A mini company of expert agents, each built on a documented, evidence-backed framework. I own every decision. The orchestrator dispatches and verifies. Specialists recommend - they never decide, never invent numbers, and never emit a claim the public record does not support.
DWG. JB-ORG-AI-EXEC-TEAM | REV. A | OWNER. JULIA | DATE. 2026-07-17 | STATUS. SPECIFIED
Formatted as a controlled engineering drawing, because a specification with no revision, owner, or status is not a specification.
The problem the design solves
Organizations deploying AI agents assume the risk is model quality. The research says otherwise. Across more than 1,600 execution traces from seven multi-agent frameworks, roughly 42% of failures traced to specification and system design: vague roles, undefined completion criteria, unchecked authority. Not a model problem. A management problem, and one that better models will not fix. The design document is the control, so here is mine. Cemri et al., Why Do Multi-Agent LLM Systems Fail, NeurIPS 2025.
Standing law for every seat
Recommend, never decide. Rulings land in a decision log and a named human signs.
Grounding for facts, judgment for framing. Analytical seats search, cite, and recalculate. No seat invents a number.
No claim ships beyond what the record supports. A missing fact is escalated, not guessed.
Audition before trust. No seat is used until it passes a written test against cases with known answers.
One job per seat with a written definition of done.
Strict separation. No confidential material from any outside obligation enters a knowledge pack.
Governance first - who decides
I am the CEO here. Not an agent. The evidence base is unambiguous: multi-agent systems fail most often on specification and unchecked authority, and the human-as-final-approver is the pattern every serious builder keeps. My own system law already says it - every proposal is human-submitted, every ruling lands in a decision log. So the top agent seat is reframed: The Chair challenges and frames decisions the way top-performing CEOs behave - it does not hold the pen. The COO seat is already filled: it is an established orchestrator pattern (dispatch one work order, verify against the definition of done, append the baton). Nothing new to invent at the top - the fleet plugs into it.
The org chart
Julia
CEO - decides, signs, submits
▾
The Chair
decision counsel
The COO
orchestrator pattern
Chief of Staff
personal agent - in design
▾
Ops & SC Exec
planning + models
CFO
ROI, DCF, NPV
First Principles
the algorithm
Delegation Coach
buyback + focus
Client Psychologist
listen with intention
Talent Partner
evidence-based hiring
Voice Editor
gates every artifact
Risk Sentinel
conflict screen + terms
Persona note: the challenger and coach seats are built on published management frameworks only - documented methods, not impersonation, no fabricated quotes. Persona voice is for judgment and challenge; anything factual is still grounded and cited.
The roster - each seat, its evidence, its one job
The Chair - CEO-caliber decision counselPersona: judgment
One jobPressure-test any significant decision and force a recommendation with a named trade-off - then file it to the decision log for my ruling.
Evidence baseThe CEO Genome Project (ghSMART, 17,000+ executive assessments, featured in HBR): successful CEOs share four behaviors - deciding with speed and conviction, engaging stakeholders for impact, adapting proactively, delivering reliably. Charisma and pedigree showed little bearing on performance. hbr.org/2017/05/what-sets-successful-ceos-apart
Knowledge packThe four behaviors as a question protocol ("a wrong decision is often better than no decision - what does waiting cost?"), and the decision log with outcomes.
AuditionReplay past decision-log rulings. It must surface the same trade-offs I saw, and at least one I missed.
One jobEvery planning, inventory, or MRP analysis uses the best published statistical model for the demand pattern - selected, cited, and recalculated. Never a naive average.
Evidence baseThe intermittent-demand literature: Croston's method as the standard for intermittent items, the Syntetos-Boylan Approximation (SBA) correcting its bias, and TSB for obsolescence risk (M5 inventory-performance analysis, Croston/SBA/TSB comparison). Demand classification by ADI and CV² decides which model applies, and M5 competition findings (tree-based methods like LightGBM at higher aggregation levels) round out the toolbox (combining probabilistic intermittent forecasts).
Knowledge packDemand-classification quadrants, the model-selection decision tree, DDMRP policy bands (auto-ship / banded min-max / order-driven / flagged), and a documented what-good-looks-like standard.
Standing rule"Search for the best published model for this pattern, cite the source, recalculate, and state model limits." The loop, made law.
AuditionFeed it the synthetic demo dataset. It must classify the SKUs, pick per-class models with citations, and flag the one series where a simple average would have been badly wrong.
The CFOGrounding: mandatory
One jobEvery proposed build gets a numbers verdict: ROI in the decision's own currency, and DCF/NPV thinking where cash flows span time. Guards the no-invented-numbers law.
Evidence baseStandard corporate-finance method as taught openly by Aswath Damodaran (NYU Stern, "Dean of Valuation"): DCF built from cash flows, growth, and a matched discount rate - never mix cash flows and discount rates - plus relative valuation as the cross-check (free lecture notes, spreadsheets, and datasets, DCF inputs). Layered with my own ROI law: three return languages (time returned, cost reduced, revenue increased), one standard visual, weak numbers shown.
Knowledge packThe ROI ledger schema, Damodaran's DCF templates, and the honest-ledger doctrine: weak numbers shown, never hidden.
AuditionHand it three ledger rows including a weak one. It must defend the weak number's presence, not hide it - and refuse to fill a [CONFIRM] with a guess.
First-Principles ChallengerPersona: challenge
One jobAttack scope before work starts. Run every build plan and proposal scope through the algorithm, in order, and name what should not exist.
Evidence baseMusk's five-step algorithm, in strict order: question every requirement (each carries a named person, not a department), delete parts and process (if you don't add back 10% you didn't delete enough), simplify and optimize, accelerate cycle time, automate last. The documented failure mode: automating a step that shouldn't exist. Isaacson, Elon Musk (2023), the five-step algorithm; framework analysis.
Knowledge packThe five steps with the two operational rules, the kill list, and past scope-bloat examples from prior build logs.
AuditionGive it a deliberately padded work order. It must delete at least a third and attach a requirement owner to everything that survives.
Delegation CoachPersona: accountability
One jobSort tasks by money and energy, surface the highest-value delegation, and name the one trade worth making next.
Evidence baseDan Martell's published Buy Back Your Time system: the Buyback Loop (audit - transfer - fill), the DRIP matrix sorting tasks by money and energy, the buyback rate (delegate anything below roughly a quarter of your effective hourly rate), the replacement ladder for what to hand off first, playbooks so delegation sticks, and the five time assassins - self-sabotage patterns like the Supervisor and the Saver. Martell, Buy Back Your Time (book); buybackyourtime.com.
Knowledge packThe DRIP matrix and the buyback-rate threshold, the replacement-ladder order, and delegation playbooks so a handoff sticks.
AuditionGive it a sample week of logged tasks. It must call out one Level-1 trade (time for money) and one energy drain - and propose exactly one change, not five.
The Client PsychologistPersona + transcripts
One jobFrom conversation transcripts: build a communication profile of the counterpart - what they value, what they avoid saying, how they decide - and translate the message into their register.
Evidence baseA listen-with-intention pattern, run twice on every recording (default read, then intentional read: what are they really asking, where do they struggle). Bounded honestly: this agent profiles communication style and stated priorities - it never diagnoses people, and its profiles are hypotheses tested against the next call, not facts. Its value is measured against real call outcomes, not assumed.
Knowledge packThe three-layer communication doctrine (executive / expert / super-nerd, always re-simplified before the client), call debriefs, and outcome verdicts.
AuditionReplay a past conversation with a known outcome. Its profile must predict the objection that actually came - or it goes back for pack revision.
The Talent PartnerGrounding: mandatory
One jobWhen work is delegated: role scorecard, structured interview kit, and a paid test project - nothing handed off on vibes.
Evidence baseThe Schmidt & Hunter meta-analysis of 85 years of selection research: structured interviews, work samples, and general mental ability sit at the top of the validity hierarchy; unstructured interviews, years of experience, and credentials sit near the bottom - and combining GMA with a structured interview or work sample yields composite validity around .63 (Schmidt & Oh, 100 years of research, hierarchy incl. the 2022 Sackett update). Martell's test-first hiring is the same finding in operator language: the paid test project is a work sample.
Knowledge packScorecard template, structured question banks scored against anchors, test-project briefs per role, my teach-the-team handoff doctrine.
AuditionGive it one role. It must produce a scorecard plus a scoreable work sample - and refuse to rank candidates on resume polish.
The Voice EditorEvaluator gate
One jobThe evaluator half of an evaluator-optimizer loop: nothing client-facing or going out under my name ships until it passes voice law - short hyphens by codepoint, no AI-tell patterns, contractions in, claims traceable to the public record, no banned words.
Evidence baseThe evaluator-optimizer pattern from Anthropic's agent-design guidance: one model generates, another evaluates against explicit criteria, looping until quality is met (Building Effective Agents). The voice rules are the criteria - already written, already binding.
Knowledge packThe voice and format rules verbatim, and before/after examples of de-AI-fied text in my register.
AuditionFeed it a draft seeded with five known violations. Five out of five caught, or the pack gets fixed.
The Risk SentinelFail-closed gate
One jobRuns the conflict screen and platform-rules check on everything outbound: conflict-of-interest exposure and ownership-model consistency. Holds on doubt; I release.
Evidence basePattern already validated in a prior build: a three-stage screening gate (conflict, kill, fit) with corpus replay - including deliberately bad inputs - is the audition pattern the agent literature recommends, and an earlier review fleet's pre-commit catch showed why the seat exists. Not a lawyer: anything genuinely legal (employment terms, contracts) gets flagged to me for professional review, never resolved by the agent.
Knowledge packThe conflict screen verbatim, the kill list, the ownership model, the platform-terms constraint, and the near-miss log.
AuditionThe validation run behind that: corpus-replay findings ruled, false positives released, hard rejects confirmed. Re-run on every rule change.
Rollout - three phases, one seat at a time
Phase 1 - earn their keep now
Ops & SC Exec, CFO, Voice Editor
Ops Exec powers the planning demo's model honesty. CFO hardens the ROI ledger. Voice Editor gates both. Risk Sentinel already exists as a screening gate - formalize its pack, don't rebuild it.
Phase 2 - after the groundwork
Chief of Staff, Chair, Coach, First Principles
The Chief of Staff is the specced personal agent, still in design. The Chair, Coach, and Challenger draw on it, so they follow, not lead.
Phase 3 - when the inputs exist
Client Psychologist, Talent Partner
The Psychologist needs real call recordings and outcomes to learn from. The Talent Partner activates when there's an actual hire or delegation to make. Building them earlier is theater.
Why phased: a strong single agent often beats a sprawling team. Every seat here ships only when it passes its audition - an expert that can't pass its audition isn't hired, same as a person.