The Agent Orchestration Stack: Fable, Astra, Sol, and ApexGenius
An operating model that matches each model to the work its capability and token cost justify, with one governed connector into Salesforce, Jira, HubSpot, and Google that every agent shares.
The complaint is usage. The cause is shape.
The complaint coming up everywhere right now is how fast usage gets burned. Limits hit early, premium tokens disappear on routine work, and the answer most people reach for is a bigger plan. Building the first iteration of ApexGenius taught a different lesson about cost efficiency, multi-agent architecture, and selecting the best model or agent for each task: usage burns fastest when one chat carries everything.
That one chat ends up holding half-formed ideas, the plan, the research, the implementation, credentials for three different apps, and the decision about whether any of it should ship. The thread gets long, it re-reads its own history on every turn, context slips, and the agent starts doing things nobody asked for because it can no longer tell what was decided from what was floated.
A person who thinks out loud makes this worse. Ten ideas go in to find the two worth doing, and a single overloaded conversation treats all ten as instructions. A smarter model in that one window does not fix it. Giving the work a shape does, and the same shape is what brings the token bill down.
This post covers the two halves of that shape, which turn out to be inseparable: how work splits across models by what each one is capable of and what its tokens cost, and how every agent reaches Salesforce, Jira, HubSpot, and Google Workspace through ApexGenius, one governed connector, instead of carrying its own keys.
What it costs when there is no shape
Three things go wrong, and they compound. Work gets duplicated or contradicts itself, because two efforts were never scoped against each other and both believed they owned the same file. Evidence gets scattered, so the honest answer to "did it work" is buried in a different window from the one being read. And access sprawls: every agent that touches Salesforce or Google gets its own key, its own connection, and its own idea of what it is allowed to do.
The last one is the quiet one. It stays invisible until something is wrong, and by then nobody can say which agent did what, under which credential, or whether anyone approved it.
There is a fourth cost that is easy to miss: tokens. An overloaded thread re-reads its own history on every turn, and a frontier model spends premium tokens on work a cheaper model would finish just as well. Shape fixes that too.
The operating model
Every piece of the shape below exists to answer one of the failures above.
A human directs. The operator talks naturally, changes their mind, and throws out ideas. They are also the persistent decision-maker. Scope, design, and release each have a gate, and a person stands at it.
Astra holds the main conversation. It remembers the whole thread, breaks the work into pieces, delegates them, watches progress, pulls the evidence back together, and reports. Astra is the only agent that talks to the operator about everything.
Fable architects and designs. When a piece of work needs structure before anyone builds it, Fable produces the architecture and the design, and that output becomes the packet the workers build from. A design at this stage is a design. It is not implemented.
Bounded workers research and implement. Each one gets a concrete assignment that does not overlap with anyone else. Research workers read within limits and return cited evidence. Implementation workers build one scoped thing from an approved packet. They do not wander.
Evidence returns to the thread. Whatever a worker produces comes back to the main conversation, where Astra integrates it and the operator can see it in one place. Nothing lives only in a side window.
Clear roles, bounded work, evidence in one thread, and a human at the gates that matter.
Give every agent in your operating model one governed door into Salesforce, Jira, and Google.
Where each model earns its tokens
Every model call has a capability and a price. The split below sends each kind of work to the model whose capability the work needs, at the lowest effort that still gets it right, so premium tokens go to judgment and cheaper tokens go to volume. The split is a working choice from one production setup, stated as such, and it will move as the models do.
Depth and judgment
Claude for architecture, synthesis, and design. Fable takes the work that needs the whole picture: architecture, long-form synthesis of research, and visual and interface design. Those are the calls where a wrong answer is expensive, so the higher token cost buys something. Opus takes the deepest reasoning, review, and the hardest worker tasks. Sonnet handles fast, focused execution when a Claude worker fits the job.
- Architecture and the packet the workers build from
- Long-form synthesis of research and evidence
- Visual and interface design, reviewed before build
- Opus for the hardest reasoning, review, and complex worker tasks
- Sonnet for fast, focused execution
- Maximum effort only for hard architecture or review calls
Bounded execution
Codex for bounded implementation, tests, and verification. Sol workers take concrete assignments from an approved packet, build them, write the tests, and verify their own work. They finish a scoped thing and report exactly what they did, and they do it at a token cost that makes four in parallel sensible.
- Concrete, non-overlapping implementation assignments
- Tests written and run inside the assignment
- Verification and a plain report of what was done
- Lower or medium effort for routine work
Maximum reasoning effort only where it earns its keep. The highest effort settings are reserved for decisions that are actually hard. Routine work runs at lower or medium effort. Most work is routine, and treating it as hard only makes it slower and more expensive.
Parallel only when the work does not overlap. Four workers running at once is useful when each one owns a separate assignment and none of them touch the same files. The moment two assignments overlap, parallel becomes a merge problem, so the work gets split until it does not.
Explicit handoffs, nothing implied. What moves between agents is a brief, an approved artifact, or evidence. Never a partial memory of a long thread. If a worker needs context, it gets it written down, which also keeps the worker's context window small and its tokens cheap.
No paying twice for the same work. No two models are asked to solve the same problem so the answers can be compared. Each piece of work has one owner, and every other model is doing something else.
The gates stay human. Scope, design, and release are still a person's decisions, whichever model did the work. The split changes who does the work. It does not change who decides.
Spend premium tokens on judgment and let every model share the same connected apps.
A human decides scope · design · release. Same gates whichever model did the work.
Why each model holds the role it holds
These are role choices from one production workflow, made on observed fit and token cost. They are stated as working choices, and they are revisited when the models change.
Fable and Claude: architecture, synthesis, and visual design
Fable holds architecture, synthesis, and visual design because that work needs the whole picture held at once, the tradeoffs kept visible, and the output turned into a coherent design another agent can build from. That is the work where the premium capability changes the result, so it is the work that justifies the premium token. It is a role assignment based on fit and cost, and it is revisited when the fit or the cost changes.
Astra, Sol, and Codex: coordination, bounded implementation, tests, and verification
Astra holds the implementation thread and coordinates the assignments. Sol workers take bounded, non-overlapping pieces from the approved packet, implement them in parallel where that is safe, write the tests, and return verification. Concrete work with a clear spec does not need the most expensive reasoning, and running several workers at once is only affordable because each one is cheap and bounded. The same human gates for scope, design, and release apply regardless.
Whichever model holds the role, it reaches your apps through one connection.
Inside the actual workspace
This is what the setup looks like on screen, because a diagram is a claim and a screenshot is at least a state. These are captures of the workspace, not proof of outcomes: nothing in them shows a deployment, a provider accepting anything, a completed OAuth flow, or a successful action in a business app.




Where ApexGenius fits: plug in once, hand it to every agent
Everything above is about roles, effort, and evidence. The part it does not solve on its own is access. Every agent in the stack needs to reach the systems where the work actually lives: Salesforce records and metadata, Jira stories, HubSpot contacts and deals, documents in Google Drive, a calendar, an inbox. If Astra, Fable, Opus, Sonnet, and four Sol workers each hold their own credentials and their own connections, that is the sprawl problem again, multiplied by eight.
ApexGenius is one governed connector to Salesforce, Jira, HubSpot, and Google that Astra, Fable, Opus, Sonnet, and Sol all share. Plug in once, hand it to every agent. No model holds app credentials. An agent asks the gateway what it can do for this account, gets back only the operations that are connected, healthy, and enabled by policy, and calls them through that one route. Account scope and tool policy are applied at the gateway, not remembered by each agent. The shared skills for go-to-market, RevOps, product, and marketing ops sit behind the same connector, so a Sol worker and Fable work the same systems the same approved way, and nothing gets rebuilt per model.
A connected app is not the same thing as an approved action. A person still decides whether a consequential write should run. Authorization means the gateway can reach the app. It does not mean anything has been done, and it does not mean anyone has agreed to it.
That is the whole differentiation, kept deliberately narrow. Plug in once, hand it to every agent, instead of a different door for every model.
The boundaries that hold
None of this works if the states blur. These lines get said out loud in the thread whenever an agent gets ahead of itself.
- Local work is not production.
- A connected or authorized app is not the same as a successful action.
- A design is not implemented.
- A prepared Jira story is not completed work.
- OAuth authorization is separate from provider validation, and both are separate from the first successful classified read or action.
- Nothing ships without human approval.
Those lines are the reason the evidence loop matters. A worker reporting that a thing is done means nothing until a person can see what state it is actually in.
Where the agents plug in
The operating model above is yours to copy. The part that keeps it governed is the connector every agent shares. ApexGenius sits between your AI and the apps where the work lives, so Claude, Codex, ChatGPT, or any MCP client reaches the same accounts through one route, with account scope, tool policy, and approvals applied once.
Today that route reaches Salesforce, Jira, Confluence, Gmail, Google Drive, Docs, Sheets, Slides, and Calendar, Apollo, Upwork, YouTube, n8n, and your own MCP servers. Connect one app or all of them. Every agent you run gets the same door, and nothing writes without your approval.
Connect your AI to the apps it needs. 7 days free · No card required.
Connect it free