The Agent Orchestration Stack: Fable, Astra, Sol, and ApexGenius

An operating model that matches each model to the work its capability and token cost justify, with one governed connector into Salesforce, Jira, HubSpot, and Google that every agent shares.

Connect it freeBrowse connectors
How work moves through the stack. A human directs and decides. Astra holds the one conversation and delegates. Fable architects and designs, with research workers feeding it evidence. Four GPT-5.6 Sol workers implement bounded, non-overlapping assignments in parallel. Evidence returns to the thread, and ApexGenius is the one governed connector every agent uses into Salesforce, Jira, HubSpot, Google, and shared skills.

The complaint is usage. The cause is shape.

The complaint coming up everywhere right now is how fast usage gets burned. Limits hit early, premium tokens disappear on routine work, and the answer most people reach for is a bigger plan. Building the first iteration of ApexGenius taught a different lesson about cost efficiency, multi-agent architecture, and selecting the best model or agent for each task: usage burns fastest when one chat carries everything.

That one chat ends up holding half-formed ideas, the plan, the research, the implementation, credentials for three different apps, and the decision about whether any of it should ship. The thread gets long, it re-reads its own history on every turn, context slips, and the agent starts doing things nobody asked for because it can no longer tell what was decided from what was floated.

A person who thinks out loud makes this worse. Ten ideas go in to find the two worth doing, and a single overloaded conversation treats all ten as instructions. A smarter model in that one window does not fix it. Giving the work a shape does, and the same shape is what brings the token bill down.

This post covers the two halves of that shape, which turn out to be inseparable: how work splits across models by what each one is capable of and what its tokens cost, and how every agent reaches Salesforce, Jira, HubSpot, and Google Workspace through ApexGenius, one governed connector, instead of carrying its own keys.

What it costs when there is no shape

Three things go wrong, and they compound. Work gets duplicated or contradicts itself, because two efforts were never scoped against each other and both believed they owned the same file. Evidence gets scattered, so the honest answer to "did it work" is buried in a different window from the one being read. And access sprawls: every agent that touches Salesforce or Google gets its own key, its own connection, and its own idea of what it is allowed to do.

The last one is the quiet one. It stays invisible until something is wrong, and by then nobody can say which agent did what, under which credential, or whether anyone approved it.

There is a fourth cost that is easy to miss: tokens. An overloaded thread re-reads its own history on every turn, and a frontier model spends premium tokens on work a cheaper model would finish just as well. Shape fixes that too.

The operating model

Every piece of the shape below exists to answer one of the failures above.

A human directs. The operator talks naturally, changes their mind, and throws out ideas. They are also the persistent decision-maker. Scope, design, and release each have a gate, and a person stands at it.

Astra holds the main conversation. It remembers the whole thread, breaks the work into pieces, delegates them, watches progress, pulls the evidence back together, and reports. Astra is the only agent that talks to the operator about everything.

Fable architects and designs. When a piece of work needs structure before anyone builds it, Fable produces the architecture and the design, and that output becomes the packet the workers build from. A design at this stage is a design. It is not implemented.

Bounded workers research and implement. Each one gets a concrete assignment that does not overlap with anyone else. Research workers read within limits and return cited evidence. Implementation workers build one scoped thing from an approved packet. They do not wander.

Evidence returns to the thread. Whatever a worker produces comes back to the main conversation, where Astra integrates it and the operator can see it in one place. Nothing lives only in a side window.

Clear roles, bounded work, evidence in one thread, and a human at the gates that matter.

Give every agent in your operating model one governed door into Salesforce, Jira, and Google.

Connect it free7 days free · No card required

Where each model earns its tokens

Every model call has a capability and a price. The split below sends each kind of work to the model whose capability the work needs, at the lowest effort that still gets it right, so premium tokens go to judgment and cheaper tokens go to volume. The split is a working choice from one production setup, stated as such, and it will move as the models do.

Claude · Fable

Depth and judgment

Claude for architecture, synthesis, and design. Fable takes the work that needs the whole picture: architecture, long-form synthesis of research, and visual and interface design. Those are the calls where a wrong answer is expensive, so the higher token cost buys something. Opus takes the deepest reasoning, review, and the hardest worker tasks. Sonnet handles fast, focused execution when a Claude worker fits the job.

  • Architecture and the packet the workers build from
  • Long-form synthesis of research and evidence
  • Visual and interface design, reviewed before build
  • Opus for the hardest reasoning, review, and complex worker tasks
  • Sonnet for fast, focused execution
  • Maximum effort only for hard architecture or review calls
Codex · GPT-5.6 Sol

Bounded execution

Codex for bounded implementation, tests, and verification. Sol workers take concrete assignments from an approved packet, build them, write the tests, and verify their own work. They finish a scoped thing and report exactly what they did, and they do it at a token cost that makes four in parallel sensible.

  • Concrete, non-overlapping implementation assignments
  • Tests written and run inside the assignment
  • Verification and a plain report of what was done
  • Lower or medium effort for routine work
Maximum effortArchitecture choices with real tradeoffs and reviews where being wrong is expensive. Fable, rarely.
Medium effortDesign passes, synthesis, and implementation with some ambiguity. Fable, Opus, or Sol, depending on the work.
Low effortRoutine implementation, tests, verification, file recovery, captures. Sol workers or Sonnet, most of the time.

Maximum reasoning effort only where it earns its keep. The highest effort settings are reserved for decisions that are actually hard. Routine work runs at lower or medium effort. Most work is routine, and treating it as hard only makes it slower and more expensive.

Parallel only when the work does not overlap. Four workers running at once is useful when each one owns a separate assignment and none of them touch the same files. The moment two assignments overlap, parallel becomes a merge problem, so the work gets split until it does not.

Explicit handoffs, nothing implied. What moves between agents is a brief, an approved artifact, or evidence. Never a partial memory of a long thread. If a worker needs context, it gets it written down, which also keeps the worker's context window small and its tokens cheap.

No paying twice for the same work. No two models are asked to solve the same problem so the answers can be compared. Each piece of work has one owner, and every other model is doing something else.

The gates stay human. Scope, design, and release are still a person's decisions, whichever model did the work. The split changes who does the work. It does not change who decides.

Spend premium tokens on judgment and let every model share the same connected apps.

Connect it free7 days free · No card required

A human decides scope · design · release. Same gates whichever model did the work.

Why each model holds the role it holds

These are role choices from one production workflow, made on observed fit and token cost. They are stated as working choices, and they are revisited when the models change.

Fable · Claude

Fable and Claude: architecture, synthesis, and visual design

Fable holds architecture, synthesis, and visual design because that work needs the whole picture held at once, the tradeoffs kept visible, and the output turned into a coherent design another agent can build from. That is the work where the premium capability changes the result, so it is the work that justifies the premium token. It is a role assignment based on fit and cost, and it is revisited when the fit or the cost changes.

Astra · Sol · Codex

Astra, Sol, and Codex: coordination, bounded implementation, tests, and verification

Astra holds the implementation thread and coordinates the assignments. Sol workers take bounded, non-overlapping pieces from the approved packet, implement them in parallel where that is safe, write the tests, and return verification. Concrete work with a clear spec does not need the most expensive reasoning, and running several workers at once is only affordable because each one is cheap and bounded. The same human gates for scope, design, and release apply regardless.

Whichever model holds the role, it reaches your apps through one connection.

Connect it free7 days free · No card required

Inside the actual workspace

This is what the setup looks like on screen, because a diagram is a claim and a screenshot is at least a state. These are captures of the workspace, not proof of outcomes: nothing in them shows a deployment, a provider accepting anything, a completed OAuth flow, or a successful action in a business app.

Wide screenshot of a Mac desktop with three regions: on the left a Claude Cowork window above a tmux grid of four Claude Code panes running Opus and Sonnet workers, in the center a browser showing the orchestration diagram page, and on the right a ChatGPT window above a Codex grid of four Astra panes at low and medium effort.
The whole desk on one screen. Left: Claude Cowork above the Claude worker grid, four panes with Opus and Sonnet workers online and waiting for delegated tasks. Center: the reference orchestration page open locally. Right: the Codex grid, four Astra panes with effort set per pane. It is a layout of who is working where, not a result.
Screenshot of the Claude Cowork app with a task list on the left and a design canvas on the right showing desktop, mobile and social artboards.
Fable on the design lead. Claude Cowork with this article's native design canvas open beside the conversation. Design work in progress. Nothing here is implemented or published.
Screenshot of a dark terminal window titled claude-grid with three empty tmux panes.
The Claude worker grid, between assignments. A tmux grid reserved for Claude workers, Opus and Sonnet. The panes are empty because no bounded assignment is running. That is the honest idle state, not a failure.
Screenshot of a dark terminal window titled codex-grid with two panes, each showing a gpt-6-astra session at low or medium effort in the ApexGenius directory with full access.
The Codex grid where Astra and the Sol workers run. Two Codex panes inside the ApexGenius repo, both running Astra, one at low effort and one at medium. Effort is set per pane, per task, which is where the token budget actually gets spent or saved. "Full access" here means local workspace access, not an approved action in any business app.

Where ApexGenius fits: plug in once, hand it to every agent

Everything above is about roles, effort, and evidence. The part it does not solve on its own is access. Every agent in the stack needs to reach the systems where the work actually lives: Salesforce records and metadata, Jira stories, HubSpot contacts and deals, documents in Google Drive, a calendar, an inbox. If Astra, Fable, Opus, Sonnet, and four Sol workers each hold their own credentials and their own connections, that is the sprawl problem again, multiplied by eight.

ApexGenius is one governed connector to Salesforce, Jira, HubSpot, and Google that Astra, Fable, Opus, Sonnet, and Sol all share. Plug in once, hand it to every agent. No model holds app credentials. An agent asks the gateway what it can do for this account, gets back only the operations that are connected, healthy, and enabled by policy, and calls them through that one route. Account scope and tool policy are applied at the gateway, not remembered by each agent. The shared skills for go-to-market, RevOps, product, and marketing ops sit behind the same connector, so a Sol worker and Fable work the same systems the same approved way, and nothing gets rebuilt per model.

A connected app is not the same thing as an approved action. A person still decides whether a consequential write should run. Authorization means the gateway can reach the app. It does not mean anything has been done, and it does not mean anyone has agreed to it.

Astra, Fable, Opus, Sonnet, and the Sol workers reach Salesforce, Jira, HubSpot, Google, and shared skills through ApexGenius only. One governed connector with account scope, tool policy, and approvals; no direct agent-to-app connections.

That is the whole differentiation, kept deliberately narrow. Plug in once, hand it to every agent, instead of a different door for every model.

Connect SalesforceConnect Google WorkspaceSee all connectors

The boundaries that hold

None of this works if the states blur. These lines get said out loud in the thread whenever an agent gets ahead of itself.

Evidence boundaries
  • Local work is not production.
  • A connected or authorized app is not the same as a successful action.
  • A design is not implemented.
  • A prepared Jira story is not completed work.
  • OAuth authorization is separate from provider validation, and both are separate from the first successful classified read or action.
  • Nothing ships without human approval.

Those lines are the reason the evidence loop matters. A worker reporting that a thing is done means nothing until a person can see what state it is actually in.

Where the agents plug in

The operating model above is yours to copy. The part that keeps it governed is the connector every agent shares. ApexGenius sits between your AI and the apps where the work lives, so Claude, Codex, ChatGPT, or any MCP client reaches the same accounts through one route, with account scope, tool policy, and approvals applied once.

Today that route reaches Salesforce, Jira, Confluence, Gmail, Google Drive, Docs, Sheets, Slides, and Calendar, Apollo, Upwork, YouTube, n8n, and your own MCP servers. Connect one app or all of them. Every agent you run gets the same door, and nothing writes without your approval.

See all connectors

Connect your AI to the apps it needs. 7 days free · No card required.

Connect it free