In short
In a multi-agent architecture, the orchestrator does not only delegate tasks — it also chooses the model that executes them. Using a powerful model to read a file and extract ten facts is wasteful; using a lightweight model to detect contradictions between sources produces insufficient results. Cognitive dispatch is the decision to assign each sub-task to the model whose reasoning level exactly matches what the task requires.
The problem: one model is not enough
In plain terms : using Opus for 100% of tasks costs on the order of $15 per million output tokens [order of magnitude estimate]; using Haiku costs on the order of $1.25 per million tokens [order of magnitude estimate]. A pipeline that dispatches 80% of traffic to Haiku cuts the bill by roughly 5× to 10× with no measurable loss on mechanical tasks.
A multi-agent system can easily deploy dozens of sub-agents in parallel. If every agent runs on the most powerful available model, cost explodes without quality improving proportionally. Conversely, uniforming on a lightweight model reduces cost but introduces errors on tasks that require deep reasoning.
The problem is not technical — APIs allow model selection per call. It is conceptual: the orchestrator must categorize each task according to what it demands cognitively, not according to how important it seems.
An extraction task (read a document, identify the dates, name them) requires no reasoning; it requires precision and speed. An analysis task (compare two positions, detect a contradiction, formulate a recommendation) requires depth. The distinction is clear in theory; it is often blurred in practice because orchestrators tend to overestimate the difficulty of their own sub-tasks.
This overestimation bias is systemic: pipeline architects tend to dispatch all tasks on the model they use themselves to design the pipeline, out of familiarity. The result is a costly system where 80% of sub-agents do work that the cheapest model would have handled just as well.
Cognitive dispatch: principle and grid
In plain terms : three task classes, three model tiers. Extraction → lightweight. Analysis → intermediate. Orchestration/long context → powerful. Each step costs 4× to 12× more than the previous one [order of magnitude estimate] — the choice must therefore be justified, not automatic.
Cognitive dispatch rests on a simple taxonomy: classify each sub-task by the type of operation it requires, then associate that type with a model tier.
| Operation type | Required tier | Examples |
|---|---|---|
| Extraction, summarization, categorization | Lightweight | Read a file and extract N facts, list entities, sort items |
| Analysis, synthesis, drafting, contradiction detection | Intermediate | Compare sources, recommend, design a structure, write |
| Complex coordination, long sessions (>100K tokens), high-context tasks | Powerful | Orchestration, context isolation, multi-step reasoning over large corpus |
This grid is a starting point, not an absolute rule. The boundary between extraction and analysis is sometimes blurry: summarizing a dense text is closer to analysis than extraction. In ambiguous cases, moving up a tier is less risky than moving down.
Haiku, Sonnet, Opus: practical application
In plain terms : Haiku to read and extract, Sonnet to analyze and draft, Opus to orchestrate and reason over heavy corpora. The matrix below crosses context and recommendation — it is the operational grid to keep in sight when designing a Pieuvre.
| Context | Recommendation | Why |
|---|---|---|
| Mechanical extraction, lint, scan, deterministic parsing | Haiku effort:low (thinking off) | Minimal cost (~$1.25/1M output tokens [order of magnitude estimate]), low latency, no reasoning required. |
| Standard drafting, single-source analysis, narrative summary | Sonnet effort:medium | Capability/cost balance (~$3/1M output tokens [order of magnitude estimate]), required analytical quality. |
| Multi-source synthesis, strategic audit, contradiction detection | Sonnet/Opus effort:high, thinking on | Multi-step reasoning justifies the extra cost; thinking mode adds ~30% output tokens [order of magnitude estimate]. |
| Pieuvre orchestrator, session > 100K tokens, long context | Opus effort:high, thinking on | System coherence ceiling; cost (~$15/1M output tokens [order of magnitude estimate]) offset by avoiding coordination errors. |
| Consolidation agent (N outputs → synthesis) | Sonnet effort:high minimum | Analytical task disguised as aggregation — Haiku produces flat syntheses. |
Claude offers three models that illustrate this hierarchy well. Haiku is the lightweight model: fast, low-cost, suited for repetitive processing tasks. Sonnet is the intermediate model: it covers the vast majority of analytical tasks in a multi-agent system. Opus is the powerful model: relevant for the orchestrator itself or for complex coordination tasks, particularly when the active context exceeds 100K tokens.
In practice, on corpora of several dozen sources, Haiku efficiently handles extraction (read, categorize, list) while Sonnet is necessary for in-depth documentary research — where analytical depth determines the quality of the final synthesis.
Opus as a sub-agent remains rare. Its main advantage manifests in very long contexts, where lighter models lose coherence. For most current architectures, the orchestrator is the only role that justifies Opus by default.
There is an intermediate case often overlooked: consolidation agents. When an agent must aggregate the outputs of ten Haiku agents and produce a coherent synthesis, it is doing analysis — not extraction. The dispatcher must treat the consolidation agent as an analytical agent, even if technically it only reads Markdown files produced by the others.
What the literature documents
Work on the efficiency of multi-agent systems converges on one principle: the quality of a pipeline does not depend on the most powerful model it uses, but on the fit between each task and the model handling it [UNVERIFIED — practical consensus not yet formalized in a single benchmark]. Cost-quality studies on GPT-4 vs GPT-3.5 in agent architectures show that selective (rather than global) model replacement is the approach that minimizes degradation while reducing costs.
Research on hierarchical agents also documents the importance of the model in orchestration roles: an orchestrator weaker than its sub-agents produces incoherent pipelines, where local agents do good work but the final integration fails. The orchestration model tier sets the coherence ceiling for the entire system.
Implementing dispatch: practical decisions
In plain terms : three sequential decisions, in that order. 1. List every sub-task. 2. Classify them (extraction / analysis / coordination). 3. Assign a model per class. Not before.
Cognitive dispatch is not a concept to apply intuitively when coding each agent call. It is structured upfront, during pipeline decomposition.
Classify tasks before choosing models
The useful sequence: first list all sub-tasks in the pipeline, then classify them according to the grid (extraction / analysis / coordination), then assign a model per class. Models are chosen only after this classification — not before.
A quick heuristic: if the task can be described by a single verb (read, list, extract, count, sort, categorize), it is extraction. If it requires two or more verbs with a logical link (read and deduce, compare and recommend, detect and explain), it is analysis.
Treat consolidation agents as analytical agents
Agents that aggregate other agents’ outputs deserve particular attention. Their task seems simple (read some Markdown, write a summary) but it is analytical: they must detect inconsistencies between sources, arbitrate contradictions, and maintain overall coherence. Dispatching a consolidation agent on Haiku means entrusting the pipeline’s most complex reasoning to the least capable model.
Define the output contract before dispatch
Each sub-task must have a defined output format before the model is chosen. That format (list of N facts, table with M columns, text of X paragraphs) constrains the necessary model tier. A highly structured output with strict constraints tolerates a lighter model better than a free-form output expecting a narrative synthesis.
Common mistakes and their consequences
In plain terms : six symmetric mistakes. Under-dispatching degrades quality; over-dispatching inflates the bill. The cost of a dispatch error rarely shows on an isolated agent — it accumulates at the scale of a batch pipeline.
Under-using a powerful model. Using Haiku for a cross-analysis task produces shallow outputs — the sub-agent lists points without detecting tensions. The error is not always immediately visible; it manifests during consolidation, when the orchestrator receives flat deliverables it must compensate for with extra work.
Over-using a powerful model. Using Sonnet or Opus to read a hundred files and extract one date each costs between 5× and 20× the price of equivalent Haiku processing, with no measurable improvement in result. At the scale of a batch production pipeline (for example 1000 extractions per day), the cost difference reaches several tens to hundreds of dollars per day [order of magnitude estimate], significant and recurring.
Dispatching on perceived importance. The most frequent mistake: choosing the model based on the strategic importance of the task rather than its cognitive nature. Extracting critical data is still extraction — it does not need Sonnet because the data matters.
Ignoring context length. Long sessions degrade coherence in lighter models before powerful ones. For tasks that require maintaining an active context of several tens of thousands of tokens, dropping one model tier can produce incoherent outputs by the end of the session.
Not defining the error criterion in advance. Without defining what constitutes an insufficient deliverable for each sub-task type, the orchestrator has no signal to detect a dispatch error. Defining the verification criterion before launching agents allows quick comparison of outputs and adjustment.
Mixing tasks in a single agent prompt. If a sub-agent receives a prompt asking it both to extract facts and deduce implications from them, the model tier cannot be optimized: you must either over-dispatch (Sonnet for extraction) or under-dispatch (Haiku for inference). Cognitive dispatch requires decomposing multi-task prompts into single-task prompts before assigning the model.
Key takeaways
- Cognitive dispatch means aligning model tier with task nature, not its importance.
- Three operation types cover the majority of cases: extraction (lightweight model), analysis/synthesis (intermediate model), coordination/long context (powerful model).
- Documentary research, even when it seems “simple,” falls under analysis — an intermediate model is the minimum.
- Dispatch errors are rarely visible on an isolated agent; they emerge during consolidation, when insufficient deliverables accumulate.
- Defining the verification criterion per sub-task type before launching agents is the most effective preventive measure.