In brief

Routing introduces conditional logic into an agent’s execution flow: instead of following a fixed path, the system evaluates context and dynamically selects the most appropriate action, tool, or sub-agent. It is the mechanism that transforms a static executor into an adaptive system. Four main mechanisms coexist — LLM-based, embedding-based, rule-based, and ML classifier — each with different cost, speed, and flexibility profiles.


Why routing is structurally important

A linear pipeline (prompt chaining) processes all requests the same way. As soon as a system must handle heterogeneous requests — a billing question, a technical inquiry, an incident report — rigidity becomes a problem. Routing solves this limitation by adding a conditional decision layer.

Gulli (2025) puts it this way: “Routing introduces conditional logic into an agent’s operational framework, enabling a shift from a fixed execution path to a model where the agent dynamically evaluates specific criteria to select from a set of possible subsequent actions.”

Routing is not a surface-level optimization. It is an architectural requirement for any system that must respond to variable inputs with differentiated processing.


The four mechanisms

1. LLM-based routing

The language model itself analyzes the input and produces a route identifier or routing instruction. Advantage: fine-grained natural language understanding, handling of ambiguous formulations. Drawback: latency and cost at each routing decision, non-deterministic behavior.

Typical use: intent classification on complex, ambiguous, or multilingual requests where an explicit rule is difficult to write.

2. Embedding-based routing

The request is converted into an embedding vector, then compared by cosine similarity to embeddings representing each available route or capability. The route whose embedding is semantically closest is selected.

Advantage: semantic flexibility without an LLM call per request. Drawback: quality depends on the embedding model and the construction of reference embeddings. Sensitive to formulations far from training examples.

Typical use: systems with a large number of routes, where semantic similarity is a relevant criterion.

3. Rule-based routing

Predefined rules (if-else, switch, keyword regex, structured data checks) determine routing. This is the fastest and most deterministic mechanism.

Advantage: execution speed, near-zero cost, entirely predictable and auditable behavior. Drawback: rigidity — every new variant requires an explicit rule, and unforeseen formulations are not covered.

Typical use: categorization based on strict and stable criteria (document type, detected language, numeric value range).

4. ML classifier

A specialized discriminative model, trained on a labeled corpus, makes the routing decision. The logic is encoded in the model’s weights, not in a prompt. Gulli mentions using LLM-generated synthetic data to build this training corpus: the LLM participates in the offline phase, not in real-time inference.

Advantage: fast inference, high performance on domains covered by training, low marginal cost at scale. Drawback: requires a labeled corpus, retraining if the domain evolves, less adaptable than LLM-based for out-of-distribution cases.

Typical use: high-volume systems where latency and per-request cost are critical.

In short: the 4 mechanisms range from fastest/most rigid (rule-based, < 1 ms per decision) to most flexible/expensive (LLM-based, 100-500 ms per decision). None is “best” — they correspond to different constraints. In large-scale production, you often combine: rule-based as first pass for obvious cases, LLM-based as fallback for ambiguous cases.


Input routing vs. state-based routing

This distinction is often overlooked but changes the architecture.

Input routing: the decision is made solely from the incoming request. This is the case for the four mechanisms described above. The system evaluates what arrives and selects a path.

State-based routing: the decision depends on the accumulated global state of the system, not just the last input. This is the paradigm of stateful graphs (LangGraph). Gulli notes: “LangGraph’s state-based graph architecture is particularly well-suited for complex routing scenarios where decisions are contingent upon the accumulated state of the entire system.”

A concrete example: a support agent that has already tried two solutions without success should not propose a third solution of the same type — state-based routing allows it to read the session history and escalate to a human. Input-only routing would simply see the last question and ignore the accumulated context.


Implementation patterns

Coordinator-Delegate

A coordinator agent receives all requests and routes them to specialized sub-agents (e.g., billing agent, technical agent, general support agent). Each sub-agent has discrete capabilities and a context adapted to its domain.

This pattern is the most common in multi-agent architectures. It implies that the coordinator must be reliable: a routing error sends the request to the wrong specialist.

Multi-level routing

Routing applies at multiple levels of the operation: initial request classification, intermediate decisions in the processing chain, tool selection within a sub-routine. Each level can use a different mechanism depending on its speed and accuracy constraints.

Cognitive dispatch

A special case of routing: directing a task to the language model whose size is appropriate for the type of cognition required. A simple extraction goes to a lightweight model (fast, inexpensive); contradiction analysis goes to a heavier model. The routing decision encodes a resource policy.


Choosing the right routing mechanism

ContextRecommendationWhy
Stable categorisation, strict criteria (detected language, format)Rule-based< 1 ms per decision, full auditability, zero cost
100+ routes, semantic similarity relevantEmbedding-basedNo LLM call per request, scalable, keeps its flexibility
High volume, stable domain, corpus availableML classifierFast inference, low marginal cost, strong in-domain performance
Ambiguous queries, unforeseen phrasingsLLM-basedFine-grained understanding, handles nuance — accept the cost and latency
Contextual decisions (multi-turn session, escalation)State-based routing (LangGraph)Takes the whole history into account, not only the last input

Strengths and limitations

DimensionStrengthsLimitations
LLM-basedNuanced understanding, zero initial configurationLatency, cost, non-deterministic
Embedding-basedSemantic flexibility, scalableQuality depends on reference embeddings
Rule-basedSpeed, full auditabilityRigid, manual maintenance
ML classifierFast, low marginal costRequires corpus, retraining if drift
State-based routingContextually coherent decisionsIncreased complexity, harder to debug

The choice of mechanism is not absolute: production systems often combine multiple levels. A fast rule-based filter eliminates trivial cases before passing ambiguous ones to the LLM.


Key takeaways

  • Routing is what transforms a linear pipeline into an adaptive system: without it, all inputs receive the same treatment regardless of their nature.
  • Four mechanisms coexist with different profiles — rule-based (deterministic and fast), embedding-based (semantic flexibility), ML classifier (scalability), LLM-based (fine-grained understanding) — and can be combined.
  • The distinction between input routing and state-based routing is architectural: the latter requires maintaining an explicit global state (LangGraph paradigm) and enables coherent decisions across an entire session.
  • Each mechanism has its limits: rule-based breaks on unforeseen cases, LLM-based is costly, ML classifier drifts if the domain evolves.
  • Routing is not an optimization layer added after the fact — it is a design decision that determines the system’s ability to handle input variability in production.