In short

Guardrails are the set of control mechanisms that frame the behavior of an LLM agent in production. They are organized into five layers: input validation, output filtering, behavioral constraints, tool restrictions, and checkpoint/rollback. Without these layers, an autonomous agent is unpredictable and potentially dangerous. With them, it becomes reliable and deployable.

In short: think of a fortress — not a prison, a system of concentric defenses. The drawbridge filters who enters (input validation). The moat protects what leaves (output filtering). Internal regulations limit even the occupants’ behavior (behavioral constraints). Armories restrict who touches the weapons (tool restrictions). And the archive logs everything to allow a return to a previous state if needed (checkpoint/rollback). The aim is not to prevent occupation: it is to make it sustainable.


The problem: the unpredictability of autonomous agents

An LLM agent in production does not only handle benign queries. It receives adversarial inputs designed to divert it from its instructions (jailbreaking), generates outputs that may violate legal or ethical rules, and sometimes has access to tools capable of irreversible actions — modifying a database, sending an email, triggering a command.

The problem is not the power of the model. It is the absence of a control structure around it. “Without proper guardrails, an AI system may be unconstrained, unpredictable, and potentially hazardous.” (Gulli, 2025, Ch.18)

Guardrails address this problem through a systemic approach: several independent lines of defense, each targeting a distinct risk vector.


The five layers of defense

Layer 1 — Input validation

Before a request reaches the main model, it passes through an input filter (input validation/sanitization). This filter detects jailbreak attempts, prompt injections, out-of-domain content, and structurally malformed requests.

A common technique is dual-model screening: a lightweight, fast model (for example Gemini Flash) handles the request on the front line. If it passes this initial filter, it is forwarded to the main model. This approach preserves performance while reducing the attack surface.

Layer 2 — Output filtering

Output filtering takes place after generation, before transmission to the user. It detects toxicity, discriminatory biases, regulated content (definitive medical advice, formal legal opinions), and factually incorrect information.

In content generation systems, this layer checks compliance with legal guidelines and identifies misinformation. In educational tutors, it blocks incorrect answers and inappropriate conversations.

Layer 3 — Behavioral constraints

Behavioral constraints define what the agent can and cannot do, independently of what the model is technically capable of producing. They are mainly implemented through structured prompting.

A legal research assistant may be technically capable of producing a precise legal opinion. The behavioral constraint forbids it from doing so — not through incapacity, but through an explicit policy encoded in its system instructions.

These constraints frame without restricting capabilities: the agent remains competent in its domain, but operates within defined operational limits.

Layer 4 — Tool restrictions

The least privilege principle applies to agents exactly as it does in classic information security: grant the agent the minimal set of permissions required to accomplish its task, and nothing more.

An agent tasked with querying a database should not have access to write functions. A document summarization agent does not need network permissions. Each unnecessary permission is an additional risk vector.

This layer is implemented at the orchestration framework level: the list of available tools is defined by configuration, not by the model itself.

Layer 5 — Checkpoint and rollback

The checkpoint/rollback pattern is borrowed from software engineering: each intermediate state of an agent is validated before moving to the next step. If an anomaly is detected, the system can return to a previous state known to be correct.

“The checkpoint and rollback pattern is a perfect example of software engineering principles applied to agents: each checkpoint is a validated state, a successful ‘commit’ of the agent’s work.” (Gulli, 2025, Ch.18)

This mechanism is particularly critical for long-running agents or those performing actions with side effects (writing, sending, deployment).

In short: the five layers are not redundant — they cover different risk vectors. An attacker who bypasses the input filter (jailbreak) will be stopped by behavioral constraints. If those drift, tool restrictions prevent destructive actions. And if despite everything an incorrect action passes through, rollback restores the system to a healthy state. Defense in depth is exactly that reasoning formalized.


Structured validation: Pydantic as a reliability tool

One technical point deserves attention: structured output validation. Forcing an agent to generate its responses within a predefined schema — for example via Pydantic in Python — guarantees two properties simultaneously.

First, structural compliance: the output can be consumed automatically by the next pipeline stage without defensive parsing. Second, anomaly detection: an output that does not satisfy the schema signals a behavioral deviation, even if the textual content appears correct.

from pydantic import BaseModel, Field

class AgentResponse(BaseModel):
    answer: str = Field(..., max_length=500)
    confidence: float = Field(..., ge=0.0, le=1.0)
    sources: list[str] = Field(default_factory=list)
    is_out_of_scope: bool = False

This type of schema forces the agent to explicitly declare whether it considers a request out of its scope — turning a silent drift into a detectable signal.


A dedicated guardrail: the policy model

An advanced approach consists of deploying an LLM dedicated to policy enforcement (policy-based guardrail). This model is given an explicit “Content Policy Enforcer” role: it evaluates each input or output against a set of formalized rules before it reaches or leaves the main agent.

The advantage is separation of concerns: the main agent focuses on its task, the policy model handles compliance. Business rules can evolve independently from the main model, without requiring fine-tuning.

The limit is the latency and infrastructure cost of an additional LLM call on every exchange.


Why guardrails make agents deployable

A common misunderstanding presents guardrails as a restriction of capabilities. The opposite is more accurate: an agent without guardrails cannot be deployed in production, which makes it useless outside controlled environments.

“Guardrails are not meant to restrict an agent’s capabilities but to ensure its operation is robust, trustworthy, and beneficial.” (Gulli, 2025, Ch.18)

Guardrails solve the trust problem required for deployment: legal teams, compliance teams, and end users need structural guarantees on the agent’s behavior. These guarantees cannot come from the model alone — they have to be architectural.


Limits and tensions

Guardrails are not free. Each layer adds latency. Input filters can produce false positives and reject legitimate requests. Overly strict Pydantic schemas can cause validation failures on outputs that are correct but slightly unexpected.

There is also a tension between robustness and flexibility: behavioral constraints that are too rigid reduce the agent’s usefulness in edge cases. Fine-tuning each layer is empirical work, not an initial one-time configuration.

Finally, classic guardrails do not protect against all sophisticated adversarial attacks. Jailbreak techniques evolve. A guardrail system must be maintained, audited, and updated — it is an operational discipline, not a one-off deployment.


What to remember

  • Guardrails form a multi-layered defense: input validation, output filtering, behavioral constraints, tool restrictions, checkpoint/rollback.
  • The least privilege principle applies to agents as it does to any system: each unnecessary permission is a risk vector.
  • Structured validation (Pydantic) turns silent drifts into detectable signals.
  • An LLM dedicated to policy enforcement separates business logic from compliance logic.
  • Guardrails do not limit the agent’s capabilities: they are the necessary condition for its deployment in production.