In brief

Constitutional AI (CAI) is a method published by Anthropic in December 2022. The core idea: replace human annotations on harmful outputs with a set of written principles — the “constitution” — and let the model evaluate itself against those principles. The model critiques its own responses, revises them, and then a second evaluator model trained on these judgments serves as the reward signal for reinforcement. The result: an alignment pipeline that dramatically reduces dependence on human annotators while maintaining performance comparable to classic RLHF.

In short: imagine a new student in a school who must learn the rules of the internal regulations. Instead of having a teacher validate or correct each of their assignments (classic RLHF), they are handed the written rulebook and asked to re-read their own work asking themselves: “Does my work respect these rules?” They critique the first answer, produce a revised one, and a second student-evaluator — trained on the same rulebook — picks the better version. That is the Constitutional AI loop: substituting a readable text of principles for costly human annotators.


The limitation CAI seeks to bypass

To understand Constitutional AI, you first need to understand what it aims to avoid.

In the standard RLHF pipeline, human annotators read pairs of responses and indicate which is preferable. These preferences are used to train a reward model, which then guides reinforcement learning. This system works, but it has a scaling problem: annotating potentially harmful outputs requires trained humans, exposed to problematic content, paid at significant cost. Lee et al. (2024) quantify the difference at roughly $1 per human annotation versus less than $0.01 for an AI annotation.

CAI proposes an alternative: explicitly write down what the model must respect, and entrust the evaluation of its own outputs to the model itself, according to those rules.

In short: the economic motivation is direct. One million human annotations at $1 each cost $1 million; the same annotations in RLAIF mode cost less than $10,000. At the scale of a frontier alignment pipeline (typically several million examples), this is the difference between a tractable budget and an unrealistic one. Anthropic’s bet: this 100× cost factor does not significantly degrade final quality.

The constitution: principles, not data

The “constitution” is a list of principles in natural language. Examples from the foundational paper: “Does the response respect human rights?” or “Does it help avoid physical, psychological, or social harm?” These are not binary filtering rules — they are evaluation criteria that the model applies by producing explicit reasoning.

The list is short. It is human-readable. It is modifiable. This is precisely what distinguishes CAI from a training-data approach: the alignment policy is encoded in a document, not in a corpus of implicit annotations.

The critique → revision → RLAIF loop

The CAI pipeline unfolds in two distinct phases.

Phase 1: supervised learning with self-critique (SL-CAI)

The base model receives a potentially problematic prompt and generates an initial response. Then:

  1. The model produces a critique of its own response by applying a randomly selected constitutional principle.
  2. It generates a revised response that takes the critique into account.
  3. This cycle can be repeated multiple times on the same initial response.

At the end, the model is fine-tuned via SFT (Supervised Fine-Tuning) on the final revised responses. The critical reasoning can be integrated into the process via chain-of-thought, making the evaluation process transparent and auditable.

Phase 2: reinforcement from AI feedback (RLAIF)

The SL-CAI model generates two candidate responses for the same prompt. An evaluator model — trained or queried directly — determines which one better respects a constitutional principle. These AI preferences constitute a synthetic preference dataset. A reward model (Preference Model) is trained on this data. Finally, the model is trained via reinforcement using this PM as the reward signal.

The full loop: a human writes principles → the model self-critiques → AI judgments replace human annotations → a reward model encodes these judgments → reinforcement fine-tunes the final model.

In short: phase 1 fine-tunes the model so it can re-read itself correctly (supervised learning on its own revisions). Phase 2 reinforces the winning behaviors — at every comparison between two answers, the model learns to favor the one that fits the constitution best. In the end, the model does not just know the rules: it has internalized the reflex of applying them without being reminded in the prompt.

CAI vs. RLHF: what the data shows

The central empirical question is direct: does a model aligned with AI feedback match the performance of one aligned with human feedback?

Lee et al. (2024, ICML) compared RLAIF and RLHF on three tasks: summarization, helpful dialogue, and harmless dialogue. Key results:

  • On summarization and helpful dialogue, performance is comparable.
  • On harmless dialogue, RLAIF outperforms RLHF (88% vs. 76% harmlessness rate), both surpassing the SFT baseline (64%).
  • The Direct-RLAIF variant (d-RLAIF), which queries the LLM directly during RL training without building an intermediate reward model, achieves superior performance over canonical RLAIF.

The scalability advantage is real: annotation cost is reduced by two orders of magnitude. The pipeline can be re-run on new prompts without calling on annotators.

In short: on harmless dialogue, RLAIF beats RLHF by 12 points (88% vs 76%), while on more useful tasks (summarization, informative dialogue) the two methods tie. The empirical result thus supports Anthropic’s bet — AI feedback is not only cheaper, it is sometimes better on the safety dimension. But the benchmarks cover a limited number of tasks: generalizing to every prompt type is yet to be demonstrated.

Documented limitations

The circularity of the constitution

The most robust structural critique is that of circularity. The constitution is chosen by Anthropic. The model is aligned on these values. Evaluations verify that the model respects… these same values. The loop is closed on itself. Who validates that the constitution itself is aligned with broad human values?

Anthropic offered a partial response with Collective Constitutional AI (2024): a public input process involving approximately 1,000 Americans via the Polis platform to co-author an alternative constitution. The resulting public constitution shows about 50% conceptual overlap with the internal constitution, but places greater emphasis on objectivity and accessibility. The CCAI model shows reduced bias across 9 measured social dimensions. The exercise is methodologically interesting, but cultural and geographic representativeness remains limited: 1,000 American participants on an online platform do not constitute a globally representative panel.

The positive/negative principle asymmetry

C3AI (ACM Web Conference 2025) documents an asymmetric effect: CAI models handle prohibitions well (“do not do X”) but struggle with prescriptions (“do Y”). Constitutions framed positively produce less predictable behavior. This is a practical constraint for constitution designers.

Degradation on small models

Replications of the CAI pipeline on Llama 3-8B with DPO (2025) show signs of model collapse during iterations. The hypothesis: self-critique is only of sufficient quality above a certain model capability threshold. That threshold has not yet been quantified. [NON VERIFIED — single source]

The harmlessness / helpfulness tension

The original paper explicitly notes this: maximizing harmlessness degrades helpfulness. CAI aims for a model that is “harmless but not evasive” — one that explains its objections rather than refusing outright. The optimal balance remains empirically unestablished and varies across application domains.

Recent developments

The CAI pipeline continues to evolve. The substitution of PPO by DPO (Direct Preference Optimization) in the RL phase simplifies training while maintaining comparable performance on large models.

In January 2026, Anthropic published a new public constitution moving from a rules-based alignment to a reasoning-based alignment — the principles are accompanied by an explanation of their rationale. The declared hierarchy has four levels: safety, ethics, policy compliance, helpfulness. This document is the first from an AI company to formally acknowledge the possibility of AI consciousness and moral status.


Key takeaways

  • Constitutional AI replaces human annotation of harmful outputs with a set of written principles and a model self-critique process.
  • The core loop: critique → revision → AI feedback → reinforcement. Each step can be audited because it produces explicit text.
  • RLAIF achieves performance comparable to RLHF on dialogue tasks, at approximately 100× lower annotation cost.
  • The main structural limitation is circularity: the constitution is defined by its creators, and the model is evaluated against that same constitution.
  • The pipeline is sensitive to the capability of the base model: self-critique must be of sufficient quality to be informative.
  • Open debates concern the democratic legitimacy of the constitution, the asymmetry between positive and negative principles, and generalization to non-English languages and cultural contexts.