In brief
Agent0 (Xia et al., 2025) is an autonomous training framework for LLM agents: two agents co-evolve from the same base model without ever using human-annotated data. A “curriculum” agent generates progressively harder problems; an “executor” agent learns to solve them with a Python code interpreter. On the Qwen3-8B-Base model, Agent0 achieves +18% in mathematical reasoning and +24% in general reasoning compared to the base model, outperforming several competing methods.
In short: imagine a chess coach and a student who share the same starting brain. The coach invents increasingly twisted positions, the student trains to solve them with an abacus (the Python interpreter). Over the course of games, the coach raises the difficulty to stay just above the student’s level — not too easy, not impossible. No human master ever set a problem: the pair rises on its own.
The problem it solves
Classical training of LLM agents via reinforcement learning (RLHF, RLVR) depends on large-scale annotated data: expensive, slow to produce, and bounded by what humans already know how to describe. Existing self-improvement approaches partially work around this constraint, but run into two limits:
- Complexity ceiling: tasks generated by the model rarely exceed its current capacity — the curriculum stagnates.
- Single-turn interactions: most methods do not capture the multi-step nature of real problems.
Agent0 attacks both problems simultaneously through symbiotic co-evolution and the integration of external tools.
Architecture: two agents, one loop
The two agents are initialized from the same base LLM (πbase). They then share an iterative loop:
The curriculum agent
Its role: generate problems situated exactly at the boundary of the executor’s current capabilities. It is trained by GRPO with a composite reward signal:
- Uncertainty reward (Runc): maximized when the executor answers correctly on ~50% of attempts — neither too easy nor impossible.
- Tool use reward (Rtool): proportional to the number of interpreter calls in the executor’s responses — incentivizes generating tasks that require code.
- Repetition penalty (Rrep): reduces the reward if tasks generated in the batch are too similar to each other (measured by BLEU).
The executor agent
Its role: solve the generated problems. It has access to a sandboxed Python interpreter: it can write code, receive the execution result, correct, and iterate before giving its final answer.
The executor’s training relies on ADPO (Ambiguity-Dynamic Policy Optimization), a GRPO variant designed for the case where pseudo-labels come from majority voting — and are therefore potentially noisy:
- Ambiguity advantage scaling: the gradient is downweighted for tasks where the vote is most uncertain (unreliable pseudo-label).
- Dynamic trust region: for ambiguous tasks, the clipping ceiling is relaxed, allowing larger updates to explore out-of-distribution solutions.
The filtered dataset
After each curriculum update, the executor generates k=10 responses per candidate task. Only tasks whose self-consistency score (proportion of concordant responses) falls within the window [0.25; 0.75] are retained for training — neither trivial nor unsolvable.
In short: the mechanics rest on three simple rules. The coach invents “boundary” problems (the student must succeed once out of two times). The coach earns more points if the problems force the student to code. And the coach loses points if the problems look too similar to each other. These three signals are enough to produce a difficulty that climbs without plateauing.
What distinguishes Agent0 from other zero-data approaches
| Method | External data | Tool | Multi-turn |
|---|---|---|---|
| R-Zero | None | No | No |
| Absolute Zero | None | Verification only | No |
| SPIRAL | None | No | Yes (zero-sum game) |
| Socratic-Zero | External API (OpenAI) | No | Yes |
| Agent0 | None | Execution + curriculum | Yes |
The key difference from Absolute Zero (which also uses a code executor): Agent0 explicitly rewards the curriculum for generating tasks that require the tool, not just for verifying answers with it. Without the Rtool reward, performance drops by 7.2% on average according to the ablation study.
Not to be confused with HyperAgents
HyperAgents (Darwin Gödel Machine) operates at the meta-level of improvement: the agent can modify the way it improves itself, recursively modifying its own improvement code. The level of abstraction is higher, and the method assumes an already capable agent.
Agent0 operates at the bootstrapping level: starting from a raw LLM without any specialized training, and progressively building capabilities through a self-generated curriculum. The originality lies in structured co-evolution and the integration of the tool as a vector of progress — not in meta-cognitive recursivity.
The two approaches are complementary: Agent0 addresses “how to start from scratch”, HyperAgents addresses “how to go further once capabilities are already in place”.
Empirical results
On Qwen3-8B-Base (3 co-evolution iterations):
- MATH: 78.0 → 82.4 (+4.4 pts)
- AIME25: 16.7 → 24.8 (+8.1 pts)
- MMLU-Pro: 51.8 → 63.4 (+11.6 pts)
- BBEH: 8.6 → 13.7 (+5.1 pts)
The progression is monotone across the three iterations: the curriculum agent generates tasks of objectively increasing difficulty (the success rate of an executor frozen at iteration 1 drops from 64% to 51% on datasets generated at iterations 2 and 3). The average number of tool calls per task increases from 1.65 to 2.60 between iterations 1 and 3.
Gains transfer to general tasks not seen in training (SuperGPQA, MMLU-Pro, BBEH), suggesting that the multi-step reasoning acquired on math generalizes.
In short: the coach and the student don’t just progress — they invent an increasingly demanding school. An external examiner (an agent frozen at iteration 1) drops from 64% to 51% success when presented with the problems from the following iterations. The difficulty really climbs, it’s not just a score effect on the same tests.
Limitations
- Current domain: evaluations cover exclusively mathematical and academic reasoning. Generalization to other domains (code, science, real agentic tasks) is not demonstrated.
- Pseudo-label validity: majority voting as a substitute for human labels introduces noise — ADPO mitigates the impact but does not eliminate it. On tasks with a large or ambiguous answer space, the method could reinforce systematic errors.
- Co-evolution cost: each iteration requires sequentially training both agents. The computational complexity is significantly higher than standard training.
- External judge: evaluation of executor responses uses GPT-4o as a judge for certain benchmarks — an external dependency that the training framework itself avoids, but which remains present at evaluation time [NOT VERIFIED for all benchmarks].
Key takeaways
- Agent0 co-evolves two agents from a raw LLM without any annotated data: a curriculum that generates tasks, an executor that solves them with a Python interpreter.
- The key is not having a tool, but explicitly rewarding the generation of tasks that require that tool — which distinguishes Agent0 from its competitors.
- ADPO replaces standard GRPO to handle the noise of pseudo-labels from majority voting.
- Gains (+18% math, +24% general on Qwen3-8B) hold across 3 iterations and transfer to out-of-distribution tasks.
- The main limitation remains the tested domain: only formal reasoning, not open agentic tasks.