In short
HACRL (Heterogeneous Agent Collaborative Reinforcement Learning) is a training framework in which multiple LLMs of different sizes or families learn simultaneously on the same tasks by sharing their reasoning trajectories. The associated algorithm, HACPO, achieves on average +3.3% over the GSPO baseline while consuming half as many rollouts. Gains are concentrated on difficult tasks: +13.7% on AIME2025 and +12.5% on AMC23 for the Qwen3-4B model.
In short: imagine two students of different levels lending each other their mock exam drafts. The strong one sees in the weaker one’s copy paths they would not have explored. The weaker one sees in the stronger one’s copy trajectories they would not have found alone. Both progress simultaneously, and the pile of copies to grade is cut in half. HACRL does the same with LLMs: verified rollouts are shared between models of different sizes, and each one benefits.
The problem: training heterogeneous agents without waste
LLM reinforcement training relies on rollout generation: the model produces responses, a verifier scores them, and the gradient is computed from these evaluations. This process is costly — each rollout consumes compute.
A naive approach to improving a set of models would be to train each one separately, or to multiply the number of rollouts per agent. HACRL starts from a different observation: if two models work on the same task, the verified rollouts of the first contain information that the second can exploit, and vice versa. A correct rollout from a strong model shows a weak model a trajectory it would not have found on its own. An incorrect rollout from a weak model can signal to a strong model regions of the response space it does not naturally explore.
Sharing these trajectories between heterogeneous agents reduces the total cost by a factor of N (for N agents) without loss of coverage.
HACPO: four mechanisms for stable transfer
Sharing rollouts between different models creates a technical problem: trajectories generated by agent j are off-policy for agent k. Using these trajectories directly in the gradient computation introduces a potentially destabilizing bias.
HACPO (the concrete implementation of HACRL) addresses this problem with four distinct mechanisms:
1. Weighting by relative capability
A coefficient ω_t^(k,j) = P̂_t^(k) / P̂_t^(j) weights received rollouts according to the current performance ratio between the two agents, estimated over a sliding window of K=5 batches. Without this coefficient, the performance of the Qwen3-1.7B-Base agent drops by 3.1% on average (ablation Table 3).
2. Per-agent advantage estimation
Each agent maintains its own reward baseline μ̂_t^(k), distinct from the others. Using a shared baseline without weighting degrades performance by 1.5% to 2.8% depending on the agent (ablation Table 2). This separation guarantees that the estimated advantage remains unbiased for each model individually (Theorem 4.1 of the paper).
3. Exponential importance sampling
The sequential importance ratio s_{t,i}^(k,j) = [π^(k)(y^(j)) / π^(j)_old(y^(j))]^(1/|y|) corrects the distribution shift between the policies of the two agents. The parameter α controls the aggressiveness of this correction; its optimal value varies by agent pair tested (α=3.0 for Qwen3-8B-Base, α=1.0 for Qwen3-4B-Base in the same experiment).
4. Asymmetric per-step clipping
Clipping bounds s_{t,i}^(k,j) ∈ [1.0 − δ, 1.0] tighten progressively over intra-batch updates. The upper bound is fixed at 1.0: the model cannot “over-exploit” positively the trajectories of another agent. Without this mechanism, the ablation (Figure 4) shows severe training instability.
In short: the four mechanisms address the same problem — when a student copies from a neighbor, you have to prevent them from taking everything blindly. Weighting corrects level differences (the too-strong neighbor will not dictate everything). The separate baseline keeps the evaluation clean for each student. Importance sampling adjusts for different reasoning methods (Llama and Qwen do not think alike). Asymmetric clipping bounds excessive enthusiasm. Without these filters, training diverges.
Results on mathematical reasoning
All experiments use 7,500 MATH questions as training data, with G=8 rollouts per prompt. Evaluations cover seven benchmarks: MATH-500, MATH, GSM8K, AIME2025, AMC23, Minerva and OlympiadBench.
Homogeneous pair (same family, different sizes): Qwen3-1.7B-Base + Qwen3-4B-Base
- The 1.7B model goes from 0.467 to 0.493 AVG (+5.6%)
- The 4B model goes from 0.578 to 0.601 AVG (+4.0%)
- Both agents improve simultaneously
Specialized pair: Qwen3-4B-Base + Qwen3-4B-Instruct
- The 4B-Base model goes from 0.684 to 0.755 AVG (+10.4%)
- The 4B-Instruct model goes from 0.795 to 0.813 AVG (+1.8%)
- The strongest agent also benefits from the transfer, including from a less performant base model
Cross-family pair: Qwen3-4B-Base + Llama3.2-3B-Instruct
- Qwen3-4B-Base: from 0.578 to 0.597 AVG (+2.0%)
- Llama3.2-3B-Instruct: from 0.351 to 0.390 AVG (+4.0%)
- Different tokenizers and parameter spaces — transfer remains positive for both
The comparison with GSPO×2 (a single agent trained with double the rollouts) shows that HACPO performs better on the majority of benchmarks despite a lower total cost. Diversity of source produces more signal than homogeneous volume.
In short: the numbers confirm the pedagogical intuition. On AIME2025 (Olympic math benchmark), Qwen3-4B goes from 0.28 to 0.42 — a gain of +13.7% in absolute value, half the gap between a naive and an expert model. On GSM8K (easy problems), the gain falls to +0.8% — there is almost nothing to learn. The harder the problem, the more useful trajectory sharing between agents becomes.
Limits to keep in mind
Several constraints delimit the scope of these results:
- All experiments involve exactly two agents. Behavior with N>2 agents is not empirically tested. Extrapolation is not guaranteed.
- The domain is solely mathematical reasoning. The benchmarks used (MATH, GSM8K, AIME, AMC, Minerva, OlympiadBench) do not cover code, text comprehension, or factual reasoning.
- The optimal value of α (importance sampling) varies by pair. There is no universal value, which requires calibration per agent configuration.
- HACRL is a training framework, not an inference framework. Deployed agents operate entirely independently — the benefit is acquired during fine-tuning, not at inference time.
- The statistical independence assumption (Assumption D.1) underlying the theoretical guarantees is only satisfied asymptotically, not in finite regimes.
Key takeaways
- HACRL enables LLMs of different capabilities to train jointly by sharing their verified trajectories, reducing rollout cost by a factor equal to the number of agents.
- HACPO achieves +3.3% on average over GSPO with half the rollouts, and outperforms GSPO×2 (equivalent resources) on the majority of benchmarks.
- Gains are asymmetric: they are more pronounced on difficult tasks (AIME2025: +13.7%) than on easy tasks (GSM8K: +0.8%).
- Cross-agent transfer benefits both participants, including the strongest agent which receives trajectories from a weaker model.
- Experiments are limited to agent pairs on mathematical reasoning — generalization to other domains and to N>2 agents remains to be demonstrated.