In brief
RLVR (Reinforcement Learning with Verifiable Rewards) improves LLM accuracy in mathematics and code, but its binary correct/incorrect signal pushes models to converge on a single response strategy, impoverishing their exploration capacity. A recent variant — multi-answer training — asks the model to produce K candidates in a single pass and rewards the set: on the MBPP code benchmark, top-1 accuracy improves by +69% and token consumption drops by 54%. Calibration, measured by Brier score, goes from 0.55 to 0.18.
The problem: when answering better means exploring less
RLVR: the verifiable signal shortcut
RLHF (Reinforcement Learning from Human Feedback) first trains a reward model from human annotations, then optimizes the LLM policy accordingly. This pipeline is expensive and introduces a risk: human annotators reward responses that appear confident rather than those that are factually correct, pushing models toward verbal overconfidence (Leng et al., 2024).
RLVR short-circuits this intermediate reward model. For domains where ground truth exists programmatically — mathematics, code, formal logic — an automatic verifier (compiler, test oracle, calculator) directly emits a binary signal: 1 if the answer is correct, 0 otherwise. You cannot convince a compiler that a wrong program is correct. This clean signal is one of the factors that allowed DeepSeek-R1-Zero, trained via GRPO without any supervised fine-tuning phase, to reach 71.0% on AIME 2024 versus 15.6% for the base model.
The trade-off: RLVR only works where a deterministic oracle exists. Creative writing, nuanced argumentation, any subjective task remains out of reach.
The silent collapse of diversity
In 2025, several independent works documented a paradoxical effect: RLVR training improves accuracy on a single attempt (Pass@1) while degrading the probability that at least one response among K attempts is correct (Pass@k for large K).
Yang Yue et al. (NeurIPS 2025) measured this phenomenon directly: during training, Pass@1 rises from 26.1% to 42.5% on training data, while Pass@256 continuously declines. Their explanation: RLVR does not generate new reasoning patterns — it simply makes more probable the paths already present in the base model that received positive rewards, at the expense of others. Compare with distillation, which introduces genuinely new paths from a teacher model.
Li et al. (ICLR 2026) formalize the underlying mechanism. The GRPO algorithm, like PPO, uses reverse-KL divergence as regularization. This divergence is mass-seeking: it concentrates the policy on high-probability modes and actively suppresses low-probability distributions. Their alternative, DPH-RL, substitutes mass-covering f-divergences (forward-KL, JS-divergence) that force the model to maintain broad coverage. Result: Pass@1 and Pass@k improve simultaneously, and catastrophic forgetting on out-of-domain tasks decreases.
This phenomenon is also called policy entropy collapse: by the end of training, model outputs converge on a small number of nearly identical strategies.
In short: RLVR makes the model more accurate on its first attempt, but reduces the diversity of strategies it can produce. It is like a student who becomes excellent at a single method, at the expense of all the other methods they had in their repertoire before.
The multi-answer approach: reformulating the objective
Producing a set rather than choosing a mode
The central intuition of Puri et al. (MIT CSAIL, March 2026) is a reformulation of the training objective. Instead of asking the model for the right answer, it is asked to produce K candidate responses in a single pass, covering the distribution of possible correct responses. The reward function sums the corrections over the set:
R_multi_RLVR = Σ (1 if response i ∈ correct responses)
This change in objective is more profound than it appears. In standard GRPO (DeepSeek-R1), the model already generates K=16 responses per question — but rewards remain individually binary and serve only to estimate a relative within-group advantage. Each response is evaluated independently; their diversity is not rewarded. With multi-RLVR reward, the sum of corrections forces the model to explore the solution space: it cannot “choose” to converge on a single response without sacrificing reward points on the K-1 others.
A notable side effect: joint reasoning over K hypotheses in a single pass reduces redundant reasoning chains. When a model generates K responses independently (K separate calls), it often reproduces the same reasoning preamble. In joint production, it is forced to take distinct paths.
The inference precedent: self-consistency
This principle is not without precedent. Wang et al. (ICLR 2023) had shown that at inference time, generating a diverse set of reasoning chains then aggregating by majority vote substantially improves performance: +17.9% on GSM8K, +11.0% on SVAMP, +12.2% on AQuA. Self-consistency operates without modifying the model — it post-hoc exploits the diversity the model already possesses.
Multi-answer training goes one step further: it trains the model to natively produce this diversity, integrating exploration into the objective itself rather than only exploiting it at inference time.
The results
Code: MBPP
On the MBPP Python code generation benchmark:
| Metric | Single-answer RLVR | Multi-RLVR |
|---|---|---|
| Top-1 accuracy | 0.29 | 0.49 (+69%) |
| Tokens used | 512 | 235 (-54%) |
| Brier score | 0.55 | 0.18 (RLCR-Multi) |
The +69% accuracy gain accompanied by a 54% token reduction is not a trade-off — both improve simultaneously. The explanation: the multi-answer model distributes its token budget across K short, distinct paths rather than a single long path attempting to be exhaustive.
Medical diagnosis: DDXPlus
DDXPlus is a differential diagnosis benchmark — a naturally multi-answer task since a patient may present several valid simultaneous diagnoses.
| Metric | Single-RLVR | Multi-RLVR |
|---|---|---|
| Correct diagnosis coverage | 0.62 | 0.79 |
| Unique diagnoses per question | ~4 | ~8 |
| ECE (calibration) | not reported | 0.02 (RLCR-Multi) |
The ECE of 0.02 represents near-perfect calibration: when the model declares 80% confidence on a diagnosis, it is correct in approximately 80% of cases at that confidence level. This is achieved via the RLCR variant (RL with Calibration Rewards), which adds to the reward signal a penalty proportional to the gap between declared confidence and observed correctness.
Why calibration is distinct from accuracy
A model can be accurate without being well calibrated — it predicts correctly but at confidence levels that do not correspond to its actual empirical accuracy. RLHF degrades calibration because human annotators reward responses that appear confident; the model learns to project confidence independently of factual correctness. The Llama-3 and Mistral-7B models tested by Leng et al. exhibit ECEs between 0.12 and 0.40.
Binary RLVR reproduces the same problem in a different form: Puri et al. measure systematic overconfidence at all confidence levels on models trained with simple binary reward. RLCR corrects this artifact by making declared confidence costly when it is wrong.
Which training signal, and when
| Context | Recommendation | Why |
|---|---|---|
| Subjective domain (writing, dialogue) | RLHF or DPO | No deterministic oracle is possible — human feedback is irreplaceable |
| Maths/code, oracle available, single answer | RLVR (GRPO) | Clean binary signal, no annotator cost |
| Maths/code, several valid solutions | Multi-RLVR | Preserves diversity, native exploration, +69% on MBPP |
| Differential diagnosis, calibration required | RLCR (RL with Calibration Rewards) | ECE = 0.02 — stated confidence aligned with accuracy |
| Small model (≤ 7B) to fine-tune on a hard task | DPH-RL (forward-KL/JS) | Avoids diversity collapse, preserves Pass@k |
Limitations and open questions
What is not documented
Multi-answer training has been evaluated on tasks where multiple correct answers naturally exist (code with multiple solutions, differential diagnosis). Generalization to standard mathematical benchmarks (GSM8K, AIME, MATH), where the answer is unique, is not documented in this work. It is unknown whether the gain transfers or whether new pathologies emerge (forced redundancy, false diversity).
A direct comparison between Multi-RLVR and GRPO at identical computational budget also does not exist. GRPO already generates 16 responses per question — the difference in objective is conceptually clear, but the performance gap on the same basis is not the subject of a controlled comparison in the current literature.
Multi-answer reward hacking
A specific exploitation vector for multi-answer reward is not documented but is plausible: a model could learn to generate K superficially different responses (lexically distinct formulations) but identical in their reasoning structure, maximizing reward without covering new modes. This would be a form of reward hacking adapted to the multi-answer signal. The reward hacking literature (Weng, 2024) does not treat this case specifically.
DAPO (ByteDance/Tsinghua, 2025) partially addresses the problem for standard GRPO via Dynamic Sampling: questions where the model achieves 0% or 100% success are filtered from the batch, forcing training on problems with non-zero variance. This exploration-maintenance mechanism remains to be explored in the multi-answer framework.
The cost of explicit calibration
RLCR requires the model to generate explicit numerical confidence estimates for each response. This output format constraint adds a structural obligation. It is not established that this constraint is free in performance on domains where calibration has not been explicitly trained, nor how it interacts with potential forms of format reward hacking.
Key takeaways
- Binary correct/incorrect reward improves accuracy on a single attempt but impoverishes the diversity of explored strategies — this is empirically documented by the degradation of Pass@k during training.
- Multi-answer training reformulates the objective: rewarding a set of K candidates rather than a unique response forces the model to natively explore the solution space.
- On MBPP (code), Multi-RLVR achieves +69% top-1 accuracy with -54% tokens compared to the single-answer baseline; on DDXPlus (medicine), diagnostic coverage rises from 0.62 to 0.79.
- The RLCR variant, which adds a calibration penalty via Brier score, reaches ECE = 0.02 — the model’s declared confidence aligns near-perfectly with its empirical accuracy.
- Gains are established on naturally multi-answer tasks. Generalization to single-answer tasks and robustness to reward hacking specific to the multi-answer signal remain open questions.