CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning".
Jane: The paper was written by Ambuj Mehrish and Sebastiano Vascon from Ca' Foscari University of Venice.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper summary: Tom: So we have a paper from Ca' Foscari University of Venice, by Ambuj Mehrish and Sebastiano Vascon. It's about test-time reinforcement learning, which is what happens when you let a language model keep training on a test set without any ground-truth labels. The standard trick is to sample many answers, take a majority vote, and reward the model for matching that vote. The authors argue that the vote is a brittle way to extract supervision, and they replace it with a method they call CoRE.
Jane: And it reads the same rollouts differently, rather than sampling more or adding an auxiliary model?
Tom: Exactly. The rollouts become a graph, where edges connect answers that agree, and the edge weights also look at how similar the reasoning is and how confident the model was. Then a game-theory procedure called replicator dynamics finds the dominant set of mutually supporting trajectories. From that single equilibrium you get a refined pseudo-label, a graded reward for each rollout, and a gate that decides whether a question is even worth training on.
Lu: The numbers back it up. Across seven backbones and five benchmarks — 42 model-benchmark cells with three seeds each — CoRE improves the untrained base by 21 point 7 points on average, compared with 20 point 4 for majority-vote TTRL. It beats the vote by up to 7 point 5 points when answers are contested, and it wins wherever agreement is contestable while staying inside seed noise where the majority is reliable.
Meng: What I like is that it costs nothing extra. You generate the same 64 rollouts you already needed and read them more carefully. No added model, no labels, and no extra sampling — the graph and the equilibrium are just a smarter use of what you already have.
Lalam: And here's the bigger picture. The whole training loop hangs on that consensus function, because it's the only supervision the model gets. If the vote picks a wrong label, you are actively rewarding wrong answers. Showing that the consensus operator is itself a real design choice, with theory behind it, changes how you would build these adaptation loops.
Jane: It also reaches the same final accuracy as voting in 54 to 70 percent fewer training steps. That matters if you're paying for compute.
Meng: And everything stays self-supervised — ground truth only shows up in the final evaluation, never in the training loop.
Tom: Right, and the analysis shows majority voting comes back as a special case, so switching costs very little. But the paper opens with a genuine puzzle about voting that's worth sitting with for a minute.
Page 1 of the paper: Tom: Where we left off, the whole idea is that a consensus function decides what test-time RL can learn. Page one shows why the default, majority voting, is a weak choice. There's a paradox first, and it's genuinely fun: a model trained on majority votes can end up more accurate than the vote itself.
Jane: How does that not collapse everything? If the supervision is wrong more often than the model, you'd expect the model to get worse.
Tom: The answer is that model errors tend to scatter. When a question is hard, the model fails in many different ways, so the wrong rollouts disagree with each other. Even when the majority label is wrong, most incorrect rollouts vote against it and get penalized. That partial negative signal preserves something worth learning from.
Lu: But the paper identifies the failure mode where that breaks. They call it concentrated error. A systematic mistake becomes the plurality while the correct derivation stays in the minority. Then the vote rewards the wrong trajectories and penalizes the right ones, and updating on that amplifies the error across steps.
Meng: There are two more problems with a vote. A binary reward gives the same score to a well-supported derivation and a lucky guess. And every question contributes equally, whether its rollouts form a tight consensus or a fragmented set of incompatible solutions.
Jane: Because a vote compresses each trajectory down to its final answer and each answer class down to a count. You lose the model's confidence during generation, and you lose the relational structure among reasoning paths entirely.
Lalam: So voting can't even see the difference between a cluster of solutions that agree on substance and a bunch of random guesses that happen to land on one number. The paper runs the math to show you can't fix this by counting more carefully — graph structure alone rarely beats the vote, and confidence weighting alone is also limited.
Tom: Exactly, and the two signals turn out to be complementary. That complementarity becomes the entire design of the method. It also has a fairly deep ancestry in graph clustering and game theory, which is what the next page covers.
Page 2 of the paper: Tom: Page three is where the paper lays out the intellectual toolkit. The contributions are threefold: a method, a theory for when it works, and evidence across those 42 settings. But the lineage is the interesting part — there's a whole body of work on aggregating multiple samples at inference time.
Jane: Self-consistency, from the chain-of-thought literature. You sample many reasoning paths and take the majority. The paper cites a confidence-weighted variant, CISC, which cuts the number of paths needed by more than 40 percent.
Lu: But that line operates at inference only. You pick the best answer and move on. A parallel line uses learned verifiers to score or rerank samples. CoRE's shift is that it turns the aggregation itself into a training signal for reinforcement learning.
Meng: And the clustering machinery comes from an older place. A theorem by Motzkin and Straus relates maximum cliques to maximizing a quadratic function over a simplex. Then Pavan and Pelillo generalized that idea to weighted graphs, under the name dominant sets — basically maximal cliques where the weights are allowed to vary.
Lalam: The algorithm that finds them is replicator dynamics, which borrows from evolutionary game theory. You have a population of strategies, and the ones with higher payoff grow over time. A result called the Baum-Eagon inequality guarantees this iteration keeps improving at every single step.
Tom: The paper's claim is that CoRE is the first method to use dominant-set extraction for reinforcement learning rewards. That's a genuinely new borrowing, and it fits — because a dominant set is exactly a group of trajectories that mutually support each other rather than just agreeing on a label.
Jane: And there's a subtle design choice on the RL side. GRPO normalizes rewards within a group, so any per-question scaling of rewards gets cancelled out. That's why the cohesiveness gate gets applied to the loss instead of the reward.
Lu: It also means that when the gate is always one, you recover the standard objective. So someone who wants to bolt this onto an existing pipeline doesn't have to change anything else.
Meng: So the ingredients are clear. The next page shows how you actually turn 64 raw rollouts into a graph, weight the edges, and run the dynamics.
Page 3 of the paper: Tom: We know the ingredients now; the construction is on this page. Each question gives you 64 rollouts, and those become the nodes of a graph. An edge exists between two rollouts only if they give the same final answer — symbolic equivalence, not string matching. That's the hard gate that structures everything.
Jane: But the edge weight also depends on how similar the reasoning is. They use TF-IDF vectors over character n-grams and take the cosine similarity between the reasoning texts.
Tom: Right, and that similarity only applies within the same answer class. Different answers share no edge, no matter how similar the prose looks. The formula combines the answer match with a weighted reasoning term, and a floor parameter at 0 point 1 guarantees agreeing rollouts keep a minimum connection even when their derivations are lexically very different.
Lu: Then confidence enters. Each rollout gets its mean token log-probability as a confidence score, those get converted into relative weights anchored at the most confident rollout, and the temperature is set to 0 point 25. Each edge is scaled by the geometric mean of its two endpoints' weights.
Meng: Which preserves the block structure and makes sure an edge only survives when both endpoints are confident. There's also a diagonal penalty so that a single rollout can't be chosen as the consensus on its own — a singleton has a negative internal score.
Jane: The dynamics then start from the center of the simplex, giving every node an equal chance, and iterate the replicator equations until convergence. Three lemmas guarantee this is sound: the constant shift doesn't change the maximizers, the objective rises monotonically, and no singleton can win when a genuine clique exists.
Lalam: The readouts are what feed the training. The answer class holding the most equilibrium mass becomes the pseudo-label. Each rollout gets a graded reward proportional to its affinity to that consensus, so a well-supported derivation scores higher than a marginal one. And the gate measures the overall coherence of the consensus — how much the group really agrees.
Tom: One subtlety: the rewards use the uncalibrated affinity, not the confidence-weighted one, because confidence has already shaped the equilibrium. Using it again would double-count.
Lu: And the whole thing is a strict generalization of the vote. Set kappa to one and make confidence uniform, and the extracted class is exactly the plurality.
Meng: So the recipe is complete. The question is whether it actually helps, and the first big answer comes in the experimental table for the math-specialized models.
Page 4 of the paper: Tom: We have the construction and the theory, so this page asks the empirical question. The math-specialized table covers Qwen2 point 5-Math at 1 point 5 and 7 billion parameters, plus DeepSeek-Math at 7 billion, each on six benchmarks. The family average improvement over the no-RL base is 21 point 6 points for CoRE, against 19 point 9 for majority voting and 19 point 8 for the graph-only variant.
Jane: So the gap is real but not huge — which fits the theory that CoRE wins where consensus is contestable, not everywhere. And notice the graph-only arm is no better than the vote. That matches the prediction that structure alone doesn't rescue a minority.
Lu: The individual cells tell the story. On Qwen2 point 5-Math-1 point 5B, CoRE beats voting by 5 point 3 points on MATH-500 and 5 point 0 on MATH Level 4. On the 7B model the biggest margin is 7 point 5 points on GPQA. DeepSeek-Math gains 3 point 8 points on AMC.
Meng: But look at where it doesn't win. On MATH Level 5 for the 1 point 5B model, and on MATH-500 for both 7B models, CoRE trails by at most 0 point 3 points. That's inside the seed noise, and those are settings where majority voting already sits above 85 percent or around 50, with little recoverable headroom.
Jane: There's also the out-of-domain GPQA column. For the 1 point 5B model, GPQA is both a different subject and a different format — multiple choice with four options — and the model hovers near the 25 percent random-guess floor. The paper is honest about this: the benchmark is a probe of the operating envelope, rather than an in-domain evaluation.
Lalam: The consistent pattern is what matters. The gains concentrate where a coherent, confident correct minority exists to be recovered. They vanish where the vote is already right, or where nobody in the sample actually knows the answer.
Tom: And the paper ties this back to the theory rather than leaving it as anecdote. The measured wins and losses line up with the recovery threshold from the earlier analysis. That's why the next table, with vanilla and instruct models, is a real test rather than a formality.
Page 5 of the paper: Tom: The math-specialized results looked strong, and now the instruct models put that to a harder test. On LLaMA-3 point 1-8B, CoRE beats majority voting by 4 point 2 points on MATH Level 4 and 5 point 1 points on GPQA. On Mistral-Nemo-Instruct, the margins are 2 point 2 points on MATH Level 5 and 3 point 8 on GPQA, with a family average of 19 point 1 points over the base.
Jane: What strikes me is that these are the noisier backbones. Majority voting on Mistral actually hurts accuracy on several benchmarks compared with doing nothing. That's the known failure mode of self-training — a weak model locks onto systematic errors. CoRE still manages to extract signal, because it can find a coherent correct cluster even when the plurality is wrong.
Lu: The vanilla models in the same family line up too. Qwen2 point 5-7B and Qwen3-8B in non-thinking mode show a mean improvement of 24 point 5 points, with CoRE beating the vote by 6 point 2 points on MATH Level 5
Page 6 of the paper: Tom: We've been exploring how CoRE replaces the majority ballot with a graph-based consensus, and page eleven steps back to show where that machinery actually earns its keep.
Jane: It lays out three operating regimes, and the first one is blunt: at the competence floor, there's simply nothing to recover. Weak models on out-of-domain or competition-level problems produce roll-outs with no coherent structure at all, so CoRE and majority voting both just sit at the baseline.
Tom: Right, that's the 1 point 5B math model on GPQA, or Mistral-Nemo on eyeME. When nobody in the sample can solve the question, no consensus rule can manufacture signal from noise.
Jane: At the opposite extreme, you get saturation. A strong model produces nearly unanimous roll-outs, the majority ratio approaches one, and CoRE's pseudo-label ends up matching the vote anyway.
Tom: So it gracefully degrades to majority voting exactly when majority voting is already fine. The interesting territory is the middle, where a coherent, confident correct minority exists but gets outvoted.
Jane: And that's where the two signals split the workload. Confidence weighting rescues cases where the wrong plurality is a degenerate cluster, like repeated empty boxed tokens, while the graph structure handles the harder cases where the wrong answer is fluent, popular, and confidently stated.
Tom: Page eleven also flags a genuine limitation. Confidence only helps if correct clusters are, on average, more confident than incorrect ones. If that ordering is weak or reversed, confidence adds nothing.
Jane: Which is why the method combines both signals instead of betting on one. Each one becomes informative in a different part of the envelope, and when no recoverable minority exists, CoRE just behaves like the old ballot.
Tom: So the consensus operator adapts to the situation rather than forcing one rule everywhere. That naturally raises the question of what happens when both signals fail together.
Jane: And that's exactly what the qualitative examples in the appendix dig into, with a concrete geometry question where the graph decisively beats all three baselines.
Conclusion: Jane: We set out to question the majority vote as the default consensus rule for test-time RL, and CoRE answers that by treating roll-outs as a graph and extracting a dominant set through replicator dynamics.
Tom: That single design choice gives you three upgrades over the ballot box: a pseudo-label that can overturn an incorrect plurality, a graded reward instead of a binary one, and a cohesiveness gate that down-weights fragmented questions.
Jane: The theory makes it falsifiable. Majority voting is a special case, the extraction threshold has a sharp formula, and confidence lowers that threshold multiplicatively in nats.
Tom: And the experiments match the predictions, not just the averages. The gains land on contested high-disagreement benchmarks, vanish at the competence floor, and dissolve into seed noise under saturation.
Jane: The practical angle is what wins me over. Same 64 roll-outs, no auxiliary models, no extra sampling, just a smarter read of what you already generated. Plus the sample efficiency gain, reaching the vote's plateau in about half the steps.
Tom: The deeper implication is that the consensus operator is a real design choice, not a plumbing detail. If the only supervision in adaptation comes from self-agreement, then how you define that agreement sets the ceiling on what the model can learn.
Jane: It also opens a path for self-training beyond math. The authors mention hidden-state representations instead of lexical features, and long-context reasoning models as natural extensions.
Tom: And honestly, the fact that the recovery examples include a geometry question where only the graph structure finds the correct minority, while confidence alone fails, shows the two signals are genuinely complementary rather than redundant.
Jane: We should also note the transparency: the codebase and fixed hyperparameters are released, so the whole thing is reproducible from the repository.
Tom: That's a good place to leave CoRE. We've seen how equilibrium-based consensus beats the ballot, and the next paper on our list tackles a different angle on reward design for reasoning models.
Jane: Let's turn the page and see what's coming up.
Ambuj Mehrish, Sebastiano Vascon
Ca' Foscari University of Venice
cs.AI
Submitted: 2026-08-10
Updated: 2026-08-11
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
The gist: CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Abstract On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from
Key concepts
- Test-time reinforcement learning
- A process where a language model continues training on a test set without ground-truth labels. The model samples multiple answers, and a consensus function provides pseudo-labels to reward correct behavior. The episode focuses on improving this consensus function.
- Majority voting
- The standard consensus method where the most common final answer among sampled rollouts is chosen as the pseudo-label. It is brittle because it ignores reasoning similarity and model confidence, and can reward systematic errors when wrong answers form a plurality.
- Dominant sets and replicator dynamics
- A game-theoretic approach from graph clustering. Rollouts form a graph with edges weighted by answer agreement and reasoning similarity. Replicator dynamics iteratively find a dominant set—a group of mutually supporting trajectories—which provides a refined pseudo-label, graded rewards, and a coherence gate.
- CoRE
- The proposed method. It constructs a graph from 64 rollouts, weights edges by answer match and reasoning similarity (TF-IDF cosine), scales by confidence, and runs replicator dynamics to find an equilibrium. The equilibrium yields a pseudo-label, per-rollout rewards, and a gate to decide if a question is worth training on.
Terminology
Summary
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning
Abstract
On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model’s own roll-outs, rewarding those that match the majority vote over N sampled answers. That vote discards a correct answer whenever it is a minority and scores every majority-matching roll-out identically. We replace it with CoRE (Consensus Rewards via Equilibrium): the N roll-outs form a graph whose edges combine answer agreement, reasoning similarity, and generation confidence, and replicator dynamics extract its dominant set, yielding a refined pseudo-label, a graded per-roll-out reward, and a per-question cohesiveness gate. CoRE strictly generalizes voting: majority voting is recovered as a special case; a block-value analysis gives a sharp threshold for when consensus recovers a correct minority against a larger wrong plurality; and confidence calibration provably lowers that threshold multiplicatively. Across seven backbones and five benchmarks (42 model–benchmark cells, three seeds each), CoRE improves the untrained base by +21.7 points on average versus +20.4 for majority-vote TTRL, wins wherever agreement is contestable with margins over the vote of up to +7.5 points, and reaches the voting baseline’s plateau accuracy in 54–70% fewer steps. Consensus, not counting: treating the roll-out group as a graph rather than a ballot box turns a brittle vote into a calibrated, graded, self-supervised reward at no extra roll-out cost.
Introduction
Reinforcement learning has advanced language-model reasoning, but typically requires labeled rewards. Test-time reinforcement learning (TTRL) instead derives supervision from the model’s own roll-outs on unlabeled test questions. It uses the majority answer as a pseudo-label, rewards agreeing trajectories, and updates the policy. Although TTRL permits other aggregation rules, majority voting remains the default and can yield models that outperform the vote used to train them. This apparent paradox arises because model errors often scatter across multiple answers. Even when the majority label is wrong, many incorrect roll-outs disagree with it and still receive negative rewards, preserving part of the training signal. The failure mode is concentrated error: a systematic mistake becomes the plurality while the correct derivation remains in the minority. Majority voting then rewards the wrong trajectories, penalizes the correct ones, and can amplify the error across updates.
Two additional limitations arise from the same reduction. First, a binary agreement reward assigns the same value to a well-supported derivation and a lucky guess. Second, each question contributes equally to training, regardless of whether its roll-outs form a coherent consensus or a fragmented set of incompatible solutions. Majority voting cannot distinguish these cases because it compresses each trajectory to a final answer and each answer class to a count. It ignores both the model’s confidence during generation and the relational structure among reasoning paths.
We address this limitation with CoRE: Consensus Rewards via Equilibrium. CoRE represents roll-outs as a graph whose edges connect trajectories with matching answers and weight their reasoning similarity by generation confidence. Replicator dynamics identifies a dominant set of mutually supporting trajectories, yielding a refined pseudo-label, graded roll-out rewards, and a cohesiveness gate for question-level updates. CoRE integrates directly into GRPO, requires no labels or auxiliary models, and uses the same roll-outs as majority-based test-time reinforcement learning. Majority voting remains a special case of the formulation.
CoRE’s dynamics begin from majority voting because uniform initialization assigns each answer class mass proportional to its size. A minority can therefore recover through coherence alone only if its internal support exceeds that of the majority by a factor that grows with the class-size imbalance. This predicts that graph structure rarely improves over voting by itself. Confidence-weighted voting is similarly limited: it rescales class counts but does not model support among reasoning trajectories. The two signals are complementary. Confidence calibration rescales each answer class by its members’ mean confidence, lowering the coherence needed for a correct minority to recover. Under our assumptions, this threshold decreases exponentially with the confidence gap between correct and incorrect roll-outs. Confidence can still fail when incorrect answers are systematically overconfident. Empirically, removing either confidence or graph structure reduces CoRE to near-majority-vote performance, whereas combining them yields consistent gains.
We evaluate CoRE on seven backbones from four developers, spanning math-specialized, base, and instruct models across AMC, AIME 2024, MATH-500 and its Level-4/5 subsets, and GPQA-Diamond, with three random seeds per setting. Across 42 model–benchmark pairs, CoRE improves accuracy by +21.7 points over the untrained base, compared with +20.4 for majority-vote TTRL. It is the strongest RL variant in every model family, outperforming voting by up to +7.5 points when consensus is uncertain and remaining within seed variation when the majority is reliable. CoRE also reaches the voting baseline’s final accuracy in 54–70% fewer optimization steps. These results suggest that CoRE complements majority voting by recovering supervision from smaller, coherent, and confident correct groups that vote counts can miss.
Contributions
-
Method. CoRE, an equilibrium-based consensus reward for test-time RL that turns the same N roll-outs into a refined pseudo-label, graded per roll-out rewards, and a question-level cohesiveness gate, with no auxiliary models or extra roll-outs.
-
Theory. Majority voting is a special case of CoRE. A block-value analysis gives the threshold at which a correct minority overturns an incorrect plurality and shows that confidence calibration lowers it multiplicatively.
-
Evidence. Across 42 model–benchmark settings and seven backbones, CoRE improves over the untrained base by +21.7 points on average, compared with +20.4 for TTRL. It is the strongest RL method in every model family, exceeds voting by up to +7.5 points when agreement is contested, and reaches the vote’s plateau in 54-70% fewer steps. Both its gains and failures follow the regimes predicted by the analysis.
Related Work
Test-time adaptation and test-time RL. Test-time adaptation uses unlabeled inputs to update model behavior, either through auxiliary self-supervision or entropy minimization. For LLMs, test-time scaling improves accuracy with additional inference compute, while TTRL converts majority votes over sampled roll-outs into reinforcement-learning rewards. Recent label-free methods instead optimize semantic-cluster entropy, intrinsic entropy, self-certainty, or majority-based self-training. These approaches extend earlier work on self-generated supervision, self-reward, and self-play, alongside unsupervised reasoning objectives and agreement-based rewards. Related analyses suggest that entropy dynamics and robustness to imperfect rewards partly explain why such updates can succeed. CoRE likewise trains on unlabeled test questions, but improves the consensus aggregator that produces the reward rather than modifying the adaptation loop itself.
Aggregating multiple samples at inference. Self-consistency marginalizes over reasoning paths by majority vote, and confidence-informed self-consistency (CISC) replaces uniform votes with a confidence-weighted vote, cutting the required number of paths by over 40% on average
. The accuracy of such compound inference schemes scales predictably with the number of model calls. A parallel line reweights or reranks samples with a learned verifier rather than by vote, using outcome-based signals, process- versus outcome-based feedback, and step-aware verifiers that score intermediate reasoning. These operate only at inference; CoRE instead turns aggregation into a training signal, and confidence-weighted voting is one of its ablation arms.
Dominant sets and replicator dynamics. The Motzkin–Straus theorem relates the maximum clique problem to a quadratic optimization over the simplex, later extended through regularization. Dominant sets generalize maximal cliques to weighted graphs and recover coherent clusters through replicator dynamics (RD). This framework, drawn on evolutionary game theory and the Baum–Eagon inequality, elegantly bridges together combinatorial optimization (clique search), optimization (maximization of a quadratic functional), and dynamical systems (stable points search). To our knowledge, CoRE is the first method to use dominant-set extraction to construct rewards for reinforcement learning and for consensus reaching.
Policy optimization. CoRE builds on GRPO, following policy-gradient methods from PPO and RLHF to critic-free variants. Since group normalization removes reward magnitude, CoRE applies its cohesiveness gate to the loss rather than the reward. This preserves the advantage estimate and reduces exposure to reward over-optimization.
CoRE: Method and Analysis
Problem Setup
Let D be an unlabeled test set and pitheta a policy language model. Test-time reinforcement learning adapts pitheta on D itself: for each question q, the policy samples N i.i.d. roll-outs o1,..., oN ∼ pitheta(· q), a consensus function maps the group to a pseudo-label y∗ and per-roll-out rewards Ri i=1N, and the policy is updated by group-relative policy optimization (GRPO) on those rewards. TTRL instantiates the consensus function with majority voting: y∗vote = arg maxa i: yˆi ≡ a and Ri = ⊮[yˆi ≡ y∗vote], where yˆi is the answer extracted from oi. Since this self-generated signal is the only supervision the model receives, the consensus function determines what test-time RL can learn; it is our object of study, with the sampling procedure and the policy-optimization update rule held fixed.
Each roll-out exposes three signals, all computable without model internals or external supervision: the extracted answer yˆi, compared by symbolic equivalence rather than string match; a reasoning representation hi ∈ Rd, a TF–IDF n-gram vector of the reasoning text; and a generation confidence ci = (1/oi) Σt log pitheta(oi,t q, oi,<t), the mean token log-probability. Majority voting consumes only the first, and only its mode. Throughout, the reward is computed solely from these signals; ground-truth answers are used only for final evaluation and for the post-hoc diagnostic analyses of §6, and never enter the training loop.
Consensus Reward Construction
CoRE replaces majority voting with dominant-set extraction over a roll-out graph that encodes answer agreement, reasoning similarity, and group cohesiveness. A dominant set generalizes the maximal clique concept to weighted graphs, allowing a compact and mutually supportive subset of trajectories to emerge from pairwise interactions. CoRE therefore produces three signals rather than TTRL’s two: a refined pseudo-label y∗, a graded reward Ri ∈ [0, 1], and a question-level cohesiveness gate wq ∈ [0, 1].
The roll-out graph is graph G = (V, E, omega) where V is the set of roll-out answers generated by the model and E ⊆ V × V is the set of edges connecting different nodes, weighted by the function omegai,j: (i,j) ∈ E → R≥0. The graph G is encoded into an affinity matrix A′.
Affinity graph. Given two nodes i, j ∈ V, two kernels compare the corresponding roll-outs: a hard answer kernel Kans(i,j) = ⊮[yˆi ≡ yˆj] and a soft reasoning kernel Krsn(i,j) = (1/2)(1+cos(hi, hj)) ∈ [0, 1]. They combine multiplicatively,
Aij = Kans(i,j) [kappa + (1 − kappa) Krsn(i,j)], Aii = 0,
so that roll-outs with different answers share no edge regardless of how similar their prose is: A is block-diagonal over answer classes, and reasoning similarity can strengthen ties only within a class. The floor kappa guarantees agreeing roll-outs a minimum affinity even when their derivations diverge lexically; kappa is the generalization knob, and at kappa = 1 the graph reduces to pure answer agreement.
Confidence calibration. Node weights wi = exp((ci − maxj cj)/tau) ∈ (0, 1] convert confidences to relative weights anchored at the most confident roll-out. They calibrate the edges symmetrically,
A′ij = Aij √(wi wj),
a diagonal congruence A′ = D1/2AD1/2, D = diag(w). The geometric mean is a symmetric choice. It preserves the block structure exactly, keeps an edge only when both endpoints are confident, and makes a cluster’s internal score scale linearly with its average confidence.
Avoid singletons. To avoid singletons, self-loops are added to the calibrated rollout-graph W = A′ − alphaI. The self-competition term −alpha on the diagonal penalizes small clusters and excludes singletons outright. Since the diagonal is negative, we shift W′ = W + Cee⊤ with C ≥ alpha to obtain a nonnegative affinity matrix, where e is the all-ones vector. The blocks external to the main diagonal are kept equal to 0 to avoid incoherent consensus.
Dominant Set Extraction. Given a readout graph G = (V, E, omega), a DS is extracted by maximizing a standard quadratic assignment problem x⊤Ax over the standard simplex. The optimization is performed by iterating the replicator dynamics
x(t + 1)i ← x(t)i (W′x(t))i / (x(t)⊤W′x(t))
until convergence x(t) −x(t +1) < epsilon or Tmax iterations. Where x(t)i is the i-th component of vector x at time t. Eq 3 starts from the barycenter of the simplex x(0) = e/N, leaving the same chances to each node to get extracted as part of the DS. At convergence, the dynamic system reaches a stable point, denoted by x∗. x∗ concentrates the mass on the most mutually supporting nodes; its support sigma = i: xi∗ > rho is the consensus. Three lemmas establish that the construction is sound: the nonnegativity shift is invariant on the simplex, so W and W′ share all maximizers and replicator fixed points (Lemma 1); the update (3) is a Baum–Eagon growth transform, ascending monotonically and converging to fixed points whose supports are dominant sets (Lemma 2); and for alpha > 0 no singleton is ever selected when a genuine answer clique exists (Lemma 3).
Read-outs. Three signals are read from the equilibrium:
y∗ = arg maxa Σi: yˆi≡a xi∗, Ri = (Ax∗)i / maxj(Ax∗)j, wq = x∗⊤Ax∗.
The pseudo-label is the answer with the greatest equilibrium mass, so it may differ from the count-based plurality. Each roll-out is rewarded by its affinity to this consensus using the uncalibrated matrix A: confidence already shapes x∗, and using A′ again would double-count it. The question-level gate wq measures consensus coherence and scales the loss rather than the rewards, since GRPO normalization cancels any per-question reward scaling. It therefore leaves advantage estimation, clipping, and optimization unchanged, with wq ≡ 1 recovering the standard objective. TTRL is the special case (kappa=1, w≡e, wq ≡1).
When Does Consensus Beat the Vote?
Because the answer kernel gates all edges, the affinity matrix is block-diagonal: each answer class Va (size na = Va) is a connected component, where a ∈ A is the set of possible answers. For x ∈ Δ write sa = Σi∈Va xi ≤ 1 for the mass on class a; the objective decomposes as x⊤A′x = Σa∈A s2a phia(x), where phia is the class’s internal coherence score, converging under the within-class dynamics to a limit mua (for a uniform clique of affinity c, mua = c(1 − 1/na)). Two questions arise: which class the global optimum prefers, and which class the replicator dynamics, started from the barycenter, actually select as the global maximizer lies in the class with the largest intensive score mua. The RD starts from the barycenter, allocating initial mass in proportion to class size; therefore the trajectory starts from the basin of the class with the largest extensive score namua. The initialization literally encodes the ballot, leaving mass to the other possible outcome proportionally to their class size; coherence and confidence are what can overturn the convergence to the largest class size.
Proposition 1 (Majority voting is a special case): With answer-only affinity (kappa=1) and uniform confidence (w ≡ e), the extracted class maximizes namua = na − 1 − alpha, which is strictly increasing in na: the pseudo-label y∗ equals the plurality answer whenever the largest class is unique.
Proposition 2 (Extraction threshold): Let a correct clique of size n∗ and affinity c∗ compete with a wrong clique of size m > n∗ and affinity cw (alpha=0). The dynamics started from the barycenter extract the correct class iff
c∗ > ((m−1)/(n∗−1)) cw.
Proposition 3 (Confidence multiplies coherence): For an answer clique S with base affinity c and confidence weights wi, the calibrated internal score obeys the exact identity
mu′S = c(mu¯S2 − nu¯S/S), mu¯S = meani wi, nu¯S = meani wi,
so with w¯S:= mu¯S2, the extraction threshold (5) becomes
c∗ > (w¯w/w¯∗) · ((m−1)/(n∗−1)) cw.
When correct roll-outs are more confident (w¯∗ > w¯w), the threshold drops multiplicatively. Confidence is not merely a refinement but the enabling signal. Under temperature-tau weighting, a confidence gap delta = c¯∗ − c¯w in mean token log-probability reduces the recovery threshold by e−delta/tau. Since ci uses natural logarithms, delta is measured in nats. At the gap observed in §6, the 8-vs-24 threshold falls from ≈3.3cw to below cw, turning minority recovery from unlikely to feasible. Confidence-weighted voting alone only rescales counts and cannot compare class-level reasoning coherence. The analysis therefore predicts that neither confidence nor graph structure is sufficient in isolation; their multiplicative interaction in (7) creates the recovery regime.
Predictions. The analysis makes falsifiable forecasts that §4 tests: (i) at kappa=1, w≡e, CoRE’s pseudo-labels match majority voting exactly (Prop. 1); (ii) the graph alone and confidence-weighted voting alone should each perform on par with the vote, while their combination should not (Props. 2–3); (iii) gains should concentrate where a coherent, confident correct minority exists to be recovered, vanish where the vote is already right or no correct roll-outs exist, and reverse where the confidence premise fails.
Experimental Setup
Models and benchmarks. We evaluate 7 backbones from four developers across three regimes: math-specialized Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and DeepSeek-Math-7B; vanilla Qwen2.5-7B and Qwen3-8B in non-thinking mode; and Llama-3.1-8B-Instruct and Mistral-Nemo-Instruct. Each model is adapted and evaluated on five free-form mathematical-reasoning benchmarks: AMC, AIME 2024, MATH-500, and its L4 and L5 subsets, which isolate harder, high-disagreement cases where correct minorities are more likely to be outvoted. We additionally use the 198-question GPQA-Diamond subset as an out-of-domain probe covering expert-validated, Google-proof
graduate-level physics, chemistry, and biology questions with four answer choices. Math-specialized models are OOD in both subject and format and remain near the 25% random-guess baseline, making GPQA-Diamond a probe of the operating envelope rather than in-domain ability.
Training. Following TTRL, each model is adapted on the unlabeled test benchmark. For every question, we sample N=64 roll-outs, derive a label-free reward from their consensus, and update the policy with GRPO; ground-truth answers are used only for final evaluation. We compare four reward rules under the same training loop: Majority, the standard TTRL vote and the kappa=1 special case of our operator (Prop. 1); EC, equilibrium consensus with graded rewards but no confidence; CISC, a pre-registered confidence-weighted voting control without the graph; and CoRE, the full confidence-calibrated method. We report pass@1 as mean@16 using 16 samples at temperature 0.6 and top-p 0.95, averaged over three seeds for each model–benchmark pair. Base denotes the frozen model without RL, and each Δ reports the corresponding arm minus Base.
CoRE configuration. We use one fixed setting across all models, benchmarks, and seeds: N=64 following TTRL, kappa=0.1, tau=0.25, alpha=0.1, and C=alpha. Replicator dynamics starts from the barycenter and stops at epsilon=10−6 or Tmax=200, with support threshold rho=1/(10N). Each question uses independently fitted character 4-gram TF–IDF reasoning vectors, and consensus adds negligible O(N2) computation relative to roll-out generation.
Results
Math-specialized models: CoRE improves the no-RL base by +25.0 points on average and is the best arm on every model. Against the reward that TTRL actually uses Majority vote, CoRE wins on the benchmarks where agreement is contestable and stays within seed noise elsewhere. On Qwen2.5-Math-1.5B it improves over Majority by +5.3 on MATH-500 and +5.0 on MATH-L4; on Qwen2.5-Math-7B by +7.5 on GPQA; and on DeepSeek-Math-7B by +3.8 on AMC. The gains concentrate on the harder benchmarks, where a correct but non-majority cluster exists to be recovered. The only benchmarks on which CoRE trails Majority is MATH-L5 on Qwen2.5-Math-1.5B and MATH-500 on both Qwen2.5-Math-7B and DeepSeek-Math-7B do so by at most 0.3 points, within the 0.2–0.6 across-seed standard deviation, and on models where Majority already scores above 85% and 50% so little recoverable headroom remains. Offline pseudo-label accuracy before policy updates shows CoRE outperforms majority voting at every roll-out count N, and the gap persists as N increases. The downstream gains therefore arise from a better training signal rather than optimization noise.
Vanilla and instruct models: The improvement is not specific to math-specialized backbones. On Qwen2.5-7B, CoRE beats Majority by +6.2 on MATH-L5, +2.9 on MATH-L4, and +1.2 on AMC (mean Δ vs base +24.5). On LLaMA-3.1-8B, CoRE improves over Majority by +4.2 on MATH-L4 and +5.1 on GPQA, and on Mistral-Nemo by +2.2 on MATH-L5 and +3.8 on GPQA (mean Δ vs base +19.1). Across all three model families the picture is consistent: CoRE is the best RL arm on most benchmarks and improves over the no-RL base by roughly +21.7 points on average across all 42 cells (family means +21.6, +24.5, +19.1), spanning four model developers (Qwen, DeepSeek, Meta, Mistral) and three capability regimes.
Sample efficiency: CoRE reaches TTRL’s plateau accuracy in 70% fewer steps on LLaMA-3.1-8B (AMC), 57% fewer on DeepSeek-Math-7B (AMC), and 54% fewer on Qwen2.5-Math-1.5B (MATH). Unlike the binary majority reward, which assigns the same target to every consensus member, the equilibrium reward Ri = (Ax∗)i/maxj(Ax∗)j grades each roll-out by its support within the consensus. This denser signal appears to reduce gradient variance, linking faster convergence to the final accuracy gains.
Analysis
CoRE vs. CISC. CISC retains confidence weighting but removes the graph, equilibrium, and graded reward, so the CoRE − CISC gap isolates the value of second-order structure. CoRE improves over CISC by +1.2 points on average, with family-level gains of +1.2, +0.8, and +1.5. On the contested MATH L4/L5 subsets, it wins 13 of 14 settings, with significance under both the sign test (p=0.002) and Wilcoxon signed-rank test (p=0.001). Unlike confidence voting, which weights roll-outs independently, the quadratic objective x⊤Ax captures mutual support and can recover a coherent correct minority.
Recovery region. Propositions 2–3 predict when a correct minority can be recovered. The axes measure its size disadvantage, r = m/n∗, and required coherence advantage, c∗/cw. The static optimum requires roughly equal coherence, whereas barycenter-initialized replicator dynamics raise the threshold to ≈r because they begin from majority mass. Confidence calibration lowers it by e−delta/tau, yielding ≈ 0.29r here. Model–benchmark pairs below this operative threshold are recoverable; those above it are not.
Operating regimes of consensus. The exceptions occur mainly at the competence floor, where roll-outs lack coherent structure (chi ≈ 0) and no reliable minority cluster exists beyond the majority signal. This regime appears for Qwen2.5-Math-1.5B on out-of-domain GPQA and Mistral-Nemo on AIME, where all consensus methods remain near the underlying performance floor. At the opposite extreme, saturation leaves little room for additional recovery. When a strong model produces nearly unanimous roll-outs, the majority ratio approaches one and the consensus rules become equivalent up to seed variation. The results on Qwen3-8B for GPQA and AMC are consistent with the kappa = 1 limit in Proposition 1, under which CoRE reduces to Majority. Thus, the largest gains arise in the intermediate regime where the sample set contains a coherent correct minority that is not selected by majority voting alone.
The analysis also characterizes when confidence information is useful. Proposition 3 assumes that correct clusters are, on average, assigned higher confidence than incorrect clusters. When this ordering is weak or reversed, confidence provides limited additional evidence. This motivates the combined use of confidence and graph structure in CoRE: each signal is most informative in a different part of the operating envelope, while the method naturally approaches Majority when no additional recoverable consensus is present.
Limitations
CoRE helps only where roll-outs disagree and a coherent correct cluster exists: it cannot manufacture signal at the competence floor and reduces to majority under saturation. The reasoning kernel uses a black-box lexical representation, white-box hidden states are a natural extension and we defer long-context (32k) long-reasoning-model training to future work.
Conclusion
We introduced CoRE, which replaces majority voting in test-time reinforcement learning with the dominant-set equilibrium of a confidence-calibrated roll-out graph. Using the same N roll-outs, CoRE produces a refined pseudo-label, graded rewards, and a question-level cohesiveness gate, while retaining majority voting as a special case. Our analysis identifies when a coherent correct minority can overturn an incorrect majority and shows that confidence lowers this recovery threshold multiplicatively. Across seven backbones and five benchmarks, CoRE improves over the untrained base by +21.7 points on average, compared with +20.4 for majority-vote TTRL, exceeds voting by up to +7.5 points, and reaches its final accuracy in 54–70% fewer steps. The gains concentrate where agreement is contested, while the failures occur near the predicted competence floor. These findings suggest that the consensus operator is a substantive component of self-supervised RL: modeling mutual support among roll-outs yields a denser training signal than treating them as independent votes. Future work may replace lexical reasoning features with hidden-state representations and extend CoRE to long-context reasoning models.
Improvements for AI systems
Based on the paper, here are the specific improvements you can make to AI systems and what the improved system can do:
1. Replace majority-vote rewards with equilibrium-based consensus rewards in test-time RL
-
Improvement: Instead of rewarding roll-outs that match the plurality answer, use replicator dynamics on a graph of roll-outs (edges weighted by answer agreement × reasoning similarity × confidence) to extract a dominant set. This yields a graded per-roll-out reward, a refined pseudo-label, and a question-level cohesiveness gate.
-
What the improved system can do: Recover correct answers that are in the minority but form a coherent, confident cluster—something majority voting misses. This leads to higher accuracy on contested, hard questions (e.g., +7.5 points over voting on GPQA) and faster convergence (54–70% fewer steps to reach the same accuracy).
2. Use confidence-calibrated graph structure, not just confidence-weighted voting
3. Add a question-level cohesiveness gate to scale the loss
4. Use graded rewards instead of binary agreement rewards
5. Generalize majority voting as a special case (with a tunable knob)
6. Use the theory to predict when consensus will help or fail
7. Extend the reasoning kernel to white-box hidden states (future work)
Summary of the improved AI system’s capabilities:
-
Higher accuracy on contested, hard questions (up to +7.5 points over majority-vote TTRL).
-
Faster training (54–70% fewer steps to reach the same accuracy).
-
Robustness across model families (math-specialized, vanilla, instruct) and benchmarks (AMC, AIME, MATH, GPQA).
-
Self-supervised—no labels, no auxiliary models, no extra roll-outs.
-
Theoretically grounded—you can predict when it will help and when it will not, and it gracefully degrades to majority voting when no recoverable consensus exists.
Abstract
On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs, rewarding those that match the majority vote over N sampled answers. That vote discards a correct answer whenever it is a minority and scores every majority-matching roll-out identically. We replace it with CoRE (Consensus Rewards via Equilibrium): the N roll-outs form a graph whose edges combine answer agreement, reasoning similarity, and generation confidence, and replicator dynamics extract its dominant set, yielding a refined pseudo-label, a graded per-roll-out reward, and a per-question cohesiveness gate. CoRE strictly generalizes voting: majority voting is recovered as a special case; a block-value analysis gives a sharp threshold for when consensus recovers a correct minority against a larger wrong plurality; and confidence calibration provably lowers that threshold multiplicatively. Across seven backbones and five benchmarks (42 model--benchmark cells, three seeds each), CoRE improves the untrained base by +21.7 points on average versus +20.4 for majority-vote TTRL, wins wherever agreement is contestable with margins over the vote of up to +7.5 points, and reaches the voting baseline's plateau accuracy in 54 -- 70 % fewer steps. Consensus, not counting: treating the roll-out group as a graph rather than a ballot box turns a brittle vote into a calibrated, graded, self-supervised reward at no extra roll-out cost.
Sources
- Concrete Problems in AI Safety
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Training Verifiers to Solve Math Word Problems
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Mathematical Problem Solving With the MATH Dataset
- OpenAI o1 System Card
- Mistral 7B
- Language Models (Mostly) Know What They Know
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
- A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
- Teaching Models to Express Their Uncertainty in Words
- Understanding R1-Zero-Like Training: A Critical Perspective
- Training language models to follow instructions with human feedback
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
- Maximizing Confidence Alone Improves Reasoning
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Proximal Policy Optimization Algorithms
- Can Large Reasoning Models Self-Train?
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection