Video wird geladen...
Video konnte nicht geladen werden
Can LLMs Self-Verify? Much better than you'd expect. LLMs are increasingly used as parallel reasoners, sampling many solutions at once. Choosing the right answer is the real bottleneck. We show that pairwise self-verification is a powerful primitive. Introducing V1, a framework that unifies generation and self-verification: 💡 Pairwise self-verification... show more
108,064 Aufrufe • vor 7 Monaten •via X (Twitter)
32 Kommentare

Paper: Code: Project page: Pairwise self-verification improves test-time scaling across code and math tasks.

Current self-verification (or aggregation) techniques during parallel reasoning have two key limitations: 1️⃣ Calibration collapse (pointwise scoring) Without comparative references, models often over-score plausible but incorrect solutions (e.g., giving 10/10 to wrong code). 2️⃣ Diversity collapse during self-refinement or aggregation If the model latches onto incorrect candidates, Pass@N may decrease as correct solutions get discarded during self-aggregation Pairwise self-verification addresses both: comparisons provide stronger verification signals without reducing Pass@N, and can provide orthogonal improvements in accuracy and latency compared with strong deep-thinking/aggregation methods.

V1-Infer: Pairwise self-verification via Swiss-system ranking Instead of scoring N candidates independently, we run a Swiss-system tournament: Solutions compete head-to-head, and rankings are determined by confidence-weighted win rates. Key idea: allocate verification budget to the most uncertain pairs, focusing comparisons exactly where they matter. Result: accurate rankings with far fewer comparisons, K << C(N,2).

Results across code generation and math: ⚡Code benchmarks: up to +19.6% Pass@1 across models ⚡Math benchmarks: up to +17.9% Pass@1 across models 😲On LiveCodeBench-V6, V1-Infer improves over RSA by +4-6% All gains come from self-verification alone (no external verifiers, execution feedback, self-aggregation, or ground truth). On SWE-bench Lite (300 GitHub issues, Gemini 2.5 Flash, using Mini-SWE-Agent): ✅ Vanilla: 26.3% ✅ Pointwise self-verification: 28.3% ✅ Pairwise self-verification: 33.3% (+7.0%) The verifier only sees the issue description and patch diffs, no repo context or agent trajectories. Pairwise comparisons help distinguish root-cause fixes vs surface patches.

Can we train models to be better self-verifiers? Introducing V1-PairRL: a unified RL framework that co-trains a single model as both generator and pairwise self-verifier. Key idea: generation and verification co-evolve. As the generator improves, the verifier trains on in-distribution outputs from the evolving model itself. We also address two reward hacking failure modes: 🧠 Safe-bet collapse: verifier predicts ~0.5 correctness scores to hedge → fixed with a sparsity threshold 🧠 Empty-solution loop: generator outputs trivial failures → fixed with selective pairing

V1-PairRL training results (DeepCoder framework): ✅ 7–9% better test-time scaling vs pointwise RL co-training ✅ Outperforms standard RL even when both use pairwise verification at inference ✅ +8.7% Pass@1 over standard RL without any test-time scaling V1 shows that pairwise self-verification is a trainable capability. V1-PairRL improves self-verification and makes the generation capability stronger as well.

Pairwise self-verification is also complementary to aggregation-based test-time scaling. As a proof-of-concept, we combine V1-Infer with a SoTA deep-think style algorithm (Recursive Self-Aggregation - great work by @siddarthv66 et al.) Every few RSA loops, we apply pairwise verification to filter candidates before aggregation. Result: faster convergence and lower latency than vanilla RSA. Interestingly, pointwise verification hurts in the same setup, converging slower than aggregation alone. This suggests self-verification can act as a fitness signal for evolutionary search.

If you're working on test-time scaling, parallel reasoning, or RL for LLMs, I'd love to hear your thoughts! Paper: Code: Project page: Co-led with Xiuyu @xiuyu_l Thanks to collaborators and advisors: @KushaSareen @sudomonish @sijun_tan @XiaoxiaWShirley @_junxiong_wang @AlpayAriyak @QingyangWu1 @samir_khaki @rish2k1 @LongTonyLian Yucheng @Boyiliee @alsuhr @ben_athi @KurtKeutzer Collaboration between UC Berkeley, Mila, Together AI, and NVIDIA. @UCBerkeley, @berkeley_ai, @Mila_Quebec, @togethercompute, @nvidia Models RL trained with @rllm_project, thanks to @sijun_tan and the rLLM team! Thanks to Modal (@modal, @charles_irl) for compute credits via the Modal for Academics grant.

V1 Infer is very well-known in the community though, also for self-verification, see e.g., and The training looks cools, but the overlaim for V1-infer is unfortunate. You should probably attribute properly.

Hi Jasper, thanks for sharing these related works. I was aware of the Scaling Generative Verifiers work, and recently, the ProofBench work, which are both concurrent, based on when we submitted our paper (including to the ICLR SPOT workshop, see openreview). Thanks for sharing your OPC work as well, I enjoyed reading it over the weekend. V1-Infer and pairwise self-verification are not well-known, and to the best of our knowledge, our work is the first to systematically show that pairwise self-verification actually works well. Here is where our contributions differ: - Efficacy of Pairwise Self-Verification: Recent work ( highlighted that pointwise self-verification struggles. We concurrently found the same thing. Our primary goal was to establish that pairwise self-verification is a viable primitive that overcomes these pointwise limitations. While the papers you linked provide early signals, they do not systematically analyze or prove the broader self-verification improvements established in V1. - Efficiency & integration with Test-Time Scaling: The algorithms explored in the concurrent works are also quite different. Your OPC paper correctly notes that all-pairs Bradley-Terry requires O(N^2) comparisons, making it unscalable. This O(N^2) inefficiency is likely why pairwise verification is actually not very well known or adopted in test-time scaling literature, and why it is rarely used as a baseline against methods like Recursive Self-Aggregation (RSA). Standard Knockout tournaments require O(N) but suffer from lower accuracy (and the hybrid approach in Scaling Generative Verifiers risks dropping correct candidates). V1-Infer uses continuous (1 to 10) margin-weighted scoring to target the most ambiguous pairs, achieving high accuracy with O(N) compute. This makes it practical and allows augmenting with aggregation-based test-time scaling approaches like RSA. (See attached graph comparing V1-Infer vs. O(N^2) Bradley-Terry from OPC paper which I just ran). - Systematic Scope: While the concurrent works provide signals for the specific domain of Math Olympiad proofs, V1 is the first to systematically evaluate compute-matched pairwise self-verification across math, coding and SWE tasks, using diverse models. We will add an extended discussion in our upcoming revision to properly cite and attribute these concurrent math-focused explorations. Thanks again for highlighting them!

Hmm, I disagree. Here is another more well-known paper that also does it: (specifically, their "Self GenSelect") Yes, your algorithm is different, but the framing of your paper is very much "we are the first to show this works", which is not true.

I might be wrong here, but I also dont see a comparison between your method and a simple knockout tournament (which I happen to know works better in my use case since I tried something very similar to V1-Infer before)

Good job Harman!

Very cool work! We have also released PRISM, which uses step-level correctness signals from a PRM to guide deep think inference Paper: Thread:

So cool! If I understand things correctly: PRISM: step-level correctness using a PRM gives a score for each solution before aggregation. Your PRM is the same model, and it is a step-level verifier so better than naive pointwise verification. V1: Uses self-verification on the solution population, ignoring the thinking trace and only using the post-</think> summary/code block, and uses pairwise verification to improve over pointwise verification. PRISM tries to refine the population, which sounds related to my proof-of-concept experiment combining RSA with pairwise self-verification (see tweet above). These ideas could potentially be used together, e.g., step-level pairwise verification!. There may also be better ways to use the individual step-level scores during aggregation, or even to train aggregators with RL to take step-level correctness into account. @sudomonish

@tuvllms Nice to see these two works arrive in parallel (pun intended). We noticed, occasionally PRM assigns full scores to multiple traces that lead to different final answers, a classic limitation when using the same model as both generator and verifier.

@tuvllms We addressed it with a comparator model (also the same base model as the generator 😀) that does pairwise self-verification between two fully scored traces with different final answers, similar in spirit to V1.

@tuvllms Happy to chat more about potential synergies (step-level + pairwise sounds especially promising)!

@SharmaRituraj19 @tuvllms Interesting, id be curious to also know how vanilla outcome reward models (pointwise) work compared woth prm. Yes happy to chat more about this stuff!

So like pair-programming but for AI?

close. It is more like “pairwise code review” for candidates: instead of scoring each solution in isolation, the AI compares two solutions head-to-head and rates both!

Wow, this work is really nice! What tool did you use to create the animation?

Thanks! I used claude code and it used manim

This is such a powerful insight - pairwise verification scales better than relying on a single LLM's confidence. Multiple perspectives checking each other creates a much more robust reasoning system. The future is ensemble thinking ✨

This mirrors my training: sparring beats solo kata every time. Pairwise verification forces adversarial truth-testing instead of confident hallucination. Tournament efficiency finally makes parallel reasoning actually deployable.

pairwise verification is such an elegant solution. better than asking the same model to verify itself. gets at the uncertainty problem directly
@threadreaderapp please #unroll

Verification is starting to look as important as generation.

Great work! How does V1's verifier budget scale with N empirically? If achieving a Δ Pass@1 lift at N=16 needs K verifier calls, then to get the same Δ lift at N=128, is the budget closer to 8K (linear), 14K (N log N), or 64K (quadratic)?

Great question which i havent been able to study properly since compute usage blows up quite a lot, and so i have tried only till N=32, and its hard to get this signal at scales smaller than 32.

Pairwise over pointwise is a smart framing — ranking relative quality between solutions sidesteps the calibration problems that make absolute scoring unreliable. Curious how V1-PairRL scales when the generation model starts producing near-identical candidates.

If the model starts producing identical candidates, then the verifier will be rewarded when it gives the same score to both solutions AND the correct score to both solutions as well (high if candidates are good, else low). This may result in reducing the discriminative capacity of the verifier; however, in practice, there is sufficient diversity when performing temperature sampling, and with techniques like zero-variance filtering of prompts, this problem should not occur.

