Training-free LLM Verification via Recycling Few-shot Examples
cs.LG, cs.AI
Submitted: 2025-06-08
Updated: 2026-08-31
Comments: EMNLP 2026 Main
Code: https://github.com/QwenLM/Qwen2.5-Mathhttps:
License: http://creativecommons.org/licenses/by/4.0/
The gist: Although large language models (LLMs) have achieved remarkable performance, the inherent stochasticity of their reasoning processes and varying conclusions present significant challenges.
Terminology
Abstract
Although large language models (LLMs) have achieved remarkable performance, the inherent stochasticity of their reasoning processes and varying conclusions present significant challenges. Majority voting or Best-of-N with external verifiers has been explored to mitigate this, but these approaches are limited in applicability or require additional training. To address this problem, we propose a novel framework that Recycles Few-shot examples to verify LLM outputs (ReFeri). Our key idea is to utilize the given few-shot examples not only to generate outputs, but also to evaluate the candidate outputs. Specifically, ReFeri combines a forward confidence score with a backward reconstruction penalty to select candidates that follow few-shot guidance while avoiding demonstration-specific overfitting. Experiments with three different LLMs across seven diverse tasks demonstrate that our framework significantly improves the accuracy of LLMs---achieving an average relative gain of 8.2%---through effective response selection.
Sources
- Universal Self-Consistency for Large Language Model Generation
- Training Verifiers to Solve Math Word Problems
- The Llama 3 Herd of Models
- Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Language Models (Mostly) Know What They Know
- Generative Reward Models
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Gemini: A Family of Highly Capable Multimodal Models
- Solving math word problems with process- and outcome-based feedback
- Recitation over Reasoning: How Cutting-Edge Language Models Can Fail on Elementary School-Level Reasoning Problems?
- Qwen2 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks