Training-free LLM Verification via Recycling Few-shot Examples

arXiv:2506.17251 · cs.LG, cs.AI · Submitted 2025-06-08 · Read on arXiv

cs.LG, cs.AI

Submitted: 2025-06-08

Updated: 2026-08-31

Comments: EMNLP 2026 Main

Code: https://github.com/QwenLM/Qwen2.5-Mathhttps:

License: http://creativecommons.org/licenses/by/4.0/

The gist: Although large language models (LLMs) have achieved remarkable performance, the inherent stochasticity of their reasoning processes and varying conclusions present significant challenges.

Terminology

Abstract

Although large language models (LLMs) have achieved remarkable performance, the inherent stochasticity of their reasoning processes and varying conclusions present significant challenges. Majority voting or Best-of-N with external verifiers has been explored to mitigate this, but these approaches are limited in applicability or require additional training. To address this problem, we propose a novel framework that Recycles Few-shot examples to verify LLM outputs (ReFeri). Our key idea is to utilize the given few-shot examples not only to generate outputs, but also to evaluate the candidate outputs. Specifically, ReFeri combines a forward confidence score with a backward reconstruction penalty to select candidates that follow few-shot guidance while avoiding demonstration-specific overfitting. Experiments with three different LLMs across seven diverse tasks demonstrate that our framework significantly improves the accuracy of LLMs---achieving an average relative gain of 8.2%---through effective response selection.

Sources

Related papers