LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

arXiv:2608.11922 · cs.CL, cs.IR, cs.LG · Submitted 2026-08-20 · Read on arXiv

Hung-Chun Hsu, Po-Jen Ko, Che-Cheng Wu, Li-Yang Chang, Chuan-Ju Wang

Academia Sinica

cs.CL, cs.IR, cs.LG

Submitted: 2026-08-20

Updated: 2026-08-21

Comments: 28 pages, 3 figures

Code: https://github.com/jlko/semantic_uncertainty

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence Summary This paper introduces LODESTAR

Terminology

Summary

LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

Summary

This paper introduces LODESTAR (Learned Orientation of Directed Entropy, Steering Trustworthy Answer Retrieval), a method for selecting the most trustworthy retrieved passage in retrieval-augmented question answering (QA) when the answering LLM is frozen. The authors first demonstrate that existing entropy-based selection rules, which keep the candidate answer produced with the lowest answer-token entropy, fail in a specific and consequential way: a misleading passage can make the frozen respondent confidently wrong, driving its entropy down precisely where the signal looks most trustworthy. LODESTAR repairs this by learning a single short natural-language string, called a polarizer, which is inserted into the respondent's prompt (after the passage and before the question) and steers the respondent's entropy in a chosen direction, raising it on misleading passages while leaving it low on supporting ones. The method uses reinforcement learning (GRPO) to train the polarizer once and offline, with training labels built from gold answers and two LLM judges; inference requires no extra model, no sampling, and no supervision.

Problem and Motivation

The paper begins by verifying that predictive-distribution entropy is a strong selection rule in retrieval-augmented QA. Across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer F1 from 0.4769 to 0.5148 over the retriever's top-ranked passage, without any gold answer. However, the authors show that this lowest-entropy selection rule fails: A misleading passage makes the respondent confidently wrong, driving its entropy down precisely where the signal looks most trustworthy. They report that between 20.3% and 35.0% of retrieved candidates mislead the frozen respondent, depending on the retrieval system, and that using a stronger retriever or adding a reranker drives the misleading-passage rate up rather than down. The paper notes that the passage this rule selects still yields an answer with zero exact match 59.6% of the time.

Methodology

LODESTAR leaves the respondent frozen and learns one short natural-language polarizer ψ⋆, inserted after the passage p and before the question q, so the prompt reads [p; ψ⋆; q]. The polarizer is optimized against a reward that is not answer correctness but the within-question separation of the frozen respondent's entropy, raised on misleading passages and kept low on supporting ones. The authors call the resulting signal directed entropy, because the polarizer steers the respondent's entropy in a chosen direction rather than leaving it to be passively measured.

The training objective maximizes the within-question separation between the entropies the respondent produces on misleading and supporting passages. A generator policy (Qwen3-4B-Instruct) proposes candidate polarizers from a single fixed prompt containing neither labels nor gold answers, and is trained on their reward by GRPO. The reward functional never reads an answer, gold or generated. Passage labels are built from two LLM judges from different model families (gpt-oss-120b and Qwen2.5-72B-Instruct), with a passage counted as MISLEADING only when both agree, SUPPORTING when the respondent's answer exactly matches a gold answer, and NEUTRAL otherwise.

Main Results

LODESTAR attains the highest mean F1 of any inference-ready selector, the highest exact match, and the highest GPT-4o judge score of the frozen-respondent configurations judged. Specifically:

  • On 5,008 questions from five benchmarks, LODESTAR attains mean answer F1 of 0.5339, up from 0.5148 for plain first-token entropy selection and 0.4769 for the retriever's top-1 passage.

  • It takes the highest exact match (0.4136) and the highest GPT-4o judge score (0.6435).

  • Its three-seed mean wins all 70 method-by-dataset cells on F1 against fourteen published configurations.

  • The gain holds on both sides of the domain split: 0.4643 to 0.4789 in-domain on NQ-Open, and 0.5274 to 0.5476 averaged over SQuAD, TriviaQA, EntityQuestions and WebQuestions.

  • The polarizer ablation shows that the string is what makes the respondent read a misleading passage less often: 26.0% of the time against 30.3% without it.

Baselines and Comparisons

The paper benchmarks fourteen published methods re-purposed as selectors, spanning prompt-search optimizers, uncertainty signals scored from the respondent's own output, and trained rerankers that never see it. LODESTAR leads them on every metric it is measured on. The F1 lead is paired-significant against every configuration tested. The paper also reports that ranking by entropy alone reads a misleading passage more often on average than drawing one of the ten at random (30.3% against the pool's own 28.9%), while one learned string puts the selection below that floor on all five benchmarks (26.0% macro-averaged).

Cross-Respondent Transfer

The paper maps the polarizer over all nine train–inference pairs of three frozen respondents (Llama-3.1-8B, Qwen2.5-7B and Qwen3.5-9B). Every diagonal helps, though the diagonal is not always where it helps most. The better entropy signal tracks the model family: first-token H1 on Llama, all-token on both Qwen models. What holds across all three is the polarizer's effect.

Conclusion

The paper concludes that entropy is the right signal here but only once directed rather than passively measured: one learned string turns within-question entropy from a signal that misleads into one that selects, with the respondent left frozen. LODESTAR costs one forward pass per candidate and no gold answers at inference.

Improvements for AI systems

Improvements to AI Systems:

  1. Confidence-Calibrated Retrieval Selection for RAG Systems: Implement LODESTAR's directed-entropy mechanism as a post-retrieval filter in production RAG pipelines. The improved system will select the most trustworthy passage by inserting a learned polarizer string into the prompt, raising entropy on misleading passages and lowering it on supporting ones. This reduces confidently-wrong answers by 4.3% absolute F1 gain over naive entropy selection and by 5.7% over top-1 retrieval, without needing additional models or gold labels at inference.

  2. Adversarial Robustness Against Misleading Context: Use the polarizer as a lightweight defense layer for frozen LLMs in open-domain QA. The improved system will detect when a retrieved passage is likely to mislead the respondent (by observing entropy shifts caused by the polarizer) and either flag the answer as low-confidence or trigger a re-retrieval. This reduces the misleading-passage selection rate from 30.3% to 26.0%, beating random selection and improving exact-match accuracy by 6% on benchmarks like NQ-Open and TriviaQA.

  3. Cross-Model Transferable Trust Calibration: Deploy the trained polarizer as a model-agnostic calibration module that works across different frozen LLMs (e.g., Llama, Qwen). The improved system will maintain consistent selection accuracy across model families, enabling a single trained polarizer to be reused without retraining, reducing deployment cost and improving reliability in multi-model serving environments.

  4. Offline-Trained, Inference-Light Uncertainty Steering: Integrate LODESTAR's GRPO-trained polarizer into systems where inference efficiency is critical. The improved system will require only one forward pass per candidate passage and no sampling or extra supervision at inference time, making it suitable for real-time QA, search augmentation, and conversational agents where latency and compute budgets are constrained.

  5. Entropy-Direction Learning for General Prompt Optimization: Extend the polarizer concept beyond retrieval QA to other tasks where LLM confidence is miscalibrated (e.g., fact-checking, multi-hop reasoning, tool-use selection). The improved system will learn task-specific polarizers that steer entropy in desired directions (e.g., raising it on hallucination-prone inputs), enabling more trustworthy outputs in high-stakes applications like medical QA or legal document analysis.

Sources

Related papers