Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

arXiv:2607.17524 · cs.CL, cs.LG · Submitted 2026-07-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift".

Jane: Token-Level Off-Policy Labeling (TOPL) proposes an off-policy training paradigm that reframes post-training as a token-level correctness prediction task,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: We’ve covered a lot about how "Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift" reframes post-training by focusing on predicting token correctness instead of just next-token prediction, and we touched on the interesting insights into the LoRA components. Now, let's wrap up what this all means for us moving forward.

Jane: Essentially, the paper argues that by training a model to distinguish between good and bad tokens in a response sequence, we get a natural mechanism to guide it toward generating accurate information while steering it away from hallucinations under distribution shift conditions. It’s about building intrinsic correctness into the generation process itself through this token-level supervision.

Lu: The authors show strong results on document summarization tasks across eleven datasets under distribution shift, and they suggest that the effectiveness of this token-level learning signal is what drives those strong out-of-distribution performance improvements compared to sequence-level methods <ref:2607.17524#pg0,token-level learning signal is>.

Meng: From a practical standpoint, the implication is that future AI training pipelines should incorporate these finer supervision signals if we want to see real gains in model robustness when deploying them on novel data. It suggests that simply fine-tuning on massive datasets isn't enough; we need smarter feedback mechanisms.

Lalam: If this methodology scales well beyond summarization and transfers effectively to other tasks like machine translation, it means we could start designing AI that is inherently more reliable across a much wider variety of applications, not just a single niche task. That versatility is really exciting for the long-term growth of our culture.

Tom: It really feels like this work offers a blueprint for how we can make generative AI systems inherently more trustworthy and less prone to generating convincing but incorrect information when they encounter new situations. The token-level focus seems to be the key mechanism here.

Jane: So, the main idea is that this approach provides a way to inject factual fidelity directly into the model's learning objective in a way that respects its existing structure, rather than forcing a complete overhaul of the training process.

Lu: The authors conclude by confirming that their token-level learning signal is critical for good performance, and they show how their decomposition into LoRA components helps us understand the underlying mechanism of factuality separation within the model's hidden representations.

Meng: So, we’re looking at a way to make the model’s internal representation space more structured in a way that separates truth from noise, which is something we need to keep an eye on for practical system design.

Lalam: I think this paper opens up a really promising avenue for building AI that doesn't just mimic patterns but understands the underlying structure of information, which is what makes it so valuable to us.

Conclusion: Tom: So we've seen how this paper tackles the tricky problem of making AI outputs accurate when things change in the data it sees. Jane, let's recap what we're focusing on right now with this discussion about "Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift."

Jane: Exactly, Tom. We’re zeroing in on how they shift the focus from general next-token prediction to a specific task of judging whether each individual word or token in a generated response is actually correct or corrupted.

Lu: The core idea is that by training the AI to classify tokens as good or bad based on the source document, it naturally learns to prioritize factual content over random noise when it encounters new data.

Meng: From my side, what’s most interesting for me is how they use LoRA for this; it suggests we don't need to retrain the whole massive model just to fix this kind of fidelity issue.

Lalam: This paper points toward a future where AI generation isn't just about sounding fluent, but about possessing an actual sense of what is true within the context provided.

Tom: Right, so the title itself really tells us that they’re looking at how to train this model off-policy, meaning using data from other times or sources to improve its performance here.

Jane: And the authors are doing a lot of heavy lifting by showing how this token-level supervision helps bridge the gap between what an AI *says* and what is actually *true* in the source material.

Lu: Their methodology uses controlled modifications to summarize texts, essentially creating deliberately flawed examples to teach the model what "bad" tokens look like before it ever sees them in a real application.

Meng: It’s cool that they applied this across several different backbone models; that suggests their approach isn't tied to one specific architecture, which is a big deal for real-world deployment.

Lalam: Thinking about the impact, this work could significantly boost our ability to deploy AI systems in fact-checking roles or areas where accuracy is absolutely non-negotiable.

Tom: It’s definitely a step toward building AI that can handle distribution shifts without totally breaking down into nonsense when it hits something new.

Jane: And the implication is that we might be able to create generation pipelines where factual errors are caught much earlier in the process, before they become part of a final response.

Lu: The way they analyze the internal structure using those LoRA components gives us a window into how AI actually separates concepts like fact from fiction inside its layers.

Meng: Practically, if we can reliably steer these hidden states toward factual behavior with just a few parameters, it simplifies the fine-tuning process immensely.

Lalam: Imagine an AI assistant that could confidently cite sources because its internal representation is fundamentally structured to distinguish between evidence and fabrication.

Tom: That’s a massive leap from just training it on more text; this is about teaching it *how* to think about correctness at the atomic level.

Jane: So, we're moving beyond simply making AI sound like a human writer and toward engineering systems that are fundamentally more grounded in reality.

Lu: The next big question for this paper is how these token-level rules translate when we move from summarization to complex, multi-step reasoning tasks.

Meng: I’m curious if the computational overhead of calculating those probabilities for every single token makes it feasible for real-time inference on smaller devices.

Lalam: That feasibility, coupled with the accuracy boost, is what really opens up possibilities for making AI a truly reliable tool in society, and we'll have to keep watching how researchers tackle that next challenge.

University of Southern California

cs.CL, cs.LG

Submitted: 2026-07-20

Updated: 2026-10-06

Importance score: 89/100

The gist: Token-Level Off-Policy Labeling (TOPL) proposes an off-policy training paradigm that reframes post-training as a token-level correctness prediction task, which is critical for improving factuality

Key concepts

Token-Level Correctness Prediction
This approach reframes language model training as a classification problem for every single output token. The model learns to predict whether a specific token is factually correct or corrupted based on the input document and preceding tokens, rather than just predicting the next word.
Binary Classification Objective
The core training objective is a binary cross-entropy loss. The model outputs a probability for each token being 'correct' (1) or 'corrupted' (0). This forces the model to learn fine-grained distinctions between faithful and hallucinated text segments.
Low-Rank Adaptation (LoRA)
TOPL uses LoRA to update only selected layers of the large language model instead of retraining all parameters. This allows for efficient training while focusing on learning a specific, discriminative subspace that separates factual information from noise in the model's hidden representations.

Terminology

Summary

Token-Level Off-Policy Labeling (TOPL) proposes an off-policy training paradigm that reframes post-training as a token-level correctness prediction task, which is critical for improving factuality and out-of-distribution generalization in large language models. The core finding is that by training the model to distinguish good and bad tokens, it naturally guides the model toward generating good tokens while avoiding the pitfalls of directly training on off-policy tokens.

How it works

TOPL reframes post-training as a token-level correctness prediction task instead of relying on next-token prediction. This is achieved by adopting a binary classification objective that predicts whether each response token is correct or corrupted, where corrupted tokens are those perturbed to introduce factual inconsistencies. The model outputs a probability for every token based on the input document and the preceding tokens: The model outputs a probability pk = P(zk = 1 D, t≤k) for every token tk. This prediction is generated by attaching a lightweight reward head to an intermediate layer of the language model, where the hidden states are computed as: H = FinalRMSNormf0:m(D, S) ∈ R S×Dhidden. The model is then trained using a binary cross-entropy loss over all response tokens.

Data Generation for Token-Level Supervision

The process begins by generating data for token-level supervision. Given a document D and its ground-truth summary S = (t1,..., tS), the sequence S serves as a positive example of a faithful response. To construct the token-level supervision signal, one natural approach is to generate a perturbed version of the summary, denoted as S', by introducing controlled modifications such as token insertions, deletions, or substitutions that make the content of the summary unfaithful to D. The paper leverages datasets like FAVA (Mishra et al., 2024), which provides model-generated summaries together with fine-grained edits, i.e. corruptions, to said summary that simulate typical hallucinations. For each token tk in the perturbed sequence S', a binary label zk ∈ [0, 1] is assigned: zk = 1 indicates that the token is faithful to the original summary S (i.e. it is a 'good' token), and zk = 0 indicates that the token corresponds to a hallucinated span introduced by the edits (i.e. 'bad' tokens).

Model Architecture and Training

TOPL training involves using Low-Rank Adaptation (LoRA) on selected layers of the backbone language model f, rather than updating all parameters directly. The objective is to predict the correctness score for each token based on its hidden representation: P(zk = 1 D, t≤k) = σ h(Hk), k = 1..., S, (2) where σ(·) denotes the sigmoid function. After training, the reward head is discarded and the learned LoRA update is merged into the backbone model to obtain an updated model fmerge for generation. The paper utilizes various backbone models such as Qwen3-8B, Llama-3.1-8B, and Gemma-3-4B, applying LoRA with a LoRA rank of r = 4 and a scaling hyperparameter α = 8.

Analysis of Effectiveness and Interpretation

The effectiveness of TOPL is analyzed through mechanistic interpretations connecting it to conditional steering. The authors hypothesize that TOPL with LoRA can be viewed as a form of conditional steering, where LoRA-A acts as a classifier for token correctness and LoRA-B defines the steering directions. Specifically:

  1. LoRA-A is hypothesized to capture discriminative concept directions for token correctness, effectively separating factual and nonfactual tokens in the hidden representation space. The paper shows that TOPL exhibits substantially stronger separability between good and bad tokens in the LoRA-A subspace compared to SFT.

  2. LoRA-B is interpreted as encoding steering directions associated with factuality. Interventions of the form ∆h = λ(µ+ − µ−)B are used to steer hidden states toward factual behavior, demonstrating that positive λ steers it toward factual representations, with an optimal magnitude balancing these effects.

  3. This strong LoRA-A separability is linked to generalization: methods with stronger separability in the LoRA-A subspace tend to achieve better OOD performance, suggesting that this structure helps the model rely on more stable and semantically grounded features rather than spurious correlations under distribution shift.

Generalization and Cross-Task Transfer

TOPL demonstrates strong out-of-distribution (OOD) generalization across 11 datasets on document summarization, achieving "strongest overall results on average across 11 datasets under distribution shift.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the Token-Level Off-Policy Labeling (TOPL) paradigm, along with what these improved systems can achieve:


  1. Improve Factuality and Reduce Hallucinations in Generative Models:

  2. Enable Robust Out-of-Distribution (OOD) Generalization Across Diverse Tasks:

  3. Increase Interpretability of Model Behavior via Representation Steering:

  4. Develop Unified, Fine-Grained Token-Level Training Objectives for Off-Policy Learning:

5.1 Improvement Details & Capabilities:

  1. Improve Factuality and Reduce Hallucinations in Generative Models:

Based on the core mechanism of TOPL, which trains the model to predict token correctness (factual vs. non-factual) via a binary classification head, the improved system can perform high-fidelity content generation by directly penalizing factual inconsistencies at the atomic level.

  • Specific Capability: Significantly reduce hallucinations in models like Large Language Models (LLMs) by ensuring every generated token aligns with verifiable ground truth derived from context.

  • Specific Capability: Achieve superior performance on factuality benchmarks (e.g., AggreFact, Bespoke-MiniCheck) compared to sequence-level or standard next-token prediction methods, leading to summaries and translations that are demonstrably more accurate.

  1. Enable Robust Out-of-Distribution (OOD) Generalization Across Diverse Tasks:

The paper demonstrates that TOPL achieves strong out-of-distribution performance across 11 datasets in document summarization and effectively transfers this benefit to machine translation tasks (as shown in Section 4.5).

  • Specific Capability: Deploy a single, robust fine-tuned model that maintains high factual consistency when applied to novel domains, languages, or input distributions not seen during training (i.e., strong generalization under distribution shift).

  • Specific Capability: Apply the same TOPL framework to other faithful generation tasks (like machine translation) and maintain state-of-the-art performance across different modalities and languages where token-level supervision is available.

  1. Increase Interpretability of Model Behavior via Representation Steering:

The research provides a mechanistic understanding by interpreting the LoRA components: LoRA-A acts as a classifier separating factual/non-factual tokens (conditioning), while LoRA-B functions as steering vectors that push hidden states toward factual or non-factual regions.

  • Specific Capability: Gain control over model behavior at inference time by applying targeted interventions along learned steering vectors (e.g., using the formula in Section 5.3).

  • Specific Capability: Develop a system where steering is not arbitrary, but is guided by a representation that explicitly distinguishes factual content from noise, allowing researchers to probe and manipulate the model's internal logic regarding truthfulness.

  1. Develop Unified, Fine-Grained Token-Level Training Objectives for Off-Policy Learning:

TOPL reframes post-training as a token-level correctness prediction task using an off-policy objective (binary cross-entropy loss) rather than direct next-token prediction or standard reward models.

  • Specific Capability: Create a unified, single training pipeline that is both fine-grained (token level) and computationally efficient (using LoRA), simplifying the optimization process compared to complex on-policy RL methods.

  • Specific Capability: Facilitate the development of more interpretable model updates, where learned LoRA adapters can be directly mapped to conditional steering mechanisms, providing a clearer link between training signal and final model behavior.

Sources

Related papers