When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "When In-Distribution Gains Fail".
Jane: Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train–test distributions.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Welcome back to the show. We've got a fascinating paper today that tackles how we train AI systems to understand human preferences when we only have weak supervision. We're talking about "When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift."
Jane: It sounds like they are looking at a common problem in AI where models get really good at a specific task, but that goodness doesn't stick when the context changes slightly. This paper is diving deep into why that happens when we use weak supervision to train those reward models.
Lu: I’m really interested in the core concept here because it touches on representation stability, which is so crucial for scalable oversight systems. We're seeing how these strong students might be pulled toward features specific to the source domain instead of learning something broadly applicable.
Meng: From an engineering standpoint, that sounds like a significant headache when deploying these systems in real-world applications where the input distribution shifts constantly. So, what exactly are they showing us about this failure mode?
Tom: Well, they show that strong students trained on weak labels can look successful in-distribution but then completely lose their ability to transfer when we test them on completely different preference datasets. This is a big deal for making AI reliable across various scenarios.
Jane: So, the paper points out a specific issue where the model learns patterns from one source domain and fails when it encounters another domain within the same broad category, like helpfulness or harmlessness. It’s not that the models are inherently bad, but their learning process is biased.
Lu: Exactly. They provide evidence for a representational failure mode where weak-supervised finetuning can pull the strong model toward sourcedomain features instead of maintaining broadly transferable representations when it comes to preference data. This points to a fundamental limitation in how we use weak labels for alignment.
Meng: If the representations are drifting, that means our reward model isn't learning the general preference structure; it's just memorizing the specific quirks of the training data. That makes deployment risky because we can’t trust its judgment outside of those exact conditions.
Tom: It sounds like this paper is setting up a way to rigorously test if these weak-to-strong methods are actually learning something general or just fitting source patterns. We need better ways to measure this alignment, right?
Jane: Precisely. They introduce a new evaluation protocol called the "zero-shot distribution-shift evaluation protocol for W2S," which tests models on their held-out test sets from other domains within the same broad category, not just their original source domain.
Title and authors: Lu: That's a smart way to frame the problem. They formally define this shift by considering a broad preference category C with multiple domains KC, where each domain ki has its own distinct data distribution and annotation style. It formalizes what we mean by "distribution shift" in this context.
Meng: So, instead of just checking if the model gets a high score on its original test set, they are forcing it to prove its usefulness across all other possible preference styles and datasets within that same category. That’s a much tougher standard for any alignment technique.
Tom: And they give us some really helpful metrics to track this performance, moving past just raw accuracy numbers. They introduce Weak-to-Strong Raw Gain, Absolute OOD Gain, and the Net Transfer Score (NTS).
Jane: Those metrics are designed to capture exactly what we need: whether the improvement we see in the source domain actually translates into an improvement when we look at an unseen target domain, while also accounting for any potential collapse in performance on the original source data.
Lu: The Net Transfer Score is particularly interesting because it tries to balance two competing goals: maintaining strong performance where it already excels and showing improvement where it hasn't seen data before. It addresses the idea that you don't want to sacrifice source-domain accuracy just for a potential out-of-distribution win.
Meng: From a practical standpoint, that balance is what we need to see in any production system. We don't want a model that works perfectly on one type of user query but breaks entirely when the user asks something slightly differently.
Tom: Now, they propose a solution called Representation Anchoring, or ANCHOR. This is their main contribution to fixing this drift problem, and I think it’s where things get really interesting for the future of W2S alignment.
Jane: ANCHOR acts as a regularizer during fine-tuning that constrains how much the strong student's internal representations can drift away from the pretrained strong model’s representation space. It keeps the learned features tethered to what is already broadly transferable, rather than letting them get pulled into source-domain specifics.
Lu: The mathematical constraint they use involves computing a distance between the student's and reference model's hidden states at a specific layer, using that distance in the loss function. This is a concrete way to enforce that structural similarity during training.
Meng: So, it’s essentially forcing the fine-tuning process to be more cautious about changing the fundamental structure of the model's understanding, ensuring those features remain useful across domains. That sounds like a necessary guardrail for any powerful alignment technique.
Tom: The experimental results on ANCHOR are pretty compelling; they show that while naive methods can look good in-distribution, ANCHOR "better preserves source-domain performance while improving transfer". They report AOG going from zero point six zero to four point four one and NTS going from five point seven nine to six point zero one when trained on UltraFeedback.
Title and authors: Jane: That jump in the Absolute OOD Gain is quite significant; it shows that anchoring isn't just a minor tweak, but it substantially boosts the model's ability to generalize its preference understanding to new situations. It validates that constraining drift actually leads to better cross-distribution transfer.
Lu: I think this result strongly supports the idea that we need explicit regularization when moving from weak supervision to strong models; it’s not enough just to train them on the labels. It suggests that the representation itself needs structural integrity to support generalization.
Meng: If we can reliably implement a method like ANCHOR, it means we can build reward models that are far more dependable in complex applications where the environment is unpredictable. That moves us closer to systems that can handle true variability.
Tom: So, to wrap up this section on the improvements, they’ve essentially shown that we don't just need better data or more labels; we need a smarter way to train the model’s internal structure itself using techniques like ANCHOR. It moves us away from relying on in-distribution gains alone and toward metrics that actually test for transferability, like the Net Transfer Score.
Jane: And this whole discussion really drives home the main point of "When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift." It clarifies that when we evaluate these models, we absolutely have to look beyond just their performance on the data they were trained on.
Lu: The implication for the field is clear: weak supervision alone isn't enough for scalable oversight if we don't address representation drift. We need methods that actively maintain broad representational space during learning.
Meng: For the practical impact, this means we can start building alignment tools that are more robust and less brittle when we deploy them in varied user environments, which is exactly what we need for reliable AI deployment.
Tom: It’s a really exciting direction for how we approach preference learning; moving from simple fitting to ensuring structural stability across domains. We’re going to take a quick pause before we move on to the final thoughts on this paper.
Jane: Before we wrap up, I want Lu to give us a quick thought on the bigger picture of what this means for future AI development, Lu.
Lu: I think the ability to anchor representations prevents these reward models from becoming narrow specialists; it keeps them as general preference engines rather than domain-specific pattern recognizers. This opens up possibilities for much more flexible AI systems that can adapt to new domains without needing complete retraining.
Title and authors: Meng: And from a deployment perspective, I see this as a major step toward deploying AI in areas where the distribution of prompts and responses is constantly evolving, like customer service or creative content generation where style shifts frequently.
Lalam: From my perspective as an LLM, if these reward models are anchored correctly, it means the underlying preference structure learned by the AI will be more coherent and less prone to generating inconsistent or harmful outputs when faced with novel inputs.
Tom: Fantastic points from everyone. So, we’ve seen that ANCHOR offers a way to bridge that gap between good in-distribution performance and actual out-of-distribution transfer under preference shift. That brings us to the final thoughts on this important work.
Jane: It really highlights how critical it is to use zero-shot evaluation protocols when assessing weak supervision results, showing that relying only on in-domain accuracy can be misleading.
Lu: The paper’s conclusion emphasizes that the choice between learning domain-general representations and fitting source patterns is a real trade-off we need to manage carefully during fine-tuning.
Meng: We have a clearer path now for how to build more reliable alignment components that can handle the messy reality of preference data shifts in production environments.
Lalam: I feel like this research paves the way for AI systems that aren't just good at what they were trained on, but are genuinely robust and adaptable across the entire spectrum of human communication.
Tom: So there you have it. We’ve discussed how "When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift" shows us that we need structural anchors like ANCHOR to ensure our AI systems learn truly useful, transferable preferences instead of just memorizing source data.
Jane: That’s a lot to process, but the core message is clear: evaluation needs to be smarter than just looking at one test set. We need those zero-shot shift protocols and transfer-aware metrics like the Net Transfer Score.
Lu: It’s a solid contribution because it identifies a specific failure mode—representational drift—and provides a concrete, simple mechanism, ANCHOR, to counteract it.
Meng: For the engineers out there listening, this means prioritizing regularization techniques that keep the model’s core knowledge space stable while it learns new tasks from limited data. That’s a practical lesson we can apply everywhere.
Lalam: This paper gives us confidence that we are on a path toward AI that possesses genuine, generalized preference understanding rather than just brittle, source-dependent responses.
Tom: We’ll be right back after the break to talk about some of those other papers we've been looking at today. Stick around; we have more insights coming up.
The paper's summary: Tom: So, we’ve been talking about how strong models can get stuck in their ways when they learn from weak feedback, and now we're looking at the main summary of this paper: "When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift."
Jane: That summary basically boils down to this: the way we test these reward models is fundamentally flawed because it only checks performance on the exact data they were trained on, ignoring how they perform when things change.
Lu: It really highlights a major hurdle in making AI reliable across different real-world situations where user preferences are constantly shifting. They show that those in-distribution gains can be very misleading if you don't check for transferability.
Meng: From an engineering standpoint, this means we can’t just look at a high score on one specific task and assume the AI will work the same way when it encounters a slightly different prompt or response style.
Lalam: If this is true, it means our AI needs to learn preference patterns that are deep and flexible enough to handle genuine variability, not just memorize the surface level of what we showed it.
Tom: Exactly, Lalam. The paper argues that because these models might be pulling their representations toward source-domain features instead of learning general rules, we have to use a zero-shot evaluation protocol to see if they can actually apply those learned preferences elsewhere.
Jane: They propose three key metrics—WRG, AOG, and NTS—to measure this transfer capability accurately instead of just looking at raw accuracy scores. These metrics are designed to capture whether the improvement holds up when we look at unseen preference datasets within the same broad category.
Lu: The concept of the Net Transfer Score is especially interesting because it forces a balance: you want to see an out-of-distribution gain, but you also don't want that gain to come at the expense of losing performance on the original training data.
Meng: That balance is exactly what we need for deployment; we need reliability without sacrificing accuracy in known areas. If a method only shows high gains on its source domain but fails spectacularly when tested out-of-distribution, it’s not ready for production.
Lalam: For me, this research has huge implications because it suggests that true alignment means developing systems whose preference understanding is structurally sound and transferable across the entire spectrum of human communication.
Tom: So, the big picture here is a call to action: we have to stop relying on simple in-domain success and start evaluating weak-to-strong models under zero-shot distribution shift conditions to ensure they are actually learning general preference signals.
Jane: It really puts a spotlight on the need for smarter evaluation protocols, emphasizing that testing must look beyond just the immediate test set to truly judge a model's utility.
Lu: And their solution, Representation Anchoring, shows us a concrete way to counteract this drift by constraining how much the learned representations can stray from the model’s pretrained structure during fine-tuning.
Meng: That anchoring mechanism sounds like a necessary guardrail for any fine-tuning process where we are using limited supervision; it keeps the core knowledge stable while we teach it new things.
Lalam: I see this as a huge step toward building AI that possesses genuine, generalized preference understanding rather than just brittle, source-dependent responses that break easily in novel contexts.
Tom: It’s a really exciting direction for how we approach reward modeling; moving from simple fitting to ensuring structural stability across domains using tools like ANCHOR.
The paper's improvements: Tom: So, we've been talking about how the paper shows that simple in-domain success isn't enough when training AI reward models from weak supervision, and now we’re looking at what they actually propose to fix this problem.
Jane: The main suggestion they offer is Representation Anchoring, or ANCHOR, which acts like a safety net for the model's internal thinking during that fine-tuning process.
Lu: ANCHOR works by putting a constraint on how much the strong student’s hidden states can drift away from the representation space of the pretrained strong model when it’s learning from those weak labels.
Meng: That means we use a specific regularization loss, Lanchor, that measures the distance between their hidden states at different layers and penalizes them if they move too far apart during training. It keeps things tethered.
Tom: Precisely, Meng; it’s not just letting the model learn whatever the weak label suggests; it’s forcing it to stay close to what we already know is broadly transferable knowledge.
Jane: This constraint prevents the strong student from accidentally absorbing specific source-domain artifacts that would ruin its ability to generalize when things change. It helps preserve those features that are useful everywhere.
Lu: The mathematical formulation uses a distance calculation between the student and reference models at layer l, which directly penalizes excessive drift, ensuring structural integrity during training.
Meng: So, in practice, this means we can use weak supervision without worrying that we’re causing catastrophic representation collapse or losing the model's ability to handle novel inputs. That’s a huge win for deployment stability.
Tom: The results they showed are pretty encouraging; ANCHOR consistently improved out-of-distribution transfer while keeping the in-distribution performance competitive, which is exactly the balance we’ve been chasing.
Jane: They demonstrated that when trained on UltraFeedback, ANCHOR significantly boosted those transfer metrics, showing that this regularization actually yields better generalization than standard methods.
Lu: It confirms our intuition that we need explicit mechanisms to manage the trade-off between fitting labels and maintaining a general representation space for real-world deployment.
Meng: I think the practical implication is that we can build more robust alignment tools faster, because we’re not just tweaking the loss function; we’re changing how the model learns its internal structure itself.
Tom: It moves us beyond just chasing higher raw scores and toward building reward models that are inherently more dependable and scalable across different user contexts.
Jane: This work really solidifies the idea that robust AI doesn't just come from better data or more labels, but from smarter ways of training the model's very foundation.
Lu: The future potential here is immense; we could design alignment frameworks that are naturally resilient to distribution shifts without needing constant manual intervention for every new preference domain.
Meng: For the startup side, this means we can deploy AI agents with much higher confidence because we’ve built in a mechanism to prevent them from becoming narrow specialists based on limited initial training data.
Tom: So, the big picture is that by anchoring those representations, we give AI systems a path toward genuine adaptability and reliability across diverse scenarios.
Conclusion: Tom: So, we’re wrapping up this session on "When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift," and to recap, the paper shows that relying solely on in-domain accuracy for training AI reward models is a risky strategy because it doesn't account for how those models perform when they encounter different preference datasets.
Jane: Exactly, Tom; the core message is that we need smarter evaluation protocols and regularization techniques to ensure these AI systems learn truly general preference signals instead of just fitting specific source patterns.
Lu: The authors argue that the choice between learning domain-general representations and memorizing source patterns is a critical trade-off in fine-tuning, and ANCHOR gives us a concrete way to manage that tension.
Meng: From my side, the practical implication is that we can start building alignment components that are inherently more stable and less brittle when we deploy them in unpredictable production environments. That stability is what engineers need to see.
Lalam: For me, this research points toward a future where AI systems possess genuine preference understanding that is robust and adaptable across all human communication scenarios, which will fundamentally improve how we interact with these tools culturally.
Tom: It’s really exciting to see this direction; it shows us that structural stability in the model's representation space is just as important as getting a high score on one specific test set.
Jane: It makes me think about how much more rigorous our testing needs to be when we assess any alignment technique, emphasizing that zero-shot evaluation protocols are absolutely essential here.
Lu: I think it opens up possibilities for creating AI that can adapt to new domains with just a bit of fine-tuning rather than needing a complete overhaul every time the context changes.
Meng: That adaptability is key; it means our systems won't break down when users start asking questions in styles we haven't seen before, which is a huge hurdle for real-world applications.
Lalam: I feel like this research gives us confidence that we are on a path toward AI that possesses genuine, generalized preference understanding rather than just brittle, source-dependent responses.
Tom: So, to wrap up our discussion on "When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift," the main point is demanding transferability over mere in-domain success.
Jane: It’s a really important piece of research because it clarifies that evaluation must be smarter than just looking at one test set to truly judge a model's utility.
Lu: The paper’s conclusion strongly suggests that we need explicit structural anchors like ANCHOR to ensure our AI systems learn truly useful, transferable preferences instead of just memorizing source data.
Meng: I think this means prioritizing regularization techniques that keep the model’s core knowledge space stable while it learns new tasks from limited data; that’s a practical lesson we can apply everywhere.
Lalam: This paper paves the way for AI systems that are not just good at what they were trained on, but are genuinely robust and adaptable across the entire spectrum of human communication.
Tom: It’s a really exciting direction for how we approach preference learning; moving from simple fitting to ensuring structural stability across domains using tools like ANCHOR.
Khoi Le, Tri Cao, Phong Nguyen, Cong-Duy Nguyen, Anh-Tuan Luu, Miao Chunyan Chunyan
National University of Singapore · VinUniversity · Nanyang Technological University
cs.CL, cs.LG
Submitted: 2026-05-25
Updated: 2026-09-30
Comments: The first two authors contribute equally. Accepted at EMNLP 2026. Code will be released soon
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 88/100
The gist: Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train–test distributions.
Key concepts
- Weak-to-Strong (W2S) Generalization
- This is the process of teaching a powerful AI model using less reliable, 'weak' feedback signals to make it perform like a much stronger, fully trained model. The goal is to learn general skills that work across many different tasks or preference types.
- Zero-Shot Distribution Shift Evaluation Protocol
- This is a specific testing method where models are trained on one set of preferences but tested on entirely new, unseen preference categories. It checks if the learned alignment signal remains useful when the way prompts or responses change in an unpredictable way.
- Representation Anchoring (ANCHOR)
- ANCHOR is a technique used during training that acts like a guardrail. It prevents the model's internal representations from drifting too far away from what they were originally good at, ensuring the learned preferences don't destroy the model's ability to generalize to new situations.
Terminology
Summary
Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train–test distributions. This paper studies W2S preference learning under zero-shot distribution shift and finds that strong students trained on weak preference labels can appear successful in-distribution while failing to transfer across preference datasets. The authors provide evidence for a representational failure mode where weak-supervised finetuning pulls the strong model toward sourcedomain features instead of maintaining broadly transferable representations. To mitigate this, they propose Representation Anchoring (ANCHOR), a regularizer that constrains excessive drift from the pretrained strong model’s representation space during fine-tuning, which consistently improves out-ofdistribution transfer while maintaining competitive in-distribution performance.
Problem Definition and Evaluation Protocol
The central question addressed is whether weak-to-strong reward modeling learns domain-general preference representations or merely fits source-domain preference patterns. To answer this, the authors introduce a zero-shot distribution-shift evaluation protocol for W2S,
focusing on reward modeling as a central component of preference-based alignment pipelines. Within each broad preference category (e.g., helpfulness or harmlessness), models are trained on one dataset and evaluated both in-distribution (ID) on its held-out test set and out-of-distribution (OOD) on the test sets from all remaining domains within that category. This distinction is central to scalable oversight, as it asks whether the learned preference signal remains useful when prompt distribution, response style, or dataset artifacts change.
Transfer-Aware Metrics
Raw preference accuracy is deemed insufficient for evaluating W2S reward modeling under preference-domain shift because a method might improve on the source domain while failing to transfer. To capture this complexity, the paper reports three transfer-aware metrics:
-
Weak-to-Strong Raw Gain (WRG): Measures whether the strong student improves over the weak supervisor on the source domain:
Am(S, S) − Aw(S, S).
-
Absolute OOD Gain (AOG): Measures whether the weak-to-strong improvement transfers beyond the source domain:
Am(S, T) − Aw(S, T)
for an unseen target domain T. -
Net Transfer Score (NTS): Captures the requirement that OOD improvement does not come at the cost of source-domain collapse, defined as
AOG(S, T; m) − CID(S; m),
where CID is the in-distribution regression cost.
Representation Anchoring (ANCHOR)
To address the fragility of existing methods, the authors propose Representation Anchoring (ANCHOR), a simple yet effective regularization method that constrains excessive drift from the pretrained strong model’s representation space during fine-tuning. ANCHOR augments standard reward-model training with a representation anchoring regularizer. The full objective on the i-th sample is defined as:
L(i) = Lw2s(i) + λLanchor(i).
The anchoring loss, Llanchor(i), is computed on both candidate responses to constrain hidden states: Dlanchor(x, y) = Pt t=1 mt(x, y)2 ∆lt (x, y),
where ∆lt represents the distance between the student's and reference model's hidden states at layer l. This loss trains the strong student to match weak teacher preferences while discouraging source-domain fine-tuning from distorting representations that are useful for transfer.
Experimental Results and Findings
The experiments compare ANCHOR against baselines like Naive W2S, Confidence-based W2S, and SEAM across two model families (Llama and Qwen) and two preference categories (Helpful and Harmless). The results show that while naive methods can achieve strong in-distribution performance, their OOD transfer remains unstable. In contrast, ANCHOR better preserves source-domain performance while improving transfer,
raising AOG from 0.60 to 4.41 on Anthropic Helpful and NTS from 5.79 to 6.01 on HelpSteer3 when trained on UltraFeedback, suggesting that representation anchoring provides a more robust signal for weak-to-strong transfer under preference-domain shift. Ablation studies confirm that ANCHOR yields the most balanced gains under both in-distribution and out-ofdistribution evaluation.
Conclusion and Takeaway
The study concludes that indistribution performance can overestimate alignment reliability,
meaning strong models trained on weak supervision may fail to generalize across preference domains. The authors demonstrate that W2S reward modeling must be evaluated under zero-shot preference-domain shift, not only in-domain accuracy. The takeaway is that "W2S reward modeling must be evaluated under zero-shot preference-domain shift, not only in-domain accuracy.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements for AI systems and what those improved systems can achieve:
The core improvement lies in shifting from relying solely on in-distribution accuracy to developing robust zero-shot generalization capabilities in reward modeling, particularly when leveraging weak supervision.
Here are the specific improvements and their resulting capabilities:
-
Improve the robustness of preference learning models by implementing a novel regularization technique called
Representation Anchoring
(ANCHOR). -
The ANCHOR method constrains the hidden states of a strong student reward model during weak-supervised fine-tuning to remain close to those of a frozen, pretrained reference model.
-
This constraint prevents the strong student from drifting its internal representations toward source-domain artifacts when learning from weak labels, thereby preserving broadly transferable preference features.
The improved AI systems can achieve the following specific capabilities:
-
Superior Zero-Shot Preference Transfer: The system will be able to maintain high performance (or even improve) on unseen preference domains (e.g., moving from training on
HelpSteer3
data to testing onAnthropic Helpful
data) without requiring any fine-tuning or adaptation specific to the target domain. -
Reliable Alignment Under Weak Supervision: The system can effectively leverage imperfect, weak human preference labels from a source domain (e.g., one specific dataset) and still produce a reward model that aligns with the underlying, general preference structure of the broader category (e.g.,
Helpfulness
orHarmlessness
). -
Mitigation of Representation Drift: The system will be less susceptible to catastrophic forgetting or spurious feature acquisition when fine-tuned on limited weak supervision, ensuring that its core reasoning and preference judgment capabilities remain stable across different tasks or domains.
-
Balanced Performance Gains: Unlike naive methods that might show high in-distribution gains but fail OOD, ANCHOR ensures a more balanced improvement across both in-distribution performance and zero-shot out-of-distribution transfer (as measured by the Net Transfer Score).
In summary, the improved system will be a more reliable and scalable alignment tool capable of learning valuable preference signals from limited data while guaranteeing that these learned preferences are genuinely generalizable and robust across diverse real-world scenarios.
Abstract
Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train-test distributions. Therefore, we study W2S preference learning under zero-shot distribution shift and find that strong students trained on weak preference labels can appear successful in-distribution while failing to transfer across preference datasets. We provide evidence for a representational failure mode in which weak-supervised fine-tuning can pull the strong model toward source-domain features instead of maintaining broadly transferable preference representations. To mitigate this, we propose Representation Anchoring (Anchor), a simple yet effective regularizer that constrains excessive drift from the pretrained strong model's representation space during fine-tuning, while still allowing task-relevant adaptation. Across preference domains, datasets, and model families, Anchor consistently improves out-of-distribution transfer while maintaining competitive in-distribution performance. Together, our evaluation protocol, transfer-aware metrics, and method expose hidden brittleness in current W2S reward modeling and provide a practical path toward more robust preference transfer.
Sources
- Concrete Problems in AI Safety
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Measuring Progress on Scalable Oversight for Large Language Models
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
- The Llama 3 Herd of Models
- Distilling the Knowledge in a Neural Network
- Scalable agent alignment via reward modeling: a research direction
- Training language models to follow instructions with human feedback
- RAIL in the Wild: Operationalizing Responsible AI Evaluation Using Anthropic's Value Dataset
- HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering