Auditing Proxy-Based Validation Across Text Spans
cs.LG, cs.CL
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: 63 pages, 7 figures, 38 tables. Code: https://github.com/wdi1024/rlc-audit
Code: https://github.com/wdi1024/rlc-audit
License: http://creativecommons.org/licenses/by/4.0/
The gist: Evaluation scores are often validated by their agreement with inexpensive proxy labels.
Terminology
Abstract
Evaluation scores are often validated by their agreement with inexpensive proxy labels. When the score and the proxy are computed from the same text span, however, that agreement can arise from surface evidence the two share rather than from the semantic construct the proxy is meant to represent. We make the distinction explicit by declaring the score, its span, the proxy and the target construct as a validation contract, then re-evaluating that proxy rule strictly outside the scored span. In a controlled HotpotQA correctness experiment varying only the shared text boundary, the score agrees with its proxy far better than with correctness at a 50-character prefix: the gap is +0.184, collapsing to at most +0.045 from 120 characters onward. At that short prefix the score still predicts whether the answer string appears later (AUC 0.634) while an equivalence test places its agreement with correctness at chance, so the reported proxy agreement does not establish that the score ranks correctness. On OR-Bench, suppressing each model's recurring opening templates removes most of the score's association with the refusal proxy, while matched-volume deletion removes almost none and construct agreement stays at chance. Only three of eleven external contracts support the off-span control, and none of the routing studies we sampled released the generations it needs. We therefore ask that a proxy-based validation claim declare the span each label is read from, report the construct agreement beside the proxy agreement, and release the generations that let the proxy be re-read off the scored span.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning
- A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
- What Does a Routing Oracle Measure Under Stochastic Decoding? Coupling, Scorer Choice, and Single-Commit Ceilings
- Training Verifiers to Solve Math Word Problems
- Gemma: Open Models Based on Gemini Research and Technology
- Cross-Model Disagreement as a Label-Free Correctness Signal
- The Llama 3 Herd of Models
- RouterBench: A Benchmark for Multi-LLM Routing System
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Mistral 7B
- Building to the Test: Coding Agents Deliver What You Check, Not What You Requested
- When the Defense Writes the Refusal: Auditing Keyword-Scored Evaluation of Inference-Time Defenses for Multimodal Large Language Models
- Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs
- EST-PRM: Stress-Testing Process Reward Models Before They Become Load-Bearing
- FreePRM: Training Process Reward Models Without Ground Truth Process Labels
- SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Qwen3 Technical Report
- Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks