CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation
The University of Hong Kong
cs.CL
Submitted: 2026-08-13
Updated: 2026-09-19
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation Summary This paper introduces Counterfactual Relevance for On-Policy Distillation (CROP), a method for selective
Terminology
Summary
CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation
Summary
This paper introduces Counterfactual Relevance for On-Policy Distillation (CROP), a method for selective on-policy distillation (OPD) that allocates token-level supervision based on task relevance, a dimension complementary to existing optimization-need criteria.
Motivation and Gap
Standard OPD supervises a student language model on trajectories sampled from its current policy, assigning equal credit to all response tokens. Selective OPD methods address this by allocating supervision non-uniformly based on estimated training value. However, existing criteria (e.g., uncertainty, teacher–student disagreement, teachability) primarily estimate optimization need
—whether a token is uncertain, differs from the teacher, or is teachable. The paper argues that these signals do not fully characterize supervision value because even tokens that are uncertain or discrepant may reflect input-generic response patterns. This motivates task relevance
as a complementary dimension: whether a token's learning signal depends on the semantic content of the current input.
Method
CROP operationalizes task relevance through a paraphrase-calibrated counterfactual sensitivity margin. For each source prompt, CROP constructs a validated original–paraphrase–counterfactual triplet offline:
-
The paraphrase preserves task-relevant meaning, requested task, and output constraints.
-
The counterfactual changes exactly one explicit material condition while preserving the requested task and output constraints.
The method holds the student's original rollout fixed and re-scores the same response prefixes under all three prompts. For each response position, it computes:
-
Counterfactual sensitivity: JSD between the student distribution under the original prompt and under the counterfactual prompt.
-
Paraphrase sensitivity: JSD between the student distribution under the original prompt and under the paraphrase prompt.
The CROP score is the difference: counterfactual sensitivity minus paraphrase sensitivity. Positions are ranked by this margin, and a batch-global binary mask selects the top positions under a fixed supervised-token budget. The teacher target and OPD objective remain unchanged.
Key Contributions
-
Task relevance for selective OPD: Distinguishes optimization need from task relevance and motivates sensitivity to controlled semantic contrasts as a complementary signal.
-
CROP method: Constructs and validates matched prompt triplets offline, re-scores the same response prefixes under all three prompts, and subtracts paraphrase sensitivity from counterfactual sensitivity using top-K JSD with a residual bucket. Converts scores into a batch-global hard token mask under a fixed budget.
-
End-to-end performance: Across two teacher–student settings, CROP improves aggregate performance by 1.92 and 2.96 points over the strongest non-CROP selector.
Experimental Results
The paper evaluates on two settings: Qwen3-4B → Qwen3-1.7B and Qwen3-8B (GRPO) → Qwen3-4B, using 16,594 validated DAPO-Math-17K prompts. Benchmarks include AIME24, AIME25, MATH-500, GPQA-Diamond, HumanEval, and IFEval.
-
Main results (10% budget): For Qwen3-4B → Qwen3-1.7B, CROP achieves the best average of 47.98, improving over Pure OPD by 3.11 points and over the strongest selective baseline (TIP) by 1.92 points. For Qwen3-8B (GRPO) → Qwen3-4B, CROP-ent achieves the best average of 57.48, while CROP obtains 57.13; both improve over Pure OPD by more than 1.8 points.
-
Selector behavior across entropy ranks: CROP exhibits a secondary concentration in the extreme low-entropy tail: 10.0% of its retained tokens fall in the lowest-entropy decile, compared with 3.6% for CREDIT and nearly zero for uncertainty-oriented selectors. This shows CROP is not reducible to an entropy proxy.
-
Budget sensitivity: CROP achieves the best average at 10%, 15%, and 20% budgets, while CROP-ent attains the overall peak of 47.99 at 25%. Both vary non-monotonically, showing that broadly increasing token coverage does not guarantee better performance.
-
Selector variants: CROP achieves the highest average of 47.98, exceeding CS-OPD (counterfactual change alone) by 1.98 points, PC-OPD (paraphrase sensitivity alone) by 4.89 points, and CROP-Teacher (teacher-side scoring) by 3.25 points.
-
Token-selection controls: At the same 10% budget, CROP-10% improves the average by 2.00 points over Random-10% and by 3.34 points over Bottom-10% (lowest-scoring positions).
Conclusion
CROP is a hard token selector that operationalizes task relevance through matched paraphrase and counterfactual interventions. It consistently improves aggregate performance over full-token OPD and all non-CROP selective baselines in both teacher–student settings. Component comparisons confirm that counterfactual sensitivity provides the primary signal and that paraphrase calibration improves the ranking. The optional CROP-ent variant shows uncertainty can be complementary, though its benefit is setting-dependent. The paper notes limitations: evidence is limited to mathematical triplets and Qwen-based settings, motivating evaluation across model families, domains, and intervention-generation procedures.
Improvements for AI systems
Improvements to AI systems based on this paper:
-
Input-conditional token supervision in distillation: Instead of allocating training signal uniformly or purely by uncertainty/disagreement, the student model can now be trained on tokens whose learning signal is provably tied to the semantic content of the specific input. This yields more efficient distillation, improving downstream task accuracy by 2–3 points at the same compute budget, especially in math and reasoning tasks.
-
Counterfactual sensitivity scoring for selective learning: The AI system can now identify which response positions are genuinely sensitive to changes in the task's material conditions (e.g., numbers, constraints) versus those that are generic or paraphrastic. This enables the system to focus its learning capacity on tokens where the correct output actually depends on the input's meaning, reducing overfitting to surface patterns.
-
Paraphrase-calibrated relevance ranking: By subtracting paraphrase sensitivity from counterfactual sensitivity, the system can distinguish
task-relevant
frominput-generic
uncertainty. This prevents the model from wasting supervision on tokens that are uncertain merely because of linguistic variation, leading to more robust generalization across rephrased prompts. -
Batch-global hard token masking under fixed budget: The improved system can now enforce a global token budget across a batch, selecting the most task-relevant positions for supervision. This makes training more compute-efficient and predictable, allowing the same distillation budget to yield higher performance than random or uncertainty-based token selection.
-
Complementary uncertainty integration (CROP-ent variant): The system can optionally combine task relevance with entropy-based uncertainty to capture both
what matters for this input
andwhat the model is unsure about.
This hybrid selector improves performance further in some settings, enabling adaptive distillation that leverages both signals when beneficial. -
Offline triplet construction for scalable supervision: The system can pre-construct validated original–paraphrase–counterfactual prompt triplets offline, then reuse them across training runs. This reduces online computational overhead while providing high-quality relevance signals, making selective distillation practical for large-scale training.
What the improved AI system can do specifically:
-
Distill a large teacher model into a smaller student with higher accuracy on math benchmarks (AIME, MATH-500) and better generalization to out-of-distribution reasoning tasks.
-
Achieve superior performance at low supervision budgets (10–20% of tokens), making it suitable for resource-constrained deployment.
-
Avoid wasting training signal on tokens that are uncertain but input-generic, leading to faster convergence and more stable training.
-
Provide interpretable token-level relevance scores, enabling developers to audit which parts of a response are truly task-dependent.
-
Transfer this selective distillation approach across model families and domains (with future extension), improving any teacher–student setup where task relevance can be defined via controlled semantic interventions.
Sources
- Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models
- Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
- Entropy-Aware On-Policy Distillation of Language Models
- When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
- Proximal Policy Optimization Algorithms
- From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
- A Survey of On-Policy Distillation for Large Language Models
- Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
- Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation
- On the Position Bias of On-Policy Distillation
- Trust Region On-Policy Distillation
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering