Robust Failure, Conservative Repair: Textual Knowledge Distillation from Cross-Model Failures
cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 16 pages, 2 figures. Accepted to EMNLP 2026 (Main track)
Code: https://github.com/ChicagoHAI/RFCR
License: http://creativecommons.org/licenses/by/4.0/
The gist: Failure-based textual knowledge distillation aims to discover gaps in a model's knowledge by examining its task errors.
Terminology
Abstract
Failure-based textual knowledge distillation aims to discover gaps in a model's knowledge by examining its task errors. The distilled knowledge can be useful for the reasoning of both this model ("source model") and other models. However, this transfer of knowledge may not be stable. We define a rule atom to be a standalone rule injected into a model's textual input at inference time. A rule atom can encode transferable task knowledge or model-specific reasoning patches that can confuse other models. Also, the injected rule atoms can be misapplied to unrelated cases, causing the model to incorrectly flip its answer based on irrelevant information. Building on a pipeline that distills training examples into task-specific cheat sheets that aid model reasoning, we examine when failure-derived rules can improve these cheat sheets. Our early experiment shows rule distillation from a single model's failures underperforms the baseline cheat sheet on non-source model families. This motivates Robust Failure, Conservative Repair (RFCR), a textual distillation procedure that derives rules from failures shared across models, sharpens their application boundaries using boundary cases, and abstains when no useful rule is found. On a 400-item BIG-Bench Hard task set, RFCR improves the baseline cheat sheets from 68.50% to 71.25% (+2.75 pp; 95% CI [+1.25,+4.50]) without performance degradation on previously correct cases. Ablations and cross-model diagnostics support that accuracy gains come from both new knowledge injection and strict rule-application control.
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
- Large Language Models as Optimizers
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection