Distilling Directional Verification
cs.CL, cs.AI, cs.LG
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/js-lee-AI/directional-verification
Terminology
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Structured Denoising Diffusion Models in Discrete State-Spaces
- Training Verifiers to Solve Math Word Problems
- RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
- The Llama 3 Herd of Models
- Reinforced Self-Training (ReST) for Language Modeling
- Mistral 7B
- Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs
- Sequence-Level Knowledge Distillation
- FUSE: Ensembling Verifiers with Zero Labeled Data
- DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models
- Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models
- Let's Verify Step by Step
- Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge
- Large Language Diffusion Models
- Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
- Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
- 2 OLMo 2 Furious
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering