Rubric-Aware On-Policy Self-Distillation for LLM Personalization
cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/SnowCharmQ/GRASP
Terminology
Sources
- GPT-4 Technical Report
- Training LLMs with Reinforcement Learning for Intent-Aware Personalized Question Answering
- VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions
- Rubric-based On-policy Distillation
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Entropy-Aware On-Policy Distillation of Language Models
- LongLaMP: A Benchmark for Personalized Long-form Text Generation
- EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation
- Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- A Survey of Personalized Large Language Models: Progress and Future Directions
- Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation
- Preference-Aware Rubric Learning for Personalized Evaluation
- Integrating Summarization and Retrieval for Enhanced Personalization via Large Language Models
- A Survey of On-Policy Distillation for Large Language Models
- Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
- Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Toward Personalized LLM-Powered Agents: Foundations, Evaluation, and Future Directions
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering