PolicyLong: Towards On-Policy Context Extension
cs.LG, cs.AI
Submitted: 2026-04-09
Updated: 2026-08-28
Comments: Work in progress. Correspondence to ucaswu@tencent.com or wuxing@iie.ac.cn
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
- What is Wrong with Perplexity for Long-context Language Modeling?
- Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language Model
- LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions
- NExtLong: Toward Effective Long-Context Training without Long Documents
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Reinforcement Learning via Self-Distillation
- dots.llm1 Technical Report
- EntropyLong: Effective Long-Context Training via Predictive Uncertainty
- DeepSeek-V3 Technical Report
- Lost in the Middle: How Language Models Use Long Contexts
- POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
- Self-Distillation Enables Continual Learning
- jina-embeddings-v3: Multilingual Embeddings With Task LoRA
- Effective Long-Context Scaling of Foundation Models
- HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks