Learning to Learn from Context: Synthetic Training from Perturbed Public Documents
cs.CL, cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/Tencent-Hunyuan/CL-bench
Terminology
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Kimi K3: Open Frontier Intelligence
- ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
- On the Measure of Intelligence
- Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
- CL-bench Life: Can Language Models Learn from Real-Life Context?
- CL-bench: A Benchmark for Context Learning
- Gemma 4 Technical Report
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models
- Humanity's Last Exam
- Modular Techniques for Synthetic Long-Context Data Generation in Language Model Training and Evaluation
- Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries
- OJBench: A Competition Level Code Benchmark For Large Language Models
- Measuring short-form factuality in large language models
- GLM-5: from Vibe Coding to Agentic Engineering
- LongSkywork: A Training Recipe for Efficiently Extending Context Length in Large Language Models
- Group Sequence Policy Optimization
- Instruction-Following Evaluation for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering