Generalization Dynamics of LM Pre-training
cs.CL
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- Taken out of context: On measuring situational awareness in LLMs
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- Nexus: Same Pretraining Loss, Better Downstream Generalization via Common Minima
- Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- In-context Learning and Induction Heads
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
- Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
- Olmo 3
- Simple Mechanistic Explanations for Out-Of-Context Reasoning
- STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
- Larger language models do in-context learning differently
- Language Models Learn to Mislead Humans via RLHF
- Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
- Which Attention Heads Matter for In-Context Learning?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering