Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility
Jo-Ku Cheng, Nikolaos Aletras, Marco Valentino
cs.CL, cs.AI, cs.LG
Submitted: 2026-08-04
License: http://creativecommons.org/licenses/by/4.0/
The gist: Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition.
Terminology
Abstract
Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretraining tasks, such as Dyck and procedural algorithms, rely on narrow primitives that fail to capture the expressive capacity of natural language. Moreover, prior studies remain restricted to relatively small token budgets, offering limited insight into skill emergence and representational dynamics. To address these limitations, we propose logic pre-pretraining (Logic-PPT) as a principled initialization strategy, leveraging formal derivations to impart richer structural and linguistic biases. Formal derivations require abstract mechanisms that are central to natural language, simultaneously binding variables, connecting quantifiers and relational dependencies, and composing predicate-argument structures over long contexts. Scaling our evaluation to a 100B-token regime, logic pre-pretraining substantially accelerates skill acquisition in LMs, achieving 80% accuracy on linguistic tasks with 36B fewer tokens than standard initialization, and outperforming alternative pre-pretraining baselines. Mechanistically, formal derivations induce persistent structural reorganization, distinctively characterized by a lower-rank, spectrally concentrated representation space. Crucially, we show that this internal geometry enables improved model compressibility via pruning, matching the dense baseline performance even at about 33% sparsity.
Sources
- Phi-4 Technical Report
- Mistral 7B
- Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
- Reasoning Core: A Scalable Procedural Data Generation Suite for Symbolic Pre-training and Post-Training
- Verbalizable Representations Form a Global Workspace in Language Models
- Training Language Models via Neural Cellular Automata
- What do Language Models Learn and When? The Implicit Curriculum Hypothesis
- Qwen3 Technical Report
- 2 OLMo 2 Furious
- Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
- Pretraining Induces a Reusable Spectral Basis for Downstream Task Adaptation
- Approaching Deep Learning through the Spectral Dynamics of Weights
- A Simple and Effective Pruning Approach for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering