Math for AI safety: an invitation for mathematicians
math.HO, cs.AI, cs.LG
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 33 pages, 5 figures
Code: https://github.com/ColombanD/open-source-game-theory
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- SGD learning on neural networks: leap complexity and saddle-to-saddle dynamics
- Remarks on the disproof of the unit distance conjecture
- Concentration of tempered posteriors and of their variational approximations
- Refusal in Language Models Is Mediated by a Single Direction
- Occam's razor is insufficient to infer the preferences of irrational agents
- Robust Cooperation in the Prisoner's Dilemma: Program Equilibrium via Provability Logic
- Sparse Representation of a Polytope and Recovery of Sparse Signals and Low-rank Matrices
- A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations
- Characterising Simulation-Based Program Equilibria
- Parametric Bounded L\"ob's Theorem and Robust Cooperation of Bounded Agents
- Cooperative and uncooperative institution designs: Surprises and problems in open-source game theory
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Toy Models of Superposition
- On the uniqueness and stability of dictionaries for sparse representation of noisy signals
- Logical Induction
- Sparse and spurious: dictionary learning with noise and outliers
- Cooperative Inverse Reinforcement Learning
- Progress in Formalizing Sphere Packing in Dimension 8
- When can dictionary learning uniquely recover sparse data from subsamples?
- The Bayesian Learning Rule
Related papers
- A unified interpretation of probability
- Remembering Solomon Marcus
- LLAMA LIMA: A Living Meta-Analysis on the Effects of Generative AI on Learning Mathematics
- A machine-checked proof of the Dong-Yang classification of optimal (n,4) binary codes for BSCs
- If you can distinguish, you can express: Galois theory, Stone--Weierstrass, machine learning, and linguistics
- Explanations, Prompts, and Formalizations: Arguments for New Norms in LLM-Enabled Mathematical Research