Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
cs.AI, cs.LG
Submitted: 2026-09-16
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these laws to remarkable results,
Terminology
Abstract
The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these laws to remarkable results, activating only a fraction of an enormous stored parameter bank for each token. That success is built on static pretraining data. A deployed model faces a different world, where much of the data that would make it more useful is not in its training set but in the live interaction it is currently handling, such as the facts a user supplies or the corrections they give. A conventional model cannot learn from this data, because its weights are frozen after training. Instead, the knowledge and behaviour supplied at run time are placed in the prompt, by retrieval or instruction, and re-read on every request only to be discarded once the request ends. We ask how an architecture could learn from live interaction by writing it into its weights. Taking inspiration from MoE, we propose the Infinite-Parameter LLM. A compact hypernetwork turns the data given at run time into a low-rank modulation of a shared base network, so the feed-forward weights are generated from live data rather than stored in a fixed bank. Where prior weight generators read the context once and freeze, we carry a Bayesian belief over the generator's latent code and update it online, so the effective weight is re-derived from that evolving belief as the session proceeds rather than fixed after one read. The stored footprint stays fixed, yet the weights the model can compile are effectively infinite. For the knowledge and behaviour supplied at run time, carrying them in the weights rather than the prompt is amortized in compute, frees the context window, persists across turns, and can generalise better than in-context use. We specify an evaluation protocol that tests exactly this against in-context learning and retrieval.
Sources
- Zhyper: Factorized Hypernetworks for Conditioned LLM Fine-Tuning
- Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
- Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning
- The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
- Using Fast Weights to Attend to the Recent Past
- Titans: Learning to Memorize at Test Time
- Conditional Computation in Neural Networks for faster models
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Language Models are Few-Shot Learners
- X-LoRA: Mixture of Low-Rank Adapter Experts, a Flexible Framework for Large Language Models with Applications in Protein Mechanics and Molecular Design
- Low-rank extended Kalman filtering for online learning of neural networks from streaming data
- Text-to-LoRA: Instant Transformer Adaption
- Doc-to-LoRA: Learning to Instantly Internalize Contexts
- Retraining-Free Merging of Sparse MoE via Hierarchical Clustering
- Unified Scaling Laws for Routed Language Models
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- DeepSeek-V3 Technical Report
- Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models
- Bayesian Mixture of Experts For Large Language Models
- LoRAMoE: Alleviate World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection