A Hierarchical Energy-Based Model for Multimodal Cognition
Subir Varma
q-bio.NC, cs.AI
Submitted: 2026-08-08
Updated: 2026-08-14
Comments: 48 pages, 14 figures
Project page: https://subirvarma.github.io/GeneralCognitics/2026/07/15/statmech4.html
License: http://creativecommons.org/licenses/by/4.0/
The gist: We propose IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical, energy-based model of multimodal cognition that extends a previously proposed single-modality
Terminology
Abstract
We propose IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical, energy-based model of multimodal cognition that extends a previously proposed single-modality model (LEPP) to integrate vision and language. Following the view that generative neural networks are effective theories of cognitive dynamics, analogous to how statistical mechanics relates to thermodynamics, IM-LEPP models cognition as latent states flowing through learned energy landscapes rather than as an account of neural circuitry. The architecture is a hub-and-spoke hierarchy, grounded in the controlled semantic cognition framework of Lambon Ralph et al., in which predictive-coding pipelines for visual objects, scenes, and linguistic units converge on a shared amodal hub modeled on the anterior temporal lobe. Each pipeline's own prediction is conditioned by, rather than overwritten by, the current hub state, preserving pipeline-specific identity while letting every prediction reflect the full multimodal context. We show this architecture gives a mechanistic account of attentional phenomena such as inattentional blindness and Necker-cube bistability, and that its structure recovers or motivates independently established findings in psycholinguistics, including surprisal theory, the N400/P600 ERP components, and garden-path reanalysis, alongside a falsifiable contrast with transformer language models on trajectory-sensitivity in next-word prediction. We also discuss data-efficient language acquisition relative to LLMs, outline a semantic/episodic memory subsystem, situate the model against predictive coding, the free-energy principle, JEPA, and Hierarchical Temporal Memory, and propose concrete experimental predictions to test its central claims Key Words: predictive processing; predictive coding; energy-based models; diffusion models; effective theory; computational neuroscience.
Sources
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- Revisiting Feature Prediction for Learning Visual Representations from Video
- Trajectory Dynamics in Language Model Hidden States Predict Human Processing Costs Beyond Surprisal
- Improving language models by retrieving from trillions of tokens
- Sampling with flows, diffusion and autoregressive neural networks: A spin-glass perspective
- Why Neurons Have Thousands of Synapses, A Theory of Sequence Memory in Neocortex
- Lost in the Middle: How Language Models Use Long Contexts
- Latent Diffusion for Language Generation
- Decomposition of surprisal: Unified computational model of ERP components in language processing
- Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
Related papers
- BrainWave: A Brain Signal Foundation Model for Clinical Applications
- Toward Robust, Reproducible, and Widely Accessible Intracranial Speech Brain-Computer Interfaces: A Comprehensive Narrative Review of Neural Mechanisms, Hardware, Algorithms, Evaluation, Clinical Pathways and Future Directions
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- Emergence of psychopathological computations in large language models
- NeuroAI and Beyond: Bridging Between Advances in Neuroscience and Artificial Intelligence
- Attraction to hierarchical feature memory explains orientation bias