Cross-attention encoding models reveal dynamic spatiotemporal routing across human higher visual cortex
q-bio.NC, cs.AI, cs.LG
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- Transformer brain encoders explain human high-level visual responses
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Perception Encoder: The best visual embeddings are not at the output of the network
- SAM 3: Segment Anything with Concepts
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- The Algonauts Project 2025 Challenge: How the Human Brain Makes Sense of Multimodal Movies
- Animate Your Thoughts: Decoupled Reconstruction of Dynamic Natural Vision from Slow Brain Activity
- DINOv3
- NEvo: Neural-Guided Evolutionary Video Synthesis for Dynamic Visual Selectivity
Related papers
- BrainWave: A Brain Signal Foundation Model for Clinical Applications
- Toward Robust, Reproducible, and Widely Accessible Intracranial Speech Brain-Computer Interfaces: A Comprehensive Narrative Review of Neural Mechanisms, Hardware, Algorithms, Evaluation, Clinical Pathways and Future Directions
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- Emergence of psychopathological computations in large language models
- NeuroAI and Beyond: Bridging Between Advances in Neuroscience and Artificial Intelligence
- Attraction to hierarchical feature memory explains orientation bias