Spike-driven Vision-Language-Action Model
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Spike-driven Vision-Language-Action Model".
Jane: Spike-driven VLA introduces a novel framework that enables end-to-end direct training for robotic manipulation using spiking neural networks, offering an energy-efficient alternative to conventional large Transformer models.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Welcome back to the show, everyone. Today we're diving into something really interesting that just dropped on arXiv, and I'm telling you, this paper is packed with stuff we need to talk about. We’re looking at the Spike-driven Vision-Language-Action Model.
Jane: It sounds like a big deal, Tom. What exactly is the main idea behind this research? Are they trying to do something completely new in how robots learn?
Lu: Well, what's really compelling here is that they're moving away from those massive Transformers that take up so much energy and time, which is a huge problem for real-world robotics. They’re proposing using spiking neural networks instead to make it way more efficient.
Meng: Efficiency is key for us because deploying complex AI on edge devices isn't always feasible with traditional models, so how does this approach actually tackle that?
Lalam: From my perspective, this work suggests a path toward building more culturally sensitive and resource-aware AI systems by focusing on sparse computation instead of dense operations.
Tom: Exactly. So, let's go over the summary of what they've done in the Spike-driven Vision-Language-Action Model paper. What’s the core contribution they’re highlighting?
Jane: They are proposing a framework that enables end-to-end direct training for robotic manipulation using spiking neural networks, which is a major step because it skips those complicated conversion steps you usually have to do with existing architectures.
Lu: The framework is structured around three main parts: first, they develop spiking visual and instruction encoders to capture multimodal inputs into sparse spike representations, then they introduce a Multi-Winner Spike Fusion module for cross-modal interaction, and finally, they use a Spike Action Chunking Transformer for continuous action generation.
Meng: That sounds like a lot of moving parts. Can you simplify what the Multi-Winner Spike Fusion does in practical terms? I'm curious how it handles connecting the visual input with the language instructions.
Jane: The Multi-Winner Spike Fusion module is designed to establish task-relevant cross-modal correspondence by using a bidirectional top-k winner-take-all spike routing to suppress background interference and yield fused memory. They calculate projections and create instruction-guided visual representations and visually grounded instruction representations that they concatenate into a task-conditioned memory.
Tom: That sounds like a clever way to filter out the noise, which is something we always struggle with in multimodal systems. So, what are the actual improvements this method offers over what's currently out there?
Title and authors: Lu: The paper points out that existing spiking VLAs often rely on ANN-to-SNN conversion methods or simply transfer weights from pre-trained ANNs to their spiking counterparts, but those methods usually keep the original ANN architecture and restrict task-specific design.
Jane: What they are proposing is direct training methods, like Wu et al. (two thousand eighteen) or Zhou et al. (two thousand twenty-three), which allow for task-specific spiking architectures and end-to-end learning of representations directly in the spike domain, which is a significant procedural improvement over the methods mentioned earlier.
Meng: So they are focusing on training the architecture itself rather than just adapting an existing one to spiking, right? That would save a lot of time in development.
Tom: Precisely. And they show competitive performance on benchmarks like LIBERO and Meta-World while requiring fewer parameters and lower estimated inference energy than conventional ANN-based VLA models, which is what makes this research so interesting for deployment.
Lalam: For me, the implication here is that we can start building more compact AI systems that can run reliably in environments where computational power is limited, which could really improve how we deploy helpful AI across many different contexts.
Tom: The final part of their work details the Spike Action Chunking Transformer, which takes the fused memory and current robot state to form a state-augmented memory before using spiking cross-attention to predict action chunks in parallel.
Jane: That chunking approach is clever because it allows the system to generate sequences of future actions simultaneously, which leads to more temporally coherent and efficient control policies compared to generating one action at a time.
Lu: They also detail specific components like the Spiking Visual Encoder, which uses a three-stage pyramidal architecture with convolutional mixer blocks and spiking self-attention blocks to produce visual spike tokens.
Meng: Speaking of those components, I see they've pre-trained the visual encoder on ImageNet1K before fine-tuning, which gives it a strong starting point for general vision tasks. That makes sense from an engineering standpoint for getting good initial weights quickly.
Lalam: It's interesting how they combine that large-scale pre-training with the direct training approach; it suggests that foundational visual knowledge can be leveraged effectively even when training in the spiking domain.
Tom: And their results are quite impressive, showing a seventy-two point four percent average success rate on the Meta-World benchmark and a ninety-two point four percent average success rate on LIBERO, which puts them ahead of models like SmolVLA by using only one-third of its parameters.
Title and authors: Jane: That performance metric is strong, especially considering they are using fewer parameters and lower estimated inference energy than those existing ANN-based VLA models.
Lu: They also provide a detailed theoretical energy estimation for the computation, counting synaptic operations as ACs at zero point nine pJ and continuous operations as MACs at four point six pJ, leading to an estimated computation energy of EbSpikeVLA = eACNAC + eMACNMAC.
Tom: That comparison of computational energy is what really sells the efficiency aspect, showing a substantial reduction compared to models that rely on dense MAC operations at every timestep.
Jane: So, in short, this Spike-driven Vision-Language-Action Model offers a way to build VLA systems that are both more capable and much more energy efficient than current state-of-the-art Transformer models.
Lu: It’s important to remember that the paper mentions they also explore robustness under diverse distribution shifts across seven perturbation dimensions on LIBERO-Plus, which suggests good reliability in varied real-world settings.
Meng: That robustness is what we need when we move these systems out of the lab and into actual operational environments where things are never perfectly controlled.
Lalam: I think that focus on robustness makes this model highly valuable because it means the AI can handle real-world unpredictability without needing constant retraining for every minor environmental change.
Tom: So, to wrap up, we've seen how this Spike-driven Vision-Language-Action Model uses direct training in the spike domain and novel fusion techniques to achieve strong performance with significantly lower energy costs.
Jane: It’s a really solid piece of research that shows how sparse computation can lead to high performance without the heavy computational load.
Lu: This work opens up a direction for developing task-specific spiking architectures, which is something we've been aiming for in the SNN field for longer, and it gives us concrete examples of what's possible with VLA directly trained in this paradigm.
Meng: From an engineering standpoint, the focus on end-to-end training means that the whole pipeline is optimized together from day one, which streamlines the development process significantly.
Lalam: I think this paper sets a good precedent for how we can build more compact and reliable AI systems that can be deployed where computational resources are scarce, which is exactly what the field needs right now.
Tom: Fantastic discussion, everyone. We’ve really explored how this Spike-driven Vision-Language-Action Model uses direct training in the spike domain and novel fusion techniques to achieve strong performance with significantly lower energy costs. That was a great session!
The paper's summary: Tom: So, to recap, this paper introduces a Spike-driven Vision-Language-Action Model that uses spiking neural networks for end-to-end training to create efficient robotic controllers, which is a big deal because it tackles the massive energy drain of current large Transformer models.
Jane: That’s right, Tom; essentially, they’ve built a whole system where the visual and language inputs are processed through spikes to generate actions in real time using far fewer resources than traditional methods.
Lu: What I find particularly fascinating is their architecture because it uses this specific Multi-Winner Spike Fusion module to manage the cross-modal communication, which seems much smarter than standard attention mechanisms for sparse representations.
Meng: From a practical standpoint, I'm really interested in the energy estimates they provide; if we can hit those power targets on edge devices, that opens up a whole new category of deployment possibilities for robotic systems.
Lalam: For me, the cultural implication is huge because if AI can be made this efficient and reliable for physical tasks, it means we can deploy helpful assistants in environments where computing power isn't readily available everywhere.
Tom: Exactly! And their results on benchmarks like LIBERO and Meta-World show that this efficiency doesn't come at the cost of intelligence; they’re actually achieving higher success rates than some of the existing models while using a fraction of the parameters.
Jane: It really shows how focusing on sparse spike representations lets these AI systems maintain high performance even when they are running on much more constrained hardware.
Lu: And the way they structure it for end-to-end training means you don't have to worry about complex conversion steps anymore, which simplifies the whole development pipeline immensely.
Meng: That direct training aspect is a major win for our engineering teams because it means we can tailor the architecture exactly to our specific robotics problem instead of trying to shoehorn it into a pre-existing framework.
Lalam: If this model becomes the foundation for more efficient embodied AI, it could significantly improve how we design and deploy assistive technologies across different cultures and physical settings.
Tom: This paper really proves that the spiking domain isn't just an academic curiosity; it’s a viable path for building practical, energy-aware robotic intelligence.
Jane: It’s exciting to see how they managed to balance that high level of functional performance with such a drastically reduced computational footprint.
Lu: The theoretical framework for the Spike Action Chunking Transformer is particularly elegant in how it handles continuous action generation in parallel using sparse routing instead of sequential processing.
Meng: I'm looking forward to seeing how this translates into actual control policies that are robust enough for real-world scenarios, since the paper also touched on distribution shifts.
Lalam: That robustness against environmental changes is what makes this model truly applicable; it suggests an AI that can adapt more gracefully when things aren't perfectly controlled in the lab.
Tom: It’s a powerful combination of smart fusion and efficient action chunking that sets a new baseline for what we expect from multimodal systems.
Jane: So, the big picture here is moving toward robotic AI that is not just capable, but also profoundly resource-conscious and adaptable to messy real-world conditions.
Lu: And I think the future work they suggest focusing on developing even more specialized spike encoders will be where we see the next wave of creative possibilities emerge.
The paper's improvements: Tom: So, we’ve talked about how they built this system from scratch using spiking networks for better efficiency, so now we need to look at what they suggest for future improvements and how that impacts things.
Jane: They are pointing out several areas where the current approach can be even refined, focusing heavily on making the fusion module more robust and improving the action generation process.
Lu: The paper emphasizes refining the Multi-Winner Spike Fusion with more sophisticated routing techniques to really sharpen that cross-modal correspondence, which could lead to much more nuanced understanding of complex instructions.
Meng: From an engineering standpoint, I’m listening for details on how they plan to handle those distribution shifts mentioned earlier; real-world deployment demands reliability far beyond what the benchmarks show currently.
Lalam: It’s exciting because if they can make the fusion module truly robust, it means the AI systems we build will be able to interpret human commands with much greater precision and less ambiguity in complex situations.
Tom: And on the action side, they suggest more sophisticated ways to chunk actions so that the continuous control policy becomes even more temporally coherent when generating a sequence of movements.
Jane: That chunking improvement should mean that a robot doesn't just guess the next move, but actually plans a short sequence of moves with better foresight before executing it.
Lu: Their future work seems focused on integrating these refinements with even more specialized spike encoders, which could allow the model to handle incredibly specific visual or language inputs with peak efficiency.
Meng: I’m curious if they plan to tackle the challenge of making this end-to-end training pipeline adaptable to entirely new robotic hardware architectures without requiring a complete overhaul of the training setup.
Lalam: If they can make it highly specialized, it means we could develop AI assistants that are incredibly adept at a very specific physical task, improving how we use technology in specialized fields.
Tom: It sounds like the next step involves pushing the boundaries of both perception and control to make these systems even more fine-tuned and reliable for complex manipulation tasks.
Jane: They are really showing us that there’s still a lot of room to optimize these modules, especially where visual understanding meets precise physical execution.
Lu: I think their suggestions open up avenues for creating entirely new types of multimodal AI that might operate on even sparser representations than what we see here.
Conclusion: Tom: So we’ve covered how the Spike-driven Vision-Language-Action Model uses spiking networks to create an efficient, end-to-end training framework for robotics, and we’ve discussed its energy savings and performance on benchmarks like LIBERO and Meta-World.
Jane: That was a really deep dive into how they manage the cross-modal interaction through that Multi-Winner Spike Fusion module, which is such a clever way to handle the language and vision inputs together.
Lu: I think what stands out most about this work is its success in making sparse spike representations work effectively for continuous action generation, which opens up some really creative possibilities for future embodied AI designs.
Meng: I’m still thinking about how much less computational overhead this means when we consider deploying these systems on physical hardware; that energy efficiency is something we need to keep pushing for practical robotics applications.
Lalam: For me, the most significant implication is that this work paves the way for a generation of AI assistants that can be deployed widely because they won't require massive data centers or huge computational power to run effectively in diverse environments.
Tom: Exactly! The paper on the Spike-driven Vision-Language-Action Model really shows us how we can build smarter, more resource-friendly robotic controllers right now.
Jane: It’s a fantastic piece of research that demonstrates a solid path toward building AI that is both capable and incredibly energy conscious.
Lu: I hope the next steps focus on pushing those architectural boundaries to see what other kinds of complex reasoning we can unlock with this spiking approach.
Meng: I look forward to seeing how this translates into real-world deployment constraints, because theoretical efficiency has to meet the demands of physical systems in practice.
Lalam: I really hope this research inspires a cultural shift where AI capabilities are accessible and helpful across a much wider range of human experiences.
Shuai Wang, *Malu Zhang*, Mingquan Liu, Weihui Dai, Dehao Zhang, Jieyuan Zhang, Yimeng Shan, Zijian Zhou, Yang Yang
University of Electronic Science and Technology of China
cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: Spike-driven VLA introduces a novel framework that enables end-to-end direct training for robotic manipulation using spiking neural networks, offering an energy-efficient alternative to conventional
Key concepts
- Spiking Visual Encoder (SVE)
- This component processes visual input, such as images of the robot's environment, by converting them into sparse spike tokens. It uses a three-stage architecture involving convolutional blocks for local features and self-attention blocks for global dependencies. This allows the model to capture hierarchical visual details efficiently.
- Multi-Winner Spike Fusion (MWSF)
- The MWSF module connects visual and language inputs by using a 'winner-take-all' routing mechanism over an affinity matrix. This process suppresses irrelevant background noise, allowing the system to establish sparse cross-modal correspondences between what it sees and what it is told, creating a fused memory.
- Spike Action Chunking Transformer (SpikeACT)
- This module handles decision-making and continuous control by taking the fused memory and current robot state. It uses spiking cross-attention to predict action chunks in parallel. This allows the model to generate sequences of actions efficiently, moving from high-level planning to low-level control.
- Energy Efficiency
- The framework is designed to be energy-efficient by using spiking neural networks instead of conventional Artificial Neural Networks (ANNs). The paper estimates computation energy based on synaptic operations and continuous MAC operations, showing a significantly lower estimated inference energy compared to existing VLA models.
Terminology
Summary
Spike-driven VLA introduces a novel framework that enables end-to-end direct training for robotic manipulation using spiking neural networks, offering an energy-efficient alternative to conventional large Transformer models. The core contribution is a system that integrates spiking visual and instruction encoders, a Multi-Winner Spike Fusion module for cross-modal interaction, and a Spike Action Chunking Transformer for continuous action generation. This approach demonstrates competitive performance on benchmarks like LIBERO and Meta-World while requiring fewer parameters and lower estimated inference energy than existing ANN-based VLA models.
How it works
The framework is structured into three core stages: multi-modal spike encoding, cross-modal fusion, and state-conditioned action decoding. First, the perception stage involves developing a spiking visual encoder (SVE) and a spiking instruction encoder (SIE) to capture multimodal inputs. The SVE encodes main and wrist observations into visual spike tokens, while the SIE encodes language instructions into contextual semantic spike tokens. These components are designed to capture hierarchical visual features
and preserve them in sparse spike representations.
Second, the Multi-Winner Spike Fusion (MWSF) module is introduced to establish task-relevant cross-modal correspondence. This module uses a bidirectional top-k winner-take-all spike routing to suppress background interference and yield fused memory.
It calculates visual and language projections, generates a shared crossmodal affinity matrix, and then computes instruction-guided visual representation (Cv) and visually grounded instruction representation (Cl), which are concatenated to form the task-conditioned memory Mn.
Finally, the Spike Action Chunking Transformer (SpikeACT) is proposed for multimodal decision-making and continuous control. This module takes the fused memory Mn and the current robot state Sn, encodes them into a state-augmented memory Mfn, and uses spiking cross-attention to predict action chunks in parallel. The process involves generating query features Qa from learnable action tokens Pa and key/value features KMfn, VMfn from Mfn, computing an affinity matrix Ra,n = QaK⊤n to convert affinities into a sparse binary routing matrix via WTAk(Ra,n), and finally reading out the routed features to produce the continuous action chunk Abn.
Key Components
The paper details several specific architectural components:
-
Spiking Visual Encoder (SVE): This encoder uses a three-stage pyramidal architecture consisting of convolutional mixer blocks for local features and seven spiking self-attention blocks for global spatial dependencies, designed to produce visual spike tokens Zv,n. It is pre-trained on ImageNet1K before fine-tuning.
-
Spiking Instruction Encoder (SIE): This encoder utilizes a spiking BERT backbone (SmoothSpike) pretrained via masked language modeling on a large corpus like TinyStories and OpenWebText. The instruction embeddings are projected to the shared VLA embedding dimension d, yielding Zl.
-
Multi-Winner Spike Fusion (MWSF): This module employs
bidirectional top-k winner-take-all spike routing
over the crossmodal affinity matrix Rn to suppress background responses and establish sparse correspondences, resulting in fused memory Mn. -
Spike Action Chunking Transformer (SpikeACT): This module combines the fused memory with the current robot state Sn to form Mfn, and then uses spiking cross-attention over this augmented memory to jointly predict action chunks Abn.
Performance and Efficiency
Extensive experiments on LIBERO and Meta-World demonstrate that Spike-driven VLA achieves competitive performance with fewer parameters and lower estimated inference energy than conventional VLA models.
On the Meta-World benchmark, it achieved an average success rate of 72.4% across all tasks, surpassing existing models like SmolVLA (0.45B) by 15.1% while using only 1/3 of its parameters and achieving an estimated inference energy of 11.6 mJ. On the LIBERO benchmark, the model achieved an average success rate of 92.4%. Furthermore, evaluation on LIBERO-Plus showed strong robustness under diverse distribution shifts across seven perturbation dimensions. The theoretical energy estimation indicates that synaptic operations are counted as ACs (0.9 pJ) and continuous operations as MACs (4.6 pJ), leading to an estimated computation energy of EbSpikeVLA = eACNAC + eMACNMAC.
Ablation Studies
Ablation studies confirm the importance of specific modules; for instance, replacing BERT with SIE reduced the average success rate by only 0.2%, while replacing DINOv3-B with SVE led to a 3.7% decrease, suggesting visual perception accounts for most of the performance gap. Additionally, ablation on WTA showed that its inclusion increases the overall success rate from 90.2% to 92.4%
across suites like Spatial and Long, highlighting its role in strengthening instruction-guided visual grounding.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements that can be made to existing AI systems by adopting the Spike-driven VLA (VLA) framework:
- Improved Energy Efficiency for Edge Deployment:
A key improvement is the shift from dense Multiply-Accumulate (MAC) operations to sparse Accumulate (AC) operations via Spiking Neural Networks (SNNs).
-
The improved system can perform complex Vision-Language-Action tasks on resource-constrained platforms (e.g., edge robots, mobile devices) while consuming significantly less energy.
-
Specifically, the paper estimates a theoretical energy consumption of approximately 11.6 mJ for inference on the proposed model with only 0.15B parameters, which is a fraction of conventional VLA models like π0 (which consumes 47.9 mJ).
- Enhanced Latency and Real-Time Control:
The event-driven nature of SNNs inherently reduces computational latency compared to recurrent ANN models requiring dense MAC operations at every timestep.
- The system can achieve faster decision-making loops, which is critical for high-speed robotic control where rapid feedback is necessary.
- Task-Specific Architectural Design (End-to-End Training):
The framework supports end-to-end direct training in the spike domain, bypassing the need for computationally expensive and architecture-agnostic ANN-to-SNN conversion methods.
- Improved AI systems can be trained directly on spiking architectures tailored precisely to the task, leading to more task-relevant representations that are inherently efficient.
- Superior Cross-Modal Fusion via Sparse Routing:
The Multi-Winner Spike Fusion (MWSF) module utilizes bidirectional top-k winner-take-all (WTAk) routing instead of dense Softmax normalization for cross-modal attention.
- The system can effectively suppress background interference and dilute weak, irrelevant visual evidence when language instructions are present, leading to more precise instruction-guided scene understanding.
- Robustness to Distribution Shifts in Real Environments:
The framework demonstrates strong robustness across diverse perturbation dimensions (lighting, object layout, sensor noise) on challenging benchmarks like LIBERO-Plus.
- The improved AI system can maintain high success rates (e.g., 53.2% overall success rate under noise) when deployed in real-world scenarios where environmental conditions differ significantly from the training distribution, making it more reliable for general-purpose robotic control.
- Efficient Action Generation via Chunking:
The Spike Action Chunking Transformer (SpikeACT) predicts continuous action chunks in parallel rather than sequentially, leveraging the fused memory and robot state.
- The system can generate sequences of future actions simultaneously, leading to more temporally coherent and efficient control policies that are ready for execution without requiring iterative inference steps for every single action.
- Reduced Parameter Count with Maintained Performance:
The proposed framework achieves competitive performance against large ANN-based models using significantly fewer parameters (e.g., 0.15B parameters vs. 2.25B for a comparable model on Meta-World), while achieving superior computational efficiency (1/15th the MAC operations).
- This allows for deployment of high-performance VLA systems on devices with very limited memory and computational power, which is currently impossible for state-of-the-art VLA models.
In summary, the improved AI system will be a compact, energy-efficient, end-to-end trained robotic controller capable of perceiving visual scenes and natural language instructions in real time on edge devices while maintaining high reliability in unpredictable environments.
Abstract
Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. However, most existing models rely on large Transformers, whose latency and energy costs hinder deployment on resource-constrained platforms. Through sparse event-driven computation, spiking neural networks offer a promising paradigm for high-performance and energy-efficient computing. Here, we propose the first Spike-driven VLA framework enabling end-to-end direct training for robotic manipulation, which mainly comprises three core components. First, we develop spiking visual and instruction encoders for multimodal perception, encoding visual observations and language instructions into sparse, reliable spike representations for subsequent cross-modal fusion. Then, we introduce Multi-Winner Spike Fusion for instruction-guided scene understanding, using bidirectional top- k winner-take-all spike routing to suppress background interference and yield fused memory. Finally, we propose a Spike Action Chunking Transformer that incorporates spiking cross-attention over the fused memory and the current robot state, enabling efficient end-to-end generation of continuous action chunks for robotic control. Extensive experiments on LIBERO and Meta-World demonstrate that Spike-driven VLA achieves competitive performance with fewer parameters and lower estimated inference energy than conventional VLA models. This work establishes a foundational framework for neuromorphic VLA modeling, paving the way for future advances in resource-efficient embodied intelligence.
Sources
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- EdgeVLA: Efficient Vision-Language-Action Models
- NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
- CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- SDTrack: A Baseline for Event-based Tracking via Spiking Neural Networks
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- DINOv3
- SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
- TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
- Spike-driven Transformer V2: Meta Spiking Neural Network Architecture Inspiring the Design of Next-generation Neuromorphic Chips
- SpikeZIP-TF: Conversion is All You Need for Transformer-based SNN
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
- QKFormer: Hierarchical Spiking Transformer using Q-K Attention
- SpikeGPT: Generative Pre-trained Language Model with Spiking Neural Networks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering