SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance".
Jane: The paper was written by Minghan Yang, Lan Yang, Ke Li, Honggang Zhang, Kaiyue Pang et al. from Beijing University of Posts and Telecommunications and University of Surrey.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that honestly sounds like science fiction come to life: “SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance.” Jane, I’ve got to say, the title alone got me hooked.
Jane: Oh, absolutely, Tom. And for our listeners who just tuned in, this is about using fMRI scans—basically measuring blood flow in the brain—to figure out what video someone was watching. Then, they actually rebuild that video from the brain data alone. It’s like a mind-reading movie player.
Tom: Right, and it’s from a team at Beijing University of Posts and Telecommunications, along with folks at the University of Surrey. The lead authors are Minghan Yang and Lan Yang. They’re tackling something that sounds impossible: your brain doesn’t record every single frame like a camcorder.
Jane: Exactly. So the title says “hierarchical semantic guidance.” That’s their secret sauce. Instead of trying to decode every pixel, they decode the *meaning* at three levels: what the first frame looks like, how things move, and the overall story. Then they use that to guide the video generation.
Tom: And that’s the part that gets me excited. We’re not just talking about a blurry blob. They’re reconstructing a kitten that’s actually a kitten, with motion that matches what the person watched. The implications for neuroscience and for brain-computer interfaces are huge.
Jane: For sure. And I love that they’re honest about the challenge. fMRI is slow—it measures brain activity over seconds, not milliseconds. So a video at thirty frames per second? The brain signal just can’t capture that speed. That’s why they focus on key semantic perceptions, not every frame.
Tom: So it’s not about reading minds in real time yet, but it’s a massive step toward understanding how the brain compresses visual stories. I can’t wait to see how they actually pulled this off. Let’s get into the details.
Jane: Me too. Because the title promises a lot, and from what I’ve seen, the results actually deliver.
Summary: Tom: So, Jane, we’ve got the title, we know the dream. But what does the paper actually do? Give us the elevator pitch.
Jane: Okay, so they built a system called SemVideo. The core idea is that they don’t try to decode the video directly from brain signals. Instead, they first use a big language model to turn the original video into three types of text descriptions. They call this SemMiner.
Tom: Right, and those descriptions are like a cheat sheet for the reconstruction. One is the anchor—what’s in the first frame. Another is the motion narrative—how things are moving. And the third is the holistic summary—the whole story. So they’re basically translating the video into words first.
Jane: Exactly. And then they train a decoder to turn fMRI signals into those same text embeddings. That’s the Semantic Alignment Decoder. Once they have the semantic meaning, they use another module to get the motion right, and then they feed all of that into a text-to-video model to generate the actual footage.
Tom: And the results? They tested it on two big datasets, CC2017 and HCP. On the semantic level, they beat every previous method. For example, their two-way video retrieval accuracy is zero point eight six five, which is the highest in the field. That means when you show the system two videos, it can pick the right one based on brain activity eighty-six point five percent of the time.
Jane: That’s wild. And they also improved temporal coherence. That’s the metric that checks if the motion in the reconstructed video actually matches the original. Their endpoint error, which measures motion mismatch, is the lowest among all methods. So the kitten isn’t just a static image—it’s actually crouching and turning like it did in the real video.
Tom: And here’s the kicker, Jane. They didn’t just throw a bunch of modules together. They ran ablation studies. When they removed the motion descriptions, the pixel-level quality dropped a lot. When they removed the anchor, the semantic accuracy tanked. So each piece of that hierarchical guidance is doing real work.
Jane: It’s like they built a translator that goes from brain to meaning, and then from meaning to video. And the meaning layer is what makes it work. I think that’s the big takeaway for the field—semantic guidance is the bridge.
Tom: And they even did a neuroscience check. They looked at which brain regions their model relies on. The motion part lights up areas like MT and MST, which are known for processing motion. So the model is actually using the brain’s natural structure, not just brute-forcing it. That’s a beautiful validation.
Jane: It really is. So the summary is: they decode meaning, they decode motion, and they generate video. And it works better than anything before it. Now, how did they actually improve on the previous attempts? Let’s talk about the specific innovations.
Improvements: Tom: Alright, Jane, so we know SemVideo works. But what did they actually improve over the previous state of the art? Because this isn’t the first fMRI-to-video paper out there.
Jane: Right. Previous methods like Mind-Video or NeuroClips had two big problems. One, the objects in the reconstructed video would change appearance between frames—like a face that morphs into a different person. Two, the motion was choppy or misaligned. SemVideo fixes both by adding that hierarchical semantic guidance.
Tom: And the key improvement is in how they handle motion. They built something called the Motion Adaptation Decoder. It’s got a tripartite attention fusion—that’s a fancy way of saying it looks at spatial details, temporal flow, and semantic meaning all at once. That’s new.
Jane: Exactly. And they also changed how they generate the final video. Instead of just feeding text into a video model, they first generate a clean anchor frame using the static description. Then they use that frame, plus the motion latents, to guide the rest of the video. It’s like giving the generator a solid starting point so it doesn’t drift off.
Tom: That’s a smart engineering choice. And they also introduced something called CC2017-SE, which is an extended version of the CC2017 dataset with all those semantic descriptions added. That’s a gift to the research community—they’re releasing it publicly.
Jane: And it shows in the numbers. Their fifty-way video retrieval accuracy is zero point two six four, which is way above the previous best of zero point two four six. That’s a huge jump. And on the HCP dataset, which is a completely different set of subjects, they still lead. So it’s not overfitting to one group.
Tom: But here’s what I find most impressive, Jane. They didn’t just improve the metrics. They improved the *interpretability*. They showed that the motion decoder relies on brain regions like MT and MST, which are known to process motion. That means the model is learning something biologically real, not just statistical patterns.
Jane: And they even ran a shuffle test. They took the reconstructed video, shuffled the frames randomly, and measured the motion error. The shuffled version was much worse. That proves the temporal order they reconstructed actually matches the real video—it’s not just a random sequence of pretty frames.
Tom: So the improvements are: better semantic alignment, better motion coherence, and better biological plausibility. That’s a triple win. I’d say this sets a new bar for the field.
Jane: Definitely. And it opens the door for more research into how we can use language as a bridge between brain signals and generative models. That’s a powerful idea.
Conclusion: Tom: Well, Jane, we’ve covered a lot of ground on “SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance.” Let’s wrap it up for our listeners.
Jane: Absolutely. So the big picture is this: they took fMRI data, decoded it into three levels of semantic meaning—static, motion, and holistic—and then used that to generate a video that matches what the person actually watched. It’s a complete pipeline from brain to screen.
Tom: And the results speak for themselves. They beat every previous method on semantic accuracy and temporal coherence. They even showed that their model uses brain regions that are known to process motion, which is a beautiful validation of the approach.
Jane: The implications are huge. This isn’t just about watching movies in your head. This could lead to better brain-computer interfaces for people who can’t speak or move. Imagine being able to communicate what you’re seeing or imagining just by thinking about it.
Tom: And for neuroscience, it gives us a new tool to understand how the brain compresses and represents dynamic visual stories. That’s fundamental science with real-world payoff.
Jane: So we’re saying goodbye to SemVideo, but we’re definitely keeping an eye on this research group. They’ve set a new standard, and I can’t wait to see what they do next.
Tom: Agreed. Thanks for joining us, everyone. Next time, we’ll be looking at another paper that pushes the boundaries of what’s possible. Until then, keep your minds open—literally.
Jane: See you next time!
Minghan Yang, Lan Yang, Ke Li, Honggang Zhang, Kaiyue Pang, Yizhe Song
Beijing University of Posts and Telecommunications · University of Surrey
cs.CV, cs.AI
Submitted: 2026-08-17
Updated: 2026-08-18
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 67/100
The gist: The paper introduces SemVideo, a novel fMRI-to-video reconstruction framework guided by hierarchical semantic information.
Key concepts
- SemVideo
- A system that reconstructs videos from brain activity using hierarchical semantic guidance. It works by first translating the original video into three types of text descriptions: what is in the first frame, how things move, and the overall story.
- Hierarchical Semantic Guidance
- The core method where the system decodes video meaning at three levels: what appears in the first frame (anchor), how objects move (motion narrative), and the complete story (holistic summary). This guidance is used to guide video generation.
- fMRI
- Functional Magnetic Resonance Imaging, which measures blood flow in the brain. The hosts note that fMRI is slow, measuring activity over seconds rather than milliseconds, which requires the system to focus on key semantic perceptions instead of every frame.
- Motion Adaptation Decoder
- A new module in SemVideo that uses tripartite attention fusion to handle motion. It looks at spatial details, temporal flow, and semantic meaning simultaneously to ensure the reconstructed motion matches the original video.
Terminology
Summary
The paper introduces SemVideo, a novel fMRI-to-video reconstruction framework guided by hierarchical semantic information. The authors identify two major shortcomings in existing fMRI-to-video reconstruction approaches: (i) inconsistent visual representations of salient objects across frames, leading to appearance mismatches; (ii) poor temporal coherence, resulting in motion misalignment or abrupt frame transitions.
The core of SemVideo is SemMiner, a hierarchical guidance module that constructs three levels of semantic cues from the original video stimulus: static anchor descriptions (Canchor), motion-oriented narratives (Cmotion), and holistic summaries (Choli). SemMiner is implemented as a two-stage semantic decomposition module built upon a multimodal large language model (MLLM). In the first stage, it generates a simple semantic summary Cbasic of the video, constrained to a maximum of 20 words, which serves as a rein
to prevent subsequent generation from drifting semantically. In the second stage, conditioned on the video and Cbasic, the framework elicits target-specific descriptions using carefully crafted instructions PL for each semantic objective.
Leveraging this semantic guidance, SemVideo comprises three key components:
-
Semantic Alignment Decoder (SAD): A cross-subject multi-level semantic decoder consisting of a subject-specific projector, a subject-shared mapper, and a Refineformer module. It aligns fMRI signals with CLIP-style embeddings derived from SemMiner, enabling precise semantic feature decoding while minimizing noise. The training of SAD is supervised by a combination of three objectives: LMSE, LSoftCLIP, and Lrefine.
-
Motion Adaptation Decoder (MAD): Reconstructs dynamic motion patterns using a novel tripartite attention fusion architecture that synergistically integrates spatial self-attention, temporal self-attention, and semantic-guided cross-attention mechanisms. By explicitly injecting semantic priors into the attention computation, this module effectively aligns motion latents with both spatial structures and semantic actions. The training objectives of MAD include a reconstruction loss and a bidirectional contrastive loss.
-
Conditional Video Render (CVR): A multi-level, staged guidance strategy for conditional video generation. Instead of relying on a single semantic cue, it progressively conditions the generation on three sources: static semantics (anchor descriptions), dynamic semantics (motion-oriented narratives), and holistic video-level semantics (holistic summaries). The decoded motion latents are passed through a pretrained VAE decoder to produce motion frames, the decoded anchor feature is combined with the first motion frame and fed into a text-to-image model to yield the initial reconstructed frame, and finally a pretrained text-to-video generator is steered jointly by the holistic semantic guidance, the anchor frame, and the motion-frame sequence.
The authors also introduce CC2017-SE, a semantic extension of the CC2017 dataset, applying SemMiner to generate three types of semantic descriptions for 4,320 training videos and 1,200 test videos.
Experiments are conducted on two publicly available datasets: CC2017 and a subset of the HCP 7T dataset. The evaluation covers three dimensions: semantic-level metrics (N-way top-1 retrieval in 2-way/50-way settings at frame/video level, VIFI-score), pixel-level metrics (SSIM, PSNR, Hue-pcc), and spatiotemporal-level metrics (CLIP-PCC, EPE). SemVideo achieves state-of-the-art performance on 8 out of 10 metrics on the CC2017 dataset, with a 2-way-V score of 0.865, a 50-way-V score of 0.264, a VIFI-score of 0.608, hue-pcc of 0.849, CLIP similarity of 0.526, and the lowest EPE of 4.788. On the HCP dataset, SemVideo also demonstrates strong generalization ability.
Ablation studies confirm the indispensability of each semantic component: removal of Cmotion causes the most precipitous decline in pixel-level and ST-level metrics; absence of Canchor significantly impairs semantic-level fidelity; and Choli proves integral for overall performance. Notably, removing Cmotion also results in a notable drop in semantic metrics, demonstrating that the motion latent space is endowed with semantic understanding.
Shuffle tests validate that the predicted inter-frame order robustly matches the ground truth, with a significantly low P-value for the full model that rises substantially when Cmotion is removed and spikes to 0.84 when the entire MAD is ablated, confirming that motion information is indeed decoded from MAD and fMRI-derived Cmotion.
Neuroscience interpretability analysis using ROI-wise visualization reveals that the anchor component predominantly draws from higher-level visual cortices, the motion component aligns with established motion-processing regions including MT, MST, and TPOJ (with strong activations in V1 supporting parallel non-hierarchical pathways), and the holistic component exhibits a distributed pattern of importance across both visual and motion-related areas. Similar activation distributions across subjects demonstrate stable feature capture.
Improvements for AI systems
Based on the paper SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance,
here are specific improvements I can implement in AI systems, along with the resulting capabilities.
Improvement: I will implement the two-stage, hierarchical semantic decomposition strategy (SemMiner) as a general-purpose video understanding module. Instead of generating a single, generic caption, the system will be prompted to produce three distinct, complementary descriptions: a static anchor (first-frame detail), a motion-oriented narrative (actions, dynamics, and transitions), and a holistic summary (integrating both). The key is the rein
mechanism: first generating a short, constrained summary (max 20 words) to anchor the model, then using that summary as a condition to generate the three detailed descriptions. This prevents semantic drift and hallucination.
What the improved AI system can do:
-
Generate richer, more structured video annotations: For any input video, it will output a multi-perspective description set, not just a single caption. This is far more useful for downstream tasks like text-to-video generation, video retrieval, and video question answering.
-
Reduce hallucination in video captioning: By constraining the initial summary and using it as a guide, the system is less likely to generate false or irrelevant details about the video content.
-
Provide fine-grained control for generative models: The anchor description can be used to condition a text-to-image model for a consistent first frame, the motion narrative can guide a motion-specific module, and the holistic summary can steer the overall video generation. This directly addresses the
appearance mismatch
andmotion misalignment
issues common in text-to-video generation.
Improvement: I will build a new decoding pipeline that explicitly separates semantic decoding from motion decoding. The core innovation is the Motion Adaptation Decoder (MAD) with its tripartite attention fusion. This module doesn't just map brain signals to a single latent space; it uses three parallel attention mechanisms: (1) spatial self-attention to capture intra-frame structure, (2) temporal self-attention to model inter-frame dependencies, and (3) semantic-guided cross-attention that injects the decoded motion narrative (from SemMiner) directly into the latent representation. This ensures the generated motion is not just smooth but also semantically correct (e.g., crouching
vs. turning
).
Improvement: I will integrate the hierarchical semantic decoding approach into BCI systems. The system will be trained to predict not just a single thought
but a structured semantic representation (anchor, motion, holistic). Furthermore, I will use the ROI-wise importance visualization technique (as described in Section 4.4) as a diagnostic tool to understand which brain regions the model is relying on for each type of semantic information.
Improvement: I will adopt the paper's three-level evaluation protocol (Semantic, Pixel, Spatiotemporal) as a standard for evaluating any generative video model. This includes the N-way top-1 accuracy for semantic fidelity, VIFI-score for overall alignment, SSIM/PSNR/Hue-PCC for pixel-level accuracy, and CLIP-PCC/EPE for temporal coherence and motion accuracy.
Sources
- Qwen2.5-VL Technical Report
- CaRiNG: Learning Temporal Causal Representation under Non-Invertible Generation Process
- A Survey of fMRI to Image Reconstruction
- Imagen Video: High Definition Video Generation with Diffusion Models
- A Penny for Your (visual) Thoughts: Self-Supervised Reconstruction of Natural Movies from Brain Activity
- Decoupled Weight Decay Regularization
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Make-A-Video: Text-to-Video Generation without Text-Video Data
- Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates
- NeuroCine: Decoding Vivid Video Sequences from Human Brain Activties
- GIT: A Generative Image-to-text Transformer for Vision and Language
- UniBrain: A Unified Model for Cross-Subject Brain Decoding
- xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
- SynMind: Reducing Semantic Hallucination in fMRI-Based Image Reconstruction
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models