Chain-of-Evidence Multimodal Reasoning for Few-shot Temporal Action Localization

arXiv:2504.13460 · cs.CV, cs.AI · Submitted 2026-08-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Chain-of-Evidence Multimodal Reasoning for Few-shot Temporal Action Localization".

Jane: The paper was written by Mengshi Qi, Hongwei Ji, Wulian Yun, Xianlin Zhang and Huadong Ma from Beijing University of Posts and Telecommunications.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're diving into a fresh arXiv paper that's got me genuinely excited. It's called "Chain-of-Evidence Multimodal Reasoning for Few-shot Temporal Action Localization." Jane, that title is a mouthful, but it's packed with meaning.

Jane: It really is, Tom. Let's break it down for our listeners. "Temporal action localization" is basically teaching a computer to find the exact start and end times of an action in a video. Think of it like finding the moment a basketball player jumps to make a shot, not just knowing the whole video is about basketball.

Tom: And the "few-shot" part is the real magic here. Usually, these systems need thousands of examples to learn. But this paper is about getting the job done with just a handful of examples, like showing the computer one or two videos of someone spiking a volleyball and then asking it to find that action in a brand new, completely different video.

Jane: Exactly. And that's where the "multimodal reasoning" comes in. The paper argues that if you only look at the pixels, you can get confused. Two videos might look almost identical, but the action is different. So, they bring in text to help.

Tom: Right, so instead of just watching the video, they also read a description of what's happening. That text gives the computer semantic clues that the visuals alone just can't provide. It's like having a friend whisper the plot to you while you're watching a silent film.

Jane: That's a perfect analogy. And the "Chain-of-Evidence" part is about making that text smarter. Instead of just saying "a man jumps," the text explains the sequence: "he runs, he jumps, he hits the ball, he lands." It's a logical chain of events.

Tom: And that chain is what helps the model understand the cause and effect, the flow of the action. It's not just a static snapshot; it's a story. I love that they're using large language models to generate this kind of structured narrative for the videos.

Jane: It's a clever way to leverage the reasoning power of LLMs for a task that's traditionally been purely visual. The authors are from Beijing University of Posts and Telecommunications, and they've clearly put a lot of thought into how to bridge that gap.

Tom: They really have. And the implications are huge. Think about video search, surveillance, sports analytics. If we can find actions with just a few examples, we can build tools that adapt to new situations on the fly.

Jane: We're not just talking about finding a specific action either. This could help us understand entire scenes, like detecting a car accident or a moment of violence in a crowd, without needing to pre-train on every possible scenario.

Tom: That's the exciting part. This isn't just an incremental step; it's a new way of thinking about how to teach machines to see and understand time-based events. We'll get into the nitty-gritty of how they actually do it next.

Jane: Stay with us, because we're about to open the hood on this fascinating method.

Summary: Tom: So, Jane, we've established that "Chain-of-Evidence Multimodal Reasoning for Few-shot Temporal Action Localization" is about finding actions in videos with minimal examples. But how do they actually pull it off?

Jane: Well, Tom, the core idea is a two-pronged attack. First, they use a pre-trained vision-language model, like VideoChat, to look at the support videos—the few examples you give it—and generate two kinds of text. One is a simple caption, and the other is their special "Chain-of-Evidence" text.

Tom: And that's where the LLM, like DeepSeek-R1, comes in. It takes the captions and the video descriptions and builds a logical narrative. It doesn't just say "a woman falls." It says "a woman is walking, she trips on an object, which causes her to fall, and then the dog runs to her." It's a causal chain.

Jane: Exactly. That chain is the key. It captures the temporal dependencies and the causal relationships between actions. That's something a simple caption just can't do. A caption is a snapshot; the CoE text is a story.

Tom: So they have these great text descriptions. But how does that help with localization? The model still needs to find the action in the query video, the one it's never seen before.

Jane: That's where their "Semantic-Aware Text-Visual Alignment" module comes in. They take the visual features from the query video and the support videos, and they also take the text features. Then they align them.

Tom: So it's not just aligning video to video, which is what most other methods do. They're aligning the query video to the support video, but they're also aligning it to the support text. That text acts like a guide.

Jane: Precisely. The text helps disambiguate. If the query video has a background that looks a lot like the foreground action, the text can help the model focus on the actual action. It's like having a highlighter on the important parts.

Tom: And they don't just use the text once. They have this whole pipeline where they align the support video with the support text first, creating a richer, text-infused representation. Then they align that with the query video. It's a multi-step process.

Jane: It's a very thorough approach. They also have this "Semantic-Temporal Pyramid Encoder" that looks at the video at different scales, both in time and in semantic meaning, to capture both the fine details and the big picture.

Tom: Right, so they're not just looking at the whole video as one blob. They're looking at short snippets, longer sequences, and everything in between. This helps them understand the action's structure.

Jane: And the results speak for themselves. On the ActivityNet1 point 3 dataset, they got a significant boost in performance, like a four percent improvement in the multi-instance five-shot scenario. On THUMOS14, they saw even bigger gains, around twelve percent.

Tom: Those are huge numbers. It really shows that this multimodal approach is working. They're not just adding text for the sake of it; they're using it to fundamentally improve the model's understanding.

Jane: And they even created their own dataset, the Human-related Anomaly Localization dataset, to test it on more practical scenarios like detecting fights or falls. That's a big deal for real-world applications.

Tom: It is. They're pushing the field forward not just with a new method, but with new data to challenge it. Next, we'll dig into the specific improvements they made over existing methods.

Jane: And why their "Chain-of-Evidence" is such a game-changer.

Improvements: Tom: We're back with "Chain-of-Evidence Multimodal Reasoning for Few-shot Temporal Action Localization." Jane, we've talked about the overall idea, but what are the specific improvements this paper brings to the table?

Jane: The biggest one, Tom, is the shift from visual-only to multimodal. Most few-shot TAL methods just look at the video pixels. They try to align the query and support videos based on how similar the frames look. But that's fragile.

Tom: Because two different actions can look very similar visually. Like a volleyball spike and a tennis serve. They're both a person jumping and hitting something. The pixels are almost the same.

Jane: Exactly. So the paper's first improvement is to bring in text as an additional signal. That text provides semantic clarity that pixels just can't. It's the difference between seeing a blur of motion and knowing it's a "jump serve."

Tom: And their "Chain-of-Evidence" text is a specific improvement over just using any text. It's not just a caption like "a person plays volleyball." It's a logical sequence of events.

Jane: Right. And that sequence is what helps the model understand the *temporal dependencies*. The model can learn that a "run" usually precedes a "jump," and a "jump" often leads to a "hit." This structured knowledge is far more informative than a static description.

Tom: So they're not just adding text; they're adding *reasoned* text. They're using the LLM's ability to reason about cause and effect to create a better training signal. That's a clever way to leverage the power of these models.

Jane: It is. And they also improved the alignment process itself. Instead of just aligning query video to support video, they align query video to a support representation that has already been fused with the text.

Tom: So the support set isn't just a video anymore; it's a video-plus-text package. And the query is matched against that richer package. That's a much more robust way to find commonalities.

Jane: And they also have that "Semantic-Temporal Pyramid Encoder." That's another improvement. It allows the model to look at the video at different temporal scales, from a single snippet to a longer sequence, and also to find semantically similar snippets within a layer.

Tom: That's like looking at a building from different distances. You see the bricks, then the windows, then the whole structure. Each scale gives you different information. And by doing this, they capture both the fine-grained motion and the overall action structure.

Jane: And all of these improvements work together. The pyramid encoder gives better features, the CoE text gives better semantics, and the alignment module puts it all together. It's a well-engineered system.

Tom: The ablation studies in the paper really show this. When they remove the text, performance drops. When they remove the pyramid encoder, it drops again. Each piece is contributing.

Jane: It's a holistic approach. They're not just tweaking one thing; they're rethinking the entire pipeline from feature extraction to alignment.

Tom: And that holistic thinking is what leads to those state-of-the-art results. They're not just beating the old methods by a little bit; they're setting a new bar for what's possible in few-shot TAL.

Jane: Absolutely. And it makes you wonder, what's next? If this works so well for actions, could it work for other video understanding tasks?

Tom: That's a great question, and we'll touch on that in our conclusion.

Conclusion: Tom: And that brings us to the end of our discussion on "Chain-of-Evidence Multimodal Reasoning for Few-shot Temporal Action Localization." Jane, what a ride.

Jane: It really was, Tom. We started with a complex title and we've unpacked a method that could genuinely change how we approach video understanding. The core idea is so elegant: use text to guide the visual model.

Tom: And not just any text, but a structured, causal narrative. That's the "Chain-of-Evidence" part. It gives the model a roadmap for understanding the action's progression, not just a static label.

Jane: The results are hard to argue with. They showed significant improvements on standard benchmarks like ActivityNet1 point 3 and THUMOS14, and they even created a new dataset to test on more practical, human-related anomalies.

Tom: That's what I find so exciting. This isn't just an academic exercise. The ability to localize actions with a few examples has real-world implications for security, sports analytics, and even content moderation.

Jane: And the methodology is sound. They've carefully designed each component, from the feature encoder to the alignment module, and they've shown through ablation studies that every piece matters.

Tom: For me, the biggest takeaway is the power of combining different types of AI. They're using vision models, language models, and reasoning models together to solve a problem that none of them could solve alone. It's a glimpse into the future of AI.

Jane: It really is. It's about building systems that understand the world the way we do, by seeing, reading, and reasoning about what's happening. It's a fantastic contribution.

Tom: We'll be sad to see this paper go, but we're excited to see what comes next from this team and the field as a whole. Thanks for joining us, everyone.

Jane: And as always, keep your eyes on the arXiv. The future is being written there every single day. See you next time!

Mengshi Qi, Hongwei Ji, Wulian Yun, Xianlin Zhang, Huadong Ma

Beijing University of Posts and Telecommunications

cs.CV, cs.AI

Submitted: 2026-08-14

Updated: 2026-08-18

Code: https://github.com/MICLAB-BUPT/VAL-VLM

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 63/100

The gist: This paper proposes a new few-shot temporal action localization (TAL) method that leverages multimodal information, specifically textual semantic information, to improve localization performance.

Key concepts

Temporal action localization
This is the task of teaching a computer to find the exact start and end times of an action within a video. It focuses on pinpointing when a specific event occurs, such as finding the moment a basketball player jumps to shoot, rather than just knowing what is in the video.
Few-shot
This refers to learning how to localize actions using only a small number of examples, typically one or two videos. This contrasts with traditional systems that require thousands of examples for learning, making the method more efficient for new situations.
Chain-of-Evidence (CoE) Text
Instead of simple captions, this text provides a logical sequence of events, explaining cause and effect. For example, it describes actions in order: 'he runs, he jumps, he hits the ball,' which helps the model understand the flow and temporal dependencies of the action.
Multimodal Reasoning
This involves using both visual data (pixels from videos) and textual data (descriptions) to solve a problem. The paper argues that combining visuals with text provides semantic clues that help overcome confusion when looking at videos alone.

Terminology

Summary

This paper proposes a new few-shot temporal action localization (TAL) method that leverages multimodal information, specifically textual semantic information, to improve localization performance. The authors argue that existing few-shot TAL methods rely solely on video-level information and neglect textual information, which can provide valuable semantic support. To address this, they introduce a framework with two key components: a Chain-of-Evidence (CoE) reasoning method and a semantic-aware text-visual alignment module.

The CoE reasoning method is designed to progressively guide the Vision Language Model (VLM) and Large Language Model (LLM) to generate CoE text descriptions for videos. This text is intended to capture more variance of action than visual features and to express the temporal dependencies and causal relationships between actions at the textual level. The semantic-aware text-visual alignment module is designed to align the query and support videos at different levels by leveraging the generated text.

The authors also introduce a new benchmark, the Human-related Anomaly Localization Dataset (HAL), which contains 12 types of anomalies, 1,159 videos and more than 2,543,000 frames in total. They designed an automated CoE reasoning pipeline to generate CoE texts for this dataset, which are richer in logic and more clearly structured than conventional textual data.

The method is evaluated on ActivityNet1.3, THUMOS14, and the newly collected HAL dataset. The experimental results demonstrate that the proposed method significantly outperforms existing methods in single-instance and multi-instance scenarios. Specifically, the authors report improvements of about 4% on the ActivityNet1.3 dataset and 12% on the THUMOS14 dataset under the multi-instance 5-shot scenario compared to the other state-of-the-art methods. The contributions are summarized as introducing a new few-shot learning method for TAL, designing a novel CoE reasoning method, collecting and annotating the first benchmark for human-related anomaly localization, and achieving state-of-the-art performance on public benchmarks.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with the resulting capabilities.

  1. Integrate a Chain-of-Evidence (CoE) Text Generation Module. I will add a two-stage pipeline to the system. First, a Vision-Language Model (VLM) will generate detailed video descriptions and then extract the main action events. Second, a Large Language Model (LLM) will be prompted to synthesize these into a structured, logical narrative that explicitly connects events with causal links (e.g., "Event A causes Event B, which leads to Event C"). This module will be used to pre-process all support videos.

  2. Implement a Semantic-Temporal Pyramid Encoder (STPE). I will replace the standard single-scale feature extractor with the STPE. This module will create a multi-scale temporal pyramid of video features and apply two types of attention: a temporal attention to model long-range dependencies across scales, and a semantic attention that dynamically selects the most similar features within a scale to enhance discrimination between foreground actions and background.

  3. Develop a Semantic-Aware Text-Visual Alignment Module. I will modify the alignment mechanism to be multimodal. Instead of only aligning query and support videos, the system will:

  • Fuse the support video features with the generated CoE text features.

  • Calculate two alignment maps: one for video-video and one for video-text.

  • Combine these maps via element-wise multiplication, using the text-based map to correct errors in the video-based map, especially for visually similar but semantically different content.

The improved system will be a few-shot, multimodal temporal action localizer with the following specific capabilities:

  • Robust Few-Shot Localization: It can accurately identify the start and end times of novel action categories in untrimmed videos using only 1 or 5 labeled support videos, without any fine-tuning on the new categories.

  • Semantic Disambiguation: It can distinguish between visually similar foreground actions and background scenes (e.g., volleyball spiking vs. playing volleyball) by leveraging the semantic information in the generated CoE text, which a purely visual system would misclassify.

  • Causal and Temporal Understanding: It can localize complex actions that have a logical progression or causal relationships (e.g., a person tripping, falling, and then a dog reacting). The CoE text provides a structured narrative that guides the model to understand the sequence of sub-events, leading to more precise boundary detection.

  • Generalization to Anomaly Detection: It can be applied to new, practical domains like human-related anomaly localization (e.g., fighting, robbery, riots). The system can use the causal reasoning in the CoE text to understand the context and logic of an anomalous event, improving its ability to find the exact anomalous segments in long, untrimmed surveillance or user-generated videos.

  • Superior Performance on Public Benchmarks: The system will achieve state-of-the-art results, with expected improvements of 4% mAP on ActivityNet1.3 and 12% mAP on THUMOS14 in the multi-instance 5-shot setting compared to existing methods.

Sources

Related papers