Uncertainty-Weighted Fusion of Image and Synthetic Event for Video Anomaly Detection
cs.CV
Submitted: 2025-05-05
Updated: 2026-09-21
Code: https://github.com/EavnJeong/IEF-VAD
Terminology
Sources
- Flamingo: a Visual Language Model for Few-Shot Learning
- Recent Event Camera Innovations: A Survey
- One Model, Multiple Modalities: A Sparsely Activated Approach for Text, Sound, Image, Video and Code
- PaLM-E: An Embodied Multimodal Language Model
- Convolutional Transformer based Dual Discriminator Generative Adversarial Networks for Video Anomaly Detection
- Cross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection
- Multimodal Chain-of-Thought Reasoning in Language Models
- Multi-modal Sensor Fusion for Auto Driving Perception: A Survey
- Expanding Event Modality Applications through a Robust CLIP-Based Encoder
- Multi-Modal Anomaly Detection for Unstructured and Uncertain Environments
- Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
- Decoupled Weight Decay Regularization
- Video Anomaly Detection and Explanation via Large Language Models
- Fusion framework and multimodality for the Laplacian approximation of Bayesian neural networks
- Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
- MIMIC-IT: Multi-Modal In-Context Instruction Tuning
- Uncertainty aware audiovisual activity recognition using deep Bayesian variational inference
- EventCLIP: Adapting CLIP for Event-based Object Recognition
- VERA: Explainable Video Anomaly Detection via Verbalized Learning of Vision-Language Models
- Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models