EM 2Mem: Event-Centric Multimodal Memory for Large Language Models
cs.CL, cs.AI, cs.LG, cs.MM
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/zjunlp/LightMem
Terminology
Sources
- Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
- LightMem-Ego: Your AI Memory for Everyday Life
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
- VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
- From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
- OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
- HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
- Omni-SimpleMem: Autoresearch-Guided Discovery of Lifelong Multimodal Agent Memory
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
- VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
- PersonaVLM: Long-Term Personalized Multimodal LLMs
- OpenAI GPT-5 System Card
- Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
- Tarsier: Recipes for Training and Evaluating Large Video Description Models
- Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning
- Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering