EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory
cs.CV
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/deepseek-ai/deepseek-harness
Project page: https://zzzmyyzeng.github.io/EpiCon
Terminology
Sources
- BabyVision: Visual Reasoning Beyond Language
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks
- MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
- MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
- LowPowAR: Power-Constrained Tone Mapping for Augmented Reality
- Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning
- MemGPT: Towards LLMs as Operating Systems
- Vision2Code: A Multi-Domain Benchmark for Evaluating Image-to-Code Generation
- Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving
- Hierarchical multimodal transformers for Multi-Page DocVQA
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Agent Workflow Memory
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- OpenSkill: Open-World Self-Evolution for LLM Agents
- WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
- Aurora: Unified Video Editing with a Tool-Using Agent
- MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models