ME-VLM: A Unified VLM for Embodied Cognition and Agent Coordination
cs.CV
Submitted: 2026-09-21
Updated: 2026-09-22
Code: https://github.com/MachEmbodied/ME-VLM
Project page: https://machembodied.com/ME-Brain/ME-VLM.html
Terminology
Sources
- Kimi K3: Open Frontier Intelligence
- GLM-5: from Vibe Coding to Agentic Engineering
- Mach-Mind-4-Flash Technical Report
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
- Hy-Embodied-VLM-1.0: Efficient Physical-World Agents
- MiMo-Embodied: X-Embodied Foundation Model Technical Report
- Vesta: A Generalist Embodied Reasoning Model
- Towards the Harness of Embodied Agents
- Guava: Distilling Frontier VLMs into a Compact Agent through a Robotic Manipulation Harness
- RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration
- PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
- EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
- A Pragmatic VLA Foundation Model
- From Foundation to Application: Improving VLA Models in Practice
- OpenVLA: An Open-Source Vision-Language-Action Model
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Robotic Control via Embodied Chain-of-Thought Reasoning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models