3D-MoE: Towards Spatial Intelligence with Mixture-of-Experts for 3D Reasoning and Action Generation
cs.CL, cs.CV, cs.RO
Submitted: 2025-01-28
Updated: 2026-09-20
Comments: Accepted to EMNLP 2026. 20 pages
Code: https://github.com/hpcaitech/Open-Sora
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- OpenVLA: An Open-Source Vision-Language-Action Model
- Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
- MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
- Visual Instruction Tuning
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
- Mixtral of Experts
- 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
- Latte: Latent Diffusion Transformer for Video Generation
- LoHoRavens: A Long-Horizon Language-Conditioned Benchmark for Robotic Tabletop Manipulation
- Mixture-of-Experts Meets Instruction Tuning:A Winning Combination for Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering