AgentGarten: Code Worlds for Evolving Agents
cs.CV
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/MirroS-Lab/AgentGarten
Project page: https://mirros-lab.github.io/agent-garten
Terminology
Sources
- Genie: Generative Interactive Environments
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
- Infinite Worlds with Versatile Interactions
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- One-step Diffusion with Distribution Matching Distillation
- Improved Distribution Matching Distillation for Fast Image Synthesis
- Self Gradient Forcing: Native Long Video Extrapolation
- ViPE: Video Pose Engine for 3D Geometric Perception
- Depth Anything 3: Recovering the Visual Space from Any Views
- Image Generators are Generalist Vision Learners
- SAM 3: Segment Anything with Concepts
- SAM 3D: 3Dfy Anything in Images
- Cosmos 3: Omnimodal World Models for Physical AI
- Wan: Open and Advanced Large-Scale Video Generative Models
- NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation
- World Simulation with Video Foundation Models for Physical AI
- Diffusion Adversarial Post-Training for One-Step Video Generation
- Holodeck: Language Guided Generation of 3D Embodied AI Environments
- SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models