SceneTeract: Probing and Improving Agent-Aware Activity Reasoning in 3D Indoor Scenes
cs.CV
Submitted: 2026-03-31
Updated: 2026-09-17
Code: https://github.com/PhyScene/PhyScene
Project page: https://sceneteract.github.io
Terminology
Sources
- WorldScore: A Unified Evaluation Benchmark for World Generation
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Demystifying MMD GANs
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- AI2-THOR: An Interactive 3D Environment for Visual AI
- Ministral 3
- SpatialLM: Training Large Language Models for Structured Indoor Modeling
- Human-Aware 3D Scene Generation with Spatially-constrained Diffusion Models
- LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D Scans
- RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
- Gemma 3 Technical Report
- Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
- Function2Scene: 3D Indoor Scene Layout from Functional Specifications
- SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
- Qwen3 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- SceneEval: Evaluating Semantic Coherence in Text-Conditioned 3D Indoor Scene Synthesis
- Gemini: A Family of Highly Capable Multimodal Models
- Viser: Imperative, Web-based 3D Visualization in Python
- MobiAgent: A Systematic Framework for Customizable Mobile Agents
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models