SimVerity: When Does Simulated Agent Success Survive Physical Deployment?
cs.AI
Submitted: 2026-08-25
Updated: 2026-08-25
Terminology
Sources
- HomeFlow: A Data Flywheel for Smart Home Agent Training with Verifiable Simulation
- Simulating the Resident: Generating Executable Smart Home Schedules via LLM Personas
- AI Agents That Matter
- Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents
- SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
- WorldEval: World Model as Real-World Robot Policies Evaluator
- Validate the Dream Before You Trust Its Verdict: Admissibility for World-Model Simulators
- WorldGym: World Model as An Environment for Policy Evaluation
- SAGE: Smart home Agent with Grounded Execution
- S5-HES Agent: Society 5.0-driven Agentic Framework to Democratize Smart Home Environment Simulation
- Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
- Predicting LLM Safety Before Release by Simulating Deployment
- When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection