TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces
cs.AI, cs.CL
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/ZhishanQ/TraceDance
Terminology
Sources
- SWE-chat: Coding Agent Interactions From Real Users in the Wild
- Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
- Seven simple steps for log analysis in AI systems
- Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
- GLM-5: from Vibe Coding to Agentic Engineering
- Seed1.5-VL Technical Report
- ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
- BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
- LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
- REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
- Kimi K3: Open Frontier Intelligence
- Evaluating whether AI models would sabotage AI safety research
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- MiniMax Sparse Attention
- Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?
- OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories
- RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions
- Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection