Agent-Editing World Model: Rethinking World Modeling for LLM Agents
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/AweAI-Team/AweAgent
Terminology
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- ClawGym: A Scalable Framework for Building Effective Claw Agents
- Toward Autonomous Long-Horizon Engineering for ML Research
- BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
- Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
- CWM: An Open-Weights LLM for Research on Code Generation with World Models
- GLM-5: from Vibe Coding to Agentic Engineering
- World Modelling Improves Language Model Agents
- DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
- Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Discriminative World Models for Web Agents
- Simulating Environments with Reasoning Models for Agent Training
- CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- World-Model-Augmented Web Agents with Action Correction
- ECHO: Terminal Agents Learn World Models for Free
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering