GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis
cs.CL
Submitted: 2026-09-30
Updated: 2026-10-04
Code: https://github.com/nousresearch/hermes-agent
Terminology
Sources
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
- Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
- Toward Scalable Terminal Task Synthesis via Skill Graphs
- CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents
- R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- OpenThoughts-Agent: Data Recipes for Agentic Models
- TaskCraft: Automated Generation of Agentic Tasks
- Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
- APEX-Agents
- Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
- How Well Does Agent Development Reflect Real-World Work?
- AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
- SWE-smith: Scaling Data for Software Engineering Agents
- Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
- EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent
- NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs
- SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering