Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems
cs.CL, cs.LG, cs.SE
Submitted: 2026-07-09
Updated: 2026-08-24
Terminology
Sources
- Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards
- SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints
- Agentic Observability: Automated Alert Triage for Adobe E-Commerce
- Bian Que: An Agentic Framework with Flexible Skill Arrangement for Online System Operations
- APEX-SQL: Talking to the data via Agentic Exploration for Text-to-SQL
- StepFly: Agentic Troubleshooting Guide Automation for Incident Diagnosis
- SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
- SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
- TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks
- Agent Workflow Memory
- CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language Models
- Stalled, Biased, and Confused: Uncovering Reasoning Failures in LLMs for Cloud-Based Root Cause Analysis
- Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering
- Qwen3 Technical Report
- ReAct: Synergizing Reasoning and Acting in Language Models
- SOP-Agent: Empower General Purpose AI Agent with Domain-Specific SOPs
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
- GLM-5: from Vibe Coding to Agentic Engineering
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering