Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?
cs.CL
Submitted: 2026-09-30
Updated: 2026-10-08
Terminology
Sources
- Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation
- Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- Towards Autonomous Mathematics Research
- GLM-5: from Vibe Coding to Agentic Engineering
- Automated Discovery Has No Universally Superior Harness
- DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks
- LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks
- Meta-Harness: End-to-End Optimization of Model Harnesses
- In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
- A lower bound for stepsize-based acceleration of gradient descent
- IH-Benchmark: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions
- AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
- AutomationBench
- Harnesses for Inference-Time Alignment over Execution Trajectories
- Agent Workflow Memory
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering