MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 23 pages, 7 figures
Code: https://github.com/abelperry/AgentProbe
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Program Synthesis with Large Language Models
- AgentIF-OneDay: A Task-level Instruction-Following Benchmark for General AI Agents in Daily Scenarios
- Evaluating Large Language Models Trained on Code
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
- ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation
- Language Models for Code Optimization: Survey, Challenges and Future Directions
- Measuring Coding Challenge Competence With APPS
- Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
- SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
- Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI
- Gemini: A Family of Highly Capable Multimodal Models
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
- CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
- IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
- FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions
- IFEvalCode: Controlled Code Generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering