RESKILL: Explicit Failure Attribution and Structured Repair for Interactive Language Agents
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: Accepted to EMNLP 2026 (Main Conference)
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
- Uncertainty-Aware Clarification in LLM Agents with Information Gain
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Odyssey: Empowering Minecraft Agents with Open-World Skills
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support
- How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
- SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
- SkillGen: Verified Inference-Time Agent Skill Synthesis
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- SkillX: Automatically Constructing Skill Knowledge Bases for Agents
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- Large Language Models Can Self-Improve At Web Agent Tasks
- AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
- EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
- SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering