COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following".
Tom: Autonomous agents executing human instructions must operate reliably even when instructions are incomplete, and this reliability requires enforcing missing information as a precondition for action.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So we’re looking at this paper called "COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following". Basically, the main idea is that when autonomous agents follow human instructions, they have to work reliably even if those instructions are incomplete.
Jane: That’s right. The authors argue that just detecting missing information isn't enough anymore. Agents often keep going with execution even after they realize something is missing, which leads to mistakes or unsafe actions instead of stopping and asking for help first.
Lu: They propose a modular architecture called COCORELI to fix this by making sure detection and prevention are structurally linked. It’s not just about spotting the gap; it's about blocking the action until that missing piece is filled in.
Meng: So if an agent sees a null field, instead of guessing what to do, it has to pause and generate a specific question to get that information before moving on? That sounds like a practical way to handle real-world tasks.
Lalam: Exactly. They represent instructions as structured objects where fields start empty, and execution is strictly blocked whenever those required fields are still unresolved. This forces an explicit resolution of missing information before anything happens, preventing implicit guessing.
Tom: It’s about this structural coupling that they claim makes a big difference compared to other methods we’ve seen for handling incomplete tasks. They show this works even when the underlying model gets bigger, which is something we need to keep in mind.
Jane: That's interesting because usually, making models larger just makes them get better at guessing instead of reliably knowing what they don't know. The paper suggests that how the task structure and uncertainty are represented during execution matters more than just raw model capability, according to COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following.
Lu: They test this in a controlled environment called ENVIRONMENT which is designed specifically to stress-test things like incomplete task specifications and reconstructing old structures. This environment forces the agents to deal with novel object types and physical constraints that mimic real assembly tasks, demanding they resolve ambiguities before acting.
Paper summary: Meng: From an engineering standpoint, I’m curious about how this structure handles those dynamic environments. If the instruction changes or you encounter a new object type on the fly, how does COCORELI manage that without needing constant retraining?
Lalam: The architecture includes components like an Instruction Parser and a Builder which figure out what the parts are and what their specific properties are, like orientation or configuration. Then there’s the Discourse Module that generates those targeted clarification questions when things aren't clear, which enforces explicit resolution before execution.
Tom: So, the mechanism for handling underspecification is very concrete: you have a system that detects a null field and immediately triggers a specific clarification question to fill it in. It stops guessing entirely.
Jane: It really boils down to making the process of asking for more information an explicit, required step in the workflow rather than something optional that an agent might skip over when it’s rushed.
Lu: The results they show are quite strong, especially for complex structure construction where COCORELI achieved a high overall accuracy of seventy-eight point five seven percent, beating both the CoT baselines and the agentic baseline. Also, in testing abstraction for ToolBench API tasks, it hit one hundred on all three metrics when compared to single-LLM CoT baselines.
Meng: That’s a significant jump from those other methods. If an AI can reliably reuse workflow structures across different tasks without needing task-specific fine-tuning, that opens up a lot of possibilities for building more adaptable systems in the real world.
Lalam: And it does show cost-invariance across different task types, meaning the output size doesn't grow with how complex or large the structure is. That’s good because it keeps resource usage predictable regardless of the complexity of what you’re trying to build.
Tom: So, when we think about what this means for practical AI deployment, it suggests that building reliable systems isn't just about having a bigger brain; it’s about building a better set of rules and structures around how those brains interact with incomplete information.
Paper summary: Jane: It shifts the focus from just improving the model itself to designing the interaction layer so that ambiguity is handled systematically and safely, which I think is really important for any collaborative AI application.
Lu: The authors point out a limitation, though, which is that their setup assumes tasks can always be perfectly represented by structured schemas. Also, they note that this evaluation environment abstracts away things like perception and multimodal grounding.
Meng: That makes sense. If the system relies entirely on having a perfect structural representation beforehand, it might struggle when the input to the agent is messy sensory data instead of clean text instructions.
Lalam: And another point they raise is that their conversational component only models a narrow form of dialogue: just asking for missing task parameters. It doesn't really cover more complex social stuff like negotiation or deep reasoning about intent.
Tom: So, while COCORELI solves the problem of execution errors due to missing inputs, it stops short on modeling the richer, more nuanced human conversations we see in collaborative work today.
Jane: That means COCORELI is excellent at enforcing structural correctness in a known format, but it doesn't necessarily give us a fully social or perfectly flexible dialogue agent yet.
Lu: But the core contribution remains architectural because it enforces that link between detection and prevention, regardless of how big the underlying model is. That’s the main point they’re making about where we should focus our research effort.
Meng: So, for me, what this means practically is that we need to bake this kind of explicit precondition checking directly into our system design from the start, not just add it on as an afterthought when things go wrong.
Lalam: It suggests that reliable agent execution really benefits from having these explicit mechanisms for both asking about missing parameters and using reusable task abstractions. That's the core message of COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following.
Conclusion: Tom: So, we've been looking at COCORELI and what it does for agents following human instructions.
Jane: Basically, this paper is about making sure when an AI is doing a complex task with incomplete directions, it doesn't just guess and fail; it has to stop and ask for the missing pieces first.
Lu: It’s a modular architecture designed to link detecting what’s missing directly to blocking the action. They treat instructions like structured objects where every required piece has to be filled in before anything moves forward.
Meng: From an engineering standpoint, that structural coupling is key because it doesn't rely on the underlying model being perfect at guessing; it forces a protocol for getting the right information down.
Lalam: The core contribution here is architectural, not about making the model bigger. It shows that enforcing this structural check works no matter how small or large the brain behind it is.
Tom: So, to wrap up, COCORELI isn't just another model; it's a system built around forcing explicit clarification when information is missing in a workflow.
Jane: Exactly. It moves the focus from just improving raw intelligence to designing the interaction layer so that ambiguity leads to a request for more detail instead of an incorrect action.
Lu: They tested this in environments that simulate real-world assembly, and it held up well, showing good accuracy even when things get messy with incomplete steps.
Meng: It’s interesting how they show cost-invariance across different task types; the system doesn't suddenly blow up in size just because the instruction gets more complicated.
Lalam: The paper suggests that this explicit mechanism for checking missing parameters is necessary if we want agents to be truly reliable collaborators, not just smart guessers.
Tom: It really puts a lot of pressure on us to build these kinds of safety checks into the design from the very beginning.
Jane: And once you have that foundation, it opens up possibilities for building more robust and trustworthy AI systems in real-world settings.
Lu: Which brings us to how this approach compares to other ways of handling uncertainty in instruction following.
Swarnadeep Bhar, Omar Naim, Eleni Metheniti, Bastien Navarri, Loïc Cabannes, Morteza Ezzabady, Nicholas Asher
IRIT
cs.CL, cs.AI
Submitted: 2025-08-29
Updated: 2026-10-04
Comments: 21 pages
DOI: 10.18653/v1/2026.sigdial-1.36
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 91/100
The gist: Autonomous agents executing human instructions must operate reliably even when instructions are incomplete, and this reliability requires enforcing missing information as a precondition for action.
Key concepts
- COCORELI Architecture
- This is a modular system that structures instructions into 'structured executable objects.' It ensures that the process of finding missing information and actually executing an action are tightly linked. The core idea is to make sure you can't proceed until all necessary parameters for an action are explicitly known or resolved.
- Underspecification Handling
- When instructions have missing parts, COCORELI handles this by using typed JSON structures where fields start as null. Execution stops immediately if any required field is unresolved. Instead of guessing, the system uses a 'Discourse Module' to ask a precise question to get the missing information.
- Controlled Evaluation Environment
- The system tests agents in a specific setting called ENVIRONMENT. This environment simulates real-world challenges like incomplete tasks and evolving situations. It forces agents to resolve ambiguities and ensure instructions are fully specified before they can perform any actions, mimicking practical constraints.
Terminology
Summary
Autonomous agents executing human instructions must operate reliably even when instructions are incomplete, and this reliability requires enforcing missing information as a precondition for action.
COCORELI Architecture
The paper proposes COCORELI (Cooperative, Reconstitution & Execution of Language Instructions), a modular architecture designed to enforce execution preconditions for reliable collaborative instruction following. This architecture represents instructions as structured executable objects whose parameters may be known or missing
and ensures that detection and prevention are structurally coupled: detecting a missing parameter simultaneously blocks execution
Mechanism for Underspecification Handling
COCORELI functions by representing objects and actions as typed JSON structures whose fields are initialized to null
which are then filled with information extracted from instructions or clarification responses. Execution is strictly controlled through this mechanism: Execution is blocked whenever required fields remain unresolved; the system instead invokes the Discourse Module to generate a targeted clarification question
This design explicitly prevents implicit guessing by enforcing explicit resolution of missing information before execution
Controlled Evaluation Environment
The system is evaluated in a controlled construction environment called ENVIRONMENT, which is designed to isolate key conditions like incomplete task specification, evolving state with limited history, and reconstruction of previously observed structures
This environment features novel object types
and stricter physical constraints inspired by real-world assembly,
which necessitates that agents resolve ambiguities and ensure that instructions are sufficiently specified before executing actions
Evaluation Results and Comparison
Empirical analysis shows that COCORELI demonstrates superior reliability under underspecified instructions compared to other paradigms. For complex structure construction (Task iii), COCORELI achieved the highest overall accuracy (78.57%), outperforming both the CoT baselines and the agentic baseline
Furthermore, in testing abstraction for ToolBench API tasks, COCORELI achieved 100 on all three metrics
when compared to single-LLM CoT baselines, illustrating that explicit structural representations and external memory enable reliable workflow reuse
Efficiency and User Burden Analysis
COCORELI exhibits cost-invariance across different task types because its output size is independent of the structure’s size or complexity,
unlike CoT systems where output size grows directly with the number of parts The system's user burden is quantified by metrics like True Positive (TP) CQs,
which measure how well clarification questions correctly target missing information,
providing an objective assessment of cognitive load
Architectural Contribution
The primary contribution is architectural rather than performance-driven because the central claim is not that COCORELI outperforms larger models, but that it enforces a structural coupling between detecting missing information and executing actions
This enforcement operates independently of model scale,
meaning the guarantee applies to missing parameter cases regardless of whether the underlying model is smaller or larger
Limitations
Limitations include the assumption that tasks can be represented through structured schemas and the fact that the current benchmark abstracts away from perception and multimodal grounding Furthermore, the conversational component models only a narrow form of dialogue: clarification for resolving missing task parameters,
not addressing richer collaborative discourse phenomena like negotiation or social reasoning Finally, COCORELI does not include an explicit planning module for decomposing high-level goals into sequences of actions,
which remains an open challenge
The results suggest that clarification and structured abstraction are not merely implementation choices but necessary components for reliable collaborative execution
The paper concludes that systems lacking explicit mechanisms for detecting missing task parameters and representing reusable task structures tend to rely on implicit inference, which frequently leads to incorrect or inconsistent actions The paper demonstrates that reliable agent execution benefits from explicit mechanisms for both clarification of missing task parameters and reusable task abstractions
The gist: COCORELI enforces a structural coupling between detecting missing information and executing actions by blocking execution until required details are resolved through targeted clarification. The paper demonstrates that reliable collaborative instruction following requires architectural enforcement rather than just model capability.
How it works
The architecture decomposes the task into six specialized components: Instruction Parser, Locator, Builder, External Memory, Discourse Module, and Executor
Each component uses the same underlying LLM but operates with a specialized prompt and structured inputs,
ensuring that Execution is blocked whenever required fields remain unresolved
to prevent implicit guessing
Role of Components
The Instruction Parser identifies object types and extracts preliminary attributes, while the Builder resolves object-specific attributes such as orientation or part configuration,
and the Locator interprets spatial descriptions into candidate coordinates The Discourse Module is responsible for generating targeted clarification questions
when fields remain null, thereby enforcing explicit resolution of missing information before execution
<ref:2509.
Improvements for AI systems
-
textbf Enforcement of Execution Preconditions for Reliable Collaborative Instruction Following: Introduce COCORELI, a
modular architecture that represents task structure, tracks missing information, and blocks execution until required details are resolved through targeted clarification.
This ensuresreliable behavior requires architectural enforcement,
preventing agents from proceeding to execution even after recognizing underspecification. -
textbf Structural Coupling of Detection and Prevention: Implement a design where
detecting a missing parameter simultaneously blocks execution,
contrasting with approaches wheredetection and execution remain decoupled and actions may still be produced under uncertainty.
This shifts the burden of resolution from the LLM's inference to an explicit, deterministic mechanism. -
textbf Deterministic Cost Control and Predictable Clarification: Utilize COCORELI's architecture to ensure that
the model cost remains constant across specified and underspecified tasks,
as the system handles missing informationdeterministically without bloating output sequences.
This contrasts with CoT systems whereunderspecification penalty manifests as stochastic increases in token usage.
-
textbf Explicit Structural Representations for Abstraction: Employ COCORELI's mechanism of representing objects and actions as
typed JSON structures whose fields are initialized to null,
allowing the system to storereusable task abstractions
that can bereconstructed in new contexts by instantiating them with different parameters.
This contrasts with agentic baselines that must rely on generating complex, error-prone code for pattern induction. -
textbf Controlled Evaluation via Environment Benchmark: Evaluate systems within the ENVIRONMENT benchmark to isolate failure modes like
underspecification handling
andstructural reuse,
providing a controlled environment whereincorrect placements are not permitted and cannot be corrected post hoc.
This allows for asystematic comparison of different architectural paradigms under well-defined uncertainty.
Abstract
Autonomous agents executing human instructions must operate reliably even when instructions are incomplete. While recent approaches improve detection of missing information, detection alone is insufficient: agents often proceed to execution even after recognizing underspecification, leading to incorrect or unsafe actions. We identify this failure as arising from a lack of coupling between detection and execution, and propose that reliable behavior requires enforcing missing information as a precondition for action. We instantiate this principle in Cocoreli, a modular architecture that represents task structure, tracks missing information, and blocks execution until required details are resolved through targeted clarification. In Cocoreli, detection and prevention are structurally coupled: detecting a missing parameter simultaneously blocks execution. We evaluate Cocoreli in a controlled construction environment isolating underspecification and sequential execution. Cocoreli blocks execution under unresolved specifications by construction, eliminating hallucinated actions. In contrast, chain-of-thought, prompt-chaining, and ReAct-style reasoning may still execute under incomplete specifications despite high detection rates. The same representation supports abstraction and reuse, and generalizes to API workflow tasks on ToolBench. These results show that reliable collaborative execution requires architectural enforcement, not just model capability
Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft
- MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation
- SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive Tasks
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Large Language Model Guided Tree-of-Thought
- Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback
- The Llama 3 Herd of Models
- Analyzing limits for in-context learning
- Re-examining learning linear functions in context
- ART: Automatic multi-step reasoning and tool-use for large language models
- TALM: Tool Augmented Language Models
- Open-Ended Instructable Embodied Agents with Memory-Augmented Large Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Agentic LLM Workflows for Generating Patient-Friendly Medical Reports
- ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
- How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Can Code Language Models Learn Clarification-Seeking Behaviors?
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering