COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following

arXiv:2509.04470 · cs.CL, cs.AI · Submitted 2025-08-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following".

Tom: Autonomous agents executing human instructions must operate reliably even when instructions are incomplete, and this reliability requires enforcing missing information as a precondition for action.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So we’re looking at this paper called "COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following". Basically, the main idea is that when autonomous agents follow human instructions, they have to work reliably even if those instructions are incomplete.

Jane: That’s right. The authors argue that just detecting missing information isn't enough anymore. Agents often keep going with execution even after they realize something is missing, which leads to mistakes or unsafe actions instead of stopping and asking for help first.

Lu: They propose a modular architecture called COCORELI to fix this by making sure detection and prevention are structurally linked. It’s not just about spotting the gap; it's about blocking the action until that missing piece is filled in.

Meng: So if an agent sees a null field, instead of guessing what to do, it has to pause and generate a specific question to get that information before moving on? That sounds like a practical way to handle real-world tasks.

Lalam: Exactly. They represent instructions as structured objects where fields start empty, and execution is strictly blocked whenever those required fields are still unresolved. This forces an explicit resolution of missing information before anything happens, preventing implicit guessing.

Tom: It’s about this structural coupling that they claim makes a big difference compared to other methods we’ve seen for handling incomplete tasks. They show this works even when the underlying model gets bigger, which is something we need to keep in mind.

Jane: That's interesting because usually, making models larger just makes them get better at guessing instead of reliably knowing what they don't know. The paper suggests that how the task structure and uncertainty are represented during execution matters more than just raw model capability, according to COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following.

Lu: They test this in a controlled environment called ENVIRONMENT which is designed specifically to stress-test things like incomplete task specifications and reconstructing old structures. This environment forces the agents to deal with novel object types and physical constraints that mimic real assembly tasks, demanding they resolve ambiguities before acting.

Paper summary: Meng: From an engineering standpoint, I’m curious about how this structure handles those dynamic environments. If the instruction changes or you encounter a new object type on the fly, how does COCORELI manage that without needing constant retraining?

Lalam: The architecture includes components like an Instruction Parser and a Builder which figure out what the parts are and what their specific properties are, like orientation or configuration. Then there’s the Discourse Module that generates those targeted clarification questions when things aren't clear, which enforces explicit resolution before execution.

Tom: So, the mechanism for handling underspecification is very concrete: you have a system that detects a null field and immediately triggers a specific clarification question to fill it in. It stops guessing entirely.

Jane: It really boils down to making the process of asking for more information an explicit, required step in the workflow rather than something optional that an agent might skip over when it’s rushed.

Lu: The results they show are quite strong, especially for complex structure construction where COCORELI achieved a high overall accuracy of seventy-eight point five seven percent, beating both the CoT baselines and the agentic baseline. Also, in testing abstraction for ToolBench API tasks, it hit one hundred on all three metrics when compared to single-LLM CoT baselines.

Meng: That’s a significant jump from those other methods. If an AI can reliably reuse workflow structures across different tasks without needing task-specific fine-tuning, that opens up a lot of possibilities for building more adaptable systems in the real world.

Lalam: And it does show cost-invariance across different task types, meaning the output size doesn't grow with how complex or large the structure is. That’s good because it keeps resource usage predictable regardless of the complexity of what you’re trying to build.

Tom: So, when we think about what this means for practical AI deployment, it suggests that building reliable systems isn't just about having a bigger brain; it’s about building a better set of rules and structures around how those brains interact with incomplete information.

Paper summary: Jane: It shifts the focus from just improving the model itself to designing the interaction layer so that ambiguity is handled systematically and safely, which I think is really important for any collaborative AI application.

Lu: The authors point out a limitation, though, which is that their setup assumes tasks can always be perfectly represented by structured schemas. Also, they note that this evaluation environment abstracts away things like perception and multimodal grounding.

Meng: That makes sense. If the system relies entirely on having a perfect structural representation beforehand, it might struggle when the input to the agent is messy sensory data instead of clean text instructions.

Lalam: And another point they raise is that their conversational component only models a narrow form of dialogue: just asking for missing task parameters. It doesn't really cover more complex social stuff like negotiation or deep reasoning about intent.

Tom: So, while COCORELI solves the problem of execution errors due to missing inputs, it stops short on modeling the richer, more nuanced human conversations we see in collaborative work today.

Jane: That means COCORELI is excellent at enforcing structural correctness in a known format, but it doesn't necessarily give us a fully social or perfectly flexible dialogue agent yet.

Lu: But the core contribution remains architectural because it enforces that link between detection and prevention, regardless of how big the underlying model is. That’s the main point they’re making about where we should focus our research effort.

Meng: So, for me, what this means practically is that we need to bake this kind of explicit precondition checking directly into our system design from the start, not just add it on as an afterthought when things go wrong.

Lalam: It suggests that reliable agent execution really benefits from having these explicit mechanisms for both asking about missing parameters and using reusable task abstractions. That's the core message of COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following.

Conclusion: Tom: So, we've been looking at COCORELI and what it does for agents following human instructions.

Jane: Basically, this paper is about making sure when an AI is doing a complex task with incomplete directions, it doesn't just guess and fail; it has to stop and ask for the missing pieces first.

Lu: It’s a modular architecture designed to link detecting what’s missing directly to blocking the action. They treat instructions like structured objects where every required piece has to be filled in before anything moves forward.

Meng: From an engineering standpoint, that structural coupling is key because it doesn't rely on the underlying model being perfect at guessing; it forces a protocol for getting the right information down.

Lalam: The core contribution here is architectural, not about making the model bigger. It shows that enforcing this structural check works no matter how small or large the brain behind it is.

Tom: So, to wrap up, COCORELI isn't just another model; it's a system built around forcing explicit clarification when information is missing in a workflow.

Jane: Exactly. It moves the focus from just improving raw intelligence to designing the interaction layer so that ambiguity leads to a request for more detail instead of an incorrect action.

Lu: They tested this in environments that simulate real-world assembly, and it held up well, showing good accuracy even when things get messy with incomplete steps.

Meng: It’s interesting how they show cost-invariance across different task types; the system doesn't suddenly blow up in size just because the instruction gets more complicated.

Lalam: The paper suggests that this explicit mechanism for checking missing parameters is necessary if we want agents to be truly reliable collaborators, not just smart guessers.

Tom: It really puts a lot of pressure on us to build these kinds of safety checks into the design from the very beginning.

Jane: And once you have that foundation, it opens up possibilities for building more robust and trustworthy AI systems in real-world settings.

Lu: Which brings us to how this approach compares to other ways of handling uncertainty in instruction following.

Swarnadeep Bhar, Omar Naim, Eleni Metheniti, Bastien Navarri, Loïc Cabannes, Morteza Ezzabady, Nicholas Asher

IRIT

cs.CL, cs.AI

Submitted: 2025-08-29

Updated: 2026-10-04

Comments: 21 pages

DOI: 10.18653/v1/2026.sigdial-1.36

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: Autonomous agents executing human instructions must operate reliably even when instructions are incomplete, and this reliability requires enforcing missing information as a precondition for action.

Key concepts

COCORELI Architecture
This is a modular system that structures instructions into 'structured executable objects.' It ensures that the process of finding missing information and actually executing an action are tightly linked. The core idea is to make sure you can't proceed until all necessary parameters for an action are explicitly known or resolved.
Underspecification Handling
When instructions have missing parts, COCORELI handles this by using typed JSON structures where fields start as null. Execution stops immediately if any required field is unresolved. Instead of guessing, the system uses a 'Discourse Module' to ask a precise question to get the missing information.
Controlled Evaluation Environment
The system tests agents in a specific setting called ENVIRONMENT. This environment simulates real-world challenges like incomplete tasks and evolving situations. It forces agents to resolve ambiguities and ensure instructions are fully specified before they can perform any actions, mimicking practical constraints.

Terminology

Summary

Autonomous agents executing human instructions must operate reliably even when instructions are incomplete, and this reliability requires enforcing missing information as a precondition for action.

COCORELI Architecture

The paper proposes COCORELI (Cooperative, Reconstitution & Execution of Language Instructions), a modular architecture designed to enforce execution preconditions for reliable collaborative instruction following. This architecture represents instructions as structured executable objects whose parameters may be known or missing and ensures that detection and prevention are structurally coupled: detecting a missing parameter simultaneously blocks execution

Mechanism for Underspecification Handling

COCORELI functions by representing objects and actions as typed JSON structures whose fields are initialized to null which are then filled with information extracted from instructions or clarification responses. Execution is strictly controlled through this mechanism: Execution is blocked whenever required fields remain unresolved; the system instead invokes the Discourse Module to generate a targeted clarification question This design explicitly prevents implicit guessing by enforcing explicit resolution of missing information before execution

Controlled Evaluation Environment

The system is evaluated in a controlled construction environment called ENVIRONMENT, which is designed to isolate key conditions like incomplete task specification, evolving state with limited history, and reconstruction of previously observed structures This environment features novel object types and stricter physical constraints inspired by real-world assembly, which necessitates that agents resolve ambiguities and ensure that instructions are sufficiently specified before executing actions

Evaluation Results and Comparison

Empirical analysis shows that COCORELI demonstrates superior reliability under underspecified instructions compared to other paradigms. For complex structure construction (Task iii), COCORELI achieved the highest overall accuracy (78.57%), outperforming both the CoT baselines and the agentic baseline Furthermore, in testing abstraction for ToolBench API tasks, COCORELI achieved 100 on all three metrics when compared to single-LLM CoT baselines, illustrating that explicit structural representations and external memory enable reliable workflow reuse

Efficiency and User Burden Analysis

COCORELI exhibits cost-invariance across different task types because its output size is independent of the structure’s size or complexity, unlike CoT systems where output size grows directly with the number of parts The system's user burden is quantified by metrics like True Positive (TP) CQs, which measure how well clarification questions correctly target missing information, providing an objective assessment of cognitive load

Architectural Contribution

The primary contribution is architectural rather than performance-driven because the central claim is not that COCORELI outperforms larger models, but that it enforces a structural coupling between detecting missing information and executing actions This enforcement operates independently of model scale, meaning the guarantee applies to missing parameter cases regardless of whether the underlying model is smaller or larger

Limitations

Limitations include the assumption that tasks can be represented through structured schemas and the fact that the current benchmark abstracts away from perception and multimodal grounding Furthermore, the conversational component models only a narrow form of dialogue: clarification for resolving missing task parameters, not addressing richer collaborative discourse phenomena like negotiation or social reasoning Finally, COCORELI does not include an explicit planning module for decomposing high-level goals into sequences of actions, which remains an open challenge

The results suggest that clarification and structured abstraction are not merely implementation choices but necessary components for reliable collaborative execution The paper concludes that systems lacking explicit mechanisms for detecting missing task parameters and representing reusable task structures tend to rely on implicit inference, which frequently leads to incorrect or inconsistent actions The paper demonstrates that reliable agent execution benefits from explicit mechanisms for both clarification of missing task parameters and reusable task abstractions

The gist: COCORELI enforces a structural coupling between detecting missing information and executing actions by blocking execution until required details are resolved through targeted clarification. The paper demonstrates that reliable collaborative instruction following requires architectural enforcement rather than just model capability.

How it works

The architecture decomposes the task into six specialized components: Instruction Parser, Locator, Builder, External Memory, Discourse Module, and Executor Each component uses the same underlying LLM but operates with a specialized prompt and structured inputs, ensuring that Execution is blocked whenever required fields remain unresolved to prevent implicit guessing

Role of Components

The Instruction Parser identifies object types and extracts preliminary attributes, while the Builder resolves object-specific attributes such as orientation or part configuration, and the Locator interprets spatial descriptions into candidate coordinates The Discourse Module is responsible for generating targeted clarification questions when fields remain null, thereby enforcing explicit resolution of missing information before execution<ref:2509.

Improvements for AI systems

  1. textbf Enforcement of Execution Preconditions for Reliable Collaborative Instruction Following: Introduce COCORELI, a modular architecture that represents task structure, tracks missing information, and blocks execution until required details are resolved through targeted clarification. This ensures reliable behavior requires architectural enforcement, preventing agents from proceeding to execution even after recognizing underspecification.

  2. textbf Structural Coupling of Detection and Prevention: Implement a design where detecting a missing parameter simultaneously blocks execution, contrasting with approaches where detection and execution remain decoupled and actions may still be produced under uncertainty. This shifts the burden of resolution from the LLM's inference to an explicit, deterministic mechanism.

  3. textbf Deterministic Cost Control and Predictable Clarification: Utilize COCORELI's architecture to ensure that the model cost remains constant across specified and underspecified tasks, as the system handles missing information deterministically without bloating output sequences. This contrasts with CoT systems where underspecification penalty manifests as stochastic increases in token usage.

  4. textbf Explicit Structural Representations for Abstraction: Employ COCORELI's mechanism of representing objects and actions as typed JSON structures whose fields are initialized to null, allowing the system to store reusable task abstractions that can be reconstructed in new contexts by instantiating them with different parameters. This contrasts with agentic baselines that must rely on generating complex, error-prone code for pattern induction.

  5. textbf Controlled Evaluation via Environment Benchmark: Evaluate systems within the ENVIRONMENT benchmark to isolate failure modes like underspecification handling and structural reuse, providing a controlled environment where incorrect placements are not permitted and cannot be corrected post hoc. This allows for a systematic comparison of different architectural paradigms under well-defined uncertainty.

Abstract

Autonomous agents executing human instructions must operate reliably even when instructions are incomplete. While recent approaches improve detection of missing information, detection alone is insufficient: agents often proceed to execution even after recognizing underspecification, leading to incorrect or unsafe actions. We identify this failure as arising from a lack of coupling between detection and execution, and propose that reliable behavior requires enforcing missing information as a precondition for action. We instantiate this principle in Cocoreli, a modular architecture that represents task structure, tracks missing information, and blocks execution until required details are resolved through targeted clarification. In Cocoreli, detection and prevention are structurally coupled: detecting a missing parameter simultaneously blocks execution. We evaluate Cocoreli in a controlled construction environment isolating underspecification and sequential execution. Cocoreli blocks execution under unresolved specifications by construction, eliminating hallucinated actions. In contrast, chain-of-thought, prompt-chaining, and ReAct-style reasoning may still execute under incomplete specifications despite high detection rates. The same representation supports abstraction and reuse, and generalizes to API workflow tasks on ToolBench. These results show that reliable collaborative execution requires architectural enforcement, not just model capability

Sources

Related papers