BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs
Sanjeev Manivannan
Indian Institute of Technology Madras
cs.AI, cs.CE, cs.ET
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 14 pages, 2 figures
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs Abstract Summary: Organizational decisions are co-created while evidence, constraints, and
Terminology
Summary
BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs
Abstract Summary:
Organizational decisions are co-created while evidence, constraints, and human priorities continue to change. In conventional transcript-based multi-agent systems, a human typically provides an initial problem, agents deliberate internally, and the system returns a final response. BoardroomAI instead treats the human as a persistent participant who can intervene during the discussion by challenging assumptions, modifying constraints, changing priorities, introducing evidence, or redirecting the decision process. The system operationalizes this human–agent coexistence through four components: (i) a typed decision graph whose nodes represent evidence, assumptions, constraints, claims, objections, alternatives, risks, and decisions, and whose edges represent semantic dependencies and specialist responsibility; (ii) an intervention compiler that converts each confirmed human action into explicit node- and edge-level graph updates; (iii) a dependency-aware propagation mechanism that measures the structural impact of the intervention, identifies the affected decision subgraph, preserves unaffected artifacts, and selectively reactivates the relevant domain specialists; and (iv) an evaluation framework that measures intervention impact, repair coverage, preservation, recomputation, and final decision validity. On 600 generated decision-DAG interventions, the proposed propagation mechanism matched exhaustive impact computation while inspecting only 14.59% of nodes. In a 12-case exploratory agent pilot, selective repair recomputed 62.11% of canonical nodes, preserved all gold-unaffected nodes, and produced valid updated decisions in six cases, while conservatively abstaining in the remaining six. These abstentions show that correct intervention routing does not necessarily provide sufficient context for decision synthesis. The paper therefore formalizes a decision-sufficient context closure for human-steered multi-agent deliberation. All reported results are synthetic and prototype-level.
Introduction and Problem Formulation:
The central research problem is revision routing: after a human intervention, the system must determine (i) which claims, assumptions, constraints, or decisions have changed semantic status; (ii) which generated artifacts must be revised, recomputed, or invalidated; (iii) which unaffected artifacts should be preserved without unnecessary regeneration; and (iv) which specialist agents must be reactivated to repair the affected portion of the deliberation. BoardroomAI treats the recommendation, its alternatives, and its supporting rationale as a shared co-created artifact. The human is neither a prompt provider restricted to the beginning nor an approver restricted to the end and can redirect the process, revise constraints, challenge an inference, change priorities, preserve a minority objection, or make an accountable override while the agents continue to elaborate the artifact.
The paper claims the following narrower system contributions:
-
Decision-graph representation with formal decision-graph semantics and alternative minimal justification environments for preserving independently supported claims.
-
Typed human interventions formalized as typed graph operations, including invalidation, weakening, contestation, supersession, and preference changes.
-
Selective specialist repair that propagates intervention impact through the decision graph, selectively reactivates only affected specialists, and specifies conservative fallback to full redeliberation when repair is insufficient.
-
DynaBoard evaluation, a 12-case synthetic pilot for evaluating routing independent of the final decision quality.
Motivation Example:
Consider a company evaluating the launch of a new product. Finance, marketing, legal, and operations specialists deliberate over multiple alternatives and eventually recommend a staged product launch under a 100 million budget. After the discussion reaches consensus, the decision maker reduces the budget to 40 million. While this change invalidates some earlier conclusions, others remain fully justified. A transcript-based multi-agent system either restarts the entire discussion or continues without explicitly determining which previous decisions remain valid, leading to unnecessary recomputation or inconsistent reasoning. BoardroomAI instead represents the deliberation as a dependency-aware decision graph. The intervention is propagated only through affected dependencies, invalidating the paid-acquisition launch strategy while preserving the independent regional-partnership justification and unrelated legal conclusions. Only the relevant specialists are reactivated, and the system produces an auditable change log explaining what changed, why it changed, and which decisions remain valid.
Formal State Representation:
At turn t, the deliberation state is S t = (G t, L t, M t, Π, H t, E t, B t), where G t = (V t, E t) is a typed directed multigraph; L t stores justification labels; M t is concise shared memory; Π contains agents, tools, and policies; H t is the intervention ledger; E t is an immutable evidence/provenance ledger; and B t records token, latency, and monetary budgets. A node is v = ⟨id, τ, x, s, o, c, p, t⟩: type, content, status, owner, confidence, provenance, and timestamp. Types are evidence/fact, assumption, hard or soft constraint, claim, objection, alternative, risk, question, intermediate conclusion, and decision. Status is active, weakened, contested, stale, superseded, rejected, unresolved, or accepted. Edge types include requires, supports, derives, rebuts, undercuts, supersedes, answers, and requiresHuman. Evidence nodes require a source locator, retrieval time, and content hash; model-generated claims never become evidence merely by repetition.
Minimal Justification Environments:
For each derived node v, the label L t(v) = J v 1,..., J v m contains subset-minimal alternative environments. Each J is a conjunction of supporting literals and rule identifiers; the label is their disjunction. Let ok t(J) = ∧ l∈J active t(l) ∧ ¬undercut t(J). Then v remains supportable iff ∨ J∈L t(v) ok t(J). Loss of some but not all environments weakens v; loss of all makes it stale. A rebuttal makes a claim contested unless an explicit rule or authoritative evidence invalidates its justification. This prevents majority agreement from silently deleting a minority objection.
Human Intervention as Decision-Graph Evolution:
An intervention parser returns a confirmed graph delta Δ h. Its components D+, D-, D↔, D?, and ρ encode additions, retractions, replacements, challenges, and preference/authority updates. A replacement does not merely overwrite text: it deactivates the old node, adds a new version, records a supersedes edge, and then reevaluates only those environments that reference the changed version. Directly changed nodes seed a queue. For each downstream node, the engine recomputes the semantic signature q t(v) = ⟨s t(v), J ∈ L t(v): ok t(J), a t(v), x t(v)⟩, where a t(v) is its active attack set. Only a signature change propagates. The impact set I h is the least fixed point of actual signature changes: I h = μX. D h ∪ v: ∃u ∈ X ∩ pred(v), σ t+1(v) ≠ σ t(v). The repair set R h ⊆ I h contains stale or contested generated artifacts and content whose source, constraint, or preference basis changed. Let P h = V t I h denote the preservation set. Preservation means that the validated semantic signature and provenance of an artifact do not change. The system is selective when R h ⊆ C h and C h ≪ V t, and it is safe only when every mandatory repair obligation is covered or the system falls back to full redeliberation.
Routing and Fallback:
Obligations include re-estimate, retrieve, reconcile, generate alternative, adjudicate objection, and resynthesize. Owners, expertise tags, source access, and estimated cost define a weighted set-cover instance; the router greedily maximizes uncovered risk-weighted obligations per expected cost, then adds a skeptic for contested high-impact claims. A full restart is mandatory for global objective changes, low dependency/extraction confidence, unresolved cycles, or an affected fraction above a preregistered threshold. Fallback is based on preregistered diagnostics: graph-coverage audit score below τ c, mean edge-confidence below τ e, changed-node fraction above τ f, a global objective/ontology edit, or an unstable component.
Conditional Preservation:
Assuming complete dependency annotations, deterministic pure update functions, and propagation over an acyclic graph or its strongly connected component condensation, every node outside the least-fixed-point impact set preserves its semantic signature. The proof is induction over topological order: an unchanged node has neither a direct delta nor a changed predecessor signature, so its deterministic update is unchanged. This is a conditional systems property, not a claim about fallible LLM extraction. The structural pass is O(V+E+Σ vL(v)); LLM cost remains empirical.
Decision-Sufficient Repair Packets:
Correct impact routing is necessary but not sufficient. The synthesizer must also receive enough unchanged boundary context to prove feasibility and reconstruct the declared decision rubric. The decision-sufficient closure of a repair set is defined as DSC(R h) = R h ∪ B h hard ∪ B h rubric ∪ B h alts ∪ B h risk ∪ B h source, where the five boundary sets contain, respectively, every active hard constraint relevant to a surviving alternative, every criterion and human-approved weight used by the synthesizer, the roots of surviving alternative justifications, unresolved objections and material risks, and the minimal evidence spans needed to verify included claims. The packet contains semantic summaries plus identifiers and provenance, not an unrestricted replay of the meeting. This closure is a design correction motivated by the reported pilot: the prototype selective condition supplied affected subgraphs and hand-selected boundary premises, yet abstained in six of twelve cases because the synthesizer could not establish a complete constraint/rubric basis.
Decision Synthesis:
Synthesis is constraint-first, not a vote over prose. For alternative d, the graph service evaluates executable hard predicates g(d), then computes a transparent soft-utility interval U(d) = Σ k ω k u k(d) − Σ r η r p r(d)l r(d), where weights ω k are human-approved, u k are normalized criterion scores, and each risk has probability p r, loss l r, and severity weight η r. Missing quantities remain intervals or unresolved nodes; they are not imputed by the judge. The synthesizer ranks only hard-feasible alternatives, reports sensitivity to weights and uncertain inputs, and abstains if no alternative is feasible. Agent votes are retained as provenance but have no direct coefficient in U.
Mechanism-Level Synthetic Validation:
The deterministic propagation core was implemented in Python and tested on 600 generated decision DAG interventions (150 each with 64, 128, 256, and 512 nodes). Each derived node had one to three subset-minimal alternative environments. One or two active source nodes were invalidated; 175 generated attempts with trivial or near-global impact were rejected before obtaining the fixed 600-case suite. An independent exhaustive evaluator recomputed every node and defined gold semantic impact as any change in active environment IDs. With complete annotations, selective propagation agreed with exhaustive recomputation in all 600 cases. It inspected 14.59% of nodes (95% bootstrap CI 13.47–15.85) and reused 89.96% of unaffected nodes (89.11–90.80). Reachability obtained full recall by inspecting 59.78% of nodes, but its mean precision was 10.47% (9.49–11.50). Direct updating and a single-environment label were cheaper but missed downstream or alternative-environment changes. The oracle gap (14.59% versus 5.75%) is boundary-inspection overhead. Median Python structural speed-up over full recomputation rose from 2.11× at 64 nodes to 22.08× at 512 nodes.
Missing-Dependency Stress Test:
Randomly deleting routing edges while retaining the complete graph for gold evaluation and deliberately disabling fallback (240 cases per rate) showed that only 10% deletion reduced recall to 90.55% (88.48–92.52) and exact-set recovery to 60.42%. Precision remains 100% because deletion hides paths but cannot create false ones in this monotone simulation; wrong-edge corruption remains untested. Thus the perfect complete-graph result validates an implementation conditional, while the stress test exposes dependency extraction and fallback detection as unresolved risks.
Exploratory Validation:
The validation consists of an end-to-end stress test over a controlled 12-task synthetic benchmark. Each organizational case contains four candidate decisions, supporting evidence, executable hard constraints, weighted decision criteria, and one hidden intervention that alters both the feasible solution space and the optimal decision. Ground-truth dependency annotations identify directly affected, transitively affected, revision-required, and unaffected artifacts. Four execution settings were compared: C1 employs a single iterative agent; C2 uses multiple specialist agents followed by complete re-deliberation after every intervention; C3 performs structured full re-deliberation using the canonical decision graph; C4 activates only specialists whose artifacts lie within the routed repair region. The exploratory implementation was executed using the product-facing gpt-5.6-terra interface with medium reasoning.
C1, C2, and C3 produced a valid updated decision for every evaluated task. C1 correctly identified only 44.44% of affected artifacts, whereas C2 preserved none of the gold-unaffected reasoning because every specialist was restarted. C3 achieved perfect routing accuracy but recomputed every canonical artifact, including 37.89% that should have remained unchanged. In contrast, C4 preserved every unaffected artifact and reduced specialist activations by 11.11%, recomputing only 62.11% of canonical artifacts. However, it abstained on six of the twelve tasks, producing a valid updated recommendation only when the routed repair packet contained sufficient information to reconstruct the final decision. The most important observation is that dependency-correct routing alone is insufficient for successful selective repair: C4 abstained whenever the available repair context was insufficient to establish a complete constraint–criterion basis for the synthesizer.
Design Implications:
Human steerability should update authoritative state rather than append another message. Preservation is also a co-creative capability because it protects valid alternatives, dissent, and provenance from unnecessary stochastic rewriting. Finally, routing minimality and decision sufficiency are different: a small impact set may still omit the hard constraints or rubric context needed for synthesis. Evaluation must therefore report validity, preservation, coverage, completion, and cost together.
Conclusion:
BoardroomAI models human–AI deliberation as evolution of an external decision graph: typed human edits propagate through alternative justifications, trigger obligation-based specialist repair, and preserve unaffected objections and provenance. On 600 complete generated decision DAGs, propagation exactly matched exhaustive impact sets while inspecting 14.59% of nodes; missing edges sharply degraded recovery. In the 12-case pilot, selective repair preserved all unaffected canonical nodes but abstained on half the cases. This negative result motivates decision-sufficient closure: correct routing must also supply the constraints, rubric, alternatives, risks, and evidence needed for synthesis. The contribution is therefore a falsifiable mechanism, not evidence of general organizational superiority.
Improvements for AI systems
Improvements to AI Systems Based on BoardroomAI:
-
Persistent Human-in-the-Loop Deliberation: Build AI systems that allow humans to intervene at any point during multi-agent reasoning—not just at the prompt or final approval stage. The system should accept mid-process edits (e.g., changing constraints, adding evidence, challenging assumptions, or overriding decisions) and treat these as first-class graph operations rather than appended messages.
-
Dependency-Aware Decision Graphs: Represent all reasoning artifacts (claims, assumptions, constraints, objections, alternatives, risks) as typed nodes in a directed graph with explicit semantic edges (supports, rebuts, undercuts, supersedes). This enables the system to track which conclusions depend on which premises and to compute the exact impact of any change.
-
Selective Impact Propagation: Implement a propagation engine that, upon a human intervention, computes the least fixed point of affected nodes by recomputing each node’s semantic signature (status, active justification environments, attack set). Only nodes whose signature changes are marked for repair; all others are preserved without regeneration. This avoids full re-deliberation and reduces computational cost.
-
Minimal Justification Environments: Store multiple subset-minimal alternative justification sets for each derived claim. When one premise is invalidated, the system checks whether other environments still support the claim. This preserves independently justified conclusions and prevents unnecessary invalidation.
-
Obligation-Based Specialist Routing: After identifying the affected subgraph, generate a set of repair obligations (e.g., re-estimate, retrieve new evidence, reconcile conflicts, generate alternatives, adjudicate objections). Route these obligations to the appropriate specialist agents using a weighted set-cover optimization that maximizes coverage of high-risk obligations per unit cost, while adding a skeptic agent for contested high-impact claims.
-
Decision-Sufficient Context Closure: Before synthesizing a final decision, automatically assemble a repair packet containing: (a) all active hard constraints relevant to surviving alternatives, (b) the decision rubric (criteria and human-approved weights), (c) roots of surviving alternative justifications, (d) unresolved objections and material risks, and (e) minimal evidence spans. This ensures the synthesizer has enough context to produce a valid decision, not just a correctly routed one.
-
Constraint-First Decision Synthesis: Rank alternatives only if they satisfy all executable hard predicates. Compute transparent soft-utility intervals using human-approved weights and explicit risk terms (probability × loss × severity). Do not impute missing quantities; leave them as intervals or unresolved nodes. Abstain if no alternative is feasible, and report sensitivity to weight changes.
-
Conservative Fallback with Preregistered Diagnostics: Automatically trigger full re-deliberation when: (a) the graph-coverage audit score falls below a threshold, (b) mean edge-confidence is low, (c) the affected fraction of nodes exceeds a threshold, (d) a global objective or ontology is edited, or (e) an unstable component is detected. This prevents silent degradation when dependency annotations are incomplete.
-
Auditable Change Logging: Maintain an immutable ledger of every intervention, its graph-level effects, and the propagation results. The system should output an explicit explanation of what changed, why it changed, which prior decisions remain valid, and which specialists were reactivated—enabling full auditability and trust.
-
Preservation as a First-Class Capability: Explicitly protect unaffected artifacts—including minority objections, valid alternatives, and provenance—from stochastic rewriting. This prevents unnecessary regeneration and preserves dissent, which is critical for accountable decision-making.
What the Improved AI System Can Do:
-
Handle mid-meeting changes gracefully: A user can reduce a budget, add a legal constraint, or challenge an assumption during an ongoing multi-agent deliberation, and the system will update only the affected reasoning while keeping all valid conclusions intact.
-
Produce valid decisions with minimal recomputation: It will recompute only 15% of nodes on average (versus 100% for full re-deliberation), preserving 90% of unaffected artifacts, while still matching exhaustive impact analysis exactly when dependencies are complete.
-
Abstain safely when context is insufficient: Instead of generating a plausible but unsupported recommendation, the system will explicitly state that it lacks sufficient constraint/rubric context and request the missing information or trigger full re-deliberation.
-
Provide transparent, auditable rationale: Every change is traceable to a specific intervention, with a change log showing which claims were invalidated, weakened, or preserved, and why.
-
Respect human authority and dissent: It will not silently delete a minority objection when a majority agrees; rebuttals make claims contested unless authoritative evidence overrides them, and human overrides are recorded as explicit graph operations.
-
Scale to complex organizational decisions: The dependency-aware propagation is efficient (22× speedup at 512 nodes versus full recomputation) and can handle multiple alternatives, risks, and weighted criteria in a single coherent decision graph.
Abstract
Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve. In conventional transcript-based multi-agent systems, humans typically provide an initial problem, agents deliberate internally, and the system returns a final response. BoardroomAI instead treats the human as a persistent participant who can intervene by challenging assumptions, modifying constraints, changing priorities, introducing evidence, or redirecting the decision process. We operationalize this human--agent coexistence through four components: (i) a typed decision graph representing evidence, assumptions, constraints, claims, objections, alternatives, risks, decisions, semantic dependencies, and specialist responsibility; (ii) an intervention compiler that converts confirmed human actions into explicit graph updates; (iii) dependency-aware propagation that identifies affected subgraphs, preserves unaffected artifacts, and selectively reactivates relevant specialists; and (iv) an evaluation framework measuring intervention impact, repair coverage, preservation, recomputation, and decision validity. Across 600 generated decision-DAG interventions, propagation matched exhaustive impact computation while inspecting only 14.59% of nodes. In a 12-case exploratory pilot, selective repair recomputed 62.11% of canonical nodes, preserved all gold-unaffected nodes, and produced valid updated decisions in six cases while abstaining in the remaining six. These abstentions show that correct intervention routing may still provide insufficient context for synthesis, motivating a decision-sufficient context closure for human-steered multi-agent deliberation. All results are synthetic and prototype-level.
Sources
- Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations
- An Ontology of Co-Creative AI Systems
- From Debate to Deliberation: Structured Collective Reasoning with Typed Epistemic Acts
- Towards AI as Colleagues: Multi-Agent System Improves Structured Ideation Processes
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection