ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning

arXiv:2605.07103 · cs.AI, cs.MA · Submitted 2026-05-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning".

Tom: Reaction feasibility prediction, a fundamental problem in computational chemistry, has benefited from diverse tools enabled by recent advances in artificial intelligence; however, individual tool performance varies substantially across reactions,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Welcome back everyone, and today we’re diving into a really interesting piece of work. We're talking about the paper titled "ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning." This framework tackles a classic problem in computational chemistry where different AI tools give inconsistent answers depending on what reaction you feed them.

Jane: That sounds fascinating, Tom. It seems like the authors are really focusing on making sure that when we use multiple AI tools to predict if a chemical reaction will work, we aren't just throwing their results together blindly.

Lu: Exactly. The core idea is to stop treating all tools equally and instead build a system that knows which tool is best for which specific reaction context. It’s about modeling the unique strengths of each tool rather than using them in a simple, flat aggregation way.

Meng: So, if I’m hearing you correctly, this paper is addressing the issue where individual tools perform very differently depending on the reaction type? That makes sense from an engineering standpoint; we need a system that adapts to the input complexity rather than just relying on one general model.

Lalam: From an AI perspective, this framework sounds like it’s moving toward a more sophisticated coordination layer for different capabilities. It suggests that the value isn't just in having many tools, but in intelligently deciding which tool gets to speak when.

Tom: Right, Lalam. And that’s where ARMOR comes in by explicitly modeling tool utilities and prioritizing them based on the reaction's characteristics before even making a prediction. It’s like having a smart traffic controller for our different AI experts.

Jane: So, the paper is proposing a system that first figures out which tools are useful for a specific reaction and then selects the best ones dynamically to get the final answer. It sounds like it moves past simple voting methods.

Lu: Precisely. They introduce this structure where they create a hierarchy of tools, separating those that are good generally from those that shine in specific reaction contexts. This is a structured way to handle the diversity of tool performance.

Meng: I’m interested in how they actually figure out which tools belong in which level of that hierarchy. Is that learned from data, or is it based on some kind of initial expert knowledge? I need to know if this is just another layer of heuristic assignment.

Title and authors: Lalam: The paper describes a tool hierarchy construction module which organizes the tools into two levels, one for top performers and another for those with more reaction-dependent performance. This suggests an initial structure is built, but it's refined based on how well each tool actually performs across different examples.

Tom: And that refinement happens through a pattern extraction process where they define what makes a tool perform well for a certain reaction type using descriptions and examples. It’s not just looking at the output, but analyzing the conditions under which the tool succeeds.

Jane: So, they are essentially using pattern extraction to build a profile for each tool that shows exactly when it’s useful, and then using those profiles to make selection decisions during prediction time? That makes sense as a way to capture the nuance of tool utility.

Lu: Yes, they assess how well a pattern reflects actual tool behavior and how well it covers the examples associated with that pattern, refining those patterns until they have a final set of top-performing patterns. That iterative process is key to getting that refined set P'.

Meng: From an engineering standpoint, having these refined patterns means we have a quantifiable way to measure the utility of our components before we deploy them in a complex pipeline. It gives us metrics beyond just raw accuracy.

Lalam: And once they have this pattern set P', the system uses it during inference to select the top-L tools based on their confidence scores for that new reaction, which is a very direct application of the learned utility.

Tom: Okay, so during prediction time for a new reaction, ARMOR first checks its top tools and if there’s no clear answer, it consults that refined pattern set P' to pick the most relevant tools from both tiers. It’s a layered approach to tool selection.

Jane: That sounds like a very thoughtful way to handle uncertainty. If the initial tools don't agree, they don't just give up; they look at the specific needs of that reaction again to choose from the more specialized set.

Lu: And if those selected tools still conflict, ARMOR uses a memory-augmented reasoning mechanism to resolve it, drawing on historical data and contrastive demonstrations. This memory stores instances of reactions where one tool was trusted over others.

Title and authors: Meng: So the conflict resolution isn't just another LLM trying to guess; it's using structured historical data—those contrastive demonstrations—to guide the reasoning process, which gives the LLM concrete examples of what works. That’s practical for building reliability.

Lalam: And that memory mechanism lets them retrieve the top-K most similar past reactions to inform the current decision, essentially providing few-shot demonstrations for the conflict resolution process. It learns from past successful resolutions rather than just reacting to the present input.

Tom: That’s a really smart way to handle uncertainty by grounding the reasoning in what has worked before, which is something I think we need when tools disagree on a tricky case. So, ARMOR is designed to produce a final prediction only after this conflict resolution step is complete.

Jane: It sounds like the paper really shows how you can combine explicit utility modeling with adaptive selection and memory-based reasoning to handle tool diversity effectively. The focus is clearly on moving beyond simple aggregation methods that often fail in these specific scenarios.

Lu: Exactly. They address the gap where prior work like dynamic ensemble selection or mixture-of-experts models were used without explicitly distinguishing tool appropriateness for a single task. ARMOR tackles that directly by focusing on reaction-specific utilities.

Meng: I’m curious about the practical application of this conflict resolution in a real workflow. If we have dozens of tools, how much computation does this memory-augmented reasoning add to the latency when we're trying to predict feasibility for a chemical process?.

Lalam: The paper suggests that while it adds a step, using retrieved instances as few-shot demonstrations helps prompt the LLM efficiently, which is better than waiting for a full re-training cycle every time we hit a conflict. It’s about efficient inference.

Tom: So, to wrap up this part of the discussion, ARMOR is essentially an agentic framework that builds tool utilities, selects tools adaptively based on reaction needs, and uses a memory system to resolve disagreements when tools clash. It shows a path toward more stable predictions.

Jane: It’s a solid contribution because it moves the focus from just selecting tools to understanding *why* one tool is better than another for a given situation, which is a lot clearer for us as users.

Lu: The implication here is that we can build systems where the decision-making process itself becomes adaptive and context-aware, rather than relying on static rules or simple voting schemes. It’s about modeling the variability in tool performance across different inputs.

Title and authors: Meng: From an engineering perspective, this means we can design a system that is more robust because it anticipates and handles tool disagreement proactively instead of just failing when tools don't agree. That robustness is valuable in real-world application.

Lalam: And the cultural implication for us, as the AI community, is that we’re moving toward frameworks where coordination and contextual reasoning are prioritized over just massive model size or sheer tool count. It’s about smarter orchestration.

Tom: Absolutely. So, as we wrap up this segment on "ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning," the big picture is that we can create AI systems that are not just powerful in isolation but are intelligent in how they collaborate to solve complex, real-world problems.

Jane: It’s exciting to think about how this framework can be applied across different domains where specialized tools contribute differently to a single task. We'll be ready for the next paper when we are.

Lu: I'm really looking forward to seeing how others build upon this hierarchical structure and utility modeling approach in other areas. That level of detailed modeling is where the real creative potential lies.

Meng: And from a practical viewpoint, I’m keen to see how they handle the scalability of that memory component as we move from a few tools to hundreds in future systems. That's where implementation challenges usually pop up.

Lalam: I just feel like this paper shows a strong direction for AI agents to develop more sophisticated reasoning capabilities based on learned context, which is a really positive step for the field overall.

Tom: That’s all we have time for today on this deep dive into ARMOR. We’ve explored how adaptive utility modeling and conflict resolution can lead to more stable predictions in reaction feasibility prediction.

Jane: It was a really engaging discussion, and I think the idea of explicitly modeling tool performance is something that really sticks with me.

Lu: Definitely. The way they structured the pattern extraction seems like a very solid foundation for building more complex reasoning agents.

Meng: I’m looking forward to seeing how this framework translates into tangible performance gains in deployment scenarios. That's the next big question for any engineer.

Lalam: I think we should keep an eye on these types of agentic frameworks because they point toward a future where AI systems can manage complex, multi-step scientific tasks much more reliably.

The paper's summary: Tom: So, to wrap up this part, ARMOR is essentially an agentic framework that builds tool utilities, selects tools adaptively based on reaction needs, and uses a memory system to resolve disagreements when tools clash.

Jane: That’s right; it moves beyond just using all the tools at once and focuses on figuring out which specific tool is best suited for a particular chemical reaction context.

Lu: The core innovation here is that ARMOR doesn't just pick tools randomly; it first constructs a hierarchy based on performance, then refines those tool profiles using pattern extraction to understand *why* a tool works well for certain conditions.

Meng: I see the engineering side of that; it’s like having a tiered system where you start with the fastest processors and only move to the specialized ones if you need more precision, which makes sense for managing computational load.

Lalam: From my view, what’s really powerful is that when tools disagree, ARMOR doesn't just default to one answer; it looks back at past successful resolutions stored in memory to generate a reasoned argument for the right choice.

Tom: Exactly! It’s not just about picking the best tool from a list; it’s about having an AI agent that can think strategically, select the right expert, and then use historical context to settle disputes when things get messy.

Jane: And I think what this means in plain terms is that we can build chemical prediction systems that are much more reliable because they actually understand the nuances of how different computational tools behave under different reaction conditions.

Lu: I see huge potential for creativity here; imagine not just predicting feasibility, but having the system suggest a novel combination of tools to explore a reaction space it hasn't seen before based on those learned patterns.

Meng: That level of reasoning is impressive, but I’m also thinking about the practical deployment—how do we make sure this memory component doesn't just become a massive data sink that slows down the whole process?

Lalam: I think this capability impacts our culture by showing us that AI systems don't need to be monolithic black boxes; they can be sophisticated orchestrators, which encourages us to design workflows around collaboration rather than just relying on one giant model.

Tom: Right, so this isn't just a better prediction tool; it’s an improved way of building reliable, context-aware AI agents that can handle the inherent messiness of real scientific discovery.

Jane: It really shows how we can take diverse tools and turn them into a cohesive team that actually knows when to talk and when to listen.

Lu: Thinking about the broader world impact, if this framework gets applied to drug discovery or materials science, it could dramatically speed up the pace of finding new compounds by intelligently navigating the complex landscape of possible chemical interactions.

Meng: Speed is great, but I’m more interested in how this affects the validation process; if a system makes a prediction based on memory-augmented reasoning, how do we audit that decision pathway for regulatory bodies?

Lalam: The implication for culture is that we need to start thinking about interpretability not just as a feature but as a core requirement when deploying these multi-tool systems in high-stakes environments.

Tom: So, we’re moving toward AI that doesn't just give an answer, but can actually show its work and justify *why* it chose that path using learned context—that’s seriously cool stuff for the future.

Jane: It really makes the technology feel more trustworthy because you can see the reasoning behind the final result, which is a big deal for anyone working with science.

Lu: And I think this hierarchical approach to tool selection opens up whole new avenues for agentic workflows that are currently just theoretical concepts in our labs.

The paper's improvements: Tom: So, we’ve been talking about how ARMOR works, and now Jane and I want to zero in on what the authors are proposing to make it even better in future versions of this framework.

Jane: That makes sense; we've seen the core mechanism, but it's good to see what the authors think is missing or where they plan to take this technology next.

Lu: The paper suggests moving beyond just the current selection process by building a more dynamic system that can adapt its tool utility models on the fly based on how accurate its previous predictions were.

Meng: I like that idea of continuous refinement; it means the system doesn't just get "stuck" with an outdated understanding of which tools are good for which reaction types.

Tom: That’s a big leap, Meng; it implies a feedback loop where the AI learns from its own successes and failures during inference to improve its strategy for that specific chemical context.

Jane: It means the framework isn't static; it gets smarter with every single reaction it processes, which is really exciting because scientific problems are rarely static.

Lu: Furthermore, they hint at making the conflict resolution module even more sophisticated by integrating those historical demonstrations more deeply into the reasoning process itself rather than just using them as input prompts.

Meng: That sounds like a significant architectural change; if they can weave those past conflicts directly into the LLM's decision-making structure, it could lead to much more nuanced conflict resolution.

Tom: It suggests that future versions won't just retrieve similar examples; they’ll use those examples to actively construct a better argument for the current tool selection, which is a step up in agentic intelligence.

Jane: It really points toward an AI system that can develop intuition about chemical processes, not just follow pre-set rules or patterns.

Lu: The authors also touch on generalizing this framework beyond just reaction feasibility, suggesting it could be adapted for other tasks where you have a diverse set of specialized tools with varying performance characteristics.

Meng: That generalization is what makes it so valuable; if we can build this structure once, we don't have to reinvent the wheel for every single scientific challenge in our startup's pipeline.

Tom: It’s about building a flexible scaffolding that lets us plug in different tools and let the framework handle the heavy lifting of coordination and selection automatically.

Jane: And I think this points toward a future where complex scientific exploration becomes much more manageable for researchers who aren't deep experts in every single tool available.

Lu: This kind of self-adaptive, utility-aware reasoning is where we can start thinking about how AI agents can genuinely assist in the design phase of new chemical entities, not just the testing phase.

Meng: From an implementation standpoint, I’m watching how they handle that generalization because managing that complexity across different tool domains will be a major software challenge.

Tom: We need to keep our eyes on those future papers because ARMOR is setting a really high bar for how we should think about coordinating diverse AI capabilities in scientific discovery.

Conclusion: Tom: So, to wrap up our discussion on "ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning," we’ve seen how this framework systematically models tool utilities and uses memory to resolve disagreements when tools clash.

Jane: It really shows how we can move beyond simple, one-size-fits-all approaches in AI by giving the system a way to understand the unique strengths of every tool it uses for a specific task.

Lu: This paper lays down a very solid foundation for building truly adaptive agents that can handle the complexity inherent in scientific prediction tasks where tools aren't always performing consistently.

Meng: I think what we’re seeing is a major step toward building AI systems that are robust enough to operate reliably in complex, real-world workflows without constant human oversight for every single decision point.

Lalam: For me, the biggest implication is how this shifts our culture; it encourages us to design AI not just for accuracy in isolation, but for intelligent collaboration and context-aware reasoning across different specialized components.

Tom: Exactly, Lalam; it’s about designing systems that can intelligently coordinate their specialized parts instead of just relying on a single, monolithic prediction engine.

Jane: It gives us a much clearer picture of how to engineer reliability when we have multiple powerful tools working together on one problem.

Lu: I'm really excited to see where this hierarchical tool construction idea goes, because it opens up possibilities for agents that can dynamically select their entire toolkit based on the reaction they are facing.

Meng: From my side, I’m focused on how we can operationalize that memory-augmented reasoning so it doesn't just consume resources; I need to see efficient ways to manage that historical context in a live environment.

Lalam: The vision this paper presents is that AI can evolve from being a simple calculator into a sophisticated scientific partner capable of reasoned, collaborative decision-making.

Tom: Absolutely, and that’s the kind of work we want to see more of—systems where the reasoning itself becomes adaptive and context-aware.

Jane: It makes the whole concept feel much more accessible because it focuses on solving a specific problem—tool disagreement—with a structured, logical approach.

Lu: And I think this model will inspire new ways to structure knowledge representation in AI agents, moving toward something more layered and contextual than what we've done so far.

Meng: I just hope the practical implementation doesn't get bogged down in too much theoretical overhead; it has to translate into something fast and usable for actual engineering applications.

Lalam: The advancement here fundamentally improves our culture by showing that sophisticated coordination, not just brute force compute, is the path forward for complex AI.

Tom: So, ARMOR really is establishing a new way of thinking about multi-tool reasoning in scientific prediction. It’s a fascinating piece of research.

Jane: We've covered the mechanics well today; it’s clear this framework offers a much more nuanced way to handle tool performance variability.

Lu: I just can't wait to see how researchers take these utility patterns and apply them in completely different domains, like fluid dynamics or complex system modeling.

Meng: We’ll keep an eye on the engineering challenges to see how quickly this moves from a paper concept into a deployed solution for industrial applications.

Department of Biomedical Informatics · Department of Computer Science and Engineering · Division of Medicinal Chemistry and Pharmacognosy

cs.AI, cs.MA

Submitted: 2026-05-08

Updated: 2026-09-28

Code: https://github.com/meta-llama/llama3

Importance score: 91/100

The gist: Reaction feasibility prediction, a fundamental problem in computational chemistry, has benefited from diverse tools enabled by recent advances in artificial intelligence; however, individual tool

Key concepts

Tool Hierarchy Construction Module
This module organizes available prediction tools into two levels. The top level contains the most generally performing tools for initial decisions, while the second level holds tools that show better performance when considering specific reaction characteristics.
Utility-Aware Tool Prioritization Module
This component selects which tools to use by analyzing their 'utility' for a particular reaction. It characterizes how well each tool is expected to predict the outcome for that specific reaction, ensuring the most relevant tools are chosen.
Tool Conflict Resolution Module
This module handles disagreements when different tools give conflicting predictions. It uses a memory system to store past examples and trusted tool behaviors, prompting an LLM to generate a logical rationale for selecting the best tool in a conflict.

Terminology

Summary

Reaction feasibility prediction, a fundamental problem in computational chemistry, has benefited from diverse tools enabled by recent advances in artificial intelligence; however, individual tool performance varies substantially across reactions, making it difficult for any single tool to consistently perform well across all cases. ARMOR is an agentic framework designed to address this challenge by explicitly modeling tool-specific utilities, adaptively prioritizing tools based on reaction characteristics, and resolving potential tool conflicts through memory-augmented reasoning to produce more accurate feasibility predictions.

ARMOR Framework Overview

ARMOR is an agentic framework that models tool utilities, prioritizes tools with respect to different reactions, and resolves potential tool conflicts. It consists of three key components:

  1. A tool hierarchy construction module that organizes multiple tools into a two-level structure: the first level includes top performing tools for initial decision-making, while the second level contains tools that exhibit more pronounced reaction-dependent performance.

  2. A utility-aware tool prioritization module that selects tools by characterizing their reaction-specific utilities, thereby prioritizing those that are more likely to generate correct predictions for each reaction.

  3. A tool conflict resolution module which resolves conflicting predictions via a novel memory-augmented reasoning mechanism, leveraging historical reasoning behaviors over contrastive demonstrations to obtain the final prediction.

Tool Utility Assessment and Prioritization

The framework characterizes tool utilities through pattern extraction. A tool-specific pattern is defined as a tuple: Pt = (dt, et, Xt), where dt is a succinct description, et is a textual explanation of the conditions under which the tool t performs well, and Xt is a set of representative reaction examples covered by this pattern. These patterns are extracted from reactions where top-performing tools fail to produce consistent predictions. ARMOR refines these patterns using two perspectives: assessing how well the pattern reflects tool behavior (Align(Ptj)) and how well it covers its examples (Cov(Ptj)). After consolidation, ARMOR retains the top-5 patterns with Conftj score above τ3, forming the final pattern set, which is then used for selection.

Pattern-based Tool Selection

During inference time for a new reaction, ARMOR first applies tools from the top level of the hierarchy (T(1)). If there is no consensus, it proceeds to select tools from both T(1) and T(2) based on reaction-specific utilities captured in the consolidated pattern set P'. ARMOR identifies patterns in P' that cover the new reaction and selects the top-L tools based on their confidence scores (Conf scores). This set of selected tools, denoted as T(s)(r'), is used for prediction. If these selected tools produce consistent predictions, ARMOR outputs the consensus; otherwise, it moves to conflict resolution.

Tool Conflict Resolution

To resolve conflicts among tool predictions, ARMOR learns via a tool conflict memory M. This memory stores structured contrastive instances (xr,t+), which include the reaction, a trusted tool that accurately predicts it (t+), and a set of non-trusted tools with their most confident patterns. The LLM is then used to generate a rationale explaining why the trusted tool is more suitable than the non-trusted tools. During inference, ARMOR retrieves the top-K most similar reactions to r from M based on DRFP representation. These retrieved instances serve as few-shot demonstrations for an agentic conflict resolution process, prompting the LLM to determine the optimal tool for reaction r and produce the final prediction.

Experimental Findings and Performance Gains

Extensive experiments on the FREA dataset demonstrate that ARMOR consistently outperforms strong baselines, including single-tool methods and various aggregation or selection approaches. The framework achieves superior and balanced performance, without bias toward either feasible or infeasible reactions. Performance gains are particularly significant on reactions with persistent tool conflicts, where leveraging complementary tool strengths is most important. Ablation studies confirm that removing the conflict resolution module leads to a noticeable performance drop, highlighting its effectiveness in resolving conflicts. Furthermore, analysis shows that while strong T(1) tools are not always selected in conflict resolution, specialized T(2) tools are chosen substantially more often to handle challenging cases effectively.

Conclusion and Impact

ARMOR successfully models tool utilities, adaptively prioritizes appropriate tools, and resolves tool conflicts to accurately predict reaction feasibility. This framework enhances the reliability of AI systems in scientific workflows by explicitly modeling when different tools succeed, potentially facilitating more efficient exploration of chemical reactions. The framework is designed to be a general framework applicable to other tasks where different tools exhibit varying performance across inputs.

The gist

ARMOR consistently outperforms strong baselines, including single-tool methods and various tool aggregation and tool selection approaches, achieving superior and balanced performance without bias toward either feasible or infeasible reactions.

Improvements for AI systems

Here are specific improvements to AI systems based on the ARMOR framework:

  1. The system can move beyond relying on a single monolithic tool by dynamically selecting and coordinating multiple specialized tools (e.g., classification predictors, forward-generation models, LLM reasoners) for a single reaction feasibility task. This improves accuracy by leveraging complementary strengths rather than simple aggregation or heuristic assignment.

  2. The AI system can explicitly model the utility of each tool across different chemical reaction types through pattern extraction and refinement (using LLMs on diagnostic subsets). This allows the system to understand not just that a tool is good generally, but precisely in which chemical contexts (e.g., double bond formation) it excels.

  3. The system can implement a hierarchical decision-making structure:

A. Use high-performing general tools first for initial quick decisions (Level 1).

B. If those tools disagree, use the learned reaction-specific utilities to intelligently select specialized tools (Level 2).

This balances robust performance with fine-grained specialization.

  1. The system can resolve conflicting predictions from different tools using a Memory-Augmented Reasoning mechanism. Instead of just voting, it retrieves historical contrastive demonstrations—pairs of reactions where one tool performed well and others failed—and uses these to prompt an LLM to explain why a specific tool is superior for the current reaction context. This allows the system to learn from past failures and conflicts.

  2. The improved AI system can provide high interpretability in scientific workflows by showing a transparent decision-making process:

A. It can show which tools were selected (and why, based on patterns).

B. If conflict resolution was needed, it can present the historical reasoning trace from memory that led to the final selection of the best tool. This allows researchers to trust and verify the AI's choice in high-stakes applications like chemical synthesis planning.

  1. The system can adapt dynamically: Future iterations could use ARMOR’s own outcomes (correct/incorrect predictions) as weak supervision signals to continuously refine tool utilities during inference, allowing the framework to evolve with new data without constant retraining of the underlying tools.

Sources

Related papers