TO-Agents: A Multi-Agent AI Framework for Subjective Preference-Guided Topology Optimization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "TO-Agents: A Multi-Agent AI Framework for Subjective Preference-Guided Topology Optimization".
Jane: The paper was written by Isabella A. Stewart, Hongrui Chen and Faez Ahmed from Department of Civil and Environmental Engineering, Massachusetts Institute of Technology and Department of Mechanical Engineering, Massachusetts Institute of Technology.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper discussion segment 1: Tom: We’ve just touched upon the broad implications, but let’s talk about the title and why "Subjective Preference-Guided" is such a big deal in this context. The authors are essentially arguing that traditional optimization is too rigid for designers who want a specific look or feel.
Jane: Right, and it’s not just about making something strong; it's about the *kind* of strength. The paper focuses on guiding the design toward things like hierarchically branched structures inspired by natural tree morphologies, which is a very subjective goal.
Meng: I want to know if this approach scales; are we talking small components, or can we build a whole system that relies on this? The paper suggests it handles both a cantilever beam and even a phone-stand product design.
Lu: That wide range of application is impressive, Lu sees the potential for applying these methods to almost any field where aesthetic goals meet functional constraints. It isn's not just an academic exercise; it has real-world versatility.
Lalam: This is about elevating the human experience, Lalam believes that by letting us input natural language like "I want a more branch-like design," we are creating a more intuitive bridge between our imagination and the machine.
Tom: It’s clear that this isn't just about getting one good solution; it’s about guiding the entire journey toward a specific preference, which is what makes TO-Agents so much more than just a standard optimization run.
Jane: And since we are focusing on subjective preference, how do we know if the AI is actually listening? Does the system have a way to objectively measure that success?
Meng: That’s where the AI Judge comes in, which is critical for understanding if this system can be trusted to consistently hit those complex design goals.
Lu: The concept of having an independent judge agent provides a robust mechanism for checking alignment, Lu sees it as essential proof that the subjective preference is being tracked.
Lalam: It’s a feedback loop that ensures our cultural values are met, Lalam believes that the judgment process validates the alignment between human intent and technological output.
Tom: That's a great way to put it—we have both a goal and a mechanism to measure its achievement, which sets us up perfectly for looking at how the framework actually puts all these pieces together in Segment three.
Paper discussion segment 2: Tom: We’ve established that TO-Agents is designed to be guided by subjective preference, but now we need to break down the actual mechanics of this multi-agent pipeline. It's not one program; it's a coordinated team of specialized agents working together.
Jane: The process starts with the human providing a natural language description, and then the Pydantic Agent takes over by converting that verbose description into validated, structured JSON for feeding into the solver.
Meng: I find this initial translation step really important; it ensures that even if the human input is loose, all necessary variables are correctly interpreted before running a high-stakes simulation.
Lu: The process flows through PyFANTOM, which is the core topology optimization solver, but Lu notes that the output isn's not just a density field; it’s rendered into images for subsequent reasoning.
Lalam: This visual representation is key because Lalam sees that it allows the agents to perceive and understand the three dee structure in a way they can use for iterative refinement.
Tom: And once we have that visual output, the Vision Agent steps in, which is where things get really interesting. It looks at the images and determines how to tweak parameters based on the human’s feedback.
Jane: That's not just tweaking random variables; it's interpreting what the human means by "make it more tree-like" and translating that into specific instructions for the Vision Agent.
Meng: The next crucial part of this pipeline is the AI Judge Agent, which scores every single revision to give a neutral assessment of whether the design is improving or not.
Lu: This judge provides objective data points on how well the subjective goal is being met, Lu sees it as providing the necessary external validation for internal agentic success.
Lalam: The entire flow suggests that Lalam believes we are automating the whole complex cycle—from initial idea to visual critique—which is a profound shift in how we design.
Tom: It’s a whole sequence of specialized roles, from input translation to visual perception and critical scoring, which is why it's so impressive. Let's see how this entire system uses its own history to get better at the next stage.
Paper discussion segment 3: Tom: We’ve seen how the agents interact in a single run, but the real power of TO-Agents is in its ability to learn and refine over time. The paper highlights that this framework isn't just running once; it’s undergoing iterative refinement.
Jane: One of the best things is that the system can recover from mistakes, which is crucial because when a design fails, trying to manually correct it usually means starting over entirely.
Meng: I was struck by how they use historical data; instead of just guessing, the agents look back at previous successful runs to guide their current parameter changes and adapt their strategy.
Lu: That concept of developing a "strategy" rather than executing a fixed plan is what Lu sees as the future—the agentic behavior is sophisticated reasoning over time, not just brute force calculation.
Lalam: We are witnessing a new form of autonomous learning, Lalam believes that the AI is not just following instructions but evolving its approach based on cultural feedback.
Tom: This self-correction ability leads to impressive success rates—specifically sixty percent of trials meet the human’s preference, and they achieved this much faster than a non-guided pipeline.
Jane: That rapid learning is partly because, even if the AI Judge scores a design poorly, the system can pivot and find a better path instead of just giving up on that specific idea.
Meng: The agents aren't just guessing; they’re using their accumulated knowledge of which levers—like the SIMP penalty or volume fraction—are most effective at fine-tuning to improve the structure.
Lu: It’s about developing an evolving strategy, Lu sees that this allows for a level of abstract planning in AI that was previously thought impossible in design tools.
Lalam: This is how we are seeing AI move beyond simple execution, Lalam believes it's becoming a truly sophisticated partner in the way we conceive of objects.
Tom: The ability to learn from its own history and adapt makes the entire system feel much more robust, and that’s what brings us to wrapping things up with the final results.
Conclusion: Tom: We’ve seen how TO-Agents takes a subjective idea and guided it through a full, iterative process, from start to finish. It's a remarkable demonstration of automated design journey.
Jane: The overall implication is that we no longer have to manually translate our subjective ideas into rigid solver settings; the AI handles the heavy lifting of translating intent into actionable code.
Meng: From a practical standpoint, I think this means we can rapidly explore complex shapes for manufacturing, which would have taken months of tedious trial and error before manual iteration.
Lu: It’s also a powerful proof that AI is capable of managing these long-horizon tasks without being explicitly told every single parameter change at scale.
Lalam: We should be excited about how this allows designs to not just meet functional requirements, Lalam believes it ensures that the technology serves our cultural goals for human interaction with objects.
Tom: And we must acknowledge the limitations, like overshooting or selective memory, which is part of a realistic look at any autonomous system.
Jane: It’s a realistic look at how AI handles its own errors; even when it's very smart, it doesn' can make mistakes and might not always follow constraints perfectly.
Lu: We are looking forward to future work on better inter-agent reasoning and addressing how the agents handle growing conversational context.
Meng: I think the biggest immediate impact is making the design process faster and more reliable for manufacturers, too, especially with that end-to-end prototyping capability they showed.
Lalam: To wrap up our discussion on this groundbreaking work, we're celebrating TO-Agents: A Multi-Agent AI Framework for Subjective Preference-Guided Topology Optimization.
Tom: It’s a remarkable achievement, and I think we have a lot of exciting things to talk about in the next paper as we move forward.
Jane: It’s certainly a powerful concept, Tom, making the entire design process much more efficient for our listeners.
Isabella A. Stewart, Hongrui Chen, Faez Ahmed
Department of Civil and Environmental Engineering, Massachusetts Institute of Technology · Department of Mechanical Engineering, Massachusetts Institute of Technology
cs.AI
Submitted: 2026-08-22
Updated: 2026-08-25
Comments: Accepted for publication in the Proceedings of the ASME 2026 International Design Engineering Technical Conferences (IDETC2026)
Code: https://github.com/ahnobari/pyFANTOM
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 85/100
The gist: The paper presents TO-Agents, a multi-agent AI framework designed to bridge the gap between human qualitative design intent and the technical requirements of topology optimization (TO).
Key concepts
- Topology Optimization
- A design method that determines the optimal material distribution within a structure to maximize performance (like strength) while minimizing weight. TO-Agents guides this process based on subjective human desires, not just rigid functional constraints.
- Multi-Agent AI Framework
- A system where multiple specialized AI agents work together in a coordinated pipeline. In TO-Agents, these agents handle distinct tasks—such as interpreting natural language or scoring revisions—to complete a complex design process.
- Subjective Preference-Guided
- The ability of the AI to optimize designs based on non-quantifiable human desires, such as achieving a 'tree-like' aesthetic. This moves beyond simple functional requirements by incorporating subjective artistic goals.
- AI Judge Agent
- A critical component that provides objective scoring and assessment during the design process. It validates whether the AI's revisions are successfully meeting the complex subjective goals set by the human user.
Terminology
Summary
The paper presents TO-Agents, a multi-agent AI framework designed to bridge the gap between human qualitative design intent and the technical requirements of topology optimization (TO). The core challenge addressed is that designers typically must manually translate qualitative intent, such as desired visual style, product experience, or manufacturability into solver settings that are not directly tied to those preferences.
The TO-Agents Framework and Methodology:
TO-Agents functions as a multi-agent AI pipeline that connects natural-language design intent with iterative topology optimization. The process is highly structured and involves several specialized agents:
-
Problem Formulation: A Pydantic agent interprets the human designer's verbose technical natural language description of the problem, transforming it into
validated, structured JSON
which provides a reliable interface for passing the problem formulation to the optimizer. -
Optimization: The TO Optimization Agent invokes an open-source solver (PyFANTOM) to execute optimization based on generated parameters. The resulting topology is then rendered in 3D and captured from six orthogonal viewpoints, with color mapping applied to aid spatial reasoning.
-
Human Preference Feedback: Following the initial run, the human designer provides one qualitative feedback describing a failure of a specific aesthetic requirement (e,g.,
Improve the design by making it have more branch-like, complexity like a tree
). -
Agentic Deep Reasoning and Refinement: The Vision Agent takes this feedback and analyzes the
full chat history,
including previous parameters and rendered 3D images. It then determines which TO parameters should be adjusted to meet the designer’s objective, operating autonomously without explicit definitions of parameter effects. -
AI Judge Evaluation: An independent AI Judge Agent evaluates both the initial and revised designs, scoring them (1–5) based on whether
the human designer’s criterion of a more branching, intricate structure has been achieved.
This judge is trained independently from the vision-language model to ensure unbiased evaluation. -
Iterative Refinement: The Judge’s feedback is fed back into the conversation history, and the Vision Agent proceeds with further revisions. The system is limited to five total iterations:
one original topology optimization followed by four revision cycles.
-
Production: Finally, a Manufacturing Agent selects the highest-scoring structure and performs post-processing (e.g., adding support structures for a phone stand) to convert the density field into a functional polygonal mesh suitable for 3D printing.
Case Studies and Results:
The framework was evaluated on two long-horizon design tasks:
-
Cantilever Beam Benchmark: A classic benchmark used in TO studies.
-
Phone-Stand Product Design: A real-world product where
human-preference design and aesthetics play a more prominent role.
In both case studies, the designer specified an aesthetic preference for hierarchically branched structures inspired by natural tree morphologies.
The system performed four revision cycles across ten independent replicates.
The results demonstrate significant success:
-
Performance:
TO-Agents produces at least one preference-aligned design in 60% of trials for each case study, corresponding to up to 6× more successful trials than an ablated pipeline without visual or historical feedback.
-
Designer Shift: The framework allows designers to shift
from low-level parameter tuning toward higher-level specification of form and function.
Observed Behaviors and Failure Modes:
The agentic pipeline reveals several collective behaviors:
-
History-Conditioned Adaptation: Agents use the growing history of prior designs, parameter settings, and AI-judge feedback to
progressively learn how to manipulate TO parameters toward the human designer’s qualitative objective.
-
Strategy Discovery: Agents can discover creative multi-parameter strategies. For example, in one replicate of the phone stand study, an agent
simultaneously raise f from 0.05 to 0.15, drop rmin from 1.5 to 1.0... and halve E,
achieving a high score despite initial failures in amonolithic shell
design. -
Recovery: The agents show the capacity for recovery when their decisions diverge from the intended objective, leveraging history to
reason backward and adapt their decisions.
However, the study also identified several critical failure modes:
-
Overshooting: Pushing parameters too far, leading to a collapse in design quality.
-
Selective Memory: The agent may ignore or override constraints when it believes doing so will help achieve its goal.
-
Misplaced Tools: Using tools or strategies that are not available to the agent.
-
Incorrect Parameter Reasoning: Making mechanically incorrect assumptions, such as reasoning that
increasing the volume fraction should enable more branching
when reducing it would be more effective.
Improvements for AI systems
As a diligent and fastidious AI researcher, I have analyzed the TO-Agents framework. The system demonstrates significant potential, particularly in its ability to manage long-horizon, qualitative design goals through an iterative feedback loop. However, as demonstrated by the failure cases (Figures 14–16) and the limitations discussed in Section 3.4, several critical points of failure exist where stochastic behavior and lack of structured reasoning could lead to costly engineering errors.
To elevate this framework from a promising proof-of-concept to a reliable, autonomous engineering tool, I propose the following specific improvements:
The Problem: The current system suffers from selective memory
and an inability to reliably retrieve backward-pointing facts
(i.e., why a previous success is lost). The agent lacks a formal, structured memory beyond the raw chat history.
The Improvement: Integrate a Retrieval-Augmented Generation (RAG) layer into the Vision Agent's pipeline. This system will automatically parse and store key design events (parameter changes, resulting scores, and corresponding visual features) into a structured knowledge graph. Before generating a revision, the agent must query this graph to identify:
-
Optimal Sub-Goals: Which specific parameter combinations led to peak performance in previous successful revisions (e.g.,
High f combined with high p was highly effective for branching
). -
Failure Signatures: Which parameter changes consistently led to undesirable outcomes (e.g.,
Increasing r min by x, y, z resulted in a loss of complexity
).
What the Improved System Can Do: The Vision Agent will no longer rely solely on visual intuition but will make its decisions based on an evidence-based historical record, dramatically reducing stochasticity and ensuring that successful strategies are not lost during subsequent iterations.
The Problem: The agents sometimes violate explicit constraints (e.g., ignoring the r min 1.5 rule) or make physically unsound mechanical adjustments (e.g., halving Young's Modulus, which is a uniform scalar multiplier on the gradient).
The Improvement: Insert a Physics Validation Agent immediately after the Vision Agent and before the Topology Optimization Agent. This agent will be hard-coded with the fundamental laws of Finite Element Analysis (FEA) and structural mechanics. Its function is to:
-
Validate Parameter Sanity: Check if suggested parameter changes fall within physically meaningful ranges for E 0 or nu.
-
Enforce Hard Constraints: Automatically reject any proposed modification that violates the defined constraints (e.g, forcing the Vision Agent to re-run its logic if r min < 1.5).
-
Correct Mechanical Errors: If a suggested change is mechanistically incorrect (e.g., reducing E 0 without compensating for load path changes), it must flag the suggestion and provide a corrected parameter range, preventing costly
intuitive
failures.
What the Improved System Can Do: The system will become inherently reliable. It will prevent the agent from pursuing physically impossible or counter-productive paths, ensuring that success
is based on valid engineering principles, not just high scores in a flawed simulation.
The Problem: The agents can hallucinate tools or suggest functions that do not exist in the PyFANTOM API (e.g., attempting to implement a p-continuation schedule
).
The Improvement: Implement a Tool-Use Validation Layer. Before the Topology Optimization Agent accepts any input, it must pass through an automated schema checker. This layer will compare the agent’s suggested parameter modifications against the known, fixed API of PyFANTOM. If the agent suggests a function or variable not supported by the current toolset, it is rejected and flagged for retraining/prompt refinement.
What the Improved System Can Do: The system will eliminate hallucinated code and function calls, ensuring that every action taken by an agent has a deterministic effect within the defined computational environment.
The Problem: The current system is purely reactive—it only refines based on the immediate feedback of the AI Judge. It can double down
on a single strategy, even if that strategy is suboptimal, or fail to explore entirely new design spaces.
The Improvement: Introduce a Strategic Exploration Agent. This meta-agent sits above the iterative loop and is tasked with managing the search space. Instead of just reacting to the best score, it will periodically introduce exploratory moves
(e.g, radically changing f or beta) that are not driven by recent success but by a statistical assessment of unvisited regions in the parameter space (e.g., using concepts from Bayesian Optimization).
What the Improved System Can Do: The system will move beyond local optima. It will proactively seek out novel, highly complex topologies that might not be immediately visible to the current visual feedback loop, allowing for truly groundbreaking design solutions rather than just iterative refinements of a single idea.
By implementing these changes, we transition TO-Agents from a system that reactively optimizes to one that strategically engineers. The improved system will possess physical rigor, historical memory, and proactive exploration, ensuring that its design iterations are both highly efficient and mechanically sound, thus justifying the high stakes of autonomous engineering design.
Sources
- GenCAD: Image-Conditioned Computer-Aided Design Generation with Transformer-Based Contrastive Representation and Diffusion Priors
- BikeBench: A Bicycle Design Benchmark for Generative Models with Objectives and Constraints
- Large Language Models Are Human-Level Prompt Engineers
- Attention Is All You Need
- A Survey of Large Language Models
- Emergent Abilities of Large Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Evaluating Large Language Models in Scientific Discovery
- Higher-Order Knowledge Representations for Agentic Scientific Reasoning
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- ReAct: Synergizing Reasoning and Acting in Language Models
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- Exploration of LLM Multi-Agent Application Implementation Based on LangGraph+CrewAI
- Agent AI with LangGraph: A Modular Framework for Enhancing Machine Translation Using Large Language Models
- GraphAgents: Knowledge Graph-Guided Agentic AI for Cross-Domain Materials Design
- Robin: A multi-agent system for automating scientific discovery
- MechAgents: Large language model multi-agent collaborations can solve mechanics problems, generate new data, and integrate knowledge
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Self-Preference Bias in LLM-as-a-Judge
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection