Programming Manufacturing Robots with Imperfect AI: LLMs as Tuning Experts for FDM Print Configuration Selection

arXiv:2603.22118 · cs.RO · Submitted 2026-03-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Programming Manufacturing Robots with Imperfect AI".

Dev: We investigate how manufacturing robots can utilize imperfect AI, specifically Large Language Models (LLMs), to acquire process expertise by treating them as tuning experts within an evidence-driven optimization loop.

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: We’ve just discussed how Ekta U. Samani and Christopher G. Atkeson tackled the issue of using Large Language Models as tuning experts for FDM print configuration selection. The core idea is to use this imperfect AI within a closed-loop optimization process to find better print settings based on evidence from actual prints.

Dev: They focus on treating the LLM as a specialized decision module, not the final authority, embedding it into a Bayesian optimization loop where it receives structured diagnostics and suggests corrective actions. This shifts the role of the AI from being an oracle to a constrained expert advisor.

Taro: I think it's interesting that they framed this so modularly; it allows for swapping out different parts of the system, which is important when we have different types of manufacturing processes we need to apply this framework to.

Rosa: Exactly, and that modularity means the core interface between the evaluator, the LLM guidance generator, and the compiler can remain stable even if we change how we define what constitutes a good print or how diagnostics are gathered for a new process.

Dev: That structure is important because it separates the learning mechanism—the optimization loop—from the reasoning engine—the LLM's suggestion generation, which helps us control complexity.

Taro: If we think about real-world applications, this means we can build systems that adapt their process expertise based on what they’ve learned from historical data without needing a complete re-training for every new setup.

Rosa: That’s the practical implication; we are moving toward acquiring process knowledge incrementally through interaction rather than relying solely on massive pre-training datasets that might not cover all edge cases.

Dev: And this iterative acquisition approach addresses the issue of slow convergence in optimization problems by providing targeted guidance at each step, which is something we need when loop rates matter.

Taro: I wonder how this modular separation helps when dealing with complex, time-dependent constraints that might pop up as the robot moves through the space while it’s printing.

Rosa: That complexity is where we see its strength; by keeping the evaluation and guidance steps distinct, we can handle dynamic feedback more explicitly than if everything were fused into one monolithic AI model.

Dev: So, in short, they are proposing a system where imperfect AI contributes specialized knowledge within a structured loop to improve physical outcomes systematically.

The paper's summary: Rosa: To summarize what the paper is doing with "Programming Manufacturing Robots with Imperfect AI: LLMs as Tuning Experts for FDM Print Configuration Selection," they are using fused deposition modeling as their case study to show how robots can acquire process expertise by using imperfect AI.

Dev: They use FDM three dee printing because it's a process where the print configuration has a strong effect on the final output quality, making it a good test case for this kind of evidence-driven learning loop <ref:2603.22118#pg0>.

Taro: The summary emphasizes that novice users often rely on defaults or generic AI recommendations, which aren't reliable for meeting specific objectives, setting up the problem they are trying to solve.

Rosa: Right, and their approach is to embed an LLM inside a Bayesian optimization loop where it acts as the tuning expert receiving structured diagnostics and proposing natural language adjustments.

Dev: The core mechanism involves an approximate evaluator that scores configurations and returns those structured diagnostics—like feasibility vetoes and risk penalties—which then feed into the LLM guidance generator.

Taro: This means the LLM isn't just guessing what to do next; it’s being guided by concrete, structured data about where the print is succeeding or failing.

Rosa: Precisely, and the LLM then proposes corrective actions based on those diagnostics, which are then compiled into machine-actionable guidance for optimization.

Dev: So instead of just asking the AI "what should I do?" it’s being told, "the surface roughness is too high here," and the LLM figures out what parameter change to suggest.

Taro: That moves the AI from vague advice to specific, actionable instructions that fit directly into the control structure of a robot.

Rosa: It really is about turning imperfect reasoning into a more controlled, evidence-based refinement process for manufacturing tasks. This paper lays out the architecture clearly for how this works in practice.

Dev: The architecture seems designed to handle the uncertainty inherent in physical systems by explicitly modeling the potential failures through those vetoes and penalties.

The paper's improvements: Rosa: Now let's talk about what they suggest as improvements; they focus heavily on how this system can be made more effective, especially concerning guidance quality. They demonstrate that in-context examples really help improve the LLM guidance quality, which is a key finding.

Dev: That's significant because it shows that simply prompting the LLM with a few examples of good corrective actions helps it propose much better adjustments than just giving it a blank slate to start with.

Taro: So we can essentially give the AI context on what kind of corrections are actually useful, which helps narrow down the search space for the optimization loop dramatically.

Rosa: Absolutely; increasing that context improves performance on sixty-two percent of objects, meaning we get better guidance much more often when we provide examples to the LLM.

Dev: And they also found that increasing the action budget helps iteration speed; allowing two actions per iteration boosts the win-rate to zero point nine two zero against a no-guidance variant, which is pretty close to what you might see from handcrafted guidance <ref:2603.22118#pg1>.

Taro: That tells me that even though we’re using an LLM, we still need some control over the search process—we can’t just let it wander aimlessly; we have to manage its exploration.

Rosa: And they also showed that tailoring the guidance quality is important, so providing examples beats not providing any context on sixty-two percent of objects tested.

Dev: So the improvement isn't just about having a better AI, but about designing the interface between the AI and our optimization loop to maximize its utility within those constraints.

Taro: I think this emphasizes that we need to focus our efforts on providing high-quality input diagnostics so that whatever guidance mechanism we use can be more effective.

Rosa: That leads us to think about how we can automate the creation of these diagnostic inputs, which is where the real engineering challenge lies for applying this research widely.

Dev: It points toward building better sensors and faster surrogates that give us richer diagnostic data, which feeds directly into making this entire loop more robust.

Conclusion: Rosa: So to wrap up our discussion on "Programming Manufacturing Robots with Imperfect AI: LLMs as Tuning Experts for FDM Print Configuration Selection," the main conclusion is that LLMs are much better used as constrained decision modules inside evidence-driven optimization loops than as end-to-end oracles.

Dev: That means they are excellent at finding the best configuration most often, achieving zero percent likely-to-fail cases on seventy-eight percent of objects, which outperforms generic AI recommendations significantly <ref:2603.22118#pg0,0% likely-to-fail cases>.

Taro: The implication for autonomy is that we can achieve high reliability in manufacturing tasks by leveraging this method to iteratively refine settings based on real print feedback rather than just guessing the right starting point.

Rosa: It confirms that we can use imperfect AI to build robust and interpretable control pipelines where the AI’s decisions are grounded in process diagnostics, which is a huge step for building trustworthy robotic systems.

Dev: Overall, it shows that when you combine structured evaluation with a Bayesian optimization loop with LLM guidance, you get very high performance in configuration selection without sacrificing safety too much during the search.

Taro: I just want to add that this moves us toward systems where the AI isn't just executing a command but is actively participating in the refinement of that command based on real-time process data.

Rosa: That’s a solid summary of how this paper uses LLMs as tuning experts in FDM print configuration selection, and it definitely opens up some exciting avenues for future work, like moving toward multi-objective optimization later on.

Dev: We're excited to see what the next iteration looks like because we need to keep pushing the loop rate and latency down if we want this to move from lab success to real-time manufacturing control.

Robotics Institute at Carnegie Mellon University

cs.RO

Submitted: 2026-03-23

Updated: 2026-10-04

Comments: Accepted at IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026

Project page: https://sheffieldml.github.io/GPyOpt

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: We investigate how manufacturing robots can utilize imperfect AI, specifically Large Language Models (LLMs), to acquire process expertise by treating them as tuning experts within an evidence-driven

Key concepts

Evidence-Driven Optimization Loop
This is a closed system where a robot tests print settings, evaluates the results against quality metrics (like printability and defects), and uses that data to intelligently select better settings. The loop repeats until the best possible configuration is found, guided by structured feedback.
LLM as Constrained Decision Module
Instead of asking an LLM to choose everything, this approach restricts the LLM's role. It receives structured diagnostics from the robot and prompts it to propose only one specific corrective action at a time. This limits the LLM's scope, making its advice highly focused and actionable for tuning.
Feasibility Vetoes
These are checks within the optimization process that immediately disqualify a print configuration if it is physically impossible or guaranteed to fail based on initial diagnostics. If a configuration triggers a veto, it is set to an infinite penalty, ensuring the optimization loop never wastes time exploring invalid settings.
Scalarization of Objectives
This technique simplifies complex goals—like balancing quality, print time, and cost—into a single numerical score. While useful for optimization algorithms, this method collapses different trade-offs into one number. The paper notes this can obscure the full range of optimal solutions.

Terminology

Summary

We investigate how manufacturing robots can utilize imperfect AI, specifically Large Language Models (LLMs), to acquire process expertise by treating them as tuning experts within an evidence-driven optimization loop. The core finding demonstrates that embedding an LLM as a constrained decision module significantly improves FDM print configuration selection compared to relying on default settings or end-to-end AI recommendations, achieving the best configurations on 78% of objects with zero likely-to-fail cases.

The gist

LLMs can contribute more effectively as constrained decision modules inside evidence-driven optimization loops than as end-to-end oracles for print configuration selection.

How it works

The paper presents a modular closed-loop approach that treats an LLM as a source of tuning expertise embedded within a Bayesian optimization loop. This process begins by casting print configuration selection as a diagnosis-driven optimization problem where each candidate is evaluated to produce a scalar objective and structured diagnostics on printability and quality issues.

  1. The Evaluator computes two primary outputs: (i) a scalar score for optimization, defined by the objective function:

Obj(x) = wt t(x)/tr + t(x)/tr + wc c(x)/cr + c(x)/cr + wqQ(x), where Q(x) is an approximate quality penalty computed from bounded diagnostic penalties pi(x). (1)

  1. The Evaluator also returns feasibility vetoes, which declare a configuration infeasible if any veto triggers, setting Q(x) = +∞. If feasible, it summarizes pi(x) within groups and converts each group score into a clipped, normalized excess ek(x) in [0, 1]. The quality penalty is then aggregated as: Q(x) = max k ek(x)+λ / (3 − max k ek(x)). (2)

  2. The LLM Guidance Generator takes these structured diagnostics and proposes corrective actions based on a list of admissible changes to the print parameters, which are restricted to a small set of high-leverage parameters. The LLM is prompted to perform a constrained decision: it identifies exactly one primary issue to address next and proposes corrective actions.

  3. The LLM Guidance Compiler converts these natural-language adjustments into two artifacts for the optimization loop: (i) a differentiable soft-violation score V(x) in [0, 1] measuring how strongly configuration x contradicts the proposed changes, and (ii) an implicated-parameter set I, which enables hard constraints by freezing I c while optimizing over I.

Evaluation and Comparison

The method is tested on 100 single-component parts from the Thingi10k dataset. The tunable configuration s(x) consists of a discrete build orientation and 13 print parameters. Baselines include default parameters, heuristic reorientation, and configurations suggested by chat-based AI models (e.g., ChatGPT 5.2 Thinking).

The results show that the LLM-guided optimization finds the best configuration most often with 0% likely-to-fail cases. Across the 100 objects, our method achieves the best objective value (lowest) among all listed approaches on 78% of objects and is within 1% / 5% of the lowest on 82% / 90%. In contrast, chat-based AI model recommendations are rarely best and exhibit 15% likely-to-fail cases.

Guidance Refinement

Ablation studies confirm that guidance quality is critical. The paper demonstrates that in-context examples improve guidance quality, with prompting with examples beating prompting without examples on 62% of objects. Furthermore, increasing the per-iteration action budget improves performance; allowing two actions per iteration yields a win-rate of 0.920 against the no-guidance variant, nearing handcrafted guidance.

Design Philosophy and Limitations

The method is designed as a modular loop that cleanly separates evaluation, guidance, and optimization. It is not printer- or material-specific but requires an evaluator calibrated to the target process that provides reliable feasibility vetoes and bounded, issue-level diagnostics. A key limitation acknowledged is that the approach collapse[s] heterogeneous objectives and defect mechanisms into a scalar, which obscures the Pareto structure over quality, time, and cost. Furthermore, tuning only a subset of print parameters means some valid interventions may be out of scope.

Conclusion

The study concludes that LLMs can contribute more effectively as constrained decision modules inside evidence-driven optimization loops rather than as end-to-end oracles for FDM print configuration selection. Future work plans include extending the guidance generator to use multimodal evidence and moving from scalarization to multi-objective optimization that returns a Pareto set.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed this paper, Programming Manufacturing Robots with Imperfect AI: LLMs as Tuning Experts for FDM Print Configuration Selection. The core innovation lies in using an LLM not as an end-to-end decision-maker, but as a constrained, evidence-driven decision module within a closed-loop Bayesian optimization framework.

Here are the specific improvements that can be made to AI systems based on this research, and what the resulting improved system can achieve:


The primary improvement is shifting LLMs from being end-to-end oracles (which are unreliable) to functioning as high-quality, constrained tuning experts within a robust, evidence-driven optimization pipeline.

Here are five specific, actionable improvements:

  1. ​​​‌‌The LLM Guidance Compiler and Soft/Hard Constraint Mapping:

  2. ​​​‌‌The resulting system can generate machine-actionable guidance for Bayesian optimization by translating natural language suggestions into precise mathematical constraints (soft violation scores, differentiable soft-violation scores, and hard constraint sets). This allows the LLM to guide the search space efficiently without requiring it to solve the entire complex optimization problem itself.

  3. ​​​‌‌The Approximate Evaluator with Structured Diagnostics:

  4. ​​​‌‌The system can rapidly score candidate print configurations by providing structured diagnostics (scalar objective, feasibility vetoes, and issue-level penalties) derived from toolpath evidence rather than relying on slow or expensive high-fidelity simulations. This allows for the evaluation of thousands of configurations in a fraction of the time required by traditional methods.

  5. ​​​‌‌The Modular Closed-Loop Architecture:

  6. ​​​‌‌The system can be decoupled and adapted to any manufacturing process (e.g., CNC machining, chemical synthesis) by swapping out the specific Evaluator module (to match the new process's diagnostic needs) while keeping the LLM/Compiler interface intact. This makes it a universal framework for using imperfect AI in robot programming across diverse domains.

  7. ​​​‌‌The Hybrid Acquisition Strategy (Soft Guidance and Hard Constraints):

  8. ​​​‌‌The system can perform iterative, highly efficient optimization by simultaneously biasing the search toward expert-suggested regions (soft guidance) while strictly limiting the next step to modify only the implicated parameters (hard constraints). This prevents catastrophic failures during exploration while ensuring rapid convergence on high-quality solutions.

The overall improved AI system can achieve the following:

  1. ​​​‌‌Maximize Print Quality Under Constraints: The system can consistently select print configurations that achieve near-optimal quality metrics (e.g., 78% of objects with 0% likely-to-fail cases), significantly outperforming novice methods and generic, single-shot AI recommendations which exhibit a high likelihood of failure (15%).

2.​​​‌‌Reduce Iterative Search Time: By using the LLM as a diagnostic filter rather than a full solver, the system can achieve superior sample efficiency—finding the best configuration in fewer iterations compared to unguided optimization methods.

3.​​​‌‌Enable Expert-Level Decision Making in Manufacturing: The system allows manufacturing robots to move beyond it printed to achieve specific, nuanced objectives (e.g., prioritizing surface finish over build time) by leveraging the LLM's ability to reason over prior run evidence, effectively embedding expert knowledge into the robot's control loop.

4.​​​‌‌Create Robust and Interpretable Control Pipelines: Because the guidance is compiled into machine-actionable guidance, the system provides a transparent audit trail of why a specific configuration was chosen at each step, which is crucial for safety and debugging in high-stakes manufacturing environments.

Sources

Related papers