Programming Manufacturing Robots with Imperfect AI: LLMs as Tuning Experts for FDM Print Configuration Selection
summary
The gist
We investigate how manufacturing robots can utilize imperfect AI, specifically Large Language Models (LLMs), to acquire process expertise by treating them as tuning experts within an evidence-driven
In short
The research investigates using Large Language Models (LLMs) as constrained decision modules within a Bayesian optimization loop to improve manufacturing robot FDM print configuration selection. By treating the LLM as a tuning expert that suggests corrective actions based on structured diagnostic feedback, the method found it significantly outperformed default settings and end-to-end AI recommendations, achieving near-perfect configurations on 78% of objects.
Key concepts
- Evidence-Driven Optimization Loop
- This is a closed system where a robot tests print settings, evaluates the results against quality metrics (like printability and defects), and uses that data to intelligently select better settings. The loop repeats until the best possible configuration is found, guided by structured feedback.
- LLM as Constrained Decision Module
- Instead of asking an LLM to choose everything, this approach restricts the LLM's role. It receives structured diagnostics from the robot and prompts it to propose only one specific corrective action at a time. This limits the LLM's scope, making its advice highly focused and actionable for tuning.
- Feasibility Vetoes
- These are checks within the optimization process that immediately disqualify a print configuration if it is physically impossible or guaranteed to fail based on initial diagnostics. If a configuration triggers a veto, it is set to an infinite penalty, ensuring the optimization loop never wastes time exploring invalid settings.
- Scalarization of Objectives
- This technique simplifies complex goals—like balancing quality, print time, and cost—into a single numerical score. While useful for optimization algorithms, this method collapses different trade-offs into one number. The paper notes this can obscure the full range of optimal solutions.
Terminology used across episodes
This episode discusses
- Programming Manufacturing Robots with Imperfect AI: LLMs as Tuning Experts for FDM Print Configuration Selection · Paper Radio
- FDM-Bench: A Comprehensive Benchmark for Evaluating Large Language Models in Additive Manufacturing Tasks
- Towards Foundational AI Models for Additive Manufacturing: Language Models for G-Code Debugging, Manipulation, and Comprehension
- LLM-3D Print: Large Language Models To Monitor and Control 3D Printing
- Large Language Models to Enhance Bayesian Optimization
- LLINBO: Trustworthy LLM-in-the-Loop Bayesian Optimization
- LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?
- Thingi10K: A Dataset of 10,000 3D-Printing Models
The paper
Programming Manufacturing Robots with Imperfect AI: LLMs as Tuning Experts for FDM Print Configuration Selection · Read on arXiv
Robotics Institute at Carnegie Mellon University
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Programming Manufacturing Robots with Imperfect AI".
Dev: We investigate how manufacturing robots can utilize imperfect AI, specifically Large Language Models (LLMs), to acquire process expertise by treating them as tuning experts within an evidence-driven optimization loop.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: We’ve just discussed how Ekta U. Samani and Christopher G. Atkeson tackled the issue of using Large Language Models as tuning experts for FDM print configuration selection. The core idea is to use this imperfect AI within a closed-loop optimization process to find better print settings based on evidence from actual prints.
Dev: They focus on treating the LLM as a specialized decision module, not the final authority, embedding it into a Bayesian optimization loop where it receives structured diagnostics and suggests corrective actions. This shifts the role of the AI from being an oracle to a constrained expert advisor.
Taro: I think it's interesting that they framed this so modularly; it allows for swapping out different parts of the system, which is important when we have different types of manufacturing processes we need to apply this framework to.
Rosa: Exactly, and that modularity means the core interface between the evaluator, the LLM guidance generator, and the compiler can remain stable even if we change how we define what constitutes a good print or how diagnostics are gathered for a new process.
Dev: That structure is important because it separates the learning mechanism—the optimization loop—from the reasoning engine—the LLM's suggestion generation, which helps us control complexity.
Taro: If we think about real-world applications, this means we can build systems that adapt their process expertise based on what they’ve learned from historical data without needing a complete re-training for every new setup.
Rosa: That’s the practical implication; we are moving toward acquiring process knowledge incrementally through interaction rather than relying solely on massive pre-training datasets that might not cover all edge cases.
Dev: And this iterative acquisition approach addresses the issue of slow convergence in optimization problems by providing targeted guidance at each step, which is something we need when loop rates matter.
Taro: I wonder how this modular separation helps when dealing with complex, time-dependent constraints that might pop up as the robot moves through the space while it’s printing.
Rosa: That complexity is where we see its strength; by keeping the evaluation and guidance steps distinct, we can handle dynamic feedback more explicitly than if everything were fused into one monolithic AI model.
Dev: So, in short, they are proposing a system where imperfect AI contributes specialized knowledge within a structured loop to improve physical outcomes systematically.
The paper's summary: Rosa: To summarize what the paper is doing with "Programming Manufacturing Robots with Imperfect AI: LLMs as Tuning Experts for FDM Print Configuration Selection," they are using fused deposition modeling as their case study to show how robots can acquire process expertise by using imperfect AI.
Dev: They use FDM three dee printing because it's a process where the print configuration has a strong effect on the final output quality, making it a good test case for this kind of evidence-driven learning loop <ref:2603.22118#pg0>.
Taro: The summary emphasizes that novice users often rely on defaults or generic AI recommendations, which aren't reliable for meeting specific objectives, setting up the problem they are trying to solve.
Rosa: Right, and their approach is to embed an LLM inside a Bayesian optimization loop where it acts as the tuning expert receiving structured diagnostics and proposing natural language adjustments.
Dev: The core mechanism involves an approximate evaluator that scores configurations and returns those structured diagnostics—like feasibility vetoes and risk penalties—which then feed into the LLM guidance generator.
Taro: This means the LLM isn't just guessing what to do next; it’s being guided by concrete, structured data about where the print is succeeding or failing.
Rosa: Precisely, and the LLM then proposes corrective actions based on those diagnostics, which are then compiled into machine-actionable guidance for optimization.
Dev: So instead of just asking the AI "what should I do?" it’s being told, "the surface roughness is too high here," and the LLM figures out what parameter change to suggest.
Taro: That moves the AI from vague advice to specific, actionable instructions that fit directly into the control structure of a robot.
Rosa: It really is about turning imperfect reasoning into a more controlled, evidence-based refinement process for manufacturing tasks. This paper lays out the architecture clearly for how this works in practice.
Dev: The architecture seems designed to handle the uncertainty inherent in physical systems by explicitly modeling the potential failures through those vetoes and penalties.
The paper's improvements: Rosa: Now let's talk about what they suggest as improvements; they focus heavily on how this system can be made more effective, especially concerning guidance quality. They demonstrate that in-context examples really help improve the LLM guidance quality, which is a key finding.
Dev: That's significant because it shows that simply prompting the LLM with a few examples of good corrective actions helps it propose much better adjustments than just giving it a blank slate to start with.
Taro: So we can essentially give the AI context on what kind of corrections are actually useful, which helps narrow down the search space for the optimization loop dramatically.
Rosa: Absolutely; increasing that context improves performance on sixty-two percent of objects, meaning we get better guidance much more often when we provide examples to the LLM.
Dev: And they also found that increasing the action budget helps iteration speed; allowing two actions per iteration boosts the win-rate to zero point nine two zero against a no-guidance variant, which is pretty close to what you might see from handcrafted guidance <ref:2603.22118#pg1>.
Taro: That tells me that even though we’re using an LLM, we still need some control over the search process—we can’t just let it wander aimlessly; we have to manage its exploration.
Rosa: And they also showed that tailoring the guidance quality is important, so providing examples beats not providing any context on sixty-two percent of objects tested.
Dev: So the improvement isn't just about having a better AI, but about designing the interface between the AI and our optimization loop to maximize its utility within those constraints.
Taro: I think this emphasizes that we need to focus our efforts on providing high-quality input diagnostics so that whatever guidance mechanism we use can be more effective.
Rosa: That leads us to think about how we can automate the creation of these diagnostic inputs, which is where the real engineering challenge lies for applying this research widely.
Dev: It points toward building better sensors and faster surrogates that give us richer diagnostic data, which feeds directly into making this entire loop more robust.
Conclusion: Rosa: So to wrap up our discussion on "Programming Manufacturing Robots with Imperfect AI: LLMs as Tuning Experts for FDM Print Configuration Selection," the main conclusion is that LLMs are much better used as constrained decision modules inside evidence-driven optimization loops than as end-to-end oracles.
Dev: That means they are excellent at finding the best configuration most often, achieving zero percent likely-to-fail cases on seventy-eight percent of objects, which outperforms generic AI recommendations significantly <ref:2603.22118#pg0,0% likely-to-fail cases>.
Taro: The implication for autonomy is that we can achieve high reliability in manufacturing tasks by leveraging this method to iteratively refine settings based on real print feedback rather than just guessing the right starting point.
Rosa: It confirms that we can use imperfect AI to build robust and interpretable control pipelines where the AI’s decisions are grounded in process diagnostics, which is a huge step for building trustworthy robotic systems.
Dev: Overall, it shows that when you combine structured evaluation with a Bayesian optimization loop with LLM guidance, you get very high performance in configuration selection without sacrificing safety too much during the search.
Taro: I just want to add that this moves us toward systems where the AI isn't just executing a command but is actively participating in the refinement of that command based on real-time process data.
Rosa: That’s a solid summary of how this paper uses LLMs as tuning experts in FDM print configuration selection, and it definitely opens up some exciting avenues for future work, like moving toward multi-objective optimization later on.
Dev: We're excited to see what the next iteration looks like because we need to keep pushing the loop rate and latency down if we want this to move from lab success to real-time manufacturing control.
More episodes
- 2610.12154-Stochastic Distribution Network Reconfiguration under Load Uncertainty
- 2607.00148-3D Point World Models: Point Completion Enables More Accurate Dynamics Learning
- 2607.02403-ACID: Action Consistency via Inverse Dynamics for Planning with World Models
- 2510.26623-A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation
- 2406.13267-The Kinetics Observer: A Tightly Coupled Estimator for Legged Robots
- 2511.02147-Census-Based Population Autonomy For Distributed Robotic Teaming
- 2603.08260-Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning
- 2602.14032-RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
- 2602.15397-ActionCodec: What Makes for Good Action Tokenizers
- 2607.01819-Koopman operator theory: fundamentals, control, and applications