Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization
Khanh Quan Pham, Majid Kundroo, Geunwoo Ban, Seongho Bae, Taehong Kim
Chungbuk National University · NPS Company Limited
cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 21 pages, 8 figures. Accepted manuscript published in Journal of Intelligent Manufacturing (2025). DOI: 10.1007/s10845-025-02619-z
Journal ref: Journal of Intelligent Manufacturing 37 (2025) 1813-1828
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper presents the Reinforcement Learning for Laser Cutting (RL2C) algorithm, a Q-learning-based method with an epsilon-greedy policy designed to optimize laser cutting parameters for optical
Terminology
Summary
This paper presents the Reinforcement Learning for Laser Cutting (RL2C) algorithm, a Q-learning-based method with an epsilon-greedy policy designed to optimize laser cutting parameters for optical films. The algorithm dynamically optimizes cutting parameters such as focal length and laser power beam, significantly reducing taper size and film wastage. RL2C incorporates a dynamic environment space adaptability mechanism to adapt to new states encountered during the learning process over multiple batches of experiments. Experimental results demonstrate that RL2C requires fewer steps and less time to find optimal cutting parameters compared to various RL-based optimization methods. Specifically, RL2C reduces the number of optimization steps by up to 12.5% and processing time by up to 81.8% compared to existing methods. This study demonstrates the potential of RL in industrial laser-cutting processes by improving cut quality, reducing time and film wastage, and minimizing manual interventions.
The paper addresses the challenge of achieving high accuracy in laser-based cutting of optical films, which requires careful tuning of parameters such as focal length and laser power beam, adjusted according to the specific properties of each film type. Traditional trial-and-error methods are slow and inaccurate. The proposed RL2C algorithm uses Q-learning with an epsilon-greedy policy to dynamically optimize cutting parameters. The algorithm operates across three critical stages: focus, power, and taper optimization. RL2C enables the agent to update its knowledge over and over as it meets more materials and cutting conditions, enhancing the algorithm's efficiency and effectiveness, overcoming the shortcomings of static optimization methods and enabling real-time decision-making in industrial settings.
The key contributions of the paper are: (1) the RL2C Algorithm, a novel Q-learning-based algorithm designed to optimize laser cutting parameters for optical films, integrating a dynamic environment space adaptability mechanism; (2) Industrial Application, highlighting the practical application of the RL2C algorithm in an industrial setting, specifically for laser cutting machines; and (3) Comparative Analysis, presenting a thorough comparison between RL2C and baseline optimization techniques, such as RL-based Bayesian Optimization, RL-based PSO, and Random Search.
The paper describes the material film structure and input dataset configuration. There are three types of material films: Material A, Material A1, and Material B, with each material film having 70 lots for data collection. Each lot represents a unique set of experimental conditions. The data are categorized into three subfolders: focus, power, and taper, corresponding to the three optimization stages. Material A and A1 consist of six layers with a total thickness of approximately 290 µm, while Material B has a simpler structure of three layers with a total thickness of approximately 230 µm.
The optimization process is formulated as a sequential process with three stages. Stage 1, Focus Optimization, aims to identify the focus value that minimizes the line width. Stage 2, Power Optimization, determines the power value that ensures the laser cut meets a specific brightness threshold. Stage 3, Taper Optimization, minimizes the taper size using the optimal focus and power values from the previous stages. The overall objective is formulated as a multi-stage process where the combined objective is to minimize line width, brightness difference, and taper size, subject to constraints defined for each stage.
The RL2C algorithm uses Q-learning to train an agent to identify the best setting for high-quality cutting for various material films and lots. The agent takes as input hyperparameters including learning rate α, discount factor γ, exploration-exploitation rate ε, and the rate of decay of the exploration parameter. The output is the optimal state, optimal policy, and optimal Q-table. The algorithm begins with the initialization of the environment, which determines the states and actions of the laser-cutting process. The Q-table is populated with zero values or small random positive values for each state-action pair. The agent uses the ε-greedy policy to manage the trade-off between exploration and exploitation. After each action, the Q-values are updated using the Bellman equation. The RL2C algorithm is developed to work for both static and dynamic environments. The Update Dynamic Q table mechanism allows the agent to modify the Q-table in response to new states during training across multiple lots. RL2C maintains a separate Q-table for each stage: one for focus optimization, one for power optimization, and one for taper optimization.
The reward functions for each stage are defined as follows. For Stage 1 (Focus Optimization), the reward is-200 for error states, 100/line width - (step × 0.1) for optimal states, and-1/line width - (step × 0.1) for normal states. For Stage 2 (Power Optimization), the reward is-5000 for error states, (brightness × 50) - (step × 2) for optimal states, and-brightness - (step × 2) for normal states. For Stage 3 (Taper Optimization), the reward is-200 for error states, 100/taper - (step × 0.1) for optimal states, and-1/taper - (step × 0.1) for normal states.
The experimental setup uses a total of 210 lots, with 90 for training and 120 for testing. The agent is trained on each material sequentially, starting with Material A, followed by A1, and then B, using the ε-greedy policy across these lots. For each lot, the agent runs 100 episodes, with a maximum of 100 steps per episode. During training, the ε value starts at 1.0 and gradually decreases over the training lots, reaching 0.1. Once training is complete, the agent is evaluated on the remaining 120 testing lots with a fixed ε of 0.1.
The results show that RL2C consistently outperforms baseline methods in terms of step efficiency across all three optimization stages. In Stage 1, RL2C requires an average of 5.1 steps during training compared to 5.4 for RL-Bayesian and 25.3 for Random Search. In Stage 2 and Stage 3, RL2C maintains superior step efficiency with averages of 1.6 and 3.8 steps, respectively. For time efficiency, RL2C shows significant improvement over baseline methods. In the training phase, RL2C completes Stage 1 with an average time of 0.03 seconds, compared to 4.77 seconds for RL-Bayesian and 0.11 seconds for RL-PSO. In the testing phase, RL2C maintains superior time efficiency, completing all stages in approximately 0.02 seconds on average.
The paper concludes that RL2C emerges as a highly efficient and scalable solution for laser cutting parameter optimization, demonstrating optimal performance in both static and dynamic environments. A key strength of RL2C lies in its ability to handle transitions between materials efficiently, adapting rapidly with minimal computational cost. Future work will explore extending RL2C to additional manufacturing processes and integrating it with other advanced learning frameworks to further enhance its applicability and scalability.
Improvements for AI systems
Improvements to AI Systems:
-
Dynamic Environment Adaptability via Incremental Q-Table Expansion: Implement a mechanism where the AI system can automatically expand its state-action space when encountering novel inputs (e.g., new material properties or machine conditions) without requiring full retraining. This enables continuous learning across heterogeneous tasks, reducing catastrophic forgetting and improving generalization to unseen industrial scenarios.
-
Multi-Stage Sequential Reward Decomposition: Adopt the three-stage optimization structure (focus → power → taper) to break complex, multi-objective problems into sub-tasks with tailored reward functions. This improves sample efficiency and convergence speed, as the AI can learn hierarchical policies where each stage’s output constrains the next, reducing search space dimensionality.
-
Adaptive Exploration-Exploitation Scheduling with Decay: Use a time-varying epsilon-greedy policy that starts with high exploration (ε=1.0) and decays to a low exploitation rate (ε=0.1) across training batches. This allows the AI to rapidly discover viable parameter regions early and refine them later, cutting optimization steps by up to 12.5% and processing time by 81.8% compared to static exploration methods.
-
Reward Shaping for Error Avoidance and Efficiency: Integrate asymmetric penalties for error states (e.g., -200 to-5000) versus small step penalties (-0.1 to-2 per step) to prioritize safety and feasibility while discouraging excessive iterations. This teaches the AI to avoid invalid configurations quickly, reducing wasted computation and manual intervention.
-
Cross-Lot Transfer Learning via Shared Q-Tables per Stage: Maintain separate Q-tables for each optimization stage but reuse them across different material lots, enabling rapid adaptation to new batches with minimal retraining. This mimics the paper’s training on 90 lots and testing on 120, showing the AI can transfer knowledge between similar tasks, improving scalability to large-scale production.
-
Real-Time Decision-Making with Low Latency: Design the AI to output optimal parameters in under 0.03 seconds per stage (as demonstrated), making it suitable for real-time control in manufacturing. This requires lightweight Q-learning updates (Bellman equation) rather than heavy neural network inference, enabling deployment on edge devices with limited computational resources.
-
Constraint-Aware Multi-Objective Optimization: Formulate the AI’s objective as a sequential constrained optimization (minimize line width, brightness difference, taper size) rather than a single scalar reward. This allows the system to explicitly handle trade-offs between quality metrics, ensuring the final solution satisfies all industrial thresholds (e.g., brightness threshold) simultaneously.
What the Improved AI System Can Do:
-
Self-Adapt to New Materials and Conditions: Automatically adjust its parameter search space when encountering novel film types or machine drift, without human recalibration, by expanding its Q-table dynamically.
-
Optimize Complex Manufacturing Processes in Real-Time: Find near-optimal cutting parameters (focal length, power) in milliseconds, enabling on-the-fly adjustments during production to maintain consistent quality.
-
Reduce Waste and Rework: Minimize taper size and film wastage by converging to optimal settings in fewer steps (average 5.1 steps for focus, 1.6 for power, 3.8 for taper), directly lowering material costs.
-
Handle Multi-Stage Sequential Tasks Efficiently: Break down complex problems into sub-problems with clear reward structures, allowing the AI to solve each stage independently while respecting dependencies, improving overall success rates.
-
Transfer Knowledge Across Batches and Products: Leverage prior learning from one material lot to accelerate optimization on new lots, reducing setup time and enabling rapid prototyping or small-batch production.
-
Operate in Resource-Constrained Environments: Run on embedded systems or PLCs with minimal memory (Q-tables) and compute (Bellman updates), making it deployable in existing industrial machinery without major hardware upgrades.
-
Provide Safe Exploration: Avoid dangerous or error-prone parameter settings by heavily penalizing them, ensuring the AI never proposes settings that could damage equipment or produce defective cuts.
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection