Agentic AI for Scalable and Robust Optical Systems Control

arXiv:2602.20144 · eess.SY, cs.AI, cs.NI, cs.SY · Submitted 2026-02-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Agentic AI for Scalable and Robust Optical Systems Control".

Dev: We present AgentOptics, an agentic AI framework for high-fidelity, autonomous optical system control built upon the model context protocol (MCP).

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Well, Dev, I'm really interested in what this paper on "Agentic AI for Scalable and Robust Optical Systems Control" is trying to achieve with AgentOptics; it sounds like a big step toward making optical systems easier to manage.

Dev: It does sound significant, Rosa. Essentially, AgentOptics is an agentic framework that lets you control various optical devices using just natural language tasks instead of writing tons of device-specific code every time.

Taro: I wonder how this system handles situations where things go wrong in the real world, especially when we're dealing with complex interactions between different parts of the optical network.

Rosa: That's exactly what I was thinking; I want to know if this control works reliably outside of a perfectly controlled lab environment and for how long before we have to worry about it breaking down.

Dev: The paper tests its robustness by putting it through a four hundred ten-task benchmark that covers things like request understanding, role-dependent responses, and handling linguistic variations in the commands.

Taro: That's important because when the world misbehaves, we need a system that doesn't just follow instructions perfectly but can actually adapt and figure out what to do next.

Rosa: So, if I give it a command like "check the status of the ROADM link," it interprets that and then executes the right steps across all those different devices.

Dev: Exactly, and that's where the structure tool abstraction layer comes in; it lets AgentOptics map those high-level natural language requests to specific, standardized actions on things like a Lumen-tum ROADM or an APEX Technologies OSA.

Taro: I’m thinking about the implications for real-world autonomy; if this framework can handle misbehavior, could we see it managing network issues autonomously without constant human intervention?

Rosa: That’s a huge question for me—can this actually manage a DWDM link provisioning task end-to-end, or is it just good at single commands?

Dev: It demonstrates success across various levels of complexity; for instance, the experimental results show AgentOptics achieves an average task success rate of eighty-seven point seven percent to ninety-nine point zero percent when using online LLMs on that comprehensive benchmark.

Taro: Compared to the code generation baselines they tested, which hit only up to a fifty point zero percent success rate across those complex tasks, this performance difference is substantial for real-world deployment scenarios.

Rosa: It really puts the capability of interpreting nuanced requests and coordinating multistep actions into perspective for optical engineers like myself.

Dev: Furthermore, the framework allows us to evaluate both commercial online LLMs and locally hosted open-source models on a Dell server setup to see what configuration gives us the best balance of accuracy and operational cost.

Taro: The implications for deployment are huge, especially since they tested it with both high-powered commercial models and smaller, locally deployed ones without quantization for comparison.

Rosa: It sounds like the core idea is moving optical control away from tedious manual scripting toward a more intelligent system that understands the intent behind the request.

Dev: Precisely; it’s about replacing manual protocol handling with an agentic workflow where the AI selects and executes tools based on semantic similarity to the user's natural language input.

Taro: Thinking about future work, I'm curious if this agentic approach can be extended to handle dynamic, real-time system adjustments rather than just static configuration tasks.

Rosa: That would be the next frontier for me—moving from setting up a link to actively optimizing its performance based on live telemetry.

Dev: The paper also shows how AgentOptics can enable closed-loop optimization, which means it doesn't just set a parameter and leave it; it monitors the result and makes adjustments autonomously.

Taro: If we look at the broader impact, this suggests that complex optical infrastructure could be managed with a level of agility that was previously unattainable because of the inherent complexity of those devices.

Rosa: It certainly feels like it moves us closer to systems where we can manage massive optical networks with far less manual effort.

Dev: So, to wrap up on this Agentic AI for Scalable and Robust Optical Systems Control paper, we've seen that AgentOptics provides a framework for high-fidelity, autonomous control by using sixty-four standardized tools across eight representative devices to handle natural language commands.

Taro: I think the biggest implication is that it gives us a path toward systems where autonomy isn't just about executing pre-defined scripts but about true system-level orchestration and adaptation to unexpected events.

Rosa: It definitely moves the conversation past just controlling individual components toward managing entire optical systems in a more fluid, responsive way.

Dev: And from an engineering standpoint, the comparison against code generation baselines shows that for multi-step coordination, this agentic approach is significantly more reliable than generating custom scripts from scratch.

Taro: I'm excited to see how this framework evolves beyond the four hundred ten benchmark tasks and into environments where those automated responses need to be even more reactive.

Rosa: It’s an interesting piece of work, really showing how agentic AI can bridge the gap between human intent and the intricate operational reality of optical hardware.

Dev: Indeed, it lays a solid foundation for building systems that can handle the necessary complexity for modern network control tasks.

The paper's summary: Rosa: So, to recap, this paper introduces AgentOptics as an agentic framework that uses a model context protocol to let natural language tasks control optical systems across different devices through standardized tools.

Dev: That's the core idea, Rosa; it’s taking high-level commands and translating them into concrete actions across a whole suite of hardware, which is pretty neat for keeping things organized.

Taro: I'm curious about how this translates to real-world scenarios; does it actually handle the messy parts where things aren't behaving exactly as expected in a lab setting?

Rosa: That’s my main concern, Taro; I want to know if this system can operate reliably outside of a perfectly controlled environment and for what kind of duration before we have to start worrying about its stability.

Dev: The benchmark they used, with its four hundred ten tasks, is designed specifically to test that robustness against linguistic variations and multi-step coordination failures.

Taro: I'm pushing on the misbehavior aspect; when the network throws a curveball or a component fails unexpectedly, what’s the system’s actual response when it can’t just follow the script?

Rosa: The paper shows it can handle complex workflows, like provisioning and optimizing a channel simultaneously, which is much more practical than just single commands.

Dev: I've looked at the latency aspects; they test both commercial online LLMs and local open-source deployments to see how the execution loop rate holds up under different computational loads.

Taro: The success rates they report are really telling; seeing eighty-seven percent to ninety-nine percent across those complex tasks compared to the code generation baselines is a big data point for autonomy researchers.

Rosa: That performance gap between AgentOptics and code generation, especially with triple-action tasks, suggests that this structure abstraction layer is doing something fundamentally different in terms of planning.

Dev: Exactly; it’s not just generating code; it’s selecting the right tool based on semantic similarity to the user's request during the reasoning process.

Taro: If this framework can manage things like DWDM provisioning and link polarization stabilization autonomously, what does that mean for a network operator who has to react in seconds?

Rosa: It means we could see system-level orchestration and closed-loop optimization where the AI monitors performance and makes dynamic adjustments without constant human intervention.

Dev: That moves us beyond simple control into true autonomous system management, which is a major step toward proactive maintenance rather than reactive troubleshooting.

Taro: The implications for infrastructure are huge; imagine an ARoF link carrying 5G fronthaul traffic where the AI autonomously optimizes bias voltage based on real-time telemetry and detects faults using DAS data.

Rosa: It’s exciting to think about how this applies beyond just setting up a connection, into actively optimizing performance in real time, which is something I’ve been hoping to see in field robotics applications.

Dev: The paper also touches on the deployment configurations, showing that the choice between a powerful commercial LLM and a well-tuned local model has direct consequences for both success rate and operational cost.

Taro: Considering these results, where do you see this technology heading next, especially concerning handling those truly novel or unforeseen system failures?

The paper's improvements: Rosa: So, to recap, the paper lays out several specific improvements for AgentOptics that really push its capabilities beyond just basic control.

Dev: Right, they focus on making the system more robust by using structured abstraction through MCP tools to handle multi-vendor optical networks without needing task-specific code generation.

Taro: That sounds promising for autonomy; if it can abstract away device protocols, it should be better equipped to handle those complex, multi-step workflows we talked about earlier.

Rosa: They also address superior handling of linguistic variation and ambiguity in commands, which means the AI gets much better at understanding what a user *means* even when the phrasing is messy.

Dev: I’m interested in the error detection part; they suggest implementing advanced diagnostic capabilities to analyze execution logs and pinpoint exactly where an orchestration error occurred during a multi-tool sequence.

Taro: That targeted debugging capability is crucial for autonomy; if the system knows *why* it failed, it can actually recover instead of just failing again on the same mistake.

Rosa: And finally, they propose implementing closed-loop optimization, which means the system moves past setting static parameters and starts dynamically adjusting link performance based on live telemetry.

Dev: That closed-loop aspect is where we get real efficiency gains; it’s about performing dynamic, iterative optimization of things like ARoF transmitter bias voltage to maximize SNR or minimize bit error rate in real time.

Taro: If the AI can detect physical faults, like a potential fiber cut using DAS monitoring with LLM-assisted event interpretation, that opens up a whole new level of proactive system management.

Rosa: It sounds like these improvements are really about giving the AI more agency—moving it from being a reactive tool to an autonomous manager of the entire optical system.

Dev: And I’m thinking about the impact on loop rates; if this closed-loop optimization happens quickly enough, we could see performance stabilizing much faster than what's currently achievable with manual tuning.

Taro: The paper flags one limitation that they admit: the method relies heavily on a well-defined set of standardized tools and schemas to work effectively, so extending it to completely novel hardware without that abstraction might be challenging.

Rosa: That’s a fair caveat; the success hinges on those sixty-four standardized primitives, meaning future work needs to focus on expanding that toolset even further.

Dev: So, the focus for next steps is clearly twofold: improving the toolset's breadth and enhancing the system's ability to manage truly dynamic, closed-loop optimization scenarios.

Conclusion: Rosa: To wrap up, we've seen how AgentOptics uses an MCP-based framework to turn natural language into high-fidelity control for complex optical systems.

Dev: It’s a powerful demonstration of how structured abstraction lets AI manage heterogeneous hardware with a much higher success rate than traditional code generation.

Taro: The ability to handle those complex, multi-action workflows autonomously is what really interests me for future autonomy research; it suggests a pathway to more adaptive systems in challenging environments.

Rosa: I’m still wondering about the practical deployment; does this system stay reliable when you take it out of the lab and put it into a real network scenario?

Dev: That's a big question, Rosa; the paper tests both online and local LLM deployments to gauge how those performance metrics hold up under different computational constraints.

Taro: For autonomy, I think the key is whether it can handle those unexpected misbehaves we discussed earlier without just crashing or making a poor recovery choice.

Rosa: The authors also highlighted that while the success rates are high, they have to rely on that standardized toolset for its effectiveness, which means expanding those tools will be necessary for broader application.

Dev: Exactly; the paper's conclusion points toward the need to focus future work on extending that sixty-four-tool abstraction layer and improving error diagnosis within those complex execution sequences.

Taro: If we can solve the problem of robust, multi-step coordination like this, it opens up possibilities for much more intelligent network management systems overall.

Rosa: It’s definitely an exciting piece of work demonstrating how agentic AI is moving closer to managing entire optical infrastructures in a fluid way.

Dev: And from an engineering standpoint, the focus on loop rate and latency during those optimization steps shows they're thinking about real-time performance, which is critical for any operational system.

Taro: I’m looking forward to seeing how this framework integrates with other autonomy research areas we've been discussing, like resilience in power distribution networks or grid coordination.

Rosa: Well, that covers the main points of Agentic AI for Scalable and Robust Optical Systems Control; it shows a strong direction for intelligent optical control.

Duke University · NEC Laboratories America · Axiomatic AI

eess.SY, cs.AI, cs.NI, cs.SY

Submitted: 2026-02-23

Updated: 2026-09-29

DOI: 10.1109/ACCESS.2026.3736374

Code: https://github.com/functions-lab/AgentOptics

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: We present AgentOptics, an agentic AI framework for high-fidelity, autonomous optical system control built upon the model context protocol (MCP).

Key concepts

AgentOptics
An agentic AI framework that uses a model context protocol (MCP) to allow users to control various optical devices using natural language tasks. It translates high-level commands into specific actions across different hardware.
Structured Abstraction Layer
A layer within AgentOptics that maps high-level natural language requests to standardized actions on specific optical devices, such as a Lumen-tum ROADM or an APEX Technologies OSA. This allows the system to handle multi-vendor networks without needing device-specific code.
Closed-Loop Optimization
A capability where the AI monitors the results of a setting and makes autonomous adjustments to link performance based on live telemetry. This moves control beyond static configuration into dynamic, iterative optimization for efficiency.
Toolset Abstraction Layer
The framework relies on a set of sixty-four standardized tools and schemas. The effectiveness of AgentOptics depends on this abstraction, which allows the AI to select and execute the right tool based on semantic similarity to the user's natural language input.

Terminology

Summary

We present AgentOptics, an agentic AI framework for high-fidelity, autonomous optical system control built upon the model context protocol (MCP). AgentOptics interprets natural language tasks and executes protocol-compliant actions on heterogeneous optical devices through a structure tool abstraction layer. We implement 64 standardized MCP tools spanning eight representative optical devices and construct a comprehensive 410-task benchmark to evaluate the performance of AgentOptics across request understanding, role-dependent responses, multistep coordination, robustness to linguistic variation, and errorhandling capability. We evaluate two deployment configurations–integrating either commercial online large language models (LLMs) or locally hosted open-source LLMs–and compare against LLM-based code generation baselines. Experimental results demonstrate that AgentOptics achieves 87.7%–99.0% average task success rates, significantly outperforming code generation approaches with up to only 50% success rate. We further validate the broader applicability of AgentOptics through five representative case studies that extend beyond accurate device control to enable system-level orchestration and monitoring, as well as closed-loop optimization. These case studies include dense wavelength division multiplexing (DWDM) link provisioning and coordinated performance monitoring of coherent 400 GbE and analog radio-over-fiber (ARoF) channels, autonomous characterization and bias optimization of a wideband ARoF link carrying 5G fronthaul traffic, multi-span channel provisioning and signal launch power optimization, closed-loop fiber link polarization stabilization, and distributed acoustic sensing (DAS)-based fiber monitoring with LLM-assisted event interpretation and detection.

AgentOptics is an MCP-based framework that defines a formal client-server architecture connecting LLM hosts to external tools and services through well-defined schemas and execution semantics. The workflow involves the user issuing a natural language task, which is received by an MCP client embedded within the host application. The client forwards the task to the LLM, which interprets user intent using domain knowledge and selects relevant MCP server(s) based on its published server description. Subsequently, the client retrieves tool descriptions and function definitions from selected servers and supplies them to the LLM, which determines the most suitable tool by evaluating semantic similarity between user task and tool metadata. The selected tool is then executed by the MCP server, invoking device-specific APIs, monitoring task completion, and returning results to the MCP client. Finally, the MCP client relays results back to the LLM for a human-readable natural language response.

The framework implements 64 standardized MCP tools across eight representative optical devices: (i) Lumen-tum ROADM (10 tools), (ii) Lumentum 400 GbE CFP2-DCO (6 tools), (iii) OptiLab LT-12-EM ARoF TX (6 tools), (iv) APEX Technologies OSA (26 tools), (v) Calient S320 Optical Circuit Switch 4, and (vi) DiCon MEMS 32×32 Optical Switch 2. For each device, an MCP server is implemented exposing a set of tools supporting core operations, including device setup, parameter control, status monitoring, and connection reconfiguration.

The system implements AgentOptics interacting with two types of LLM serving:

  1. AgentOptics-Online: An MCP client sends tasks to online LLMs (e.g., GPT-4o mini, GPT-5, Deepseek-V3, Claude Haiku 3.5, and Claude Sonnet 4.5).

  2. AgentOptics-Local: Hosted on a local Dell PowerEdge R750 server with a 64-core Intel Xeon Gold 6548N CPU @2.6 GHz and an NVIDIA 40 GB A100 GPU, utilizing Qwen models with parameter sizes of 0.4 B, 8 B, and 12 B deployed without quantization using the vLLM inference framework.

The evaluation benchmark consists of a comprehensive set of carefully designed user tasks (410 expanded tasks) covering single-, dual-, and triple-action invocations across multiple devices, including five representative task variants: Paraphrasing, Non-sequitur, Error tests, Chain measures multi-step reasoning and state consistency, and Role evaluates contextual and role-based instruction following.

The comparison against the CodeGen baseline shows that AgentOptics consistently achieves the highest success rates across all task complexities. Specifically, AgentOptics-Online attains near-perfect performance (98.8%–100.0%) for single-action tasks, 99.3%–100.0% for dual-action tasks, and 97.0%–100.0% for triple-action tasks, whereas the CodeGen baseline exhibits substantially lower success rates across all task complexities (e.g., only up to 50.0%).

Improvements for AI systems

Here are specific improvements to existing AI systems based on the AgentOptics framework, along with what those improved systems can accomplish:


  1. Improvements in Control Reliability and Robustness via Structured Abstraction (MCP)

  2. Enhancement of Multi-Step Coordination and Complex Task Execution Capabilities

  3. Superior Handling of Linguistic Variation and Ambiguity in User Commands

  4. Robust Error Detection and Diagnosis for Heterogeneous Systems

  5. Implementation of Closed-Loop, Autonomous System Optimization

Specific capabilities enabled by these improvements:

  1. A system can reliably manage complex, multi-vendor optical networks (e.g., DWDM provisioning) by abstracting device-specific protocols into standardized MCP tools, ensuring that high-level natural language intent translates into correct sequences of actions across ROADMs, 400G transceivers, and OSAs without requiring task-specific code generation.

  2. The improved system can perform sophisticated, multi-action workflows—such as simultaneously provisioning a new channel and optimizing its launch power to minimize pre-FEC BER—by autonomously coordinating tool invocation order and managing state consistency across sequential steps, significantly outperforming code generation baselines that fail at triple-action tasks.

  3. The AI system can interpret highly varied or ambiguous natural language inputs (e.g., paraphrasing, non-sequitur) with high fidelity, correctly selecting the appropriate MCP tools from 64 standardized primitives to execute the intended action, maintaining near-perfect success rates (98.8%–100.0%) across complex linguistic challenges compared to current LLM-based code generation approaches (up to 50% success rate).

  4. The system can provide advanced diagnostic capabilities by analyzing execution logs and using reasoning to identify tool orchestration errors, such as incorrect tool naming or missing required arguments, allowing it to pinpoint exactly where a failure occurred in a multi-tool sequence (e.g., distinguishing between calling the wrong read vs. set operation), leading to highly targeted debugging of the agent's planning process.

  5. The AI system can enable true closed-loop automation by performing dynamic, iterative optimization of link performance based on real-time telemetry (e.g., autonomously adjusting ARoF transmitter bias voltage to maximize SNR/BER) and proactively detecting physical faults (e.g., identifying potential fiber cut events from DAS waterfall plots using prompt engineering), moving beyond simple control into autonomous system management.

Sources

Related papers