LLM-hRIC: LLM-empowered Hierarchical RAN Intelligent Control for O-RAN

arXiv:2504.18062 · cs.NI, cs.AI · Submitted 2026-02-27 · Read on arXiv

Lingyan Bao, Sinwoong Yun, Jemin Lee, Tony Q.S. Quek

Yonsei University · Electronics and Telecommunications Research Institute · Singapore University of Technology and Design

cs.NI, cs.AI

Submitted: 2026-02-27

Comments: To appear in IEEE Communications Magazine

Journal ref: IEEE Communications Magazine, pp. 1-7, 2026

DOI: 10.1109/MCOM.001.2500315

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 46/100

The gist: The paper introduces the "LLM-empowered hierarchical RIC (LLM-hRIC) framework to improve the collaboration between RICs in O-RAN." The research addresses a "critical real-world problem" which is

Terminology

Summary

The paper introduces the LLM-empowered hierarchical RIC (LLM-hRIC) framework to improve the collaboration between RICs in O-RAN. The research addresses a critical real-world problem which is bridging the ‘coordination gap’ between the strategic, longterm control of the non-RT RIC and the tactical, real-time execution of the near-RT RIC. This gap is driven by challenges such as traditional ML techniques struggle to process the diverse and complex information in O-RAN, the high computational complexity hinders near-RT decision-making, and the lack of domain-specific finetuning.

The proposed LLM-hRIC framework consists of two primary layers: 1) the LLM-empowered non-RT RIC (i.e., upper layer control) and 2) the RL-empowered near-RT RIC (i.e., lower layer control). In this architecture, The LLM-empowered non-RT RIC acts as a guider, offering a strategic guidance to the near-real-time RIC (near-RT RIC) using global network information, while The RL-empowered near-RT RIC acts as an implementer, combining this guidance with local real-time data to make near-RT decisions.

The operational process of the LLM-empowered non-RT RIC involves several steps:

  1. Data Integration: The non-RT RIC ingests multi-modal data from gNB and near-RT RIC, primarily via the O1 interface, including long-term time-series metrics, unstructured textual data, topology data, and advanced visual data.

  2. Input Validation and Prompt Construction: Data undergoes a validation stage to handle potentially noisy or corrupted measurements and is structured into a unified, context-rich format.

  3. Environment Understanding and Actionable Insights Generation: The LLM analyzes and understands the context... and generate the text response, that includes the strategic guidance y (e.g., transmit power allocation strategies).

  4. Guidance Extraction and Verification: The output is validated for its conformance to predefined formats and constraints, with a fallback mechanism to provide a default policy if validation fails.

  5. Guidance Transfer to the near-RT RIC: The guidance vector y is then sent to the near-RT RIC via the A1 interface.

The RL-empowered near-RT RIC follows a two-step process:

  1. Data Integration: It ingests the strategic guidance from the non-RT RIC (via the A1 interface) and the observed local information from the gNB (via the E2 interface).

  2. Policy generation: The RL-based xApp then analyzes these inputs to generate an optimal, fine-grained policy that is sent to the gNB for execution.

To verify the framework, the authors present a use case based on IAB networks for power allocation in the integrated access and backhaul (IAB) networks. The objective is to maximize the total throughput T by optimizing the power allocation of MBSs to each SBSs under the total power constraint on the MBS. The training of the RL-empowered near-RT RIC is conducted through three sequential phases:

  • Phase 1 (LLM-guided Exploration): The power allocation policy p m is determined by adding noise to the LLM’s initial guidance policy p om, i.e., p m = p om + noise.

  • Phase 2 (Policy Blending): "The power allocation policy p m is determined by a weighted combination of the LLM-provided policy p om and the RL-generated policy p dm, i.e., p m = w p om + (1 - w) p dm, where the blending coefficient w decays gradually from 1 to 0 over time."

  • Phase 3 (Self-directed Decision): The power allocation policy p m is determined solely by the RL agent, i.e., p m = p dm.

Experimental results demonstrate the feasibility of the framework. Regarding latency, the LLM inference latency is less than 1.5 seconds, confirming its decision time is well within the operational budget required for the LLM-hRIC (2s), and the DDPG inference time is less than 0.07 ms, which is significantly faster than the 10 ms requirement for the most demanding near-real-time control. In terms of performance, the proposed LLM-hRIC framework achieves higher total throughput compared to the baseline methods (DLN, DCN, and EPA). The results also indicate that the decaying w is critical for performance, as it ensures a smooth transition from guided exploration to self-directed optimization, which is necessary for stable convergence to the highest-performing policy.

Finally, the paper identifies several open challenges of the LLM-hRIC framework for O-RAN:

  • Leveraging Multi-Model Data in LLM-hRIC: Challenges include the lack of robust support for multi-modal inputs in the wireless communication domain and the need for designing an effective multi-modal RAG engine.

  • Fine-Tuning of LLMs: Issues include collecting or effectively utilizing the limited available data on wireless communication, generating high-quality instruction data, and implementing a privacy-preserving architecture.

  • Systematic Collaboration between RICs: Challenges involve the temporal mismatch between the large time scale of non-RT RICs and the different small time scales of near-RT RICs, as well as the need for task-aware, customized guidance.

  • Timely Inference: The need to balance model size, inference speed, and accuracy through technologies like model pruning, quantization, and knowledge distillation.

Improvements for AI systems

1. Multi-modal Wireless Retrieval-Augmented Generation (RAG) Engine

  • The Improvement: Implement a specialized RAG engine that utilizes vector databases to index not just text, but also graph-based topology data, time-series patterns, and spectral heatmaps.

  • What the system can do: Instead of the LLM merely processing current snapshots, it can perform historical context queries. For example, if a sudden drop in throughput occurs, the system can instantly retrieve similar topological configurations and signal interference patterns from the past to provide more accurate, context-aware strategic guidance.

2. LLM-driven Reward Shaping for Hierarchical Reinforcement Learning (HRL)

  • The Improvement: Shift the LLM’s output from a static guidance vector (y) to a dynamic reward function (R strategic) that is injected into the near-RT RIC’s RL training loop.

  • What the system can do: This resolves the temporal mismatch. The LLM defines the goal (e.g., prioritize energy efficiency over peak throughput for the next 10 minutes), and the RL agent optimizes its own fine-grained, millisecond-level actions to maximize that specific reward, allowing for much more flexible and sophisticated long-term/short-term alignment.

3. Digital Twin-driven Synthetic Instruction Tuning

  • The Improvement: Utilize a high-fidelity Digital Twin of the O-RAN environment to generate massive amounts of synthetic State-Action-Reasoning triplets.

  • What the system can do: This overcomes the lack of domain-specific data. The system can simulate rare, high-impact network events (e.g., sudden massive user influx, hardware failure, or extreme interference) to train the LLM on how to generate strategic guidance for scenarios that haven't occurred in the real world yet, making the network highly resilient.

4. Edge-Native Knowledge Distillation (Teacher-Student Architecture)

  • The Improvement: Use the large-scale LLM in the non-RT RIC as a Teacher to distill its strategic reasoning capabilities into a highly compressed, quantized Micro-LLM (Student) deployed closer to the edge.

  • What the system can do: This drastically reduces the 1.5s inference latency. The system can provide near-strategic guidance in sub-100ms intervals, allowing the non-RT RIC to react to network fluctuations much more frequently without the massive computational overhead of the full-scale model.

5. Closed-Loop Critic Feedback Mechanism

  • The Improvement: Implement a feedback loop where the near-RT RIC sends Execution Success Reports (comparing the intended policy vs. actual network outcome) back to the non-RT RIC.

  • What the system can do: The LLM acts as a Critic that analyzes why a strategic guidance failed (e.g., The power allocation was too aggressive for the current interference levels). It then performs self-correction, automatically refining its prompting strategy or internal logic to ensure future guidance is more implementable by the RL agent.

Sources

Related papers