EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction
Danyu Li, Ling Zhou, Rubing Huang, Xian Zhong, Bin Zou, Kui Jiang
Macau University of Science and Technology · Macau University of Science and Technology Zhuhai Research Institute · Wuhan University of Technology · Jiangsu University · Harbin Institute of Technology
cs.LG, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 50/100
The gist: EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction Abstract RNA-Protein Interactions (RPIs) are critical for regulating cellular functions.
Terminology
Summary
EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction
Abstract
RNA-Protein Interactions (RPIs) are critical for regulating cellular functions. While traditional wet-lab experiments for RPI detection are costly and time-consuming, Deep Learning (DL) methods provide an efficient computational alternative for RPI Prediction (RPIP). In particular, Graph Neural Networks (GNNs) are promising, as they naturally model RPI networks. However, existing GNN-based methods often rely on homogeneous graphs or predefined meta-paths, which limit their ability to handle data sparsity and to generalize to cold-start scenarios involving unknown molecules. To address these limitations, we propose Edge Generation-guided Relation-aware Learning (EGRL), a novel framework with several key components: implicit meta-path learning to capture relational semantics without handcrafted paths; a multi-relation-aware attention mechanism for adaptive fusion of interaction patterns; a graph generator that predicts potential (soft
) edges to support cold-start nodes; and a multi-feature fusion predictor for final interaction scoring. EGRL is jointly trained with a primary task loss and an auxiliary generator loss. Comprehensive evaluations on four benchmark datasets demonstrate that EGRL achieves competitive overall performance. More importantly, it exhibits superior generalization in cold-start settings, achieving an Area Under the Receiver Operating Characteristic curve (AUROC) of 0.867 and an Area Under the Precision-Recall curve (AUPR) of 0.861 on unknown molecules, corresponding to improvements of 8.6% in AUROC and 5.0% in AUPR over prior state-of-the-art methods. The code will be released soon.
Keywords: RNA-Protein Interaction Prediction, Graph Neural Networks, Implicit Meta-path Learning, Multi-relational Graph Modeling, Cold-start Generalization
1. Introduction
The RNA-Protein Interactions (RPIs) play a vital role in various cellular processes, including gene regulation and protein synthesis, with significant implications for disease prediction, functional annotation of biomolecules, and drug design. Traditional biological experiments for RPI detection are often expensive and time-consuming, thereby motivating the development of computational approaches.
Early computational methods for RPI Prediction (RPIP) mainly relied on traditional machine learning models, such as support vector machines and random forests, using hand-crafted features extracted from sequence or structure information. Although these methods achieved some success, their performance was limited by the quality of manual feature engineering and their inability to capture complex nonlinear interaction patterns. Subsequently, Deep Learning (DL) has gained increasing attention in RPIP due to its ability to automatically learn complex features and interaction patterns from biological data. DL techniques, especially Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs), have been widely adopted for RPIP. CNNs extract local motif patterns from sequences, while RNNs and their variants model sequential dependencies. However, these methods typically treat each RNA-protein pair independently and ignore the global relational structure among multiple RNAs and proteins.
In recent years, Graph Neural Networks (GNNs) have emerged as a powerful framework for modeling RPIs. GNNs naturally represent biological entities as nodes and their interactions as edges in a graph, enabling the modeling of relational structure and dependencies within complex biological systems. By recursively aggregating information from neighboring nodes, GNNs learn enriched node representations that encode both intrinsic features and topological context. This capability makes them well-suited for analyzing diverse biological networks, including RNA-protein interaction graphs, protein-protein interaction graphs, and RNA-RNA interaction graphs.
However, conventional GNN-based methods still face notable limitations in RPIP. Many approaches rely on homogeneous graph constructions, which oversimplify the inherent heterogeneity of biological systems, where nodes (e.g., RNAs and proteins) and interactions (e.g., binding and functional association) are of distinct types. Although some studies incorporate predefined meta-paths to model higher-order relations, such designs require substantial domain expertise and may fail to capture all relevant interaction patterns. These limitations restrict model expressiveness, hinder generalization under sparse data, and limit performance in predicting interactions involving uncharacterized molecules (i.e., the cold-start problem).
To address these challenges, we propose Edge Generation-guided Relation-aware Learning (EGRL), a novel heterogeneous GNN framework that integrates implicit meta-path modeling, a graph generator for cold-start nodes, and multi-relational Graph Attention Network (GAT) layers with multi-feature fusion. The main contributions are summarized as follows: (1) Implicit meta-path learning: We propose an implicit meta-path learner that automatically captures the importance of different relation types without manual design, dynamically integrating relation-level semantics into node representations. (2) Graph generator for cold-start nodes: We introduce a jointly trained graph generator that predicts soft edges for unseen RNAs/proteins based on their sequence features, enabling effective cold-start interaction prediction. (3) Multi-relational GAT with multi-feature fusion: We design a multi-relational GAT that processes each edge type independently, along with a predictor that combines concatenation, element-wise product, and absolute difference of node embeddings for enhanced interaction modeling. (4) Competitive performance and effective cold-start generalization: Extensive experiments on four benchmark datasets demonstrate that EGRL achieves competitive performance against state-of-the-art methods and exhibits strong generalization under molecule hold-out and sequence-cluster-based cold-start settings.
2. Related Work
Traditional machine learning methods for RPIP typically rely on manually extracted features. For example, RPI-Pred utilizes sequence and structural information with a support vector machine classifier. Despite their interpretability and low computational cost, these methods suffer from limited expressive power, as hand-crafted features may fail to capture complex interaction patterns.
Deep learning approaches have significantly advanced RPIP by automatically learning hierarchical representations. CNNs extract local motif patterns from RNA or protein sequences treated as one-dimensional signals. For instance, MCNN predicts RNA-protein binding sites by integrating multiple CNNs, each processing RNA sequence segments of different lengths. RNNs, particularly long short-term memory (LSTM) models, capture long-range dependencies. Some methods further combine CNNs and RNNs to leverage both local and sequential patterns. In addition, RPICapsuleGAN, a generative adversarial network-based model, employs adversarial training to learn data distributions and enhance feature representations for RPIP.
In recent years, Graph Neural Networks (GNNs) have become a mainstream approach for RPIP and have achieved remarkable performance. The GNN paradigm aligns naturally with RNA-protein interaction modeling, where RNAs and proteins are represented as nodes and interactions as edges. Consequently, RPIP can be formulated as a link prediction problem. GNNs learn node representations by aggregating neighborhood information, enabling effective prediction of potential interactions. Several GNN-based methods have been proposed for RPIP. LPICGAE incorporates graph autoencoders to infer potential interactions. LncPNet designs a network embedding approach for lncRNA-protein interaction prediction. RPI-GGCN integrates gated graph convolutional networks with a co-regularized variational autoencoder, leveraging Gated Recurrent Units and Graph Convolutional Networks (GCNs) to extract topological information. DeepPN employs parallel CNN and GCN layers to capture hidden features of RNA sequences for RPIP. LPI-KCGCN utilizes a two-layer GCN to learn latent representations of RNAs and proteins and subsequently predicts their interaction scores.
However, many GNN-based methods treat the RPI graph as homogeneous and primarily focus on enhanced node feature learning, which oversimplifies the intrinsic heterogeneity of biological systems. In practice, RPI graphs are heterogeneous, involving multiple types of nodes and interactions. To address this, heterogeneous graph learning methods have been explored. For example, BiHo-GNN integrates homogeneous and heterogeneous network features through bipartite graph embedding, coupled with a mutual optimization strategy. Meta-paths have also been widely adopted to capture higher-order semantic relations in heterogeneous graphs, such as in RNA-disease association prediction. An explicit multilevel meta-path aggregation graph embedding model has been proposed for miRNA-disease association prediction. These studies suggest that meta-path-based feature learning is a viable strategy for modeling heterogeneous graph structures.
Graph Attention Networks (GATs) have been further introduced to enhance representation learning in heterogeneous interaction graphs. GATs adaptively aggregate information from neighboring nodes by learning attention coefficients, and have been shown to outperform conventional sequence-based models such as CNNs and bidirectional LSTM networks. NABind integrates sequence and structural descriptors with an attention mechanism to aggregate edge information. RPIembeddor employs a GAT-based framework to incorporate structural and functional features for RPI classification. Unlike self-attention in Transformers, which relies on positional encoding, graph attention computes importance based on node features and graph connectivity, making it more suitable for graph-structured data. In RPIP, modeling the importance of interaction patterns between nodes is more meaningful than relying on absolute positional information. Therefore, GAT provides a natural and effective framework for RPI modeling.
Based on the above observations, we propose EGRL. To the best of our knowledge, this is the first study that integrates implicit meta-path learning, a graph generator for cold-start nodes, and a multi-relational GAT with multi-feature fusion for RPIP.
3. Motivation
RPIs constitute core regulatory mechanisms of cellular activities and are closely associated with the occurrence and progression of various diseases. Accurate RPI prediction is therefore of great significance for drug target discovery, disease diagnosis, and therapeutic development. In recent years, deep learning models have substantially improved the efficiency of RPIP; however, several key challenges remain:
-
Insufficient training data: Although large volumes of RNA and protein sequences are available, experimentally verified interaction data remain relatively scarce, limiting the generalization ability of supervised models.
-
Diverse interaction patterns: RPIs involve multiple types (binding, regulation, catalysis, etc.), whereas most existing methods treat them as a single homogeneous relation, thereby losing important semantic distinctions.
-
Cold-start problem: When a new RNA or protein appears without known interaction edges, conventional GNNs fail to produce meaningful embeddings, leading to degraded prediction performance.
To address these challenges, we propose EGRL, which integrates four key components: (1) an implicit meta-path learning module that automatically discovers and aggregates relation-level semantics from multiple edge types to enrich node representations, alleviating data sparsity without requiring additional labels; (2) a multi-relational GAT that computes separate attention weights for each edge type (e.g., RNA-RNA similarity, protein-protein similarity, known RPIs, and generated soft edges), thereby preserving interaction heterogeneity; (3) a graph generator that predicts probabilistic soft edges between new and existing nodes using only sequence features, trained with a pseudo cold-start auxiliary loss to enable dynamic integration of unseen nodes during inference; and (4) an interaction predictor with multi-feature fusion, which captures both linear and multiplicative relationships for final interaction scoring. By integrating these components, EGRL achieves competitive performance on benchmark datasets and demonstrates strong generalization in cold-start scenarios.
4. Proposed Method
The framework consists of four key components: (1) Implicit meta-path learning, which automatically learns the importance of different relational paths; (2) Graph generator, which dynamically synthesizes soft edges for cold-start nodes; (3) Multi-relational GAT, which propagates information across multiple edge types; and (4) Interaction predictor with multi-feature fusion, which combines node embeddings for final interaction scoring.
4.1 Implicit meta-path learning
Given a heterogeneous graph with node set V = R ∪ P (RNAs and proteins) and multiple edge types R = r1, r2,..., rK, we aim to capture the semantic importance of each meta-path without manual design. Let X ∈ R(N×d) denote the initial node feature matrix (after linear projection), where N = R + P and d is the hidden dimension. For each relation type rk, we define the set of involved nodes as: N̂k = i (i, j) ∈ Ek or (j, i) ∈ Ek, where Ek denotes the edge set of type k. A relation-specific linear transformation Wk ∈ R(d×d) is applied, and the transformed features are averaged to obtain the relation representation: mk = (1/N̂k) Σ i∈N̂k Wk X[i,:]. The importance of each relation is computed via an attention network operating on concatenated representations: α = Softmax(MLP([m1 ∥... ∥ mK])), where ∥ denotes concatenation and the Multilayer Perceptron (MLP) consists of two linear layers with ReLU and dropout. The aggregated meta-path feature is: h meta = Σ k=1 K αk mk. Finally, h meta is broadcast to all nodes and added to the original features via a residual connection: X' = X + Dropout(h meta). This enriches node representations with relation-level semantics while preserving node-specific information.
4.2 Graph generator for cold-start nodes
To handle cold-start scenarios where new RNAs or proteins appear without known interactions, we introduce a graph generator that predicts probabilistic edges between new and existing nodes. The generator is a pairwise MLP that takes node features as input and outputs interaction probabilities. Let X new ∈ R(M×d) denote the embeddings of M cold-start nodes (after the same feature projection and implicit meta-path enhancement), and let X old ∈ R(N old×d) denote the embeddings of existing nodes. For each pair (i, j), the interaction probability is computed as: p ij = σ(MLP gen(X new[i,:] ∥ X old[j,:])), where ∥ denotes concatenation and σ is the sigmoid function. MLP gen consists of two hidden layers with ReLU activation followed by a linear output layer. The output is a probability matrix P ∈ [0, 1](M×N old). To incorporate these predicted interactions into the graph, we construct bidirectional soft edges with weights equal to the predicted probabilities. These soft edges are appended to the existing edge list and treated as an additional relation type in subsequent GAT layers.
Training Strategy for the Graph Generator: To ensure that the generator produces meaningful soft edges, we jointly train it with the main model using a self-supervised auxiliary loss. During each training epoch, a subset of RNA and protein nodes is randomly selected as pseudo cold-start nodes. For each such node, the generator predicts interaction probabilities with all candidate nodes, and a binary cross-entropy loss is computed against the ground-truth interaction matrix: L gen = (1/M rna) Σ i∈C rna BCE(P i,·, Y i,·) + (1/M prot) Σ j∈C prot BCE(P j,·, Y j,·), where C rna and C prot denote the pseudo cold-start RNA and protein sets, respectively, P is the predicted probability matrix, and Y is the corresponding submatrix of the training interaction matrix. This auxiliary loss is combined with the main prediction loss using a weighting factor λ gen and optimized jointly.
4.3 Multi-relational GAT
The enhanced node features X' (after meta-path learning) are fed into a multi-relational GAT to perform relation-aware message passing. We consider multiple edge types, including original hard edges (RNA-RNA, protein-protein, similarity, and known RNA-protein interactions) as well as soft edges generated for cold-start nodes (during both training and inference). For each relation type r, an independent GAT convolution is applied: h r = GATConv r(X', E r, edge attr = W r), where E r denotes the edge index of relation r, and W r is the corresponding edge weight vector (for soft edges) or omitted for hard edges. Each GATConv employs h attention heads and outputs features of dimension d. The outputs from all relations are averaged and combined with a residual connection, followed by layer normalization: X(l+1) = LayerNorm(X(l) + Dropout((1/R) Σ r=1 R h r)), where X(0) = X' and R denotes the set of all relation types (including soft edges). This process is repeated for L layers, enabling the model to capture both local and multi-hop dependencies.
4.4 Interaction predictor with multi-feature fusion
After the final GAT layer, we obtain embeddings for RNAs and proteins: h R ∈ R(R×d) and h P ∈ R(P×d). For a given RNA-protein pair (r, p), we construct a feature vector: f cross = h R[r] ∥ h P[p] ∥ h R[r] ⊙ h P[p] ∥ h R[r] − h P[p] ∈ R(4d), where ⊙ denotes element-wise product. This formulation captures both linear and multiplicative relationships. The feature vector is then passed through a two-layer predictor network with layer normalization and ReLU activation: ŷ = σ(W2 · ReLU(LayerNorm(W1 f cross + b1)) + b2), where σ is the sigmoid function, W1 ∈ R(d×4d) and W2 ∈ R(1×d) are weight matrices, and b1, b2 are bias vectors. The output ŷ ∈ [0, 1] denotes the predicted interaction probability.
4.5 Joint training objective
EGRL is trained by minimizing a weighted binary cross-entropy loss for the main prediction task, combined with a generator auxiliary loss. Let y i ∈ [0, 1] denote the ground-truth label for the i-th RNA-protein pair. The main loss is defined as: L main = −(1/N) Σ i=1 N [w p · y i log ŷ i + (1 − y i) log(1 − ŷ i)], where w p is the weight assigned to positive samples. In practice, w p = β · neg/pos, where β further emphasizes positive interactions. The total loss is given by: L total = L main + λ gen L gen, where λ gen is a hyperparameter controlling the contribution of the generator loss. This joint optimization encourages the generator to learn meaningful soft edges from node features, enabling effective interaction prediction for unseen nodes during inference.
5. Experimental Setup
5.1 Research questions
To evaluate the proposed approach, we investigate the following research questions (RQs):
-
RQ1: How do the core architectural components of EGRL affect performance? We conduct an ablation study with five configurations: Full (all modules enabled); -A (w/o ImplicitMetaPath); -B (w/o MultiRelationalGAT); -C (w/o GraphGenerator); -D (w/o Multi-featureFusion).
-
RQ2: How do key hyperparameters influence the performance of EGRL? We evaluate the robustness of the model with respect to key hyperparameters: number of k-nearest neighbours (k ∈ 3, 5, 10, 15, 20), generator loss weight λ gen ∈ 0.02, 0.05, 0.1, 0.2, and pseudo cold-start node ratio ∈ 0.01, 0.05, 0.1, 0.15.
-
RQ3: How does EGRL compare with state-of-the-art methods? We compare EGRL with representative RPIP methods on benchmark datasets, including RLF-LPI, RPI-SAN, RPITER, IPMiner, and NPI-GNN.
-
RQ4: Can EGRL generalize to completely unseen RNA and protein families? We conduct two evaluations: Experiment 1 (Molecule hold-out) and Experiment 2 (Sequence-cluster-based partitioning using CD-HIT with identity thresholds of 80% for RNA and 40% for protein).
5.2 Datasets
We utilize four benchmark datasets: RPI369, RPI1807, RPI2241, and NPInter2. RPI369, RPI1807, and RPI2241 are derived from RNA-protein complexes collected from the RNA-Protein Interaction Database (PRIDB) and the Protein Data Bank (PDB). RPI369 is a subset of RPI2241, excluding interactions involving ribosomal proteins or ribosomal RNA. The RPI1807 dataset is obtained by parsing data from the Nucleic Acid Database (NDB) and PRIDB. The NPInter2 dataset consists of experimentally validated ncRNA-protein interactions, particularly involving lncRNAs, collected from the NPInter2 database. Since RPI369, RPI2241, and NPInter2 do not provide negative samples, we follow prior work to generate an equal number of non-interacting pairs by randomly combining RNAs and proteins from positive samples. To reduce false negatives, pairs are discarded if they are highly similar to known interactions; specifically, a generated pair (R1, P1) is removed if there exists (R2, P2) such that the sequence identity between R1 and R2 exceeds 80% and that between P1 and P2 exceeds 40%.
5.3 Previous RPIP methods under comparison
We evaluate EGRL on the four datasets by comparing it with representative methods in RPIP: RLF-LPI (an ensemble framework integrating LSTM autoencoder with attention and fuzzy decision-making), RPI-SAN (a sequence-based method combining deep learning with random forest), RPITER (a multi-level deep learning framework integrating CNNs and stacked autoencoders), IPMiner (a stacked autoencoder-based approach with random forest), and NPI-GNN (an end-to-end graph neural network-based method).
5.4 Evaluation metrics
We adopt the commonly used five-fold cross-validation to evaluate model performance. Multiple metrics are employed, including Accuracy (ACC), Precision (PREC), Recall (REC), F1-score (F1), and Matthews Correlation Coefficient (MCC). In addition, threshold-independent metrics, namely the Area Under the Receiver Operating Characteristic Curve (AUROC) and the Area Under the Precision-Recall Curve (AUPR), are also reported.
5.5 Hardware and software environment
EGRL is implemented in Python using PyTorch 2.6.0+cu118 and PyTorch Geometric 2.6.1. The input features for RNA and protein nodes are obtained from pre-trained language models: RNA-FM provides 256-dimensional embeddings for RNAs, and ESM2-t33-650M produces 1280-dimensional embeddings for proteins. These features are projected into a unified 128-dimensional hidden space via linear layers. The hidden dimension of both the meta-path processor and the relation-aware network is set to 128. The model is trained using the AdamW optimizer with an initial learning rate of 1 × 10−4 and a cosine annealing scheduler. Training is performed for 300 epochs with a batch size of 256. A weighted binary cross-entropy loss is employed, where the positive sample weight is dynamically set to 10·neg/pos to address class imbalance. Gradient clipping is applied with a maximum norm of 1.0, and dropout is set to 0.2. All experiments are conducted on a single NVIDIA V100 (32GB) GPU using five-fold cross-validation.
6. Experimental Results
6.1 Answer to RQ1: Ablation study
The ablation results on the four datasets show that the core components of EGRL contribute to the final performance to different extents. Removing the implicit meta-path module (-A) decreases the Accuracy on RPI2241 from 0.925 to 0.889. For the graph generator (-C), its removal leads to only slight performance degradation, e.g., on RPI369 (Accuracy: 0.876 → 0.870, F1: 0.878 → 0.871) and RPI1807 (Accuracy: 0.968 → 0.955), suggesting limited benefit under the five-fold cross-validation setting, which does not fully reflect strict cold-start scenarios. In contrast, removing the multi-feature fusion module (-D) results in substantial degradation across all metrics. On RPI369, Accuracy drops from 0.876 to 0.849 and AUPR from 0.875 to 0.806; on RPI2241, Accuracy decreases from 0.925 to 0.859 and MCC from 0.888 to 0.820. This indicates that simple concatenation is insufficient to capture effective interactions, while the Hadamard product and absolute difference play a key role in modeling feature interactions. The multi-relational GAT module (-B) is also critical: removing it reduces AUROC on RPI369 from 0.883 to 0.833 and Accuracy on RPI1807 from 0.968 to 0.946, highlighting the importance of relation-aware modeling in heterogeneous graphs. Notably, all ablated variants contain fewer parameters than the full model. Overall, the complete EGRL consistently achieves the best performance, indicating that its effectiveness arises from the coordinated design of multiple modules.
6.2 Answer to RQ2: Impacts of hyperparameter tuning
We conduct systematic analyses on three key hyperparameters: the number of kNN neighbors (k), the generator loss weight (λ gen), and the pseudo cold-start node ratio (cold-ratio). Performance on the RPI1807 dataset remains stable across different values of k. While AUROC stays consistently high as k increases from 3 to 20, Recall improves from 0.983 to 0.988, and both F1-score and MCC show a gradual increase (with MCC peaking at 0.932 for k = 20). Precision slightly decreases at k = 10 (0.945), indicating that enlarging the neighborhood captures more positives but may introduce noise, reflecting a precision-recall trade-off. Overall, k = 5 provides a balanced configuration in terms of Accuracy (0.968) and Precision (0.956), whereas k = 20 favors higher Recall and MCC. These results indicate that the model is robust to the choice of neighborhood size under the standard evaluation setting.
In cold-start scenarios, where test nodes are absent from the training graph, the soft-edge completion mechanism of the graph generator becomes critical. On both RPI1807 and RPI2241, varying λ gen within [0.02, 0.05, 0.1, 0.2] and cold-ratio within [0.01, 0.05, 0.1, 0.15] leads to only minor AUROC fluctuations (within 0.0016 and 0.0022, respectively). Although the absolute AUROC on RPI2241 (approximately 0.95) is lower than that on RPI1807 (approximately 0.99), both datasets consistently show low sensitivity to these hyperparameters. This indicates that the generator module operates stably across a wide configuration range and effectively supports cold-start prediction. In summary, EGRL demonstrates strong robustness across both normal and cold-start scenarios.
6.3 Answer to RQ3: EGRL vs. State-of-the-art methods
EGRL achieves competitive performance across all four benchmark datasets. On RPI369, it attains the highest Accuracy, Precision, and MCC, indicating a clear advantage and strong capability in learning from limited data. On RPI1807, EGRL achieves the highest Recall (0.993), while other metrics remain close to the best-performing methods; for instance, its Precision (0.962) and AUROC (0.992) are slightly below those of IPMiner (0.978) and RPI-SAN (0.999), respectively, yet remain well balanced overall. On RPI2241, EGRL obtains the highest Accuracy (0.925) and MCC (0.888), with consistently strong performance across other metrics, demonstrating robustness on larger datasets. On NPInter2, EGRL achieves the highest Recall (0.977) and maintains competitive AUROC (0.986) and MCC (0.913), further validating its effectiveness across diverse data distributions.
6.4 Answer to RQ4: Generalization capability validation
The molecule hold-out cold-start results on the NPInter2 dataset yield an AUROC of 0.867 and an AUPR of 0.861 on the unseen test set. Although these values are lower than those obtained under five-fold cross-validation, EGRL achieves state-of-the-art performance under the same evaluation protocol. In particular, compared with the previous state-of-the-art method ZHMolGraph (AUROC: 0.798, AUPR: 0.820), EGRL improves AUROC by approximately 8.6% and AUPR by 5.0%. The molecule hold-out results show that after removing the protein (P84104) and the RNA (n342366) from NPInter2, for protein P84104, only one interacting RNA is missed, while for RNA n342366, all interacting partners are correctly predicted. Some newly predicted interactions may correspond to false positives or potentially undiscovered interactions, indicating both prediction uncertainty and exploratory capability.
The results under the more stringent sequence-cluster-based setting yield an AUROC of 0.801 and an AUPR of 0.822. Compared with the molecule hold-out setting (AUROC: 0.867, AUPR: 0.861), both metrics decrease, likely because this strategy removes clusters with high sequence similarity to the training data, thereby reducing potential information leakage and increasing task difficulty. Although false negatives and false positives exist, the majority of interactions are correctly predicted, indicating that EGRL maintains stable predictive capability even under more challenging conditions.
7. Conclusions and future work
Graph neural networks (GNNs) have emerged as a promising paradigm for RNA-protein interaction (RPI) prediction. Given the multi-level and multi-type regulatory mechanisms underlying RPIs, constructing heterogeneous RNA-protein graphs that capture both node-level and path-level relational information is essential. Such modeling provides a solid foundation for advancing AI-driven biomedical research and drug discovery.
In this work, we propose EGRL for accurate RPI prediction. EGRL performs relation-aware encoding over multi-relational graphs. Its key strengths are threefold: (1) it automatically captures interaction patterns via implicit meta-path exploration; (2) it models multiple subgraphs (RNA-RNA, protein-protein, similarity, RPI, and generator-induced soft edges) using a multi-relational GAT and integrates their representations; and (3) it incorporates a graph generator trained jointly with the backbone to enhance generalization to unseen nodes (cold-start). Extensive experiments based on five-fold cross-validation demonstrate that EGRL achieves competitive performance across four benchmark datasets. Moreover, it shows robust generalization under both molecule hold-out and sequence-cluster-based settings, indicating its ability to capture underlying network connectivity and support the discovery of potential novel interactions.
For future work, we plan to explore multimodal data integration by incorporating additional biological information, such as RNA secondary structure, protein tertiary structure, gene expression profiles, and epigenetic signals, together with sequence features. This is expected to provide more comprehensive representations and further improve prediction accuracy and reliability. We also aim to develop more transferable frameworks for RPI prediction across species and tissue types, thereby extending the applicability of artificial intelligence in bioinformatics.
Improvements for AI systems
Based on the paper, here are the specific improvements you can make to AI systems and what the improved systems can do:
1. Implicit meta-path learning for heterogeneous graph reasoning
-
Improvement: Replace handcrafted meta-paths with an attention-based module that automatically learns the importance of different relation types (e.g., RNA-RNA similarity, protein-protein similarity, known RPIs) from data.
-
Improved AI capability: The system can generalize to new domains without requiring domain experts to manually define relational paths, making it applicable to other heterogeneous networks (e.g., drug-target, gene-disease, social networks) where relation semantics are unknown or complex.
2. Graph generator for cold-start generalization
-
Improvement: Add a jointly trained pairwise MLP that predicts probabilistic
soft
edges for unseen nodes based solely on their features, with a pseudo cold-start auxiliary loss during training. -
Improved AI capability: The system can make meaningful predictions for entities with no prior interaction history (e.g., newly discovered proteins, novel drugs, or users in a recommendation system) by dynamically synthesizing plausible connections, avoiding the failure mode of standard GNNs that produce meaningless embeddings for isolated nodes.
3. Multi-relational GAT with edge-type-specific attention
-
Improvement: Process each edge type independently with separate GAT convolutions, then aggregate via averaging with residual connections and layer normalization.
-
Improved AI capability: The system preserves the heterogeneity of interaction types (e.g., binding vs. regulation vs. similarity) and learns type-specific importance weights, enabling more nuanced representation learning in complex networks—useful for multi-omics integration or multi-modal knowledge graphs.
4. Multi-feature fusion predictor (concatenation + Hadamard product + absolute difference)
-
Improvement: Combine node embeddings using concatenation, element-wise product, and absolute difference before feeding into a two-layer MLP predictor.
-
Improved AI capability: The system captures both linear and multiplicative (nonlinear) relationships between entity pairs, improving interaction scoring beyond simple dot products or concatenation—applicable to link prediction tasks where interactions are governed by complex combinatorial rules (e.g., protein-protein binding, drug synergy).
5. Joint training with auxiliary generator loss
-
Improvement: Optimize the main prediction loss and a self-supervised generator loss (with a weighting hyperparameter λ gen) simultaneously.
-
Improved AI capability: The system learns to produce useful soft edges as a byproduct of the main task, improving sample efficiency and robustness in low-data regimes—valuable for biomedical applications where labeled interaction data are scarce but unlabeled entity features are abundant.
6. Robustness to hyperparameter variation
-
Improvement: The framework shows stable performance across kNN neighborhood sizes (k=3–20), generator loss weights (0.02–0.2), and pseudo cold-start ratios (0.01–0.15).
-
Improved AI capability: The system can be deployed in production with minimal hyperparameter tuning, reducing the need for extensive grid search and making it more practical for non-expert users or automated pipelines.
7. Sequence-cluster-based evaluation for fair generalization testing
-
Improvement: Use CD-HIT clustering (80% RNA identity, 40% protein identity) to partition data and evaluate on completely unseen sequence families.
-
Improved AI capability: The system’s generalization claims are validated under stricter, more realistic conditions (avoiding information leakage from similar sequences), making it trustworthy for real-world discovery tasks where novel sequences are common (e.g., metagenomics, variant analysis).
8. Transferable architecture for multi-omics and cross-species prediction
-
Improvement: The framework’s modular design (implicit meta-path + multi-relational GAT + generator + fusion predictor) is not specific to RNA-protein interactions.
-
Improved AI capability: The system can be adapted to other heterogeneous biological networks (e.g., drug-disease, miRNA-mRNA, protein-metabolite) or even non-biological domains (e.g., recommendation systems, fraud detection) with minimal changes, by swapping input features and edge types.
9. Improved cold-start AUROC/AUPR by 8.6%/5.0% over prior SOTA
-
Improvement: Achieves AUROC 0.867 and AUPR 0.861 on unseen molecules (NPInter2 dataset), outperforming ZHMolGraph.
-
Improved AI capability: The system can reliably predict interactions for entirely new molecules, enabling early-stage drug target prioritization, functional annotation of uncharacterized proteins, and hypothesis generation for experimental validation—reducing wet-lab costs and time.
10. Stable performance under strict sequence-cluster partitioning
-
Improvement: Maintains AUROC 0.801 and AUPR 0.822 even when test clusters share <80% RNA and <40% protein identity with training data.
-
Improved AI capability: The system is robust to evolutionary divergence and can handle novel sequence families, making it suitable for cross-species or cross-tissue applications where sequence homology is low.
Summary of what the improved AI system can do:
-
Predict interactions for unseen entities (cold-start) with high accuracy, using only sequence/feature data.
-
Automatically discover relevant relational semantics without manual meta-path engineering.
-
Handle heterogeneous interaction types in a unified framework.
-
Maintain performance with minimal hyperparameter tuning.
-
Generalize to other link prediction tasks in biology and beyond.
-
Provide reliable predictions for novel molecules, accelerating drug discovery and functional genomics.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks