ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes
Ziyan Wang, Liwen Wu, Cheng Xie, Song Gao, Zhenli He, Xin Jin
Yunnan University
cs.LG, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 100/100
The gist: ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes Abstract Text-Attributed Graphs (TAGs), endowed with abundant textual content along with
Terminology
Summary
ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes
Abstract
Text-Attributed Graphs (TAGs), endowed with abundant textual content along with topological structures, have emerged as a versatile backbone for real-world anomaly detection spanning large language model security, social network moderation, and cyber threat identification. Unlike conventional Graph Anomaly Detection (GAD), which relies primarily on structural irregularities, TAG anomaly detection must jointly leverage both topological patterns and fine-grained textual semantics to capture nuanced anomalous behaviors. The current GNN-based anomaly detectors adopt holistic message-passing schemes that indiscriminately fuse structural proximity and textual semantics during propagation, leading to deep cross-modality coupling. This entanglement acts as a noise amplifier, obscuring subtle anomalous signals and directly giving rise to the Blurred-Anomaly-Boundary (BAB) issue by rendering normal-anomalous decision boundaries poorly separable. This challenge is further amplified for graph foundation models that require robust cross-domain generalization. To bridge this gap, we introduce a novel foundation model for TAG anomaly detection featuring decoupled topological and textual prototypes. Our framework constructs dual prototype banks to independently model structural normality and semantic consistency, effectively isolating anomaly cues that are otherwise diluted during coupled aggregation. Extensive experiments across 14 diverse benchmark datasets demonstrate that our method consistently achieves state-of-the-art performance in cross-domain settings. Notably, the ablation studies further corroborate the prevalence of the BAB issue in conventional coupled TAG anomaly detectors, and show that our decoupled prototype design effectively mitigates this challenge.
Introduction
Graph anomaly detection (GAD) aims to identify nodes that deviate from dominant graph patterns and has been widely applied to fraud detection, social network moderation, cybersecurity, and recommendation systems. In real-world graphs, nodes are often associated with rich textual content in addition to relational structures, such as paper abstracts, product descriptions, and user posts. Such data are commonly formulated as text-attributed graphs (TAGs), where anomalies may arise from irregular connections, subtle semantic inconsistencies, or both. Effective TAG anomaly detection therefore requires jointly exploiting textual semantics and graph topology.
Existing GAD methods are primarily designed for numerical attributes and structural irregularities. Although recent TAGAD methods employ language models to encode raw text, most still rely on holistic GNN message passing to jointly propagate textual and topological information. As illustrated in Figure 1, this coupled propagation progressively smooths or distorts subtle anomaly cues, reducing the distinction between normal and anomalous nodes and thereby giving rise to the Blurred-Anomaly-Boundary (BAB) issue.
Recent graph pre-training and prompt-learning methods have explored transferable graph representations across tasks and domains. However, the BAB issue becomes more severe in the generalist setting, where one model is trained on multiple source graphs and directly applied to unseen domains. Cross-domain variations in textual semantics and connectivity patterns make jointly propagated representations prone to retaining domain-specific neighborhood information, further weakening anomaly cues under distribution shifts and hindering zero-shot generalization.
To address this issue, we propose ProTAGAD, a prototype-based foundation model that decouples textual and topological representation learning. ProTAGAD separately learns transferable textual anomaly prototypes and topological normality prototypes without exchanging hidden representations, and combines their anomaly scores only at the decision level. This design avoids cross-modal interference while preserving the complementary anomaly evidence of both modalities.
We evaluate ProTAGAD on 14 TAG datasets under a source-to-target zero-shot protocol. ProTAGAD achieves the best AUROC on seven of eight unseen target graphs, with an average rank of 1.12, demonstrating strong cross-domain generalization. Additional analyses further reveal the complementary roles of textual and topological prototypes and verify that decoupled prototype modeling effectively alleviates the BAB issue.
Main Contributions
• We formally identify the Blurred-Anomaly-Boundary (BAB) issue in generalist TAG anomaly detection, attributing its root cause to entangled cross-modal fusion via holistic GNN message passing.
• We propose ProTAGAD, a foundation model for TAG anomaly detection that leverages decoupled textual anomaly prototypes and topological normality prototypes to isolate modality-specific anomaly evidence without cross-modal interference.
• Extensive experiments on 14 diverse TAG benchmarks achieve state-of-the-art zero-shot cross-domain performance; ablations further verify dual prototype complementarity and empirically confirm both the prevalence of the BAB issue and the efficacy of our decoupled design.
Related Work
Graph Anomaly Detection: GAD aims to identify nodes that deviate from the dominant attribute or structural patterns. Existing methods mainly include reconstruction-, self-supervised-, spectral-, affinity-, and augmentation-based approaches. Reconstruction-based methods detect anomalies through the reconstruction of attributes or structures (DOMINANT, AnomalyDAE, ComGA). Self-supervised methods learn anomaly-sensitive node-context or neighborhood patterns (CoLA, HCM-A). Spectral methods exploit high-frequency graph signals (BWGNN, GHRN), while affinity-based methods model one-class normality and suppress suspicious connections (TAM, GCTAM). Augmentation-based methods synthesize anomalies to enhance model robustness (Semi-GGAD, CAGAD). However, these methods generally follow a one-for-one paradigm with fixed numerical attributes, requiring retraining on new graphs and overlooking the fine-grained semantics of raw text.
Generalist Graph Anomaly Detection: GGAD extends conventional GAD from graph-specific learning toward cross-domain generalization, aiming to train a unified detector that can be transferred to unseen graphs. ARC introduces in-context learning for transferable anomaly detection, while UNPrompt and AnomalyGFM explore zero/few-shot detection through unified prompts and graph foundation models. Recent studies further investigate domain shifts through invariant representation learning and prototype-based knowledge transfer, such as IA-GGAD, DR-GGAD, OWLEYE, and ProMoS. However, existing GGAD methods generally focus on numerical node representations and structural distribution shifts. The semantic information contained in raw texts is not explicitly modeled, and the interaction between textual and topological anomaly evidence remains unexplored, limiting their applicability to text-attributed graphs.
Text-Attributed Graph Anomaly Detection: TAGAD identifies anomalous nodes by jointly modeling textual attributes and graph topology. Conventional methods typically encode texts into fixed-dimensional representations, which may miss subtle anomaly cues. Recent approaches improve textual modeling through contrastive learning or large language model reasoning. CMUCL captures textual and structural inconsistencies through multi-scale contrastive learning, CoLL extracts textual anomaly evidence through collaborative LLM reasoning, and TAGAD studies realistic anomaly construction and retrieval-augmented zero-shot detection. However, existing methods often entangle textual and topological information, causing semantic anomalies to be smoothed by neighborhood aggregation, while LLM-based approaches rarely model transferable structural normality. Most are also tailored to individual graphs. Instead, ProTAGAD learns decoupled textual and topological prototype banks and combines their scores for cross-domain detection on unseen graphs.
Preliminaries
Notations: Let G = (V, A, T, X) denote a text-attributed graph, where V = v1,..., vN is the node set, A ∈ 0, 1 N×N is the adjacency matrix, and T = ti N i=1 contains the raw textual attributes associated with the nodes. An entry Aij = 1 indicates an edge between vi and vj. The matrix X ∈ RN×d denotes the initial node feature matrix, whose i-th row xi ∈ Rd is the d-dimensional feature vector associated with node vi.
For each node vi, we denote its one-hop neighborhood by Ni and construct a local textual subgraph containing the text of the target node and its neighbors: Git = ti ∪ tj j ∈ Ni. The local textual subgraph enables the agent to assess both the semantic content of the target node and its consistency with neighboring texts.
Generalist TAG Anomaly Detection: Given a collection of source-domain TAGs Ttrain = G(1) train,..., G(ns) train and a disjoint collection of unseen target-domain TAGs Ttest = G(1) test,..., G(nt) test, where ns and nt denote the numbers of source and target graphs, respectively, our goal is to learn a unified anomaly detector from Ttrain and directly generalize it to Ttest. During inference, the model parameters are frozen, and no target-domain labels or additional fine-tuning are available. For each node vi in a target graph, the detector outputs an anomaly score S(vi) ∈ R, where a larger value indicates a higher likelihood of being anomalous.
Methodology
To mitigate the BAB problem, we propose ProTAGAD, which decouples textual and topological representation learning by avoiding shared cross-modal message propagation. ProTAGAD separately learns textual and topological prototypes and fuses their anomaly scores for zero-shot inference on unseen graphs.
Textual Prototype Learning
The textual module aims to extract transferable semantic anomaly patterns from raw node texts. We first employ a text encoder fΘ txt to map the textual attributes into a continuous representation space: Xt = fΘ txt(T) = [xt1,..., xtN]⊤ ∈ RN×dt, where xti denotes the textual representation of node vi.
Based on these textual representations, we employ a lightweight text anomaly probability estimator to estimate the textual anomaly probability of each node: Ŷ = fΘ prob(Xt) = [ŷ1,..., ŷN]⊤, where ŷi ∈ [0, 1] denotes the estimated textual anomaly probability of node vi.
Existing text encoders often emphasize general semantics, making fine-grained anomaly cues difficult to capture. To address this issue, we introduce the Agents-Skills-Review-Confidence mechanism. Given a local textual subgraph Git, the anomaly agent first performs a coarse assessment. The Skill Router then identifies the graph domain, activates the corresponding domain-specific skill, and conducts fine-grained analysis of whether the target text is semantically consistent with its topic or content. The reviewer further verifies the decision consistency, supporting textual evidence, and confidence reliability. After reviewing all nodes, we identify an elbow point in the anomaly confidence distribution as the threshold. Nodes with confidence scores above the threshold are assigned a pseudo-label of 1, while the remaining nodes are assigned 0. The resulting binary indicator for node vi is defined as ci = Agents-Skills-Review-Confidence(Git), where C = [c1,..., cN]⊤, with ci = 1 indicating a high-confidence textual anomaly and ci = 0 indicating low textual anomaly confidence.
These agent-derived binary indicators are used only as pseudo-labels to supervise the text anomaly probability estimator through binary cross-entropy loss: Lprob = −(1/N) Σ [ci log(ŷi) + (1 − ci) log(1 − ŷi)].
We further use these pseudo-labels to partition the node representations into anomalous and normal groups, which are then aggregated to construct the corresponding textual prototypes: pt− = (1/∥C∥1) Σ ci·xti, pt+ = (1/∥1 − C∥1) Σ (1−ci)·xti, where pt− and pt+ are obtained by averaging the textual representations of the inferred anomalous and normal nodes, respectively, thereby capturing the representative semantic patterns of the two groups.
However, constructing the prototypes alone does not explicitly constrain the positions of individual nodes in the representation space. To further enhance the separation between anomalous and normal textual patterns, we introduce a prototype alignment objective: Lalign = (1/N) Σ [ci log(1 + exp(ϕ(xti, pt+) − ϕ(xti, pt−))) + (1 − ci) log(1 + exp(ϕ(xti, pt−) − ϕ(xti, pt+)))], where ϕ(·, ·) denotes cosine similarity. For ci = 1, the objective pulls xti toward pt− and pushes it away from pt+, while for ci = 0, it pulls xti toward pt+ and pushes it away from pt−. This objective yields a more discriminative textual representation space with a clearer normal-anomalous boundary.
The probability estimation and prototype alignment objectives optimize different components of the textual module: Lprob optimizes the text anomaly probability estimator to estimate textual anomaly probabilities, while Lalign optimizes the text encoder to learn discriminative semantic representations of anomalies.
After training, the textual anomaly score of node vi is defined as S t(vi) = fΘ prob(xti) + ϕ(xti, pt−) + ∥xti − (1/N(i)) Σj∈N(i) xtj∥2 2, where the first term is the textual anomaly probability estimated by fΘ prob, the second measures the semantic affinity between the node representation and the textual anomaly prototype learned from the source graphs, and the third term captures the deviation between a node's textual representation and its local textual neighborhood. Before summation, all score terms are standardized using Z-score normalization to eliminate their scale differences.
Topological Prototype Learning
Following ProMoS, we learn two node representations to model transferable structural normality. A Graph Transformer is first trained with self-supervised contrastive learning to obtain topology-aware node representations: H = fΘ GNN(X, A) = Graph Transformer(X, A), where H = [h1,..., hN]⊤ encodes the node attributes together with their topological contexts.
To obtain a complementary representation, we further employ an MLP that takes only the node features as input and learns to approximate the topology-aware representations through knowledge distillation: H′ = fΘ MLP(X) = σ((σ(XW + b))W′ + b′). Although the MLP does not directly use the adjacency matrix, the distillation process transfers topology-related knowledge from H to H′. The two representations therefore provide complementary views for subsequent topological prototype construction.
To capture diverse structural normality patterns, we apply K-means clustering to the topology-aware representations H and use the resulting cluster centers as topological prototypes: Ps = K-means(H, K) = ps1,..., psK, where each psk represents a typical structural normality pattern in the topology-aware representation space. To ensure that the projected representation h′i preserves the structural patterns encoded in hi, we align their similarity distributions over the shared topological prototypes. Specifically, the structural consistency objective is defined as Lstr = (1/N) Σ KL(softmax(hi·(Ps)⊤) ∥ softmax(h′i·(Ps)⊤)), where KL(· ∥ ·) denotes the Kullback–Leibler divergence. This objective encourages h′i to retain the relative affinities of hi to different topological prototypes, thereby preserving its topology-aware structural semantics.
Based on these prototypes and the two complementary node representations, we define the topological anomaly score of node vi as S s(vi) = KL(softmax(hi·(Ps)⊤) ∥ softmax(h′i·(Ps)⊤)) + ∥hi − psm∥2 2, m = arg mink ∥hi − psk∥2. Specifically, the two terms measure the discrepancy between the similarity distributions of the two node representations over the topological prototypes and the deviation from the nearest topological prototype, respectively. Before summation, both score terms are standardized using Z-score normalization to eliminate scale differences.
Decoupled Dual-Prototype Anomaly Scoring
The textual and topological branches model semantic and structural anomalies in separate prototype spaces. The two branches produce the textual anomaly score and the topological anomaly score, respectively. ProTAGAD combines the two scores only at the final scoring stage, thereby avoiding the interference caused by deeply coupled representations. The overall anomaly score of node vi is defined as S(vi) = S t(vi) + S s(vi). Before final aggregation, the textual and topological anomaly scores are independently standardized using Z-score normalization to ensure comparable scales. Consequently, a node is considered more anomalous when it exhibits a large standardized deviation in either textual semantics or topological structure, resulting in a higher final anomaly score S(vi).
Experiments
Experimental Settings
Datasets: We adopt a cross-domain source/target split across 14 text-attributed graphs spanning the citation, e-commerce, web, and encyclopedia domains. We synthesize and inject realistic anomalies that mirror real-world scenarios, including off-topic papers, citation manipulation, misleading products, fraudulent co-purchases, fake encyclopedia entries, and promotional content. Building on CMUCL, we further introduce contextual and structural anomalies to construct a challenging benchmark to evaluate zero-shot cross-domain generalization. The source graphs for training are Cora, Arxiv, Children, Products, History, and WikiCS, while the unseen target graphs for testing are Citeseer, Pubmed, Grocery, Movies, Toys, Fitness, Cornell, and Texas.
Baselines: We compare ProTAGAD with 18 representative baselines: (1) GAD methods—DOMINANT, BGNN, BWGNN, GHRN, CoLA, HCM-A, TAM, Semi-GGAD, CAGAD and GCTAM; (2) GGAD methods—ARC, IA-GGAD, AnomalyGFM, UNPrompt, OWLEYE and ProMoS; (3) TAGAD methods—CMUCL and CoLL.
Implementation: We report AUROC and AUPRC as mean ± standard deviation over five random seeds. Each method is trained once on Ttrain and directly evaluated on Ttest under a pretrain-only protocol. We use BGE, GraphTransformer, and DeepSeek4-Flash as the text encoder, graph encoder, and LLM agent, respectively. The agent is used only for offline preprocessing, with its cached outputs shared across all seeds; no LLM calls are required during training or inference. This preprocessing takes approximately 395.8 seconds and costs USD 6.64.
Main Results
Table 2 summarizes the AUROC results on eight unseen target graphs. ProTAGAD achieves the best performance on seven datasets and attains the lowest average rank of 1.12, demonstrating stable zero-shot cross-domain generalization. The most substantial improvements are observed on Citeseer (+8.65%) and Pubmed (+8.17%), indicating that the textual prototype bank effectively preserves fine-grained semantic anomalies and prevents them from being weakened during topological aggregation. The consistent gains on Texas, Grocery, Movies, Toys, and Fitness further suggest that the decoupled topological prototype bank can capture transferable structural normality across different domains. The only exception is Cornell, where ProTAGAD achieves an AUROC of 80.31%, which is 0.94% lower than the best baseline. This small gap may result from the extremely limited graph size, which provides insufficiently diverse patterns for stable prototype estimation. Overall, the results show that independently modeling semantic consistency and topological normality enables ProTAGAD to preserve more discriminative anomaly cues under domain shifts. In addition, the standard deviations remain below 1.1% on all target graphs, confirming the model's stable performance across random seeds.
Ablation Study
To disentangle the contribution of each prototype module, we evaluate four variants: (i) Backbone, which removes both prototype modules; (ii) + SP, which incorporates only the structural prototype module; (iii) + TP, which incorporates only the textual prototype module; and (iv) Ours, which combines both modules. Structural prototypes consistently improve the Backbone across all target graphs, with particularly large gains on Fitness (+33.82%), Cornell (+25.44%), and Texas (+24.02%), demonstrating their ability to capture transferable structural normality. Textual prototypes also yield notable improvements on Pubmed (+11.89%) and Fitness (+9.33%), indicating their effectiveness in capturing semantic anomaly cues. Combining both modules achieves the best performance on every target graph, confirming that textual and structural prototypes provide complementary evidence for cross-domain anomaly detection.
Effect of Decoupled Prototype Modeling
To evaluate decoupled prototype modeling, we compare ProTAGAD with a coupled variant under identical settings. In the coupled variant, the textual features and original node features are first projected to the same dimensionality, separately l2-normalized, and then averaged. The fused features replace the original node features as the input to the graph encoder. Both textual and topological prototypes are constructed from the resulting fused representations, and anomaly scores are computed in this shared prototype space.
We further introduce Anomaly Boundary Separability (ABS) to quantify the BAB issue. Let f̂+(s) and f̂−(s) denote the kernel density estimates of the anomaly scores for normal and anomalous nodes, respectively. Both densities are estimated using a Gaussian kernel with a shared bandwidth determined by applying Scott's rule to the pooled scores of the two groups. We define ABS = (1/2)[DKL(f̂+(s) ∥ M) + DKL(f̂−(s) ∥ M)], where M = (1/2)(f̂+ + f̂−) denotes the mixture distribution, and DKL(· ∥ ·) represents the Kullback–Leibler divergence. A larger ABS indicates clearer separation between normal and anomalous nodes and thus a less severe BAB issue.
ProTAGAD consistently outperforms the coupled variant on all eight target graphs, increasing the average AUROC from 61.36% to 78.89% and the average ABS from 0.2260 to 0.4553. The lower ABS of the coupled variant indicates that normal and anomalous nodes exhibit more similar anomaly-score distributions, leading to weaker boundary separability and a more severe BAB issue. In contrast, decoupled prototype modeling preserves the complementary anomaly information of the two modalities, resulting in more discriminative anomaly scores and clearer normal–anomalous boundaries.
Parameter Sensitivity
We investigate the effect of the number of structural prototypes K in Eq. (12) on the Toys and Grocery datasets. Both AUROC and AUPRC gradually improve as K increases from 1 to 10, and achieve their best performance at K = 10. A small K provides insufficient prototypes to characterize the diverse structural normality patterns across graphs. In contrast, an excessively large K may fragment the normal structural distribution and introduce domain-specific noise, leading to performance degradation when K = 15. Overall, ProTAGAD remains relatively stable under different values of K, while K = 10 provides the best balance between prototype diversity and cross-domain generalization.
Efficiency Analysis
Figure 5 compares the runtime, GPU memory usage, and mean AUROC of different methods during training and inference, where bubble size denotes memory consumption. ProTAGAD lies near the upper-left region in both panels, indicating that it achieves the highest detection performance with competitive computational efficiency. It also maintains relatively low GPU memory usage during training, while its inference memory consumption is comparatively high. Overall, ProTAGAD achieves a favorable trade-off between detection performance and computational efficiency.
Conclusion
We presented ProTAGAD, a zero-shot foundation model for generalist text-attributed graph anomaly detection on unseen graphs. By decoupling textual and topological representation learning and separately learning textual and structural prototypes, ProTAGAD alleviates the Blurred-Anomaly-Boundary (BAB) issue. Across 14 TAG datasets, ProTAGAD achieves the best AUROC on seven of eight target graphs with an average rank of 1.12. Ablation studies and coupled-versus-decoupled comparisons further validate the complementarity of the two prototype branches and the effectiveness of the proposed design. Future work will explore more memory-efficient prototype learning and extend ProTAGAD to large-scale graph domains.
Improvements for AI systems
Improvements to AI systems based on this paper:
-
Decoupled multi-modal representation learning: Replace holistic message-passing that fuses textual and topological features during propagation with separate modality-specific prototype banks. This prevents cross-modal noise amplification and preserves subtle anomaly cues that are otherwise diluted. The improved system can detect anomalies that are only visible in one modality without interference from the other.
-
Dual prototype architecture for normality/anomaly modeling: Maintain separate textual prototypes (anomalous vs. normal semantic patterns) and topological prototypes (multiple structural normality clusters). The improved system can independently assess semantic consistency and structural regularity, then combine scores only at the decision level, yielding clearer decision boundaries.
-
Agent-based pseudo-labeling for text anomaly estimation: Use an Agents-Skills-Review-Confidence mechanism where an LLM agent performs coarse assessment, a skill router activates domain-specific skills for fine-grained analysis, and a reviewer verifies consistency and confidence. This produces high-quality pseudo-labels for training the text anomaly estimator without requiring manual annotation.
-
Prototype alignment objective with contrastive separation: Apply a loss function that pulls node representations toward their correct prototype (anomalous or normal) while pushing them away from the incorrect one, using cosine similarity margins. The improved system learns a more discriminative representation space with sharper normal-anomalous boundaries.
-
Knowledge distillation for topology-aware representations: Train an MLP to approximate Graph Transformer outputs via KL-divergence alignment over shared topological prototypes. This creates complementary views (one topology-aware, one feature-only) that expose structural anomalies through distributional discrepancy, enabling detection of nodes whose structure deviates from learned normality patterns.
-
Cross-domain zero-shot generalization via transferable prototypes: Learn prototypes from multiple source graphs that capture domain-invariant normality and anomaly patterns. The improved system can be directly applied to unseen target graphs without fine-tuning, achieving state-of-the-art performance across citation, e-commerce, web, and encyclopedia domains.
-
Multi-term anomaly scoring with Z-score standardization: Combine probability estimates, prototype affinity, local neighborhood deviation, and prototype-distribution divergence, with each term independently standardized. The improved system produces calibrated anomaly scores that are robust to scale differences across modalities and domains.
-
Quantitative boundary separability metric (ABS): Introduce a Kullback-Leibler divergence-based metric between normal and anomalous score distributions to measure decision boundary clarity. The improved system can evaluate and optimize for this metric, ensuring well-separated anomaly scores rather than just ranking accuracy.
Sources
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- Boost then Convolve: Gradient Boosting Meets Graph Neural Networks
- Wiki-CS: A Wikipedia-Based Benchmark for Graph Neural Networks
- Zero-shot Generalist Graph Anomaly Detection with Unified Neighborhood Prompts
- A Survey of Generalization of Graph Anomaly Detection: From Transfer Learning to Foundation Models
- LLM-Powered Text-Attributed Graph Anomaly Detection via Retrieval-Augmented Reasoning
- Generalist Graph Anomaly Detection via Prototype-Based Distillation
- Text-Attributed Graph Anomaly Detection via Multi-Scale Cross- and Uni-Modal Contrastive Learning
- When Graph meets Multimodal: Benchmarking and Meditating on Multimodal Attributed Graphs Learning
- Learning on Large-scale Text-attributed Graphs via Variational Inference
- OWLEYE: Zero-Shot Learner for Cross-Domain Graph Data Anomaly Detection
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks