Incremental Evaluation and Training in Relational Deep Learning

arXiv:2608.13023 · cs.LG, cs.DB · Submitted 2026-08-13 · Read on arXiv

Jakub Peleška, Gustav Šír

Czech Technical University in Prague

cs.LG, cs.DB

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

Terminology

Summary

Paper: arXiv:2608.13023v1 [cs.LG], 13 Aug 2026


Relational Deep Learning (RDL) models multi-tabular databases as temporal heterogeneous graphs to enable end-to-end representation learning. However, the paper identifies a critical gap: "prevailing RDL evaluation practices rely on static, single-episode dataset snapshots, overlooking the continuous, time-evolving nature of real-world databases. Consequently, current RDL benchmarks fail to capture how model performance changes as new data accumulates over time."

The authors note that databases are not static repositories; they dynamically expand and shift over time, causing models trained on historical snapshots to become incrementally obsolete. While existing frameworks like RelBench and ReDeLEx correctly use timestamps for chronological train-test splits to prevent temporal leakage, this snapshot-based evaluation assesses models in a single, isolated episode, fundamentally failing to capture how model performance degrades over time, e.g., due to temporal concept drift.

The paper introduces an incremental, multi-episodic evaluation and training paradigm that:

  1. Simulates real-world database growth by transitioning RDL benchmarking from static snapshots to continuous, horizon-shifting evaluation

  2. Empirically demonstrates that temporal concept drift is a prevalent factor in standard RDL tasks using state-of-the-art models and large-scale datasets

  3. Defines and benchmarks multiple incremental fine-tuning strategies, showing knowledge transfer is highly effective and computationally superior to from-scratch retraining in the continuous RDL setting

  4. Proposes an exponential decay-based evaluation metric that prioritizes near-future predictive accuracy, closely aligning benchmark evaluation with practical operational priorities

A relational database R is formally defined as a finite set of relations, where each relation has a heading (attributes) and a body (tuples). The relational entity graph is defined as G = (V, E, φ, ψ, τ), where:

  • Nodes V are all tuples from all relations

  • Edges E are defined by primary key–foreign key (PK-FK) relationships

  • φ maps nodes to node types (source relations)

  • ψ maps edges to edge types (PK-FK attribute pairs)

  • τ assigns timestamps to nodes

RDL models follow a four-stage pipeline: "(i) a table-level attribute encoder initializes node embedding matrices... (ii) a table-level tabular model employs tabular architectures to refine these embeddings... (iii) a graph neural model applies message-passing based on the defined edge types... and (iv) a task-specific model head maps the final node embeddings to predictions."

Tasks are formalized via a task table Ttask with tuples (v, y, t), where v is the target entity, y is the label, and t is the anchor time. Two categories exist:

  • Static tasks: predict missing values in a fixed snapshot

  • Temporal tasks (forecasting): predict future behaviors using queries Q∆w(t) with window size ∆w, requiring strict chronological splits

The paper introduces a fixed time increment ∆I ≥ ∆w representing the temporal stride between consecutive model updates. Evaluation proceeds over a sequence of monotonically increasing evaluation horizons: T∆I = ti ti = ti−1 + ∆I, ti ≤ tval, i ≥ 1. At each episode, new training examples from anchor times in (ti−1, ti] are revealed, and this strict append-only growth prevents temporal leakage, ensuring the model never accesses information beyond its current validation horizon.

Four regimes are benchmarked:

  1. From Scratch (Baseline): Trained from scratch using all available data up to tval, serving as a (computationally expensive) baseline

  2. Finetune (Cumulative): Retains weights from episode i−1 and fine-tunes on the entire historical dataset up to tval

  3. Finetune (Incremental): Updates the episode i−1 model exclusively on the newest data increment to maximize training speed and immediate responsiveness

  4. Finetune (Upsampled): A hybrid strategy fine-tuning on a distribution that upsamples the most recent increment while retaining historical data to mitigate catastrophic forgetting

The paper introduces a time-weighted exponential decay metric: given N future evaluation windows and decay rate α ∈ [0,1], each sample s in window i receives weight:

Ws,i = (1−α)(i−1) / Σⱼ Sj(1−α)(j−1)

This reduces to a uniform mean when α = 0, while α → 1 strictly isolates performance on the immediate next increment.

The evaluation uses 12 tasks (regression and binary classification) on 4 databases from RelBench:

  • rel-avito: ad-ctr, user-clicks, user-visits

  • rel-f1: driver-position, driver-dnf, driver-top3

  • rel-stack: post-votes, user-badge, user-engagement

  • rel-trial: site-success, study-adverse, study-outcome

The architecture uses tabular ResNet node encoders (via PyTorch Frame) paired with GraphSAGE with 2 message-passing layers (hidden dimensionality 128, sum aggregation). Training uses Adam (learning rate 0.001), batch size 128, maximum 2000 steps per episode, with 5 random seeds.

The paper confirms the presence of concept drift through a general trend of performance degradation over time across tasks. Figure 4 shows that From Scratch models trained in the multi-episodic setting exhibit performance degradation over time. The data analysis identifies two principal growth patterns: roughly linear (rel-avito ad-ctr, rel-f1 driver-position) and exponential (rel-stack user-engagement, rel-trial study-adverse). With the exception of rel-avito ad-ctr, all remaining tasks exhibit clear evidence of temporal concept drift.

The fine-tuning regimes consistently match or outperform the baseline, an effect most pronounced in the rel-stack user-engagement and rel-trial study-adverse tasks. After just 100 training steps (few-shot scenario), pretrained models quickly achieve performance on par with or superior to a fully trained from scratch baseline, underscoring the computational efficiency and predictive potential of continuously fine-tuned models.

Table 2 shows that except for driver-position, varying the decay factor preserves the optimal training regime rankings. However, absolute metric values shift in distinct patterns: post-votes and ad-ctr exhibit minimal variation, study-adverse shows a consistent performance drop, and driver-position improves at α = 0.3.

The paper's core findings are:

  1. We validated the widespread prevalence of temporal concept drift in standard RDL tasks, highlighting why snapshot-based testing is insufficient

  2. "Incremental fine-tuning strategies successfully mitigate this degradation. Transferring knowledge from previous training episodes consistently matched or exceeded the performance of expensive from-scratch retraining, often adapting within just a few optimization steps"

  3. Introducing a time-weighted exponential decay metric allowed for a more realistic assessment of a model's utility by properly prioritizing near-future predictive accuracy

The paper acknowledges several limitations:

  • The setup currently focuses on append-only database growth, while real-world systems involve data updates and deletions (CRUD operations)

  • Mitigating catastrophic forgetting over extended horizons remains a challenge

  • Extending this continuous evaluation to emerging monolithic RDL foundation models remains an open challenge

Improvements for AI systems

Improvements to AI Systems:

  1. Continuous Temporal Adaptation Module: Implement an incremental learning scheduler that automatically detects concept drift via performance monitoring across evaluation horizons. The system would trigger fine-tuning on new data increments only when drift exceeds a threshold, reducing unnecessary computation while maintaining accuracy.

  2. Time-Aware Loss Weighting: Integrate the exponential decay metric directly into the training objective. The AI system would assign higher loss weights to recent training samples, prioritizing near-future predictive accuracy during optimization rather than treating all historical data equally.

  3. Hybrid Memory Replay with Upsampling: Build a dynamic replay buffer that automatically upsamples the most recent data increment while retaining a stratified sample of older data. The system would adjust the upsampling ratio based on measured drift severity, balancing adaptation speed against catastrophic forgetting.

  4. Few-Shot Episode Initialization: Leverage pretrained weights from the previous episode as initialization for new tasks. The system would perform rapid adaptation within 100–200 optimization steps, reducing training time by an order of magnitude compared to from-scratch retraining, while matching or exceeding baseline accuracy.

  5. Horizon-Aware Early Stopping: Implement validation on the immediate next time window (rather than a static holdout set) to determine when to stop training each episode. This ensures the model is optimized for the specific upcoming data distribution, not a generic historical average.

  6. Drift-Adaptive Architecture Selection: Automatically choose between fine-tuning regimes (cumulative vs. incremental vs. upsampled) based on detected growth patterns (linear vs. exponential) and task-specific drift characteristics. For exponential-growth tasks, the system would favor upsampled fine-tuning; for linear-growth tasks, cumulative fine-tuning.

  7. Multi-Episode Knowledge Distillation: After each episode, distill the current model's knowledge into a compact student model that captures temporal invariants. This student model serves as a regularization target during subsequent fine-tuning, reducing catastrophic forgetting over extended horizons.

Capabilities of the Improved AI System:

  • Deploys in production environments with continuously growing databases, maintaining predictive accuracy over months or years without full retraining, automatically adapting to concept drift.

  • Reduces computational costs by 80–95% compared to periodic from-scratch retraining, while achieving equal or better performance on near-future predictions.

  • Provides calibrated confidence about its own temporal relevance, flagging when its predictions become unreliable due to drift and requesting new data or retraining.

  • Handles both linear and exponential data growth patterns gracefully, adjusting its learning strategy based on observed data dynamics.

  • Operates in few-shot regimes, adapting to new temporal distributions within minutes (100–200 optimization steps) rather than hours or days.

  • Prioritizes actionable near-term predictions (e.g., next week's user behavior, next month's clinical trial outcomes) over long-term averages, aligning with real-world operational needs.

Abstract

Relational Deep Learning (RDL) models multi-tabular databases as temporal heterogeneous graphs to enable end-to-end representation learning. However, prevailing RDL evaluation practices rely on static, single-episode dataset snapshots, overlooking the continuous, time-evolving nature of real-world databases. Consequently, current RDL benchmarks fail to capture how model performance changes as new data accumulates over time. To address this limitation, we introduce an incremental, multi-episode evaluation and training paradigm to assess and improve the temporal robustness and adaptability of state-of-the-art RDL models. Using established large-scale datasets, we examine data evolution and model training dynamics, demonstrating that temporal concept drifts occur in the majority of predictive tasks. We present multiple incremental training regimes for fine-tuning the models and demonstrate that transfer learning is both feasible and highly effective in the RDL setting. Alongside a new temporal evaluation metric that prioritizes near-future accuracy, we show that our incrementally fine-tuned models consistently outperform the standard, expensive, from-scratch trained baselines.

Sources

Related papers