LOB-ID: Evaluating Synthetic Market Data by Inception Distances

arXiv:2608.13082 · q-fin.CP, cs.AI, cs.CE · Submitted 2026-08-13 · Read on arXiv

Andreea Bacalum, Zhuohan Wang, Ollie Olby, Martin Garaj, Namid Stillman

Simudyne · King's College London

q-fin.CP, cs.AI, cs.CE

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 8 pages, 3 figures, 1 table

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: The paper introduces LOB-ID (Limit Order Book Inception Distance), an embedding-based framework for evaluating synthetic limit orderbook (LOB) data.

Terminology

Summary

The paper introduces LOB-ID (Limit Order Book Inception Distance), an embedding-based framework for evaluating synthetic limit orderbook (LOB) data. The authors argue that existing evaluation methods for generative LOB models, which typically rely on stylised facts and selected market statistics, are informative but partial, since a sequence can match a wide range of individual statistics and still fail to behave like a coherent market trajectory. These methods miss the underlying joint structure of the data, including the dependence between successive book states, between liquidity at the touch and in the deeper levels, between the bid and ask sides of the book, and between all of these and the prevailing market regime.

To address this, the authors adapt the Fréchet Inception Distance (FID) and the Monge Inception Distance (MIND) from image-generation evaluation to LOB data. They train a DeepLOB network on four months of Level-2 order-book data for five equities listed on the Hong Kong Stock Exchange (HKEX). After removing the LSTM and classification head, they use the network's Inception representation as a domain-specific feature extractor, mapping a window of T = 100 successive event-indexed Level-2 book states, x ∈ RT ×40, to a fixed 96-dimensional feature vector. FID compares Gaussian approximations of the embedding distributions, while MIND computes a scaled average of the squared 2-Wasserstein distances between their one-dimensional projections along random unit directions, using 256 random projections and a scaling factor of alpha = 3d = 288.

The paper makes four contributions. First, it introduces LOB-ID, adapting FID and MIND to LOB data. Second, it shows that MIND is consistent across nearby trading days, distinguishes between instruments, responds monotonically to controlled perturbations, and stabilises at later stages of DeepLOB training. Third, it analyses the adversarial weaknesses of FID and selected statistic-based LOB evaluations, showing that MIND remains sensitive to the loss of joint structure they miss. Fourth, it compares five generative models—Zero Intelligence (ZI), Compound Hawkes (CH), LOBGAN, LOBS5, and DiffLOB—on a common dataset using LOB-ID, stylised facts, and existing benchmark metrics.

In the stability analysis, pairwise MIND distances across all stock-day pairs in the test period show that "the off-diagonal entries within each 6 × 6 stock block are consistently dark, showing that MIND assigns low distances to different days from the same instrument. In contrast, the cross-stock blocks are substantially lighter, revealing a clear separation between instruments."

In the sensitivity analysis, the authors apply four perturbations of increasing strength epsilon to historical LOB data: a linear trend, a random walk, a price jump, and row deletion. For all four perturbations, MIND is zero when epsilon = 0 and rises steadily as the distortion becomes stronger, with no reversals.

In the adversarial attacks, the first attack targets FID by maximising the shape loss while keeping the embedding mean and covariance close to real values. The results show that FID increases from a real–real noise floor of 0.17 to 0.59, giving RFID = 3.5. MIND increases from 0.54 to 40.3, yielding RMIND = 74.5, approximately 21 times the FID ratio. The second attack rearranges deep-book values (levels 2 to 10) while preserving Level-1 data and marginal distributions. The authors find that the distributional statistics and impact-response components from LOB-Bench do not detect it, with the attack sitting 95% ± 1% below the noise floor across the 18 stock-days. However, LOB-ID, using MIND, is able to detect the attack, with the attack sitting 75% ± 17% above the noise floor.

In the checkpoint analysis, the authors evaluate MIND scores at five checkpoints during DeepLOB training. They find that as DeepLOB training approaches convergence, the MIND scores also converge towards stable plateaus across all generator configurations, and the relative ordering of the generator configurations remains unchanged across checkpoints.

In the model comparison, the authors evaluate five generative models using stylised facts, LOB-Bench distances, and LOB-ID. The results show that among non-oracle generators, "null-conditioned DiffLOB is strongest overall. It reproduces the most stylised facts, with 7.33 ± 0.69, and records the lowest L1 and LOB-ID distances, ranking second only on Wasserstein-1. LOBS5 follows, with the lowest Wasserstein-1 distance and second place under both L1 and LOB-ID. The separation is clearest under LOB-ID, which orders the models as DiffLOB (null), LOBS5, LOBGAN, CH and ZI." Future-conditioned DiffLOB, which receives realised future regime variables, attains the best values on all three distance measures and is reported as an upper reference point.

The authors conclude that LOB-ID provides a stable measure of similarity between real and synthetic LOB data, and that it captures temporal and cross-level relationships that individual statistics may miss. They note that LOB-ID provides a reusable distance-based alternative to the LOB-Bench discriminator, and that it produces an ordering consistent with the dependencies represented by the tested models. The paper acknowledges limitations, including that LOB-ID depends on the choice of embedding network and that the current embedding is constructed from Level-2 book states, it evaluates temporal and cross-level trajectory structure rather than the full message-level process.

Improvements for AI systems

Improvements to AI Systems:

  1. Generative Model Evaluation with Joint-Structure Sensitivity: Replace or augment existing evaluation metrics (e.g., LOB-Bench, stylised facts) in generative models for financial time series with LOB-ID (specifically MIND). This enables AI systems to detect subtle losses of joint dependence—between book levels, bid-ask sides, and temporal states—that individual statistics miss. The improved system can rank generative models more reliably, as demonstrated by the clear separation of DiffLOB, LOBS5, LOBGAN, CH, and ZI.

  2. Adversarial Robustness in Synthetic Data Validation: Integrate MIND-based detection into AI pipelines that validate synthetic market data. The system can now catch adversarial attacks that preserve marginal distributions and Level-1 data but rearrange deep-book liquidity—attacks that fool existing distributional statistics and impact-response benchmarks. This makes AI-based market simulation systems more secure against data poisoning or subtle manipulation.

  3. Training-Stage Monitoring for Representation Learning: Use MIND as a checkpoint-based diagnostic during the training of embedding networks (e.g., DeepLOB). The improved AI system can monitor convergence of the feature extractor itself, ensuring that the learned representations stabilise and that relative model rankings remain consistent. This provides early stopping or retraining signals when the embedding is not yet capturing stable market structure.

  4. Cross-Instrument and Cross-Day Transfer Learning: Leverage the demonstrated stability of MIND across days for the same instrument and its sensitivity to cross-stock differences. An improved AI system can use LOB-ID as a domain-adaptation metric—automatically detecting when a generative model trained on one stock fails to generalise to another, or when a model’s output drifts across trading days, enabling dynamic recalibration.

  5. Perturbation-Aware Quality Control: Apply the monotonic sensitivity of MIND to controlled perturbations (trend, random walk, price jump, row deletion) as a calibration tool. An AI system can now quantify the degree of structural distortion in synthetic LOB data with a single scalar, allowing for automated thresholding—e.g., rejecting generated data that exceeds a pre-defined MIND distance from real data, ensuring quality control in high-frequency trading backtests.

  6. Unified Evaluation for Conditional Generative Models: Use LOB-ID to compare null-conditioned vs. future-conditioned generative models (as with DiffLOB). The improved AI system can now quantify the added value of conditioning on future regime variables, providing a principled way to decide whether to include auxiliary information in generative models, based on the reduction in MIND distance to real data.

  7. Feature Extraction Reusability: Adopt the DeepLOB-based Inception representation (96-dimensional, windowed over 100 event-indexed states) as a reusable, domain-specific feature extractor for other LOB tasks—such as anomaly detection, regime classification, or reinforcement learning state representation—since it already encodes temporal and cross-level joint structure that individual statistics miss.

Sources

Related papers