Harbormaster: Evidence-Gated, Replay-Safe Maritime Anomaly Detection on AWS

arXiv:2610.00519 · cs.CR · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Harbormaster: Evidence-Gated, Replay-Safe Maritime Anomaly Detection on AWS".

Elias: Ships broadcast their positions through AIS, and those reports can be false or missing.

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So we’re moving on to the title and authors of this paper, Harbormaster: Evidence-Gated, Replay-Safe Maritime Anomaly Detection on AWS. The title itself really sums up the core technical approach they've taken. Elias I think "Evidence-Gated" immediately tells us they are prioritizing verifiable proof over just throwing a model at the data and hoping for the best.

Priya: And "Replay-Safe" is important because in any system dealing with continuous streams, you have to worry about state corruption or inconsistent results if you have to reprocess historical data later.

Nadia: Right, Priya; that replay safety is critical when the input data, like AIS reports, can be unreliable or arrive out of order. Elias And looking at the authors and where they’re based, it hints at a strong background in both large-scale cloud systems and deep learning research.

Priya: It sounds like a team that understands both the infrastructure plumbing and the statistical modeling required for this kind of complex detection.

Nadia: They definitely do, because they’ve managed to put these very different pieces—deterministic scoring, stream processing, and heavy model inference—into one cohesive production-shaped system on AWS.

Elias: That synthesis is what makes it interesting; usually you see those components handled by separate tools that have their own limitations in communication.

Priya: I’m just wondering about the specific domain focus mentioned in the title, maritime detection, and how that niche application informs their design choices.

Nadia: It shows they’re not trying to build a general-purpose anomaly detector; they are tailoring these robust mechanisms specifically for the challenges of processing real-time ship positional data.

Elias: That specificity means their assumptions about required speeds or space-time geometry must be very well calibrated for that particular domain.

Priya: I think that calibration is exactly what makes the physics-first scoring path so meaningful; it grounds the AI's decision in known physical laws before any learned parameters are even considered.

Nadia: That grounding is what prevents the system from chasing spurious correlations just because a model happened to find a pattern in noisy data.

Elias: So, we’re looking at an architecture that marries strong data governance with practical, domain-specific constraints.

The paper's summary: Nadia: To summarize the Harbormaster paper, it describes a production system on AWS built around three core rules designed to handle untrusted AIS reports reliably. Elias Essentially, they’ve created a pipeline that ensures alerts are attributable, reviewable, and recoverable even when messages repeat or models change.

Priya: They achieve this through several key design choices: running physics checks before the learned model and using an idempotent projection of PostgreSQL in DynamoDB for the authoritative registry.

Nadia: That deterministic scorer needs no learned model to return a result, while the heavy model is handled behind a scale-to-zero endpoint, and they use Kafka to carry change events from Debezium reading logical replication into PostgreSQL.

Elias: The summary also highlights their replay-safe CDC path where the write succeeds only if the item has no applied LSN or holds an older one than the current one, which is crucial for ensuring data integrity during log replay.

Priya: And they detail a five-workflow system that manages everything from initial report checking to building shipping-lane context from historical batches and submitting separate model jobs asynchronously.

Nadia: That comprehensive workflow shows they’ve accounted for the entire lifecycle of an anomaly alert, not just the detection part, which is quite thorough.

Elias: The paper really emphasizes that they're addressing maintenance debt in machine learning code by incorporating validation into the training pipeline and using a structured promotion ladder for candidate models.

Priya: So, they’ve essentially built a system where every decision point is guarded by some form of verification—whether it’s physics, an LSN check, or a model performance probe.

Nadia: That's the central theme: making the process of moving from raw data to an actionable alert transparent and auditable for human review.

The paper's improvements: Elias: Now, let’s talk about the specific improvements they suggest in Harbormaster, which are really about how they make this system production-ready through their gated promotion ladder. Nadia They detail a rigorous four-step ladder for model promotion: holdout gate requiring an AUC of at least zero point eight five and a mean CRPS of at most one point zero, followed by reward-hacking probe and shadow comparison steps.

Priya: The reward-hacking probe, blocking candidates when the reward rises while physical consistency gets worse, seems like a very clever way to prevent models from becoming overly reliant on unrealistic predictions just because they score high in some abstract metric.

Nadia: It’s a direct check against reward hacking; if the model's abstract score improves but its underlying physical plausibility drops, it gets blocked immediately, which is much stronger than just looking at accuracy. Elias That ties directly into their physics-first scoring path because they are linking the learned reward to physical constraints.

Priya: And then they have the shadow comparison step, where they check if the mean absolute score difference between the candidate and a baseline stays below a limit, which is set by the caller.

Nadia: That delta threshold, set at zero point zero five in their tests, shows a commitment to controlling how much deviation we allow before promoting something. Elias The canary weight step adds another layer by moving through weights like five twenty-five fifty and then one hundred while using a burn-rate check to manage the risk at each stage.

Priya: It sounds like they’ve created a very deliberate progression where each step builds confidence based on different forms of evidence—statistical performance, physical consistency, and comparative scoring.

Nadia: That deliberate progression is what makes their system feel much more robust than just training a single massive model and deploying it directly.

Conclusion: Elias: So to wrap up the Harbormaster paper, the main implication is that for complex, real-world data streams like maritime AIS reports, we can build systems where trust is derived from verifiable evidence rather than just accepting a black-box AI output. Nadia That means operators get alerts that are fully attributable and recoverable, which fundamentally changes how we deploy AI in operational environments.

Priya: I think the most significant part for me is that they show how combining physics constraints with rigorous model gating allows us to move toward deploying AI in safety-critical areas with a much higher degree of confidence.

Nadia: They provide a practical framework showing exactly how to handle the inherent uncertainty of noisy inputs by creating guardrails at every stage of the pipeline. Elias Overall, it’s an interesting blueprint for building resilient pipelines that prioritize data integrity and operational recoverability over sheer predictive power.

Priya: I think the idea that replaying any prefix or suffix leaves each projected key at its highest applied LSN really solidifies how much control we regain over the historical state of the system.

Nadia: Indeed, Priya; Harbormaster is a significant piece of work because it shows how to engineer production-grade AI infrastructure that respects both computational limits and real-world physical constraints.

Arun Sharma

University of Minnesota

cs.CR

Submitted: 2026-09-30

Updated: 2026-09-30

Code: https://github.com/arunshar/harbormaster-maritime-ai

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: Ships broadcast their positions through AIS, and those reports can be false or missing.

Key concepts

Physics-First Scoring
This initial stage uses basic geometry and speed calculations to filter reports before any complex AI model runs. It checks if a vessel's reported movement is physically possible based on known speeds and distances, discarding improbable jumps or excessively fast movements immediately.
Replay-Safe Change Data Capture (CDC)
The system uses PostgreSQL for authoritative data and Debezium to stream changes into Kafka. A worker ensures every update is applied idempotently, using a commit protocol that guarantees the state of the projected data matches the highest applied change log, preventing errors during reprocessing.
Evidence-Gated Promotion
A candidate machine learning model must pass a series of rigorous tests—including AUC scores and reward-hacking probes—before it can be used in production. This ladder ensures that only models demonstrating consistent performance and physical accuracy are promoted to handle real traffic.

Terminology

Summary

Ships broadcast their positions through AIS, and those reports can be false or missing. This paper describes Harbormaster, a production-shaped system on AWS that follows three rules.

How it works

Harbormaster is built around answering the design question: How can a cloud pipeline turn untrusted AIS reports into alerts that stay attributable, reviewable, and recoverable when messages repeat and when models change? The system makes three core design choices: a fast physics check runs before the heavier model, a deterministic scorer needs no learned model to return a result, and PostgreSQL holds the authoritative registry while its DynamoDB read store is an idempotent projection of the change log.

The system operates through five workflows triggered by different events. These include Workflow 1 checking each incoming vessel report, Workflow 2 saving an analyst’s review label, Workflow 3 carrying committed registry changes into serving copies, Workflow 4 building shipping-lane context from historical batches, and Workflow 5 submitting a separate model job to an asynchronous endpoint.

Physics-First Scoring

The scoring front door uses physics as its first source of evidence. This involves several checks that run before any learned model is engaged. Key components include:

  1. A required-speed check and a stream-layer physics gate that run before any learned model.

  2. The gap detector uses space-time prisms to merge candidate gaps (Section 4).

The required speed and the stream gate use closed-form functions of reported fixes and reference data. For consecutive reports A and B, the system computes the haversine distance dAB on a sphere of radius R and the required speed vreq. The plausibility pphys equals 1 when a vessel at the 4P(a)A B L = vmax (tB − tA) dAB / a = L/2 b = 1/2 p(L squared − d 2AB). The Flink job forwards a report to the scorer only when pphys ≥ 0.3. This gate delays scoring of fast jumps and drops every report whose required speed exceeds vmax/0.3, which is about 83 knots.

Replay-Safe Change Data Capture

The system employs a replay-safe CDC path using PostgreSQL 16 on RDS as the authoritative source, with Debezium reading committed changes through logical replication into Kafka. The registry worker takes three actions for each event: attempting a conditional write of a whole item to the DynamoDB read store, removing the stale Redis cache entry, and appending a row to an Iceberg audit table.

The commit protocol ensures idempotency: The worker applies every event in a batch, flushes every sink, and only then commits its Kafka offsets. The write succeeds only when the item has no applied LSN or holds an older one, as defined by the guard: APPLY(k, l, v) ⇐⇒ l⋆(k) = ⊥ ∨ l⋆(k) < l (Equation (12)). This ensures that replaying any prefix or suffix of the change log leaves each projected key at the value of its highest applied LSN.

Evidence-Gated Promotion

A candidate model must pass a rigorous evidence-gated promotion ladder before receiving traffic. The ladder has up to four steps:

  1. Holdout gate, requiring an AUC of at least 0.85, a mean CRPS of at most 1.0, and a calibration ratio inside [0.8, 1.2].

  2. Reward-hacking probe, which blocks a candidate whose reward rises while its physical consistency gets worse: block if R¯cand > R¯base ∧ ucand > ubase ∨ s¯cand < s¯base (Equation (17)).

  3. Shadow comparison, passing when the mean absolute score difference is at most a limit δsh that the caller sets (set to 0.05).

  4. Canary weight, which moves the candidate through weights 5, 25, 50, and 100 while a burn-rate check guards each step.

Scale-to-Zero Serving

The heavier model runs behind a SageMaker asynchronous inference endpoint designed for scale-to-zero. The endpoint uses two policies: Track (steering ApproximateBacklogSizePerInstance toward a setpoint of 5) and Wake (setting capacity to exactly 1 when the average H¯k of HasBacklogWithoutCapacity is at least 1 in two consecutive 60 s periods). This design includes two dimension bugs that disabled scaling paths without error, demonstrating how live checks expose latent issues.

Evaluation

The paper reports a scoped evaluation where every result names where it ran, including BOUNDED-AWS-WINDOW (private evidence), LOCAL (local stack), SYNTHETIC-DRILL (deterministic drill), and PUBLIC-REPO results. Key measurements include:

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that could be made to AI systems by incorporating Harbormaster's design principles, and what these improved systems could achieve:


  1. The Improved System: A Replay-Safe Maritime Anomaly Detection Pipeline (Harbormaster Inspired)

  2. Specific Improvements:

Ease the development of production-grade anomaly detection systems by replacing monolithic, untrustworthy models with a verifiable, evidence-gated scoring path.

  1. What the Improved System Can Do:

  2. Can process untrusted, noisy data streams (like AIS reports) in real-time while ensuring that any resulting alert is fully attributable to a specific rule or model and can be recovered perfectly after any system failure or data replay without state corruption.

  3. Can maintain a deterministic state projection of historical data (e.g., vessel positions) by enforcing strict, LSN-guarded change data capture (CDC) mechanisms, guaranteeing that replaying any prefix or suffix of the log results in the exact same derived state as the highest applied event.

  4. Can safely deploy and promote machine learning models through a rigorous, evidence-based promotion ladder (Holdout Gate → Reward-Hacking Probe → Shadow Comparison → Canary), preventing deployment of models that show performance degradation (drift) or reward hacking without passing explicit checks against physical constraints.

  5. Can incorporate physics-informed constraints directly into the scoring path, where mandatory physics checks run before any learned model inference, ensuring that anomalous events are grounded in real-world motion plausibility (e.g., required speed limits and space-time prism feasibility) before complex models even consider the report.

  6. Can automatically detect and manage data drift (input shift, calibration shift, concept drift) by comparing current input samples against a reference set using statistical measures like Population Stability Index (PSI) and Kolmogorov-Smirnov statistics, triggering alerts when model performance begins to degrade.

  7. Can operate efficiently at scale-to-zero serving capacities for heavy models by utilizing asynchronous inference endpoints and sophisticated scaling policies that dynamically manage resource allocation based on request backlog while ensuring rapid wake-up from zero instances, minimizing idle costs.

  8. Can provide high confidence in its outputs by calculating a comprehensive severity score by fusing evidence from multiple detectors (e.g., gap detection, corridor severity, speed severity) using a noisy-OR rule and an analyst review threshold, ensuring that alerts are only generated when the combined evidence meets a high confidence level for human review.

  9. Can provide transparent debugging and auditing by labeling every measurement with its precise scope (bounded window, local run, synthetic drill), allowing researchers to definitively trace any observed result back to the exact environment and input data used for validation.

Abstract

Ships broadcast their positions through the Automatic Identification System (AIS), and those reports can be false or missing. An operator who acts on an anomaly alert needs that alert to be attributable, reviewable, and recoverable after a failure. This paper describes Harbormaster, a production-shaped system on AWS that follows three rules. Physics checks run on every report before any learned model runs. The DynamoDB read store is an idempotent projection of PostgreSQL, and a guard on the log sequence number (LSN) of each change protects every write to it. A candidate model must pass a holdout gate and a shadow comparison before a canary gives it traffic, and a burn-rate check guards each canary step. The paper proves that replaying any prefix or suffix of the change log leaves each projected key at the value of its highest applied LSN. It also notes effects this result does not cover, such as repeated cache invalidations and repeated audit rows. In a bounded AWS window, a one-hour soak returned 35,999 HTTP 200 responses to 36,000 requests, with a 95th-percentile (p95) client latency of 142.751 ms. In a separate bounded AWS run of 900 s, the stream path received a burst of 400 records/s, and its consumer lag later drained to zero. Every number in the paper carries a label that says where it was measured, and the paper lists the parts of the design that were never built.

Sources

Related papers