Harbormaster: Evidence-Gated, Replay-Safe Maritime Anomaly Detection on AWS

summary

Video file (mp4)

The gist

Ships broadcast their positions through AIS, and those reports can be false or missing.

In short

Harbormaster is a system designed to detect untrusted AIS vessel reports by combining fast physics checks with a heavy machine learning model. It ensures alerts are attributable and recoverable through replay-safe change data capture, evidence-gated model promotion, and scale-to-zero serving on AWS. This addresses the challenge of processing unreliable maritime data reliably.

Key concepts

Physics-First Scoring
This initial stage uses basic geometry and speed calculations to filter reports before any complex AI model runs. It checks if a vessel's reported movement is physically possible based on known speeds and distances, discarding improbable jumps or excessively fast movements immediately.
Replay-Safe Change Data Capture (CDC)
The system uses PostgreSQL for authoritative data and Debezium to stream changes into Kafka. A worker ensures every update is applied idempotently, using a commit protocol that guarantees the state of the projected data matches the highest applied change log, preventing errors during reprocessing.
Evidence-Gated Promotion
A candidate machine learning model must pass a series of rigorous tests—including AUC scores and reward-hacking probes—before it can be used in production. This ladder ensures that only models demonstrating consistent performance and physical accuracy are promoted to handle real traffic.

Terminology used across episodes

This episode discusses

The paper

Harbormaster: Evidence-Gated, Replay-Safe Maritime Anomaly Detection on AWS · Read on arXiv

Arun Sharma

University of Minnesota

Ships broadcast their positions through the Automatic Identification System (AIS), and those reports can be false or missing. An operator who acts on an anomaly alert needs that alert to be attributable, reviewable, and recoverable after a failure. This paper describes Harbormaster, a production-shaped system on AWS that follows three rules. Physics checks run on every report before any learned model runs. The DynamoDB read store is an idempotent projection of PostgreSQL, and a guard on the log sequence number (LSN) of each change protects every write to it. A candidate model must pass a holdout gate and a shadow comparison before a canary gives it traffic, and a burn-rate check guards each canary step. The paper proves that replaying any prefix or suffix of the change log leaves each projected key at the value of its highest applied LSN. It also notes effects this result does not cover, such as repeated cache invalidations and repeated audit rows. In a bounded AWS window, a one-hour soak returned 35,999 HTTP 200 responses to 36,000 requests, with a 95th-percentile (p95) client latency of 142.751 ms. In a separate bounded AWS run of 900 s, the stream path received a burst of 400 records/s, and its consumer lag later drained to zero. Every number in the paper carries a label that says where it was measured, and the paper lists the parts of the design that were never built.

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Harbormaster: Evidence-Gated, Replay-Safe Maritime Anomaly Detection on AWS".

Elias: Ships broadcast their positions through AIS, and those reports can be false or missing.

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So we’re moving on to the title and authors of this paper, Harbormaster: Evidence-Gated, Replay-Safe Maritime Anomaly Detection on AWS. The title itself really sums up the core technical approach they've taken. Elias I think "Evidence-Gated" immediately tells us they are prioritizing verifiable proof over just throwing a model at the data and hoping for the best.

Priya: And "Replay-Safe" is important because in any system dealing with continuous streams, you have to worry about state corruption or inconsistent results if you have to reprocess historical data later.

Nadia: Right, Priya; that replay safety is critical when the input data, like AIS reports, can be unreliable or arrive out of order. Elias And looking at the authors and where they’re based, it hints at a strong background in both large-scale cloud systems and deep learning research.

Priya: It sounds like a team that understands both the infrastructure plumbing and the statistical modeling required for this kind of complex detection.

Nadia: They definitely do, because they’ve managed to put these very different pieces—deterministic scoring, stream processing, and heavy model inference—into one cohesive production-shaped system on AWS.

Elias: That synthesis is what makes it interesting; usually you see those components handled by separate tools that have their own limitations in communication.

Priya: I’m just wondering about the specific domain focus mentioned in the title, maritime detection, and how that niche application informs their design choices.

Nadia: It shows they’re not trying to build a general-purpose anomaly detector; they are tailoring these robust mechanisms specifically for the challenges of processing real-time ship positional data.

Elias: That specificity means their assumptions about required speeds or space-time geometry must be very well calibrated for that particular domain.

Priya: I think that calibration is exactly what makes the physics-first scoring path so meaningful; it grounds the AI's decision in known physical laws before any learned parameters are even considered.

Nadia: That grounding is what prevents the system from chasing spurious correlations just because a model happened to find a pattern in noisy data.

Elias: So, we’re looking at an architecture that marries strong data governance with practical, domain-specific constraints.

The paper's summary: Nadia: To summarize the Harbormaster paper, it describes a production system on AWS built around three core rules designed to handle untrusted AIS reports reliably. Elias Essentially, they’ve created a pipeline that ensures alerts are attributable, reviewable, and recoverable even when messages repeat or models change.

Priya: They achieve this through several key design choices: running physics checks before the learned model and using an idempotent projection of PostgreSQL in DynamoDB for the authoritative registry.

Nadia: That deterministic scorer needs no learned model to return a result, while the heavy model is handled behind a scale-to-zero endpoint, and they use Kafka to carry change events from Debezium reading logical replication into PostgreSQL.

Elias: The summary also highlights their replay-safe CDC path where the write succeeds only if the item has no applied LSN or holds an older one than the current one, which is crucial for ensuring data integrity during log replay.

Priya: And they detail a five-workflow system that manages everything from initial report checking to building shipping-lane context from historical batches and submitting separate model jobs asynchronously.

Nadia: That comprehensive workflow shows they’ve accounted for the entire lifecycle of an anomaly alert, not just the detection part, which is quite thorough.

Elias: The paper really emphasizes that they're addressing maintenance debt in machine learning code by incorporating validation into the training pipeline and using a structured promotion ladder for candidate models.

Priya: So, they’ve essentially built a system where every decision point is guarded by some form of verification—whether it’s physics, an LSN check, or a model performance probe.

Nadia: That's the central theme: making the process of moving from raw data to an actionable alert transparent and auditable for human review.

The paper's improvements: Elias: Now, let’s talk about the specific improvements they suggest in Harbormaster, which are really about how they make this system production-ready through their gated promotion ladder. Nadia They detail a rigorous four-step ladder for model promotion: holdout gate requiring an AUC of at least zero point eight five and a mean CRPS of at most one point zero, followed by reward-hacking probe and shadow comparison steps.

Priya: The reward-hacking probe, blocking candidates when the reward rises while physical consistency gets worse, seems like a very clever way to prevent models from becoming overly reliant on unrealistic predictions just because they score high in some abstract metric.

Nadia: It’s a direct check against reward hacking; if the model's abstract score improves but its underlying physical plausibility drops, it gets blocked immediately, which is much stronger than just looking at accuracy. Elias That ties directly into their physics-first scoring path because they are linking the learned reward to physical constraints.

Priya: And then they have the shadow comparison step, where they check if the mean absolute score difference between the candidate and a baseline stays below a limit, which is set by the caller.

Nadia: That delta threshold, set at zero point zero five in their tests, shows a commitment to controlling how much deviation we allow before promoting something. Elias The canary weight step adds another layer by moving through weights like five twenty-five fifty and then one hundred while using a burn-rate check to manage the risk at each stage.

Priya: It sounds like they’ve created a very deliberate progression where each step builds confidence based on different forms of evidence—statistical performance, physical consistency, and comparative scoring.

Nadia: That deliberate progression is what makes their system feel much more robust than just training a single massive model and deploying it directly.

Conclusion: Elias: So to wrap up the Harbormaster paper, the main implication is that for complex, real-world data streams like maritime AIS reports, we can build systems where trust is derived from verifiable evidence rather than just accepting a black-box AI output. Nadia That means operators get alerts that are fully attributable and recoverable, which fundamentally changes how we deploy AI in operational environments.

Priya: I think the most significant part for me is that they show how combining physics constraints with rigorous model gating allows us to move toward deploying AI in safety-critical areas with a much higher degree of confidence.

Nadia: They provide a practical framework showing exactly how to handle the inherent uncertainty of noisy inputs by creating guardrails at every stage of the pipeline. Elias Overall, it’s an interesting blueprint for building resilient pipelines that prioritize data integrity and operational recoverability over sheer predictive power.

Priya: I think the idea that replaying any prefix or suffix leaves each projected key at its highest applied LSN really solidifies how much control we regain over the historical state of the system.

Nadia: Indeed, Priya; Harbormaster is a significant piece of work because it shows how to engineer production-grade AI infrastructure that respects both computational limits and real-world physical constraints.

More episodes

← Home