Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept

arXiv:2510.23994 · cs.LG · Submitted 2025-10-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept".

Jane: The paper was written by Geoffery Agorku, Sarah Hernandez, Hayley Hames and Cade Wagner from University of Arkansas.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everyone. Today we’re cracking open a fresh one from the arXiv preprint server, and the title is a mouthful — "Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept." Jane, when you first saw that title, what jumped out at you?

Jane: Honestly, Tom, the phrase "proof of concept" is what got me. That tells me these researchers at the University of Arkansas are saying, hey, we’ve got an idea, we’ve tested it on a real river, and now we want the world to poke holes in it. And the idea itself is pretty clever — they’re trying to count barges without actually looking at the barges.

Tom: Right, because barges are those big rectangular cargo boxes being pushed along the Mississippi, and they don’t have their own engines or GPS trackers. The tugboat pushing them does, though — it broadcasts its position constantly through something called AIS, the Automatic Identification System.

Jane: Exactly. So the tug is like a delivery truck, and the barges are the trailers. You know where the truck is, but you don’t know how many trailers it’s hauling unless someone tells you. And right now, on most of America’s inland waterways, nobody’s telling you in real time.

Tom: And that’s a huge deal, because the paper points out that over six hundred million tons of cargo move on these rivers every year. But the people managing locks, ports, and supply chains are basically flying blind when it comes to knowing how many barges are coming their way.

Jane: So the researchers came up with a workaround. They used satellite imagery to actually see the barges and count them — that’s their ground truth. Then they matched those satellite snapshots to the tugboat’s AIS data at the exact same moment. That gives them a labeled dataset: here’s what the tug’s movement looked like, and here’s how many barges it was pushing.

Tom: And once they had that pairing, they could train machine learning models to predict barge count using only the AIS data — no satellite needed anymore. That’s the clever part, right? The satellite is just the teacher; the AIS data is what does the real work in the field.

Jane: Precisely. And the implications are big for something called Maritime Domain Awareness. If you can predict barge counts in real time from data you already have, you can help lock operators schedule lockages more efficiently, help ports plan their crane and dock assignments, and give shippers the visibility they’ve been missing.

Tom: So it’s not just an academic exercise — this could actually make the supply chain run smoother. But before we get too deep into the weeds, let’s talk about how they actually built this thing. That’s coming up next.

Summary: Tom: So we’re back with "Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept," and Jane, I want to get into the meat of the methodology. How did they actually pull this off?

Jane: Well, Tom, they focused on a two hundred forty-mile stretch of the Lower Mississippi River, from Baton Rouge down to the Gulf of Mexico. That’s the busiest waterway corridor in the country. They gathered satellite images from Planet Labs over a few months in early two thousand twenty-four and they pulled AIS data from the Marine Cadastre system.

Tom: And the satellite images are how they counted the actual barges, right?

Jane: Yes, but here’s the thing — they didn’t just have a human staring at every image. They used a pre-trained computer vision model called YOLO to automatically detect vessels and barges in the satellite scenes. Then two researchers manually verified those detections to make sure the counts were accurate. That gave them twenty-six labeled instances — twenty-six moments where they knew exactly how many barges a tug was pushing.

Tom: Twenty-six samples. That’s not a lot, is it?

Jane: It’s small, and the paper is upfront about that. But for a proof of concept, it’s enough to see whether the approach has legs. The next step was matching those satellite detections to the AIS tracks. They used a spatiotemporal matching procedure — so the AIS points had to be within two minutes of the satellite image timestamp, and the vessel’s path had to intersect with where the satellite saw the tug.

Tom: So you’ve got a tug, you know where it was, you know how many barges it had, and now you’ve got its entire AIS trajectory. What do you do with that?

Jane: That’s where the feature engineering comes in. They extracted over thirty features from each trajectory — things like average speed, speed variability, how often the vessel changed course, how much the heading wobbled, acceleration patterns, even things like the entropy of the course distribution. The idea is that a tug pushing twelve barges moves very differently than a tug pushing two barges.

Tom: And that makes sense — a bigger tow is heavier, harder to turn, slower to accelerate. So the movement signature should be different.

Jane: Exactly. Then they used a feature selection technique called Recursive Feature Elimination with Cross-Validation to figure out which of those thirty features actually mattered. And they tested six different regression models — from simple linear ones to complex ensemble methods.

Jane: The winner was the Poisson Regressor, which is a model designed for count data — perfect for something like "how many barges." It achieved a Mean Absolute Error of one point nine two barges. So on average, its prediction was off by less than two barges. That’s pretty impressive for a first attempt.

Tom: Less than two barges off, using only the tug’s movement data. That’s the kind of result that makes you sit up and pay attention. But I’m curious — which features actually mattered most? That’s the part that could really tell us something about how these vessels behave.

Improvements: Tom: We’re back with "Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept," and Jane, you mentioned the Poisson Regressor won. But what I really want to know is — what did the feature analysis reveal? What actually predicts barge count?

Jane: Great question, Tom. The two most frequently selected features across all the models were course entropy and a composite feature they called "trip duration times speed variation." Course entropy is basically a measure of how unpredictable the vessel’s heading is over the trip. A tug pushing a huge tow can’t make sharp, erratic turns — it has to plan smooth, gradual maneuvers. So its heading signal is very regular, very low entropy.

Tom: So a low-entropy course means a big tow, and a high-entropy course means a small, nimble one?

Jane: Exactly. And the second feature, trip duration multiplied by speed variation, captures something similar — longer trips with very consistent speeds suggest a heavy, stable tow that doesn’t speed up and slow down much. The paper also found that speed coefficient of variation and mean course change rate were strong predictors. It all points to the same story: bigger tows move more deliberately.

Tom: That’s a really elegant finding. But the paper also acknowledges some limitations, right? I mean, twenty-six samples is small, and they only looked at one river.

Jane: Right, and they’re honest about that. The dataset was imbalanced — they had lots of small tows, like two to four barges, but very few large ones with a dozen or more. That means the model is better at predicting small tows and struggles with the rare big ones. In their test results, you can see the model underpredicting a twelve-barge tow, for instance.

Tom: So what’s the path forward? How do they improve this?

Jane: The paper lays out a clear roadmap. First, they need more data — more samples, more river systems, more vessel types. Second, they want to incorporate environmental conditions like river current, wind, and water levels, because those obviously affect how a tug moves. A tug fighting a strong current will have a different speed profile than one moving with it, regardless of how many barges it’s pushing.

Tom: That’s a really good point. If you don’t account for the river itself, you might misread a tug’s movement.

Jane: Exactly. And third, they mention trying more advanced models — things like recurrent neural networks or graph neural networks that can capture temporal dependencies in the trajectory data. The current models treat each trip as a bag of statistics, but a sequence model could learn how the vessel’s behavior evolves over time.

Tom: So the improvements are about breadth — more data, more context, more sophisticated modeling. This is a proof of concept, but it’s pointing toward something that could actually be deployed. Let’s bring in Lu and Meng to get their take on what this means for the real world.

Lu: I’m really excited about the transferability angle. The paper mentions testing this on other rivers like the Ohio or the Columbia. If the features hold up across different environments, you’ve got a generalizable tool, not just a Mississippi-specific one.

Meng: And from an engineering standpoint, the beauty is that AIS data is already being collected everywhere. You don’t need to install new hardware on barges or tugs. You just need the software to process the data. That makes this scalable in a way that IoT trackers or additional cameras aren’t.

Tom: So the foundation is solid, but there’s real work ahead to make it robust. Let’s wrap this up in our final segment.

Conclusion: Tom: Alright, we’re closing out our discussion of "Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept." Jane, give us the final summary — what did we learn?

Jane: We learned that you can predict how many barges a tugboat is pushing using only its AIS movement data, with an average error of under two barges. The key insight is that bigger tows move differently — they have lower course entropy, steadier speeds, and more deliberate maneuvers. The satellite imagery was just the training tool; the real product runs on data we already collect.

Tom: And the implications? What does this mean for the world?

Jane: It means lock operators could know in advance whether a tow will need a single or double lockage. Ports could plan their resources before a vessel arrives. Shippers could finally get the visibility that’s been missing from inland waterway transport. And it does all this without installing a single new sensor on a barge.

Lu: The scalability is what excites me. This could be the foundation for a real-time national barge tracking system — something that doesn’t exist today.

Meng: And it’s practical. The data infrastructure is already there. This is a software problem, not a hardware problem, and that makes deployment much more feasible.

Tom: Well said. So we’ve got a clever proof of concept with a clear path forward. The team at Arkansas has given us something to build on, and we’ll be watching to see how it evolves. Thanks for joining us, everyone. We’ll be back with the next paper soon.

Geoffery Agorku, Sarah Hernandez, Hayley Hames, Cade Wagner

University of Arkansas

cs.LG

Submitted: 2025-10-28

Updated: 2026-08-18

License: http://creativecommons.org/publicdomain/zero/1.0/

Importance score: 79/100

Key concepts

AIS (Automatic Identification System)
The system used for tracking vessels, where the tugboat broadcasting its position constantly provides data on the location and movement of a barge tow. This data is essential for predicting how many barges are being pulled.
Course Entropy
A measure of how unpredictable a vessel's heading is over the trip. The study found that larger tows have lower entropy because they must follow smooth, gradual maneuvers, unlike smaller, more nimble tows.
Poisson Regressor
A statistical model used in the research designed specifically for count data. It successfully predicted the number of barges with a Mean Absolute Error of 1.92, indicating that its predictions were consistently accurate.
Spatiotemporal Matching
The process used to link satellite images (which show actual barge counts) to the AIS data. The match required the AIS points to be within two minutes of the image timestamp and for the vessel's path to intersect with where the barges were seen.

Terminology

Summary

Summary

This paper introduces a novel method to predict the number of barges being towed by a tug/towboat on inland waterways using Automatic Identification System (AIS) vessel tracking data and Machine Learning (ML). The study addresses the critical challenge that barges are non-self-propelled and unpowered, meaning they lack AIS transponders, creating a black box problem that hampers efficiency, compromises safety, and limits strategic planning across the waterway ecosystem.

The objectives of the paper are threefold: (i) to estimate the number of barges being towed by a tug/towboat with minimal error, leveraging both static vessel geometry and dynamic movement patterns derived from AIS data; (ii) to identify the most influential predictors of barge count; and (iii) to establish a foundation for future spatial generalization of barge count prediction models across diverse riverine environments.

The methodology integrates computer vision, spatiotemporal data fusion, trajectory analysis, and machine learning. Two complementary datasets are required: vessel trajectories from AIS data (including vessel IDs, locations, speeds, and directions) and satellite imagery from commercial sources (e.g., Planet Labs or Maxar). Although the aim is to use only AIS-derived features for prediction, an interim process eases manual sample labeling by applying pre-trained machine learning models (e.g., YOLO) to automatically detect vessel and barge objects from satellite imagery, which are then manually classified and matched to AIS records.

A spatiotemporal matching procedure associates vessels detected in satellite imagery with AIS records. For every satellite scene, AIS messages transmitted within two minutes before or after the image acquisition timestamp are selected. After vessels are identified in imagery using YOLO-based detection and bounding boxes are georeferenced, AIS records are grouped by vessel identifier (e.g., MMSI) and chronologically ordered to form paths. A detection is considered a valid match if its polygon intersects with an AIS trajectory and falls within the defined temporal window.

Once AIS metadata is linked to detected vessels, vessel trips are extracted by identifying stop locations using a density-based clustering algorithm with low-speed thresholds (e.g., <2 knots), minimum stop durations, and spatial clustering of AIS pings. These stops serve as origin and destination points, with each trip defined as the movement segment between two consecutive stop points. The trip that includes the satellite capture time is selected as the matching trip.

A comprehensive set of 30 AIS-derived features capturing vessel geometry, dynamic movement, and trajectory patterns was created. These features were grouped into categories such as speed and acceleration dynamics, directional stability, and operational patterns. The feature set was guided by the premise that vessels towing larger barge configurations tend to exhibit reduced maneuverability, smoother courses, and fewer directional changes. Features include vessel length, width, draft, mean/median/standard deviation of speed, interquartile range of speed, median absolute deviation of speed, maximum/minimum/range of speed, coefficient of variation of speed, percent time at low/high/optimal speed, entropy of speed distribution, mean positive/negative acceleration, standard deviation of acceleration, maximum deceleration, acceleration sign changes, standard deviation of course, entropy of course over ground, standard deviation of turn rate, total absolute course change, mean/standard deviation of difference between course over ground and heading, total trip time, distance traveled, direct straight-line distance, sinuosity index, vessel area, draft to length ratio, and various interaction/polynomial terms.

Six regression models from different algorithmic families were selected for comparison: three ensemble-based models (Random Forest Regressor, Catboost Regressor, AdaBoost Regressor), one kernel-based model (Support Vector Regressor), one regularized linear model (ElasticNet), and one classical count-based model (Poisson Regressor). Recursive Feature Elimination with Cross-Validation (RFECV) was employed for feature selection, iteratively training a model, ranking features by importance, and removing the least influential features while evaluating predictive performance via cross-validation. This procedure was applied separately across all six models, tailoring feature selection to each model's learning characteristics.

A stratified K-fold cross-validation approach was used with 2 folds due to the small dataset size (26 labeled instances). The data was stratified by the number of barges attached to each vessel so that each fold preserved the overall distribution of barge count classes. Model evaluation used Mean Absolute Error (MAE), defined as the absolute error in barge count estimation.

The case study covered a 240-mile section of the Mississippi River from Baton Rouge to the Gulf of Mexico, the nation's most active corridor for waterborne commercial traffic, with data gathered between January and April 2024. High-resolution satellite imagery (3 m) from Planet Labs was used, with a custom API processing AIS tracks and detecting spatiotemporal overlap with Planet imagery tiles. AIS data was recorded at 1-minute intervals across U.S. inland waterways, sourced from the Marine Cadastre Web Tool. Stop identification parameters included a maximum speed threshold of 1.0 knot, minimum stop duration of 60 minutes, and a spatial radius of 300m.

The final annotated dataset consisted of 26 labeled instances, each with an associated barge count, representing moving tows. Smaller barge configurations (e.g., two to four barges) were more frequently observed, while larger configurations (e.g., 12 or more barges) were less common, showing class imbalance which can influence prediction accuracy for less frequent barge counts.

Among the six regression models tested, the Poisson Regressor outperformed all others, achieving an average MAE of 1.920 barges across the two folds, using 12 of the 30 features. The Support Vector Regressor followed closely with an MAE of 1.922 barges using one feature (DUR SOGCV). AdaBoost ranked third with an MAE of 2.784 barges using two features (COG ENT, DIST KM). The Random Forest Regressor produced an MAE of 3.525 barges based on four features (ACC ZC, COG ENT, DIST KM, DUR SOGCV). The Catboost Regressor and ElasticNet exhibited the highest error levels, with MAEs of 4.148 and 4.445 barges respectively; Catboost employed 28 features, whereas ElasticNet used two (SOG PCT LOW, COG STD).

For the Poisson Regressor, in Fold 1 the model achieved an MAE of 2.482, overpredicting some zero-barge cases and underpredicting the single large tow (predicted 8.25 vs. actual 12). Fold 2 produced a lower MAE of 1.301, with most mid-range counts estimated within roughly ±1 barge of their true values. The results indicate that while the Poisson Regressor is well-calibrated for common, low-to-mid tow sizes, it still struggles with rare, high-barge outliers.

Feature importance analysis revealed that Entropy of Course Over Ground (COG ENT) and Trip Duration interacted with Speed Variation (DUR SOGCV) are jointly the most frequently selected features across all regression models. COG ENT quantifies the complexity of a vessel's heading profile; large barge tows, with their high inertia and limited turning radius, force tugboat operators to plan and execute only the most gradual, predictable maneuvers, smoothing out the heading signal, resulting in lower entropy. DUR SOGCV captures the interaction between trip duration and speed variability; longer tows with many barges require very consistent speed profiles to safely manage increased mass and drag, so trips that combine extended durations with tightly regulated speeds become strong indicators of high barge counts. Features not selected, such as maximum speed (SOG MAX) or mean positive acceleration (ACC POS MEAN), were likely excluded due to redundancy, lower predictive power, or misalignment with behavioral patterns.

The study concludes that the proposed approach provides a scalable, readily implementable method for enhancing Maritime Domain Awareness (MDA), with strong potential applications in lock scheduling, port management, and freight planning. The ability to predict the exact number of barges in an approaching tow in real-time would allow lock masters to optimize queuing, pre-position resources, and manage water levels, addressing the mismatch between modern tow sizes and legacy lock dimensions. Real-time data also provides value to public agencies like the U.S. Army Corps of Engineers, Departments of Transportation, and the Maritime Administration for planning, performance monitoring, and economic analysis.

Future work will expand the proof of concept to explore model transferability to other inland rivers with differing operational and environmental conditions, including the Ohio River (wider channel, higher traffic density), the Arkansas River (more constrained navigation), and the Columbia River (influenced by tidal patterns and hydropower infrastructure). Future work also aims to expand the dataset to include a broader range of vessel types, tow configurations, and geographies; incorporate environmental conditions such as river flow velocity and direction, wind speed and direction, water levels, turbulence, visibility conditions, channel width, and seasonal hydrologic variability; and explore more advanced deep learning architectures such as recurrent neural networks (RNNs) or graph neural networks (GNNs) to capture complex spatial and temporal dependencies.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to an AI system for barge count prediction, along with what the improved system can do:


  • Improvement: Integrate the paper's finding that course entropy (COG ENT) and trip duration × speed variation (DUR SOGCV) are the most predictive features. Build a feature extraction module that computes these 30 kinematic features from raw AIS trajectories, with emphasis on directional stability and speed consistency metrics.

  • What it can do: The system can automatically transform raw AIS pings into a rich feature vector that captures vessel maneuverability constraints imposed by tow size, enabling accurate barge count estimation without needing direct barge observation.

  • Improvement: Implement the Poisson Regressor as the core predictive model, paired with RFECV for automatic feature selection. This combination achieved the best MAE (1.92 barges) using only 12 of 30 features, outperforming ensemble and kernel-based models.

  • What it can do: The system can predict barge counts in real-time with minimal error, while automatically identifying and discarding redundant or noisy features, reducing computational overhead and improving generalization.

  • Improvement: Build a two-stage pipeline: (a) satellite imagery + AIS matching using a 2-minute temporal window and polygon-intersection spatial matching, and (b) AIS trajectory reconstruction via density-based stop clustering (speed 60 min, radius 300 m). This pipeline generates high-quality labeled training data.

  • What it can do: The system can automatically generate ground-truth labels for barge counts from satellite imagery, eliminating the need for manual annotation and enabling continuous model retraining as new data becomes available.

  • Improvement: Incorporate stratified K-fold cross-validation (specifically 2-fold given the small dataset) that preserves the distribution of barge count classes. This prevents under-representation of rare large tows (e.g., 12+ barges) during training.

  • What it can do: The system can produce reliable performance estimates even with imbalanced datasets, ensuring that predictions for uncommon but critical large tows are not systematically biased.

  • Improvement: Add a feature importance analysis layer that ranks predictors by frequency of selection across models. This provides interpretability and guides future feature engineering efforts.

  • What it can do: The system can explain why a particular barge count was predicted (e.g., low course entropy and high trip duration × speed variation indicate a large tow), making it useful for operators who need to trust and audit the model's decisions.

  • Improvement: Design the model architecture to be spatially agnostic, with features that are normalized by vessel geometry (e.g., draft-to-length ratio, speed coefficient of variation) rather than absolute values. This allows retraining on new rivers (e.g., Ohio, Arkansas, Columbia) with minimal data collection.

  • What it can do: The system can be deployed on any inland waterway with AIS coverage, adapting to local operational and environmental conditions (e.g., tidal influence, channel width, traffic density) with only a small set of labeled samples for fine-tuning.

  1. Ingest raw AIS data (1-minute interval) and satellite imagery (e.g., Planet Labs 3m resolution) for any river segment.

  2. Automatically match vessels between the two data sources using spatiotemporal fusion, generating labeled barge count samples.

  3. Reconstruct vessel trips by identifying stops and segmenting trajectories, then compute 30 kinematic features per trip.

  4. Predict barge count in real-time using the Poisson Regressor with RFE-selected features, achieving MAE < 2 barges.

  5. Provide confidence intervals and feature importance explanations for each prediction.

  6. Retrain continuously as new satellite imagery becomes available (e.g., daily updates), improving accuracy over time.

  7. Deploy across multiple rivers with minimal adaptation, enabling lock scheduling, port management, and freight planning at a national scale.

This system directly addresses the paper's stated objectives: minimal-error barge count estimation, identification of influential predictors, and a foundation for spatial generalization.

Abstract

Accurate, real-time estimation of barge quantity on inland waterways remains a critical challenge due to the non-self-propelled nature of barges and the limitations of existing monitoring systems. This study introduces a novel method to use Automatic Identification System (AIS) vessel tracking data to predict the number of barges in tow using Machine Learning (ML). To train and test the model, barge instances were manually annotated from satellite scenes across the Lower Mississippi River. Labeled images were matched to AIS vessel tracks using a spatiotemporal matching procedure. A comprehensive set of 30 AIS-derived features capturing vessel geometry, dynamic movement, and trajectory patterns were created and evaluated using Recursive Feature Elimination (RFE) to identify the most predictive variables. Six regression models, including ensemble, kernel-based, and generalized linear approaches, were trained and evaluated. The Poisson Regressor model yielded the best performance, achieving a Mean Absolute Error (MAE) of 1.92 barges using 12 of the 30 features. The feature importance analysis revealed that metrics capturing vessel maneuverability such as course entropy, speed variability and trip length were most predictive of barge count. The proposed approach provides a scalable, readily implementable method for enhancing Maritime Domain Awareness (MDA), with strong potential applications in lock scheduling, port management, and freight planning. Future work will expand the proof of concept presented here to explore model transferability to other inland rivers with differing operational and environmental conditions.

Sources

Related papers