Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept

summary

Video file (mp4)

In short

The episode discusses a proof of concept paper from the University of Arkansas detailing how to predict barge tow size using only vessel trajectory data. Researchers used AIS data and derived features to achieve an average prediction error of less than two barges. The technology promises smoother supply chains by providing real-time visibility into inland waterway traffic.

Key concepts

AIS (Automatic Identification System)
The system used for tracking vessels, where the tugboat broadcasting its position constantly provides data on the location and movement of a barge tow. This data is essential for predicting how many barges are being pulled.
Course Entropy
A measure of how unpredictable a vessel's heading is over the trip. The study found that larger tows have lower entropy because they must follow smooth, gradual maneuvers, unlike smaller, more nimble tows.
Poisson Regressor
A statistical model used in the research designed specifically for count data. It successfully predicted the number of barges with a Mean Absolute Error of 1.92, indicating that its predictions were consistently accurate.
Spatiotemporal Matching
The process used to link satellite images (which show actual barge counts) to the AIS data. The match required the AIS points to be within two minutes of the image timestamp and for the vessel's path to intersect with where the barges were seen.

Terminology used across episodes

This episode discusses

The paper

Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept · Read on arXiv

Geoffery Agorku, Sarah Hernandez, Hayley Hames, Cade Wagner

University of Arkansas

Accurate, real-time estimation of barge quantity on inland waterways remains a critical challenge due to the non-self-propelled nature of barges and the limitations of existing monitoring systems. This study introduces a novel method to use Automatic Identification System (AIS) vessel tracking data to predict the number of barges in tow using Machine Learning (ML). To train and test the model, barge instances were manually annotated from satellite scenes across the Lower Mississippi River. Labeled images were matched to AIS vessel tracks using a spatiotemporal matching procedure. A comprehensive set of 30 AIS-derived features capturing vessel geometry, dynamic movement, and trajectory patterns were created and evaluated using Recursive Feature Elimination (RFE) to identify the most predictive variables. Six regression models, including ensemble, kernel-based, and generalized linear approaches, were trained and evaluated. The Poisson Regressor model yielded the best performance, achieving a Mean Absolute Error (MAE) of 1.92 barges using 12 of the 30 features. The feature importance analysis revealed that metrics capturing vessel maneuverability such as course entropy, speed variability and trip length were most predictive of barge count. The proposed approach provides a scalable, readily implementable method for enhancing Maritime Domain Awareness (MDA), with strong potential applications in lock scheduling, port management, and freight planning. Future work will expand the proof of concept presented here to explore model transferability to other inland rivers with differing operational and environmental conditions.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept".

Jane: The paper was written by Geoffery Agorku, Sarah Hernandez, Hayley Hames and Cade Wagner from University of Arkansas.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everyone. Today we’re cracking open a fresh one from the arXiv preprint server, and the title is a mouthful — "Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept." Jane, when you first saw that title, what jumped out at you?

Jane: Honestly, Tom, the phrase "proof of concept" is what got me. That tells me these researchers at the University of Arkansas are saying, hey, we’ve got an idea, we’ve tested it on a real river, and now we want the world to poke holes in it. And the idea itself is pretty clever — they’re trying to count barges without actually looking at the barges.

Tom: Right, because barges are those big rectangular cargo boxes being pushed along the Mississippi, and they don’t have their own engines or GPS trackers. The tugboat pushing them does, though — it broadcasts its position constantly through something called AIS, the Automatic Identification System.

Jane: Exactly. So the tug is like a delivery truck, and the barges are the trailers. You know where the truck is, but you don’t know how many trailers it’s hauling unless someone tells you. And right now, on most of America’s inland waterways, nobody’s telling you in real time.

Tom: And that’s a huge deal, because the paper points out that over six hundred million tons of cargo move on these rivers every year. But the people managing locks, ports, and supply chains are basically flying blind when it comes to knowing how many barges are coming their way.

Jane: So the researchers came up with a workaround. They used satellite imagery to actually see the barges and count them — that’s their ground truth. Then they matched those satellite snapshots to the tugboat’s AIS data at the exact same moment. That gives them a labeled dataset: here’s what the tug’s movement looked like, and here’s how many barges it was pushing.

Tom: And once they had that pairing, they could train machine learning models to predict barge count using only the AIS data — no satellite needed anymore. That’s the clever part, right? The satellite is just the teacher; the AIS data is what does the real work in the field.

Jane: Precisely. And the implications are big for something called Maritime Domain Awareness. If you can predict barge counts in real time from data you already have, you can help lock operators schedule lockages more efficiently, help ports plan their crane and dock assignments, and give shippers the visibility they’ve been missing.

Tom: So it’s not just an academic exercise — this could actually make the supply chain run smoother. But before we get too deep into the weeds, let’s talk about how they actually built this thing. That’s coming up next.

Summary: Tom: So we’re back with "Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept," and Jane, I want to get into the meat of the methodology. How did they actually pull this off?

Jane: Well, Tom, they focused on a two hundred forty-mile stretch of the Lower Mississippi River, from Baton Rouge down to the Gulf of Mexico. That’s the busiest waterway corridor in the country. They gathered satellite images from Planet Labs over a few months in early two thousand twenty-four and they pulled AIS data from the Marine Cadastre system.

Tom: And the satellite images are how they counted the actual barges, right?

Jane: Yes, but here’s the thing — they didn’t just have a human staring at every image. They used a pre-trained computer vision model called YOLO to automatically detect vessels and barges in the satellite scenes. Then two researchers manually verified those detections to make sure the counts were accurate. That gave them twenty-six labeled instances — twenty-six moments where they knew exactly how many barges a tug was pushing.

Tom: Twenty-six samples. That’s not a lot, is it?

Jane: It’s small, and the paper is upfront about that. But for a proof of concept, it’s enough to see whether the approach has legs. The next step was matching those satellite detections to the AIS tracks. They used a spatiotemporal matching procedure — so the AIS points had to be within two minutes of the satellite image timestamp, and the vessel’s path had to intersect with where the satellite saw the tug.

Tom: So you’ve got a tug, you know where it was, you know how many barges it had, and now you’ve got its entire AIS trajectory. What do you do with that?

Jane: That’s where the feature engineering comes in. They extracted over thirty features from each trajectory — things like average speed, speed variability, how often the vessel changed course, how much the heading wobbled, acceleration patterns, even things like the entropy of the course distribution. The idea is that a tug pushing twelve barges moves very differently than a tug pushing two barges.

Tom: And that makes sense — a bigger tow is heavier, harder to turn, slower to accelerate. So the movement signature should be different.

Jane: Exactly. Then they used a feature selection technique called Recursive Feature Elimination with Cross-Validation to figure out which of those thirty features actually mattered. And they tested six different regression models — from simple linear ones to complex ensemble methods.

Jane: The winner was the Poisson Regressor, which is a model designed for count data — perfect for something like "how many barges." It achieved a Mean Absolute Error of one point nine two barges. So on average, its prediction was off by less than two barges. That’s pretty impressive for a first attempt.

Tom: Less than two barges off, using only the tug’s movement data. That’s the kind of result that makes you sit up and pay attention. But I’m curious — which features actually mattered most? That’s the part that could really tell us something about how these vessels behave.

Improvements: Tom: We’re back with "Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept," and Jane, you mentioned the Poisson Regressor won. But what I really want to know is — what did the feature analysis reveal? What actually predicts barge count?

Jane: Great question, Tom. The two most frequently selected features across all the models were course entropy and a composite feature they called "trip duration times speed variation." Course entropy is basically a measure of how unpredictable the vessel’s heading is over the trip. A tug pushing a huge tow can’t make sharp, erratic turns — it has to plan smooth, gradual maneuvers. So its heading signal is very regular, very low entropy.

Tom: So a low-entropy course means a big tow, and a high-entropy course means a small, nimble one?

Jane: Exactly. And the second feature, trip duration multiplied by speed variation, captures something similar — longer trips with very consistent speeds suggest a heavy, stable tow that doesn’t speed up and slow down much. The paper also found that speed coefficient of variation and mean course change rate were strong predictors. It all points to the same story: bigger tows move more deliberately.

Tom: That’s a really elegant finding. But the paper also acknowledges some limitations, right? I mean, twenty-six samples is small, and they only looked at one river.

Jane: Right, and they’re honest about that. The dataset was imbalanced — they had lots of small tows, like two to four barges, but very few large ones with a dozen or more. That means the model is better at predicting small tows and struggles with the rare big ones. In their test results, you can see the model underpredicting a twelve-barge tow, for instance.

Tom: So what’s the path forward? How do they improve this?

Jane: The paper lays out a clear roadmap. First, they need more data — more samples, more river systems, more vessel types. Second, they want to incorporate environmental conditions like river current, wind, and water levels, because those obviously affect how a tug moves. A tug fighting a strong current will have a different speed profile than one moving with it, regardless of how many barges it’s pushing.

Tom: That’s a really good point. If you don’t account for the river itself, you might misread a tug’s movement.

Jane: Exactly. And third, they mention trying more advanced models — things like recurrent neural networks or graph neural networks that can capture temporal dependencies in the trajectory data. The current models treat each trip as a bag of statistics, but a sequence model could learn how the vessel’s behavior evolves over time.

Tom: So the improvements are about breadth — more data, more context, more sophisticated modeling. This is a proof of concept, but it’s pointing toward something that could actually be deployed. Let’s bring in Lu and Meng to get their take on what this means for the real world.

Lu: I’m really excited about the transferability angle. The paper mentions testing this on other rivers like the Ohio or the Columbia. If the features hold up across different environments, you’ve got a generalizable tool, not just a Mississippi-specific one.

Meng: And from an engineering standpoint, the beauty is that AIS data is already being collected everywhere. You don’t need to install new hardware on barges or tugs. You just need the software to process the data. That makes this scalable in a way that IoT trackers or additional cameras aren’t.

Tom: So the foundation is solid, but there’s real work ahead to make it robust. Let’s wrap this up in our final segment.

Conclusion: Tom: Alright, we’re closing out our discussion of "Predicting Barge Tow Size on Inland Waterways Using Vessel Trajectory Derived Features: Proof of Concept." Jane, give us the final summary — what did we learn?

Jane: We learned that you can predict how many barges a tugboat is pushing using only its AIS movement data, with an average error of under two barges. The key insight is that bigger tows move differently — they have lower course entropy, steadier speeds, and more deliberate maneuvers. The satellite imagery was just the training tool; the real product runs on data we already collect.

Tom: And the implications? What does this mean for the world?

Jane: It means lock operators could know in advance whether a tow will need a single or double lockage. Ports could plan their resources before a vessel arrives. Shippers could finally get the visibility that’s been missing from inland waterway transport. And it does all this without installing a single new sensor on a barge.

Lu: The scalability is what excites me. This could be the foundation for a real-time national barge tracking system — something that doesn’t exist today.

Meng: And it’s practical. The data infrastructure is already there. This is a software problem, not a hardware problem, and that makes deployment much more feasible.

Tom: Well said. So we’ve got a clever proof of concept with a clear path forward. The team at Arkansas has given us something to build on, and we’ll be watching to see how it evolves. Thanks for joining us, everyone. We’ll be back with the next paper soon.

More episodes

← Home