Computational Onboard Data Management for Heterogeneous Autonomous Vehicle Streams

summary

Video file (mp4)

The gist

Autonomous vehicles generate massive, heterogeneous data streams that current logging and storage systems fail to manage efficiently, necessitating a new approach to on-board data management.

In short

AVS is a system designed to manage massive, diverse data from autonomous vehicles by combining computation with hierarchical storage. It handles high-rate streams like LiDAR and video through modality-aware compression and tiering data between fast SSDs for immediate access and slower HDDs for long-term archiving. This balances real-time ingestion with efficient, flexible querying.

Key concepts

Modality-aware reduction
This technique involves applying different compression strategies specifically tailored to the type of data being processed, such as LiDAR points or video frames. For instance, point clouds are downsampled using a voxel grid to reduce density while keeping important geometric details intact.
Hot–cold tiering
Data is organized into tiers based on access frequency and urgency. 'Hot' data resides on fast SSDs for immediate retrieval during operations, while 'cold' data is migrated to slower HDDs for long-term storage, ensuring the system balances speed and capacity.
Perceptual hashing (pHash)
Used for video data reduction, pHash identifies visually similar frames based on their content rather than exact pixel matches. This allows the system to discard redundant frames that look nearly identical, significantly reducing storage size while maintaining visual fidelity.
Metadata indexing
A lightweight layer of structured information is used to organize the stored data. By indexing object attributes and temporal slices, this allows users to perform complex queries—like retrieving only GNSS data or specific time windows—without needing to scan massive raw files.

Terminology used across episodes

This episode discusses

The paper

Computational Onboard Data Management for Heterogeneous Autonomous Vehicle Streams · Read on arXiv

University of Delaware

Autonomous vehicles continuously generate large volumes of heterogeneous sensor and state data streams. While this data supports real-time autonomy, it also enables emerging vehicle computing applications that require retained historical context. Existing in-vehicle logging stacks are largely replay- or compliance-oriented: rich sensor streams are often discarded after online processing or recorded only in limited windows for offline replay. This creates a new onboard vehicle-data operating point: transforming transient multimodal streams into compact, long-horizon, queryable history under embedded resource constraints. We present Autonomous Vehicle Storage (AVS), a computational and hierarchical in-vehicle data management system for this operating point. AVS introduces a vehicle-specialized architecture including (i) utility-aware modality preprocessing before persistence, (ii) a chunk-indexed append-only logger for sustained ingest, time-range lookup, and prefix recovery, and (iii) an SSD-HDD hot/cold hierarchy that supports direct cold queries. Evaluated on a real L4 platform with three days of traces and concurrent LiDAR, camera, and GPS logging, AVS reduces storage by 8.0-8.7x relative to raw rosbag2 from its modality and task aware preprocessing, delivers hot-tier (SSD) query Time To First Byte (TTFB) of 0.9-9.1 ms and cold-tier (HDD) query TTFB of 25.6-44.0 ms, recovers from power loss in 9.98 s with a bounded 0.24 s loss window, and sustains background archival to HDD at about 107-108 MB/s. It shows the feasibility of AVS on real in-vehicle platforms, and it also exposes key tradeoffs and open challenges for vehicle data engineering.

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Computational Onboard Data Management for Heterogeneous Autonomous Vehicle Streams".

Dev: Autonomous vehicles generate massive, heterogeneous data streams that current logging and storage systems fail to manage efficiently, necessitating a new approach to on-board data management.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So we're looking at this paper now titled "Computational Onboard Data Management for Heterogeneous Autonomous Vehicle Streams," and what they're saying is that autonomous vehicles are becoming mobile computing platforms that generate massive, diverse data streams, potentially up to fourteen terabytes per day <ref:2511.19453#pg0>.

Dev: That scale of data really puts existing logging and storage systems in a tough spot because they just can't handle the sheer volume or the different types of information coming in <ref:2511.19453#pg0>.

Taro: Exactly, we need a system that can serve both immediate real-time control needs and long-term analytics, which is what this paper seems to be tackling with its proposal for AVS <ref:2511.19453#pg0>.

Rosa: The core thesis of the paper is proposing AVS, which they frame as a computational and hierarchical storage system that co-designs computation with a specific layout, including modality-aware reduction and compression, hot–cold tiering for daily archival, and a lightweight metadata layer for indexing <ref:2511.19453#pg0>.

Dev: It sounds like the paper is arguing that current approaches are failing because they can't manage that heterogeneity or the diverse access patterns required by things like forensics versus long-term usage analysis <ref:2511.19453#pg1>.

Taro: I agree, and what matters is how AVS aims to bridge the gap between real-time control loops and those third-party applications that need historical data, which they show in Figure two illustrating the system architecture <ref:2511.19453#pg2>.

Rosa: They claim this design is grounded in system-level benchmarks covering SSD/HDD filesystems and embedded indexing, and it's validated on embedded hardware using real L4 autonomous driving traces <ref:2511.19453#pg0>.

Dev: From an engineering standpoint, the paper focuses on decoupling the storage sidecar from the main autonomy compute and operational ECUs via an Ethernet switch, which they state ensures that storage activities never interfere with safety-critical perception–planning–control loops <ref:2511.19453#pg2>.

Taro: That separation is crucial for safety, but I'm curious about what happens when the world misbehaves and we need to query that stored data immediately for forensic reconstruction or policy analysis <ref:2511.19453#pg1>.

Rosa: The paper suggests AVS supports diverse downstream uses like infrastructure analysis, safety forensics, and even long-term usage tracking through different data structures like compressed raw data and structured metadata <ref:2511.19453#pg2>.

Dev: I'm looking at the specific concepts they mention regarding data reduction, like using voxel grid downsampling for LiDAR data to maintain geometric fidelity while cutting the footprint by about four point two times compared to the original point density <ref:2511.19453#pg0>.

Taro: That level of compression is impressive when you consider the real-time constraints they mentioned, and I'm thinking about how that reduction impacts our ability to analyze complex scenarios later <ref:2511.19453#pg0>.

Rosa: They also talk about using perceptual hashing, or pHash, for image data to identify and discard visually similar frames based on content redundancy, which they showed resulted in a four point zero six times size reduction with a latency of about one point four five milliseconds per-image <ref:2511.19453#pg0>.

Dev: A millisecond latency for that kind of pruning is fast, but I'm wondering if the system can handle the ingestion rate without dropping frames when all those different modalities like high-rate LiDAR and low-rate CAN traces are flowing in simultaneously <ref:2511.19453#pg0>.

Taro: That simultaneous ingest is a big test for heterogeneity, and I want to know how robust this system is when we introduce unexpected sensor noise or data bursts during extreme driving conditions <ref:2511.19453#pg0>.

Rosa: The overall goal they set out with AVS seems to be achieving predictable real-time ingest, fast selective retrieval, and a substantial footprint reduction while operating under modest resource budgets <ref:2511.19453#pg0>.

Dev: So, the system is designed to handle that trade-off between performance in real time and the need for substantial long-term storage capacity on constrained hardware <ref:2511.19453#pg0>.

Taro: The implication for autonomy research is that we could finally move beyond just logging raw data and start using this system to generate structured, queryable insights directly from the vehicle's operation <ref:2511.19453#pg2>.

Rosa: It really feels like the paper is laying out a blueprint for how future autonomous platforms will manage their massive data lifecycle efficiently, moving away from just ephemeral loggers <ref:2511.19453#pg0>.

Dev: If this architecture proves reliable under real operational conditions, it could significantly reduce the bandwidth and storage demands on vehicle hardware going forward <ref:2511.19453#pg0>.

Taro: For the broader impact, I see this enabling better safety forensics and policy reconstruction by making historical data readily accessible and manageable <ref:2511.19453#pg1>.

Rosa: Ultimately, the authors are showing that a system co-designing computation with a hierarchical layout can deliver the necessary organization for handling heterogeneous AV streams <ref:2511.19453#pg0>.

Dev: We need to keep an eye on how they handle those metadata indices because that's where the speed of selective retrieval really lives or dies <ref:2511.19453#pg0>.

Taro: It seems like this work could help turn vehicle data from just a byproduct into a valuable, structured asset for the entire mobility ecosystem <ref:2511.19453#pg2>. This whole discussion on the computational and hierarchical storage system AVS really sets the stage for understanding how vehicles will manage their own massive data footprint going forward.

Conclusion: Rosa: So we've been looking at how this paper tackles the challenge of managing all that massive, messy data coming from autonomous vehicles.

Dev: Yeah, it really gets to the heart of why we can't just keep logging everything onto a single system anymore.

Rosa: The title itself, "Computational Onboard Data Management for Heterogeneous Autonomous Vehicle Streams," feels pretty descriptive about what they are trying to achieve.

Taro: It points directly at the core problem: managing different types of data streams on the vehicle itself in a way that's actually useful for autonomy.

Dev: I think it highlights the shift from just simple logging to needing a system that can actively process and organize that information on-board before it gets too heavy.

Rosa: Exactly, and the authors they cite seem to be really focused on building something practical for real-world vehicle deployment, not just theoretical constructs.

Taro: I’m interested in their conclusion because it should summarize how this system actually handles the trade-off between getting real-time responses and storing enough data for later analysis.

Dev: That balance is key; if they nail that without introducing unacceptable latency spikes during critical driving situations, then this could be genuinely useful for deployment.

Rosa: I think the authors are making a strong case that we need this kind of dedicated storage architecture because the sheer volume of sensor data is simply overwhelming existing methods.

Taro: It suggests a future where vehicle data isn't just dumped; it becomes an organized asset that supports everything from immediate incident investigation to long-term operational improvement.

Dev: And if they can keep the resource usage low while still supporting those diverse access patterns, then the real-time performance metrics they’re reporting matter a lot for me.

Rosa: It really paints a picture of how autonomous vehicles will start acting more like mobile computers that have to intelligently manage their own data lifecycle.

More episodes

← Home