Modeling Robotics Dataset Construction as an Artifact-Based Build Process

arXiv:2606.00162 · cs.RO, cs.CV, cs.LG · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "Modeling Robotics Dataset Construction as an Artifact-Based Build Process".

Rosa: Modeling robotics dataset construction as an artifact-based build process substantially reduces dataset update latency while maintaining a deterministic build design that supports reproducibility.

Dev: First, who's behind it and why it matters.

Paper summary: Rosa: So, we’re talking about this paper, "Modeling Robotics Dataset Construction as an Artifact-Based Build Process." Basically, the authors are tackling the messy part of getting machine learning data from robot recordings—the conversion process—by treating it like a formal build system. They claim that modeling this construction as an artifact-based build process over a dependency graph can significantly reduce how long it takes to update datasets while keeping everything totally predictable for reproducibility.

Dev: That sounds promising for reducing those iteration cycles we always struggle with in the lab, Rosa, but I'm curious about what they actually mean by "artifact-based build process" in this context. Does it mean they’re just using a standard dependency graph structure or is it something more specific to how dataset generation works?

Taro: From an autonomy research standpoint, the thesis seems to be that by formalizing the construction steps, we gain explicit dependency tracking and selective recomputation which is crucial when dealing with massive amounts of multimodal sensor data. This structured approach should make debugging much clearer when things go wrong during the pipeline execution.

Rosa: Exactly what Taro said, Dev; it’s about creating a system where every piece of processed data is an artifact that has a clear history of its inputs, which helps us manage those slow iteration cycles you mentioned. The core idea is that instead of running sequential scripts every time we tweak something, we build a dependency graph where inputs go into operations and operations produce new artifacts.

Dev: I see the structure; it sounds like they're essentially applying concepts from CI/CD systems like Bazel to dataset creation, which implies strong determinism because the rebuild decision relies on an "action digest" derived from those declared inputs and operation definitions. That deterministic nature is what we need for reliable testing.

Taro: And that digest mechanism is key because if the inputs and rules are the same, you get the exact same output artifact, which means we can cache results effectively without worrying about corrupted or stale data sneaking into our training sets. The paper lays out how this system supports selective recomputation, meaning only what actually changed needs to be rebuilt.

Rosa: It really matters because if this works outside the controlled lab environment, it could mean that researchers can rapidly generate new datasets from their raw recordings without getting bogged down in manual scripting overhead every single time they want to test a new idea. That’s the practical impact we need to keep in mind, Dev.

Paper summary: Dev: I agree that the reduction in recomputation overhead is what really moves the needle for us on the engineering side; it’s not just theoretical speed but actual time saved during development loops. The paper specifically mentions that this formulation enables dependency tracking and artifact reuse for dataset generation, which directly addresses our need for efficiency.

Taro: I think the implication here is that we can scale up our data collection and processing capabilities much more effectively because the system handles the complexity of dependencies automatically across different stages like frame decoding or trajectory extraction. This moves us closer to systems where the pipeline manages itself intelligently, even when things are complex.

Rosa: It seems like this work is about moving away from ad hoc scripts toward a structured, reproducible methodology for creating robotics data, and that's exactly what the title suggests about "Modeling Robotics Dataset Construction as an Artifact-Based Build Process." It shifts the focus from writing custom scripts to defining a formal build structure.

Dev: And looking at the paper’s claims regarding performance, they state that Bagzel substantially outperforms the sequential rosbag2nuscenes baseline in all evaluated execution modes, showing gains like up to three hundred eighty-six point two six times speedup in warm builds on a twenty point four GB dataset <ref:2606.00162#pg0,builds on a 20.4 GB dataset>. That level of performance improvement across different modes is what really catches my attention as an engineer concerned with latency and throughput.

Taro: That massive speedup suggests that for large-scale data processing, this approach provides a much more viable path than the traditional sequential methods, especially when considering how quickly we need to iterate on our autonomy models. The scalability analysis across dataset sizes from five point one GB up to twenty point four GB also shows consistent outperformance in warm and incremental modes, which is vital for long-running experimental setups <ref:2606.00162#pg0,across dataset sizes from 5.1>.

Rosa: That consistency across different data sizes, from five point one GB to twenty point four GB, really speaks to the robustness of this artifact-based approach; it doesn't seem like its efficiency drops just because we have a larger dataset to process <ref:2606.00162#pg0>. This gives me hope that this methodology can be applied reliably outside of perfectly curated lab settings too, which is my main concern as a field roboticist.

Dev: I worry about the operational reality of running this; if we push the loop rate and need near real-time data processing, how does this artifact modeling handle those strict timing constraints without introducing unacceptable latency during the build evaluation phase? That’s a critical failure mode to consider for any deployment scenario.

Paper summary: Taro: That brings up a point about autonomy under duress; if we're dealing with unpredictable world misbehavior, we need certainty in our data generation pipeline, and this deterministic execution path seems designed specifically to provide that necessary reliability when the system is under stress. The paper’s mention of integrating with the Slurm workload manager suggests it can handle distributed execution too.

Rosa: So, it sounds like this paper is showing us how to build a more resilient data infrastructure for robotics, focusing on making data generation faster and more reliable through formal structure. We’re looking at how this formal modeling impacts our ability to move from raw sensor logs to usable training data much quicker than before.

Dev: And the specific mention of Bagzel-xattr providing consistent gains, with a mean runtime reduction of five point nine percent compared to default Bagzel in the input granularity study, shows that optimizing how we manage those large inputs is also a tangible lever for improving build performance, even when we are talking about smaller variations in source files <ref:2606.00162#pg0,gains, with a mean runtime reduction of 5.9% compared to>.

Taro: It seems the implication is that this structured approach doesn't just speed things up; it fundamentally changes how we think about the pipeline itself, treating data construction as a formal engineering problem rather than just a series of manual steps. That shift in perspective is what makes it relevant for autonomy research moving forward.

Rosa: So, to wrap up this discussion on "Modeling Robotics Dataset Construction as an Artifact-Based Build Process," it seems the main message is that we can achieve reproducible dataset generation through deterministic execution by modeling the construction process as a dependency graph, and this approach substantially reduces the latency in getting new datasets ready for training.

Dev: I think the real world implication is that this technique gives us a powerful tool to drastically cut down on our engineering overhead when iterating on complex sensor data pipelines, which is something we all deal with constantly. We need to see how quickly we can integrate these custom rules for things like frame decoding and annotation processing into our existing workflows.

Taro: I think the impact is that it provides a solid foundation for building more scalable and trustworthy data infrastructure that can support autonomous systems operating in increasingly complex, real-world environments where data needs to be generated on demand with high confidence.

Rosa: It’s exciting to think about how this could translate into faster prototyping and more reliable testing cycles for the next generation of robotic systems, moving us closer to having ready-to-use datasets much sooner.

Conclusion: Rosa: So, we've been diving into this paper titled "Modeling Robotics Dataset Construction as an Artifact-Based Build Process," and now it's time to talk about what this whole thing actually means for us in the field.

Dev: Yeah, Rosa, I agree that the title itself sounds a bit academic at first glance because it frames dataset creation like a software build process, but we need to figure out how this translates into real-world performance constraints for our systems.

Taro: From an autonomy researcher's view, the core concept here is taking something usually done manually and turning it into a structured pipeline where everything has a traceable history, which is essential when we want to debug why the AI made a certain decision in a complex scenario.

Rosa: Exactly, Taro; this approach suggests that by treating dataset creation as an artifact build, we gain much better control over reproducibility than traditional scripting allows.

Dev: And that control is what interests me from an engineering standpoint; if the authors are successfully modeling dependencies and using things like action digests to determine when a rebuild is necessary, it points toward a system with much lower latency for those iterative update cycles you mentioned earlier.

Taro: I think the real implication is that we can build more trustworthy data infrastructure for autonomous systems because we're moving away from fragile, sequential workflows toward something deterministic where the output artifact is guaranteed based on its inputs.

Rosa: That makes sense; it shifts the focus from just getting a file created to managing a reliable system that can reliably produce those files repeatedly.

Dev: I'm still thinking about the practical side, Rosa; if this model is highly effective in controlled environments, how long do you think it takes for us to see real-world gains when we apply this methodology to messy, unstructured data from field tests?

Taro: That's the million-dollar question, Dev; while the paper shows excellent results on specific datasets like nuScenes, the challenge will be proving that this artifact modeling holds up when the world misbehaves and introduces unexpected variability.

Rosa: That’s a fair point; we need to see if this structure can handle more noise than what was in those controlled experiments before we can really say it's ready for our rugged field robots.

Dev: I'm optimistic, though; the fact that they showed such substantial speedups in warm and incremental builds on large datasets suggests the underlying mechanism is quite robust and doesn't rely on perfectly clean input data every single time.

Taro: That robustness is what we need; if the system can handle variability without breaking, it opens up possibilities for creating datasets that reflect real-world operational challenges rather than just pristine test scenarios.

Rosa: So, to wrap up this segment, this paper's idea is fundamentally about applying formal build systems to data generation to achieve better control and speed for robotics pipelines. Next time we discuss how they implemented those custom rules for things like frame decoding and annotation processing within the dependency graph.

Institute for Autonomous Driving, University of the Bundeswehr Munich

cs.RO, cs.CV, cs.LG

Submitted: 2026-05-29

Updated: 2026-10-07

Comments: Accepted at the 2026 IEEE 22nd International Conference on Automation Science and Engineering (CASE 2026). 7 pages, 6 figures, 2 tables. Code: https://github.com/UniBwTAS/bagzel

DOI: 10.1109/CASE69030.2026.11704392

Code: https://github.com/UniBwTAS/bagzel

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 89/100

The gist: Modeling robotics dataset construction as an artifact-based build process substantially reduces dataset update latency while maintaining a deterministic build design that supports reproducibility.

Key concepts

Artifact-based DAG formulation
This models the dataset creation process as a Directed Acyclic Graph (DAG) where nodes are either generated artifacts (the datasets themselves) or operations that create them. Edges show which inputs are needed for which operations, allowing the system to track dependencies and reuse existing outputs efficiently.
Action Digest Rebuild Decisions
To decide whether to rebuild a part of the dataset, the system calculates an 'action digest' based on all declared inputs and rules. If this digest hasn't changed since the last build, cached artifacts are reused, ensuring that identical inputs always produce identical outputs.
Server-side Digest Management (Bagzel-xattr)
This mechanism stores file digests as metadata on a central server instead of recalculating them every time. This prevents repeated client-side hashing of large files during the build evaluation phase, speeding up the process by checking stored digests first.
Custom Rules
These are specialized instructions that describe how to perform common dataset preparation tasks, such as decoding sensor data or exporting to a standard format like nuScenes. By encapsulating these steps into rules within the build graph, construction is integrated directly into the dependency tracking system.

Terminology

Summary

Modeling robotics dataset construction as an artifact-based build process substantially reduces dataset update latency while maintaining a deterministic build design that supports reproducibility.

How it works

The core idea of the work is to model robotics dataset construction as an artifact-based build process over a dependency graph, leveraging principles from established CI/CD systems like Bazel. This formulation involves defining an abstract bipartite build graph P = (V, E), where node set V is divided into artifact nodes (Va) and operation nodes (Vo). Edges in E encode declared dependencies, alternating between artifacts and operations—artifacts serve as inputs to operations, and operations produce derived artifacts. Rebuild decisions are based on an action digest computed from the declared inputs and operation definition. If the digest changes, only the affected path is recomputed; otherwise, cached artifacts are reused.

The implementation is realized through Bagzel, an open-source framework built on the Bazel build system that enables reproducible and incremental generation of nuScenes-format datasets from ROS recordings. This framework supports two primary output formats: a visual dataset format and the widely adopted nuScenes format. The process involves developing custom rules that encapsulate common dataset preparation steps, including frame decoding, trajectory extraction, annotation processing, and export to standardized dataset formats, thereby incorporating construction directly into the Bazel dependency graph.

Key Components and Mechanisms

The methodology introduces several specific mechanisms to achieve deterministic and efficient builds:

  1. Artifact-based DAG formulation: This models the build process as a Directed Acyclic Graph (DAG), allowing for dependency tracking, selective recomputation, and artifact reuse for dataset generation.

  2. Action Digest Rebuild Decisions: Rebuild decisions are triggered if the digest changes, which ensures that identical inputs and rules yield identical digests and outputs, enabling correct cache reuse.

  3. Server-side Digest Management (Bagzel-xattr): To address the bottleneck of client-side hashing for large inputs, Bagzel-xattr introduces a server-side digest management mechanism. This stores file digests as metadata on a file server; during build evaluation, the system uses this stored digest to determine if dependent steps must be executed, avoiding repeated client-side full-file hashing.

Experimental Evaluation and Results

The approach was empirically validated by comparing Bagzel and Bagzel-xattr against a sequential rosbag2nuscenes baseline across three execution modes: (i) Cold build, (ii) Warm build, and (iii) Incremental build. The primary performance metric is the end-to-end wall-clock runtime T per build execution.

(RQ1: Execution Efficiency)

The results show that Bagzel substantially outperforms the baseline in all modes, with the largest gains in warm and incremental builds. For instance, on a 20.4 GB dataset, Bagzel achieved up to 386.26× speedup in warm builds and 7.21× in incremental builds.

(RQ2: Scalability)

Scalability analysis across dataset sizes (5.1 GB to 20.4 GB) revealed that all Bagzel variants consistently outperform the rosbag2nuscenes baseline across all three build modes. Crucially, in warm and incremental modes, Bagzel runtimes remain close to one second, indicating weak dependence on dataset size within the evaluated range.

(RQ3: Input Granularity and Digest Management)

The study investigated how input granularity (varying the number of source files) and server-side digest management influence incremental runtime. The results indicated that for both methods, input granularity has only a minor effect on incremental runtime. Furthermore, Bagzel-xattr provided consistent gains over default Bagzel, with a mean reduction of 5.9% in the input granularity study.

Conclusion

The paper concludes that modeling dataset construction as an artifact-based build process is an effective strategy for reducing dataset update latency in robotics data pipelines. The findings demonstrate that this approach enables reproducible dataset generation through deterministic execution, substantially reducing recomputation overhead. While the evaluation was conducted on a single workstation, the results provide initial evidence of the benefits of incremental artifact-based execution.


The gist

Modeling robotics dataset construction as an artifact-based build process substantially reduces dataset update latency while maintaining a deterministic build design that supports reproducibility.

The core idea of the work is to model robotics dataset construction as an artifact-based build process over a dependency graph, leveraging principles from established CI/CD systems like Bazel. This formulation involves defining an abstract bipartite build graph P = (V, E), where node set V is divided into artifact nodes (Va) and operation nodes (Vo). Edges in E encode declared dependencies, alternating between artifacts and operations—artifacts serve as inputs to operations, and operations produce derived artifacts.

Improvements for AI systems

Here are the specific improvements an AI system could achieve by implementing the principles described in this research:

  1. Automatic, Deterministic Dataset Generation: The system will move from ad-hoc, error-prone scripts to a robust, artifact-based build process (Bagzel). This allows for reproducible dataset creation from raw ROS bag recordings.

  2. Massive Iterative Speedup in Training/Evaluation Loops: The system can achieve up to 386x speedup during warm builds and 7.21x speedup during incremental builds on large datasets (e.g., the 20.4 GB dataset). This drastically reduces the time spent waiting for data preprocessing, enabling much faster model iteration cycles in robotics research.

  3. Guaranteed Reproducibility: By modeling dataset construction as a Directed Acyclic Graph (DAG) with explicit dependencies and content-addressed caching, the system ensures that any given build output is deterministic. If you rerun the process with the exact same inputs and rules, you are guaranteed to get the exact same dataset artifact.

  4. Efficient Handling of Large Data Inputs: The integration of server-side digest management (Bagzel-xattr) prevents repeated full-file hashing of massive raw recordings during incremental updates. This makes iterative refinement feasible even with multi-hundred GB datasets by only recomputing what has genuinely changed, significantly reducing build time overhead.

  5. Scalable Performance Across Data Volumes: The system demonstrates superior scaling behavior compared to sequential baselines as dataset size grows (5.1 GB to 20.4 GB), especially in warm and incremental modes where runtime remains nearly constant, whereas the baseline runtime scales linearly with data size (15.50 s/GB).

  6. Optimized Input Granularity Management: The system can intelligently manage the partitioning of input data (e.g., choosing between 1, 2, or 8 bag files) to optimize incremental build performance. It identifies that server-side digest management benefits more as input granularity increases, allowing the system to adapt its caching strategy for more complex data splits.

In summary, an AI system leveraging this paper can transform robotics research workflows from slow, fragile manual data preparation into a high-throughput, deterministic pipeline capable of rapidly testing and iterating on large-scale datasets.

Related papers