TRACTOR Benchmark for Evaluating C to Rust Translators
cs.SE, cs.CR, cs.PL
Submitted: 2026-09-20
Updated: 2026-09-20
Code: https://github.com/DARPA-TRACTOR-Program/PUBLIC-Test-Corpus
Project page: https://rust-lang.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Memory-safety vulnerabilities remain a persistent source of security risk in critical software, much of which is implemented in memory-unsafe languages such as C and C++.
Terminology
Abstract
Memory-safety vulnerabilities remain a persistent source of security risk in critical software, much of which is implemented in memory-unsafe languages such as C and C++. Recent advances in programming languages, program analysis, and artificial intelligence have created new opportunities to modernize these legacy systems through automated translation to memory-safe languages such as Rust. The DARPA Translating All C to Rust (TRACTOR) program seeks to develop scalable techniques for translating large C codebases into safe, performant, and maintainable Rust. MIT Lincoln Laboratory serves as the program's independent test and evaluation organization and has developed a standardized benchmark for systematically assessing C-to-Rust translation tools. This report describes the TRACTOR benchmark, including progressively challenging test batteries and larger milestone projects, as well as the supporting evaluation infrastructure and metrics for assessing correctness, safety, idiomaticity, and performance. The benchmark and associated evaluation infrastructure are publicly available to support the broader development and evaluation of C-to-Rust translation technologies.
Sources
- SACTOR: LLM-Driven Correct and Idiomatic C to Rust Translation with Static Analysis and FFI-Based Verification
- Translating Large-Scale C Repositories to Idiomatic Rust
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties