Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language

arXiv:2608.13029 · q-bio.GN, cs.AI, cs.SE · Submitted 2026-08-13 · Read on arXiv

Johan Henriksson

Umeå University · Umeå Centre for Microbial Research · Integrated Science Lab · Science for Life Laboratory

q-bio.GN, cs.AI, cs.SE

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/henriksson-lab/paper_rust_translation1

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language The field of bioinformatics struggles with legacy code - old code that is commonly used but may no

Terminology

Summary

Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language

The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be written in an now-unfamiliar language (e.g. Perl, Fortran). This incurs maintenance cost (technical debt), but dynamically typed languages also negatively impacts the environment and fail to make use of modern hardware. Legacy code may also have security or safety problems that make it unsuited for use in clinical settings. Here we show that agentic AI, combined with static analysis, can be used to translate legacy code to the modern language Rust. We provide prompts and supporting software to aid systematic translation, and evaluate it on common software for NGS and imaging. We showcase the result on our software Bascet: Size was reduced by 80x, build time decreased by 10x, and performance of key steps improved >3x. Unix dependencies were also removed, making Bascet the only single-cell pipeline able to run on native Windows, without a container. Large-scale refactoring of bioinformatics software is thus now possible at a limited budget, enabling more complex tools to be developed.

Introduction

Bioinformatics software underpins all major biological studies today, yet funding for software is scarce, the need for software maintenance is not given enough attention, and there is a lack of workforce. While major version releases can be published anew, bug fixes typically cannot, and there is commonly little incentive for a non-author to contribute to someone else's software. This has resulted in a fragmented bioinformatics ecosystem, where it can be hard to combine software (R vs Python), be hard to run on modern computers (requiring outdated libraries), not use modern hardware efficiently, or it may no longer scale for newer large-scale datasets. When software is poorly designed (commonly due to budget or time constraints, PhDs or postdocs leaving the lab, lack of training, or change of design), issues arise that need to be circumvented when the software is used later. This is referred to as technical debt - money saved upon software conception has to be paid for later maintenance when the software is actually used (e.g., by setting up dedicated software environments or fixing the bugs). Research software has particularly high technical debt, with issues taking a long time to resolve. New challenges are now also on the horizon: (1) R, Python and Matlab consume >50x more energy than compiler or semi-compiled languages for certain workloads, making them bad for the environment. (2) Hardware is becoming scarce due to ramped up demand, requiring code to be more efficient in the coming years. (3) Computer hardware is evolving away from the use of CPUs to more efficient architectures, requiring extensive software rewriting. All of these major problems call for a rethinking of how bioinformatics research is being conducted.

Strong initiatives for improving bioinformatics code do exist; e.g., Bioconductor is an effort to increase the interoperability and quality of R-packages. While the same software may produce different results when run on different operating systems, Nextflow has narrowed this gap. However, despite the best efforts, major problems persist: It can be difficult to mix R and Python code, and the choice of language is largely based on package availability - i.e., RNA-seq analysis is almost synonymous with R while machine learning is synonymous with Python. Similarly, the library Bioformats is a cornerstone for microscopy, supporting 165 file formats (v8.5), but being written in Java, requires use of the Java virtual machine, limiting interoperability. Bioinformaticians are thus expected to learn at least two programming languages to cover most use cases. However, the increased demand for web applications also adds the need for Javascript, and high performance computing further requires knowledge of any of C, C++, Rust, Golang, Zig, or CUDA/OpenCL in the extreme case. The need to learn deeply about such a large range of topics is incompatible with the pressure to publish, and progress quickly in academia.

In this study we show how agentic AI, combined with classic static analysis, has the potential to rewrite all bioinformatics software in a single modern language. To be able to handle all use cases, this eliminates R and Python as the one language (not suited for implementing, e.g., aligners). Also, C/C++ are generally too difficult to use for most bioinformaticians due to poor tooling (e.g., no standardized build system nor docucumentation system), legacy code dependencies and unsafe language features (pointers). However, newer languages now exist. Rust was designed by Mozilla foundation to handle the requirement for performance and security in their Firefox webbrowser (https://rust-lang.org/). It has since been accepted in the Linux kernel, and since C was made popular due to Unix, Rust has by analogy the potential of becoming the next dominant language. Major companies such as Dropbox, Discord, Vercel, Cloudflare, Meta/Whatsapp have already migrated significant parts of their code to Rust, both to improve performance and security.

In summary, this study shows how Rust is able to overcome systematic ecosystem problems in bioinformatics, including version mismatches, software distribution and clinical implementation. We also demonstrate how Rust has the potential to become the only language a research group or bioinformatician ever has to master to cover the full stack of bioinformatics.

Materials and Methods

Choice and use of AI agents

A combination of Claude and Codex was used for this study. As new versions were released during translation, this study should not be seen as a comparison of specific agentic models. In the beginning of the project, it was mainly driven by Claude code (Opus 4.7). Once Codex (20% GPT5.4 and 80% GPT5.5) was deemed more suited for translation work, it became the primary choice, especially during the early critical stages of translation. At the time of writing, GPT-5.5 (Codex) is the default, followed by Claude for resolving difficult tasks. Claude was used for all GUI (graphical user interface) programming, where the screenshots of the original software was provided as references. The translations were produced over a period of 11 weeks, using a Claude Max 20x subscription (11 weeks), and two Codex Pro subscriptions (8 weeks). Up to 90% of the limit of all three licenses were used during this period.

Choice of translation targets

The software corpuses included in this translation belong primarily to three categories: (1) Microbial and NGS software that is upstream of the current version of Bascet, (2) Software that is of relevance for image analysis and spatial omics, (3) upstream library dependencies, which in particular cover compression and file format support.

Translation

Naive prompts using Claude to translate software resulted in code that lacked most features, and the replicated features had bugs or gross simplification of algorithms. To get better translation results, a systematic check list of prompts was developed along with translation. The currently recommended prompts are included in Supplemental Material. However, both Codex and Claude have picked up implicit assumptions based on our workspaces, and the prompts may thus not be as efficient for other users. Further independent testing is needed.

The core principles that were used for all later translations (the majority) is that:

  • Each original function should be translated into one Rust function

  • No additional functions should be introduced

  • The logic of each function is independently conserved into the translated function

  • The code should follow the original code origanization, including directory structure

  • Stubs of all functions and structs are filled in, with idiomatic mapping from original code names to Rust snake case (this is the Rust naming convention)

Further prompts are injected (see Supplemental material) but these principles ensure (for Codex) that the translation is principled and conservative, enabling later audit.

Most translations were roughly done according to the following idealized steps:

  1. Static analysis and scaffolding of translation (i.e., write stubs for all functions)

  2. First complete translation of the code

  3. Testing of the code on synthetic data (comparing to original code output)

  4. Testing of the code on real data (comparing to original code output)

  5. Benchmarking and speed optimization

  6. Source code specific cleanup

  7. Addition of idiomatic Rust API; deciding if some features should be optional (e.g., making it possible to omit command line interface)

Codex was used for the initial translation of all later corpuses as it was found to be more conservative, both in terms of verbatim translation of individual functions, but also in following the instructions. Because initial mistranslation compounds later code design, it is crucial that the first steps are done conservatively. Claude was later adopted to primarily resolve bugs that Codex could not find, or to find optimization opportunities.

Translation of C code

C-code translated according to our instructions results in non-idiomatic Rust code, where raw libc functions are frequently used. Such code is not safe and Rust cannot guarantee that the translation has no memory referencing errors, free-after-use errors, nor memory leaks. Such code should be made idiomatic after translation, but with care to not reduce speed or increase RSS in the process. The following is an example prompt used to translate jpegxr (31 kLOC):

/goal Make this code more rust idiomatic. This means data structures, functions etc need to rather use rust typical storage choices. No C types should remain and no C strings should remain. All malloc and free should be replaced by rust idiomatic memory allocation. Replace C strings with vec if not sure if unicode is safe. replace pointers with rust references when equivalent and performance is not lost(!). Use Option to represent possible null values. Whenever possible, replace data structures that are casted with named enums to avoid any need of casting. Ensure that the exceptions or errors that were originally handled are handled by translation. Do not introduce helper functions but keep translation 1-1 as far as possible. Refactor aggressively. Code need not compile at each step. Do not introduce helpers or shims, replace pointers with references optimistically and clean up later (let the code break).

Additional prompts were added during and after goal execution, e.g.:

  • I think TRUE and FALSE can be removed mechanically across all files quite easily

  • I checked the code and it looks like most *mut in function declarations can be replaced with &mut. A lot of errors will appear but they can be handlded in parallel

  • Relax faithful translation: functions related to memory allocation and deallocation need not be translated 1-1. There should be little need to manually free memory in Rust

Initial translation to non-idiomatic Rust (tested and benchmarked) took 10 hours and 15M tokens (Codex). Conversion to idiomatic Rust was performed in two sessions; first 8h with Claude, then 40 min Codex (700k tokens). To verify the API, prompts were given to another agent focused on Bioformats-rs: (paraphased) Is the API rust idiomatic + Refactor the API for your needs. Ensure that the API is general and not only fulfilling the needs of Bioformats. Letting agents edit across code bases is however generally not recommended as it can unintentionally collapse API separation.

Development of static analysis support tools

Early translation attempts were made solely using agentic AI. The following errors in translation occurred commonly:

  • Expressions were swapped for no reason, i.e., a**a, or branches reordered

  • The wrong variable types were used, causing overflows

  • Order of operations was changed, causing floating point math to give different answers

  • Algorithms were replaced with naive alternatives

  • When multiple similar functions were present, the wrong function was called or reused for a different purpose

  • Helper functions were introduced to try to make an existing function able to also solve other problems, cascading into alternative designs that later could not be extended

  • SIMD optimizations were omitted

To mitigate these problems, additional prompts were added to make the agents translate the code more verbatim. However, these prompts were ineffective for Claude, and mostly insufficient for Codex. A concern is that the agentic AI's are not precise enough in tracking the overall code structure. Tracking code structure is however something that classic algorithms can do efficiently and with high precision.

To implement such algorithms, static analysis tools were thus developed using Claude (vibe coded). The library Treesitter (https://github.com/tree-sitter/) was used for parsing the original code into an AST (abstract syntax tree). This in turn enabled algorithmic analysis of code complexity and call graph structure. The agents are then prompted about their location which contains agent handover instructions, enabling them to figure out how to use the tools. Manual reminders were also prompted whenever the translation progress seemed stuck.

The following tools were developed (Figure 1):

  • CCC (Code complexity comparator) compares the call graph of the original code and candidate translation. In addition, it compares functions by the presence of types, symbols (numbers and strings), cyclomatic index, and logical expressions (first order logic). CCC relies on the presence of a mapping file (Rust function vs original function) which can be generated by agentic AI, or a first order guess is made. CCC also has a graphical user interface to compare the translated code vs the original but it was not extensively used for later translations.

  • tracehash aids code instrumentation. The inputs and outputs of functions are hashed, enabling rapid comparison of deviations (assuming pure functions) and difference in how many times the functions are called (higher level logic failure). tracehash also enables output of non-hashed data for selected functions, which can be used to better understand deviations once they have been detected.

  • gdbtv (gdb-translation-verifier-rs) compares the original and translated code in parallel GDB sessions. This is similar to the idea of bisimulation, used to prove equivalence of software. Unlike the first two tools, neither Claude nor Codex reached for this tool automatically, suggesting that their training prefers their first two modes of debugging.

To ensure that the tools are useful in practice, the agentic AI (primarily Claude) was instructed to add features as needed (in the original code repository) during initial use of the tools. In particular, the file describing which original function corresponds to which Rust function was improved to handle ambiguous cases; and lifetime functions (constructors and destructors) are given different treatment to relax the 1-1 correspondence (Rust does not typically need destructors, and constructors are implemented differently).

Validation and use of static analysis support tools

During development, the agents are instructed to use the static analysis tools, and both AI agents invoke them automatically (primarily CCC and tracehash). However, use of these tools turned out insufficient to remove all translation errors.

Later translation attempts focused on improving translation efficiency. After initial scaffolding, parallel agents can be used to do the translation. The translation is performed bottom-up following the call graph extracted by CCC. Qualitatively, it appears that parallelization also made the AI agents fix translation errors faster. A behavior frequently noted is that, once an error was discovered, the AI agents would give most attention to a smaller set of functions rather than perform a broad audit. This was especially the case when searching for performance problems (i.e. the agent analysis a hot spot). However, when the translation failed to progress, the error was typically elsewhere, causing a cascade of problems. Use of parallel subagents enforced a broader search that enabled such problems to be found. As a bonus, by adressing general translation faithfulness instead of one particular error, multiple problems could be fixed in parallel.

While AI agents can easily be prompted to audit code, they will struggle to find all issues if not done systematically on large code bases. Instead, a better strategy is to perform per-file auditing. Agents may also miss problems during one audit session and this must be accounted for. Codex tended to be better at fixing problems and generated less false positives, but both Claude and Codex could be productively used for auditing. The following prompts (paraphrased) were used for systematic audit:

  1. I want to audit each file and fix problems. To ensure all files are checked properly, list all files in TOAUDIT.md along with check boxes. Each file must pass two consecutive audits without complaints before it is considered ok. I want function names to systematically map to Rust snake case. Each original function should have one rust function. Do not add helper functions beyond this. Ensure that all logic is covered and that all functions are retained.

  2. /goal fix all files according to TOAUDIT.md and update file as needed

For VTK, one of the largest codebases, initially translated by Claude, a systematic parallel audit by Codex took about 4 days. However, systematic audits did not remove all errors, as revealed by testing on real data.

Benchmarking

All benchmarks were chosen by Claude or Codex, with some degree of user input to steer the choice. The primary prompts for adding benchmarks were analogous to, e.g., Compare speed, RSS, parity to original code on realistic data, Test on larger real data. Do not simply replicate test data we already have and Test on 100M reads. Additional prompts include ensure that the comparison is fair and ensure that the JVM has been warmed up before measuring. Codex was used for the final benchmarking of each program as Claude frequently generated biased benchmarks.

Because each program has a large number of subprograms and parameters, individivudal benchmarks should not be interpreted as definite; rather, their primary purpose has been used for regression (e.g., to check if algorithms have not been translated properly). However, their overall trend across corpuses is interpreted in this study to give a qualitative idea about the potential for speed improvement across different types of source material.

All benchmarks were performed on an Intel Xeon Gold 6138 2.00GHz, 192GB RAM. Different number of thread counts and input data sizes were stressed. Real data was used for most of the final benchmarking.

Use of agentic AI for manuscript generation

Codex was used to estimate numbers cited in this text, using the following prompts. The numbers must be interpreted qualitatively (typos in prompts included as-is):

  • which software are used for analysis of microbial analysis, covering isolates and metagenomics? + give me a tool count, excluding GUIs → 61

  • I want as estimate of how many packages these tools and libraries need as upstream dependencies. run in parallel + I want transitive package count → 2,000-3,500 unique transitive software packages, excluding databases.

  • "see benchmarks.csv [this is a file listing the github repositories seen in benchmark, Figure 1]. I want to know how many lines of code, excluding unit tests and comments, of rust. developing some utility for this, then use parallel agents → 700,098 Rust LOC, excluding comments and tests".

Draft figures were generated as vector graphics using Claude, and polished using Inkscape. Agentic AI has not been used to edit the text in any way.

Results

Hybrid agentic AI translation is possible for most but not all software

Out of approximately 40 translation attempts, 35 software packages across imaging, NGS (next-generating sequencing) and upstream libraries have been driven to a state where they have been tested, benchmarked, and are ready for use by early adopters (i.e., translation bugs are still expected to be present but we already use the code in production).

The time required to translate a code base varies greatly. The latest translation, jpegxr (C++, 31k lines of code) was translated and verified in less than 20 hours using the latest prompts provided (Materials and Methods; 10% faster without targeted optimization); other programs such as BWAMEM2 required weeks where validation and optimization took most of the time. An important note is that agentic AI by default is single-threaded and does not make efficient use of modern computers. However, the current prompts explicitly instruct the AI agents to parallelize, and this is especially important when using Codex as it is less prone to automatically spawn parallel subagents. Claude spawns parallel agents if the workload is suited. To further parallelize further, the simplest way is to translate multiple codebases in parallel.

All translated packages were optimized but under the constraint that the logic for each function should be retained in its Rust counterpart (1-1 translation). Thus, most optimizations related to improving memory usage, and in particular, to avoid the need to copy data. Common Rust performance tricks are to rather take a reference to a data (instead of a copy), return a reference to a data, or to add data to a preallocated buffer instead of returning a new buffer (useful whenever the buffers can be cleared and reused). Only in rare cases did the translation use the Rust unsafe keyword (which enables safety checks to be disabled) and raw pointers (the Rust default is to instead use a safer start-to-end range as a reference, also called a fat pointer). The AI agents were unlikely to use these unsafe constructs without active prompting but also had a tendency to needlessly copy memory during the first translation pass. However, both Codex and Claude have added code to forget memory, effectively causing memory leaks; thus the user still needs to perform some type of manual audit of the code.

The performance of different translations are compared in Figure 2. In most cases, Rust was able to match or exceed the speed of the original software. The time spent optimizing the code was unevenly distributed, with highly optimized software (C/C++) requiring by far more work to reach speed parity (such as the de novo assembler SKESA, and various compression libraries). The speedup of other translations vary greatly, and since Python and R libraries both tend to embed C/C++ code, the expected speedup is highly case dependent. Java benefitted from a great speedup but Java is notoriously hard to benchmark due to its design. The Java virtual machine (JVM) interprets code by default (like R/Python) but its hotspot compiler detects code that is run repeatedly and compiles the code on the fly. The code must thus first be warmed up to give representative benchmarks. Some care was taken during benchmarking but the bottomline is that Java is not strictly comparable.

The ability to lower or retain memory usage (RSS; resident set size) varies greatly. The fastest software also tends to have the lowest RSS as memory bandwidth is increasingly a bottleneck in modern computers. Dynamically typed languages such as R/Python typically have the highest RSS overhead as they need to track the types of all variables, while the Rust compiler can omit this information during runtime. In several translations, Rust however has higher RSS than the original software. It is likely that all of these cases can be fixed but either requires (1) giving up on the 1-1 translation, or (2) using raw pointers, or (3) using more complex multithreading buffer strategies. Because this affects the possible faithfulness or safety of the code, aggressive attempts at reducing RSS further have not been performed. Finally, Java is again an outlier, where memory is also reserved for the JVM. This memory is somewhat offset for larger workloads but since the JVM both relies on garbage collection and does typing by erasure, memory overhead is largely unavoidable.

Translation of C code is feasible but the output is non-idiomatic Rust. Because the generated code contains a large number of pointers and manual memory management via libc, this code does not capture the safety guarantees of Rust. Such code can however be rewritten to better match Rust idioms using a dedicated translation pass (see Materials and Methods).

Metaprogramming remains a particular challenge for the provided translation approach. Two attempts were made at translating BLAST but neither resulted in a useful product. There are several obstacles to the approach taken in this study related to metaprogramming: (1) BLAST generates source code during the build process, (2) BLAST combines C and C++, (3) BLAST relies heavily on the C++ preprocessor to generate code, (4) BLAST also makes heavy use of templates. The concept of one Rust function per original function breaks down when a large amount of the code revolves around generating functions (i.e., metaprogramming). C++ templates also do not always map to Rust generics, as Rust generic types are required to satisfy certain traits (e.g., additivity), while C++ template types are better thought of as cut and paste. Rust generics also lack type specialization. More research is thus needed in translation strategies for metaprogramming-heavy code.

Translation to Rust removes the need for Conda and containers, and increases portability

Our work on translation was prompted by our issues in developing Zorn/Bascet - the first single-cell preprocessing software aimed at microbial metagenomics. Unlike single-cell RNA-seq analysis of eukaryotes, which has a small number of fixed stages (debarcode, align, count features), single-cell microbial analysis is more akin to genome isolate analysis, for which a large number of packages exist. We aimed to integrate them for single-cell needs but largely failed to deliver a product using established methodology, i.e., packages via Conda, wrapped in a container (Docker or apptainer). The failure was due to a mix of (1) Zorn/Bascet being a complex package and workflow manager, (2) Conda failing to resolve compatible packages, (3) Conda failing to resolve versions in reasonable time, (4) many configurations of containers being required to cover all operating systems and CPU architectures. All of these problems could be resolved by translation.

By translating all upstream dependencies, version management is now solely managed by the Rust package manager Cargo. Rust and Cargo are designed to enable conflicting versions of software to be installed in parallel, unlike R, Python and most other programming languages. This removes the need and issues related to the Conda package manager.

After translation, all upstream software of Bascet are now used as libraries, removing the fragile CLI interface in favour of statically typed Rust functions. But by using the code as libraries, only a single binary is produced, containing the whole ecosystem (Bascet + upstream software; Figure 3a). This removes the need for containers to glue the files together. Thus, we now distribute naked binaries (200mb - 80x smaller than the original container). Together with the removal of Conda, compilation time is also reduced from 20min to about 1-2 min.

Cross-compilation also enabled distribution of native binaries for Apple M3 computers, increasing performance on non-Intel hardware (containers will require hardware emulation). Finally, during translation, we also conditionally disabled or replaced Unix/Posix-specific instructions (e.g., setting of file permissions, posix threads, etc). This has made Bascet the first single-cell toolkit that can run under native Windows (i.e., no need to install WSL2 subsystem).

Translation improves performance primarily for linked libraries

Single-cell analysis requires the processing of a large number of inputs. Modern single cell datasets can encompass up to 100M cells. This amount of data requires highly efficient software and storing the data of each cell in an individual file is not feasible due to file system design. Furthermore, running command line utility on each file would have major overhead. Assuming one second/cell in startup time, this would amount to a total of 1,157 days of just software starting cost. However, many programs load databases (from HMM profiles to entire KMER databases), resulting in loading times up to minutes. It is thus not reasonable to merely call other programs using standard procedures (such as in Nextflow or Snakemake pipelines) in this context.

To remove the software startup overhead, we added high-level APIs to each package, and divided it into separate steps (Figure 3b): (1) input file parsing and loading, (2) database loading, (3) processing, and (4) saving output. Out of these, only processing must be done for each cell. Furthermore, copying of memory (from one variable to another) is an increasingly expensive operation. By adapting the API to operate on input data by reference, zero-copy strategies can sometimes be achieved. Taken together, software such as SKESA and GECCO run over 3 times faster after integration in Bascet (benchmarks part of manuscript in preparation).

Floating point handling is hazardous

Unlike in theoretical math, floating operations are not commutative, i.e., (ab)c ≠ a(bc). This made translation of some software (such as HMMER) especially difficult as the AI agents frequently did not respect order of operations. Handling of integer types (and over/underflow) was also a problem, but generally easier to work out.

The reliance on precise floating point combined with imprecise translation resulted in p-values differing to some degree from the original code. This in turn caused problems during validation, as the small errors escalated to larger ones: (1) Order of results, such as gene lists, being different. (2) Results being omitted due to p-value cut-offs.

All floating point errors are resolved in the translation but certain types of optimizations could not be performed; e.g., SIMD (Single Instruction Multiple Data) are specialized CPU instructions that can greatly improve speed for certain workloads (e.g., vector multiplications). They however reorder the operations, resulting in slightly different results for floating point math. A future possibility is to add SIMD as an optional feature, but with a warning that the results will differ from the original software.

Discussion

This study shows that large-scale translation of open source code is possible using a combination of suitable prompts, static analysis support software, and agentic AIs. This will have particular impact on research software, which has higher technical debt than other open source software. This can be expected, given its higher degree of complexity, combined with a smaller user base. Appropriate use of agentic AI can thus help catch up with the technical debt.

This study focuses on Rust as a possible single language for all bioinformatics applications. We show that performance is comparable to C/C++, and that code written in Java, Perl and Python can be translated with great speed improvements. The memory usage can also be reduced, but idiomatic Rust can also penalize memory usage, likely due to its stricter ownership model. Reducing memory further may require larger rewrites of the code, or the use of unsafe language features.

We show that use of Rust as the single language removes the need for Conda and container infrastructure, removing a major layer of complexity, and largely resolves the issue of how to handle conflicting software versions. The size of distributed software can be greatly reduced. Rust cross-compilation also supports other CPU architectures such as the Apple M3.

A major concern is the extent to which translation introduces bugs in the code. This does happen, and a new cycle of software testing will be required. However, a large number of bugs can be found by simply comparing the output of the translated software vs original code output, on a large number of varied corpuses of real data. This study found it essential to use real world data, and that the tests generated by agentic AIs are largely insufficient. A challenge is that complex software has many settings or algorithms, making comprehensive testing difficult. Because of the 1-1 correspondence during translation (one function in, one function out), it should however be possible to design specialized testing software. Some of the functions are, e.g., similar enough that even mathematical equivalence proofing could be a route forward. It also suggests that specialized translation algorithms (i.e. transpilers) might be able to handle a large amount of translation, while also reducing cost of AI and speeding up translation.

During this translation, at least 3 bugs were uncovered, suggesting that the bugs introduced may be offset by the bugs being resolved. The bugs involved use of a broken dependency (bug since long already resolved but new version not integrated), use-after-free (possible memory corruption) and an unhandled corner case. Further memory related bugs may have been resolved implicitly by removing pointer-based memory management (malloc/free). A fair assumption is that the agentic AI technology will keep maturing to the point where bugs as a whole will be net negative.

In the case of Bascet, we directly interfaced upstream software through Rust functions. Besides possibly massive speed increases, this is the only way that safety can be guaranteed across software. Current bioinformatics practice frequently makes use of untyped Unix pipes, but this approach does not guarantee that the data is compatible across software. The translation approach removes the need for pipes, and named pipes, neither of which are compatible with Windows. Native Windows compatibility is increasingly important as many users now use Ubuntu-in-windows (WSL2), making it impossible to run Linux commands from RStudio.

Every line of code is a liability. By maintaining translations rather than producing greenfield implementations, the increase in total software maintenance burden is limited. In addition, provenance is clarified, addressing the major concern about the copyright of AI-generated code. During this study, the translations for two programs (openslide, minibwa) were made to track newer versions. The 1-1 translation ensured that upstream changes were localized to well-defined spots in the translation, enabling updates taking less than one hour. However, many upstream changes were related to the build system, documentation, and other aspects not affecting the translation. This suggests that a large number of repositories could be maintained with little effort using parallel AI agents. However, this would put the bioinformatics ecosystem at risk of supply-chain attacks, as has plagued the Javascript community (via the Node package manager). A more conservative approach to updating may thus be suitable.

The translations also highlighted the challenges in precise p-value reproduction. This will become an even larger problem as bioinformatics code starts to use GPUs, as for the sake of performance, some operations are not guaranteed to occur in deterministic order. A partial solution is to ban to avoid or ban p-value cutoffs, in favour of probabilistic weighing schemes that are less sensitive to rounding errors. As p-value cutoff tuning enables p-hacking, and handling of larger data has become easier, the field is overdue in taking this step.

Finally, while this study shows the feasibility of translation, there is still ample room for better translators and verifiers. Translation guidelines have already been generated (e.g. https://rewrites.bio/), but the sudden appearance of the AI technology calls for a rethinking of bioinformatics. The ability to refactor software at massive scale enables strategic opportunities for better integration of our many software packages, resolving big challenges: Compute needs can be reduced, barriers to adoption removed, environmental impact reduced and the reliability of software increased - which is especially important as NGS is adapted in the clinic. As such, this work hopefully inspires greater future endeavours in bioinformatics.

Improvements for AI systems

Based on this paper, here are specific improvements to AI systems and what the improved systems can do:

1. Add static-analysis-guided translation constraints

  • Improve: Integrate AST parsing (e.g., Tree-sitter) into the agent’s translation loop to enforce 1:1 function mapping, preserve call graphs, and block the introduction of helper functions.

  • Improved system can: Translate legacy code (C, Perl, Java, Python) to Rust with verifiable structural equivalence, reducing silent logic errors like swapped branches or wrong variable types.

2. Implement automated equivalence verification

  • Improve: Add tools that hash function inputs/outputs (tracehash-style) and run parallel GDB sessions (bisimulation-style) to compare original vs. translated code automatically, not just on demand.

  • Improved system can: Detect translation bugs (e.g., floating-point order changes, integer overflows, omitted SIMD) during development, and flag them before real-data testing, cutting debugging time from weeks to hours.

3. Enable conservative, audit-driven translation modes

  • Improve: Add a “faithful translation” mode where the agent must (a) not add new functions, (b) preserve directory structure, (c) map names to snake case, and (d) pass two consecutive per-file audits before marking complete.

  • Improved system can: Produce auditable, maintainable translations that track upstream changes (e.g., openslide, minibwa) in under an hour, reducing technical debt and enabling parallel maintenance of many repositories.

4. Add memory-safety and idiom refactoring passes

  • Improve: After initial translation, run a dedicated pass that replaces raw pointers, malloc/free, and C strings with Rust references, Vec, and Option, while preserving performance (no needless copies).

  • Improved system can: Convert unsafe C-derived Rust into idiomatic, memory-safe Rust without regressions, eliminating use-after-free and memory leaks that agents currently introduce.

5. Implement parallel subagent orchestration for broad audits

  • Improve: When an error is found, force the agent to spawn parallel subagents to audit all related functions, not just the suspected hotspot, to avoid cascade failures from misdiagnosis.

  • Improved system can: Find and fix multiple translation errors simultaneously, reducing the “tunnel vision” problem where agents fix one symptom while missing the root cause elsewhere.

6. Add floating-point determinism checks

  • Improve: Add a verification step that compares p-values, gene lists, and result ordering between original and translated code, with tolerance settings, and flag any operation reordering that changes results.

  • Improved system can: Guarantee that translated bioinformatics tools produce statistically identical outputs, avoiding omitted results due to p-value cutoffs and enabling safe clinical deployment.

7. Support metaprogramming-heavy code translation

  • Improve: Extend translation to handle C++ templates, preprocessor-generated code, and build-time code generation (e.g., BLAST) by using a hybrid approach: static scaffolding for generated functions, then manual/agentic mapping to Rust generics or macros.

  • Improved system can: Translate complex codebases like BLAST, which currently fail, by treating metaprogramming as a separate translation stage rather than a 1:1 function mapping.

8. Provide cross-compilation and dependency-free output

  • Improve: Automatically strip Unix/Posix-specific calls (e.g., file permissions, pthreads) and generate native binaries for Windows, macOS (Apple M3), and Linux without containers.

  • Improved system can: Produce portable, single-binary bioinformatics tools that run natively on all major OSes, removing Conda/container dependencies and reducing distribution size by 80x.

9. Add API-level integration for library use

  • Improve: After translation, generate idiomatic Rust APIs that split software into stages (parsing, database loading, processing, saving) and support zero-copy references, so translated code can be used as libraries, not just CLIs.

  • Improved system can: Enable pipelines like Bascet to call upstream tools as functions, eliminating startup overhead (e.g., 1,157 days saved for 100M cells) and achieving >3x speedups by avoiding per-file process launches.

10. Implement automated regression benchmarking

  • Improve: Add a benchmarking module that automatically compares speed, RSS, and output parity on real data (not synthetic), with prompts to “warm up JVM” and “test on 100M reads” to ensure fair comparisons.

  • Improved system can: Continuously validate translations against original software, catching performance regressions and correctness drift, and guiding optimization efforts to where they matter most.

Abstract

The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be written in an now-unfamiliar language (e.g. Perl, Fortran). This incurs maintenance cost (technical debt), but dynamically typed languages also negatively impacts the environment and fail to make use of modern hardware. Legacy code may also have security or safety problems that make it unsuited for use in clinical settings. Here we show that agentic AI, combined with static analysis, can be used to translate legacy code to the modern language Rust. We provide prompts and supporting software to aid systematic translation, and evaluate it on common software for NGS and imaging. We showcase the result on our software Bascet: Size was reduced by 80x, build time decreased by 10x, and performance of key steps improved >3x. Unix dependencies were also removed, making Bascet the only single-cell pipeline able to run on native Windows, without a container. Large-scale refactoring of bioinformatics software is thus now possible at a limited budget, enabling more complex tools to be developed.

Sources

Related papers