Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFT
cs.SE, cs.AI, cs.ET, cs.LG
Submitted: 2026-08-13
Updated: 2026-08-13
Code: https://github.com/moxin-org/C2Rust
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Translating C code into safe, idiomatic Rust is a longstanding software-engineering goal because it can eliminate entire classes of memory-safety vulnerabilities while preserving the functional
Terminology
Abstract
Translating C code into safe, idiomatic Rust is a longstanding software-engineering goal because it can eliminate entire classes of memory-safety vulnerabilities while preserving the functional behavior of legacy systems. Large language models (LLMs) have shown promise for this task but typically underperform when applied off-the-shelf, since general-purpose pretraining rarely emphasizes idiomatic Rust generation, cross-language semantic equivalence, or the ability to reason about and repair compiler/runtime feedback. In this report we describe a three-stage fine-tuning curriculum applied to Qwen3-27B that is designed to progressively specialize the model for the C-to-Rust (C2Rust) translation task: (1) continued pretraining on Rust-centric corpora to strengthen the model's prior over idiomatic Rust syntax and standard-library usage; (2) supervised fine-tuning (SFT) on the microsoft/Verus Training Data dataset to instill debugging and self-repair behavior over Rust code; and (3) task-specific SFT on paired C/Rust solutions derived from LeetCode problems to teach direct semantic translation. We evaluate the resulting model using the agentic, static-analysis-guided verification framework of SACTOR, which performs structure-aware, two-phase (unidiomatic to idiomatic) translation with foreign-function-interface (FFI)-based end-to-end (E2E) testing. We report success rate, idiomaticity (Clippy lint counts, unsafe-code fraction), and failure-mode analyses, and compare our fine-tuned model against baseline Qwen3-27B and other LLMs evaluated under the same framework.
Sources
- DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation
- The Llama 3 Herd of Models
- Mercury: Ultra-Fast Language Models Based on Diffusion
- Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality
- A Survey on Diffusion Language Models
- Effective MoE-based LLM Compression by Exploiting Heterogeneous Inter-Group Experts Routing Frequency and Information Density
- DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance
- Efficient Reasoning with Hidden Thinking
- Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs
- Qwen3.5-Omni Technical Report
- EvoC2Rust: A Skeleton-guided Framework for Project-Level C-to-Rust Translation
- Dream-Coder 7B: An Open Diffusion Language Model for Code
- Dream 7B: Diffusion Large Language Models
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- 7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties