SuperCoder: Assembly Program Superoptimization with Large Language Models
summary
The gist
SuperCoder: Assembly Program Superoptimization with Large Language Models This paper investigates whether large language models (LLMs) can serve as superoptimizers, generating assembly programs that
In short
The episode details a paper titled "SuperCoder," which explores using Large Language Models for assembly program superoptimization. Researchers trained a fine-tuned model on over 8,000 programs using reinforcement learning. The study concludes that SuperCoder achieved 95% correctness and an average speedup of 1.46x over the standard GCC -O3 compiler output.
Key concepts
- Superoptimization
- This is a computer science term defined as taking an existing program and finding the absolute fastest version of it, ensuring that the optimized code performs exactly the same function as the original code.
- GCC -O3
- This represents the highest optimization level achieved by traditional compilers. It serves as the performance baseline for SuperCoder, which is designed to find optimizations that surpass what this industry-standard compiler can achieve.
- Reinforcement Learning (RL)
- A training methodology used in SuperCoder where the model learns through a reward function. The model receives rewards for speedup and zero points if the code is incorrect, forcing it to learn optimal performance while maintaining correctness.
- Test-based Validation
- A method of verifying program correctness that involves running the code against various inputs and checking if the outputs match expected results. This approach allows scaling to complex programs with loops, unlike previous methods.
Terminology used across episodes
This episode discusses
- SuperCoder: Assembly Program Superoptimization with Large Language Models · Paper Radio
- GPT-4 Technical Report
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Program Synthesis with Large Language Models
- Kevin: Multi-Turn RL for Generating CUDA Kernels
- Learning to superoptimize programs
- Evaluating Large Language Models Trained on Code
- Meta Large Language Model Compiler: Foundation Models of Compiler Optimization
- StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
- Mercury: A Code Efficiency Benchmark for Code Large Language Models
- Equivalence Checking of ML GPU Kernels
- CodeMonkeys: Scaling Test-Time Compute for Software Engineering
- Compiler generated feedback for Large Language Models
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Coding Challenge Competence With APPS
- Qwen2.5-Coder Technical Report
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- StarCoder: may the source be with you!
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
- DeepSeek-V3 Technical Report
The paper
SuperCoder: Assembly Program Superoptimization with Large Language Models · Read on arXiv
Anjiang Wei, Tarun Suresh, Huanmi Tan, Yinglun Xu, Gagandeep Singh, Ke Wang, Alex Aiken
Stanford University · University of Illinois Urbana-Champaign · Carnegie Mellon University · Nanjing University
Superoptimization is the task of transforming a program into a faster one, and ideally the very fastest possible one, while preserving its input-output behavior. In this work, we investigate whether large language models (LLMs) can serve as superoptimizers, generating assembly programs that outperform code already optimized by industry-standard compilers in end-to-end runtime. We construct the first large-scale benchmark for this problem, consisting of 8,072 assembly programs averaging 130 lines, in contrast to prior datasets restricted to 2-15 straight-line, loop-free programs. We evaluate 23 LLMs on this benchmark and find that the strongest baseline, Claude-opus-4, achieves a 51.5% test-passing rate and a 1.43x average speedup over gcc-O3. To further enhance performance, we fine-tune models with reinforcement learning, optimizing a reward function that integrates correctness and performance speedup. Starting from Qwen2.5-Coder-7B-Instruct (61.4% correctness, 1.10x speedup), the fine-tuned model SuperCoder attains 95.0% correctness and 1.46x average speedup, with additional improvement enabled by Best-of-N sampling and iterative refinement. Our results demonstrate, for the first time, that LLMs can be applied as superoptimizers for assembly programs, establishing a foundation for future research in program performance optimization beyond compiler heuristics.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "SuperCoder: Assembly Program Superoptimization with Large Language Models".
Jane: The paper was written by Anjiang Wei, Tarun Suresh, Huanmi Tan, Yinglun Xu, Gagandeep Singh et al. from Stanford University and University of Illinois Urbana-Champaign and Carnegie Mellon University and Nanjing University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everybody. Today we’re digging into a paper that’s got a pretty bold title: “SuperCoder: Assembly Program Superoptimization with Large Language Models.” Jane, when you first saw that title, what jumped out at you?
Jane: Honestly, Tom, the word “superoptimization” is what got me. It sounds like something out of a superhero movie, but it’s actually a real computer science term. It means taking a program and trying to find the absolute fastest version of it that still does exactly the same thing.
Tom: Right, and the paper is from folks at Stanford, UIUC, Carnegie Mellon, and Nanjing University. They’re basically asking: can a large language model do this better than the compilers that have been refined for decades?
Jane: And that’s the wild part. Compilers like GCC are the result of like forty years of expert tuning. The idea that an AI could look at the output of GCC’s highest optimization level and say, “I can do better,” is kind of audacious.
Tom: Audacious is the word. But they didn’t just ask the question in theory. They built a massive benchmark of over eight thousand assembly programs to test it. That’s a huge step up from the tiny, loop-free examples that superoptimization research has been stuck on for years.
Jane: Exactly. Prior datasets were like two to fifteen lines of straight-line code. These programs average a hundred and thirty lines, and they include loops, function calls, the whole deal. So this is the first time anyone has tried to scale this problem up to something that looks like real software.
Tom: And the implications are pretty big. If an LLM can find optimizations that GCC misses, that means there’s still performance on the table that traditional compiler heuristics just can’t reach.
Jane: Right. And it’s not just about making benchmarks faster. It’s about whether we can build tools that automatically squeeze more speed out of the code that runs our data centers, our phones, our cars. That’s the long-term prize here.
Tom: So stick around, because we’re going to break down how they actually pulled this off, what the results were, and whether this is going to change how we think about code optimization forever.
Jane: And we’ll get into the nitty-gritty of the reinforcement learning they used, because that’s where the magic really happens.
Summary: Tom: So, Jane, we’ve got the title. Now let’s talk about what the paper actually claims. The summary is pretty punchy: they say that a fine-tuned model called SuperCoder achieves ninety-five percent correctness and a one point four six times average speedup over GCC’s-O3 output. That’s not a typo, right?
Jane: No typo, Tom. And let me put that in plain terms. If your program takes ten seconds to run with the best compiler settings, SuperCoder can, on average, make it run in about six point eight seconds. That’s a thirty percent reduction in runtime, just by rewriting the assembly code.
Tom: And that’s after the compiler has already done its best job. The baseline they’re comparing against is GCC-O3, which is the highest optimization level most people ever use in production.
Jane: Right. And the really interesting part is that they didn’t just evaluate one model. They tested twenty-three different LLMs, from open-source ones like Qwen and Llama to commercial ones like Claude and GPT. Most of them were terrible at this task.
Tom: Terrible is putting it mildly. Some of them couldn’t even produce compilable assembly. DeepSeek-R1, which is a reasoning powerhouse, got a zero percent compilation rate because it just wrote essays about optimization instead of actually doing it.
Jane: Yeah, that’s a fascinating failure mode. But the best commercial model, Claude Opus four did manage a fifty-one point five percent test pass rate and a one point four three times speedup. So the capability exists in some models, but it’s not consistent.
Tom: And that’s where their contribution comes in. They took a smaller, open-source model, Qwen2 point 5-Coder-7B, which was already pretty decent, and they fine-tuned it with reinforcement learning. The base model had sixty-one percent correctness and a one point one times speedup. After training, it jumped to ninety-five percent correctness and one point four six times speedup.
Jane: So the reinforcement learning isn’t just a small tweak. It’s the difference between a model that occasionally gets lucky and a model that reliably produces faster, correct code. That’s a massive improvement.
Tom: And it shows that the task isn’t just about raw model size or pretraining data. It’s about training the model specifically for this objective, which is optimizing runtime while preserving correctness.
Jane: Exactly. And that’s the core insight of the whole paper. LLMs can be superoptimizers, but they need to be taught to do it. The raw capability is there, but it has to be unlocked.
Tom: So next, we should probably dig into how they built this benchmark and why it’s such a big deal compared to what came before.
Improvements: Tom: Alright, Jane, let’s talk about what this paper actually improves over the state of the art. Because it’s not just “we tried a thing and it worked.” They had to build the whole playground from scratch.
Jane: Right. The biggest improvement is the dataset itself. Prior superoptimization work was stuck on tiny, straight-line programs with no loops. This paper introduces a benchmark of eight thousand seventy-two assembly programs, averaging a hundred and thirty lines each, and they include loops, branches, and real control flow.
Tom: That’s a massive leap. And they didn’t just grab random code. They pulled from CodeNet, which is a huge corpus of competitive programming submissions. So these are real programs that solve real problems, not synthetic toy examples.
Jane: And they were careful about the test cases too. They re-ran every program on its inputs to generate correct outputs, because a lot of the original CodeNet submissions were buggy or failed some tests. That means the ground truth they’re optimizing against is actually solid.
Tom: And they report that their test suites achieve ninety-six point two percent line coverage and eighty-seven point three percent branch coverage. So when a model passes all the tests, it’s not just passing a couple of easy cases. It’s actually exercising most of the code paths.
Jane: That’s a big deal because the correctness check is the weak point of any test-based approach. You can’t prove a program is correct for all inputs, but with that level of coverage, you can be pretty confident you’re not letting broken code slip through.
Tom: And then there’s the training methodology. They used reinforcement learning with a reward function that gives zero points for incorrect code and the actual speedup for correct code. No partial credit for getting close.
Jane: That’s a really clean design. It forces the model to learn that correctness is a hard requirement, not something you can trade off for speed. And it worked. The fine-tuned model, SuperCoder, went from sixty-one percent to ninety-five percent correctness.
Tom: And they also compared PPO and GRPO, two different reinforcement learning algorithms. They got nearly identical results, which suggests the reward function is doing the heavy lifting, not the specific algorithm.
Jane: Right. And they also tried supervised fine-tuning, where you just show the model good examples and ask it to imitate them. That worked too, but not as well as reinforcement learning. The RL approach is better because it directly optimizes for the thing you care about: speedup.
Tom: So the improvements are threefold. A realistic benchmark, a training method that works, and a demonstration that the whole pipeline can beat a production compiler. That’s a solid contribution.
Jane: And it opens the door for a lot of follow-up work. But before we get ahead of ourselves, let’s look at the actual first page of the paper and see how they frame the problem.
First Page: Tom: So Jane, we’ve been talking around the paper, but let’s actually look at the first page. The abstract is where they lay out the whole thesis, and it’s pretty direct.
Jane: It is. They start by defining superoptimization as the task of transforming a program into a faster one, ideally the fastest possible, while preserving behavior. And then they ask whether LLMs can do that for assembly code that’s already been optimized by GCC.
Tom: And the key phrase there is “already been optimized.” They’re not starting from scratch. They’re starting from the output of a compiler that has been tuned for decades, and they’re trying to find the leftover performance.
Jane: Right. And the first page also makes a clear distinction between their approach and prior work. Traditional superoptimization used search algorithms and formal verification, which only works for tiny, loop-free programs. Their approach uses test-based validation, which scales to real programs with loops.
Tom: And that’s a fundamental trade-off. Formal verification gives you a proof of correctness, but it doesn’t scale. Test-based validation scales, but you’re trusting your test suite. They mitigate that with high coverage, as we discussed.
Jane: Exactly. And the first page also introduces the reward function concept. They’re training the model to maximize speedup, but only if the code is correct. That’s the core of their reinforcement learning setup.
Tom: And they mention that they evaluated twenty-three LLMs. That’s a lot of compute, but it gives a really clear picture of where the field stands. Most models can’t do this task at all, and only a few can do it well.
Jane: The best commercial model, Claude Opus four got a one point four three times speedup. But their fine-tuned model, SuperCoder, beat that with one point four six times, and it’s a seven-billion parameter model that runs locally. That’s a big deal for accessibility.
Tom: So the first page sets up the whole story. The problem is hard, the prior approaches don’t scale, and they’re proposing a new way that combines LLMs with reinforcement learning and test-based validation.
Jane: And it’s not just a theoretical proposal. They built it, they tested it, and they showed it works. That’s what makes this paper exciting.
Tom: So we’ve covered the title, the summary, the improvements, and the first page. Let’s bring in some other voices to talk about what this means for the future.
Conclusion: Tom: Alright, we’ve spent a lot of time on “SuperCoder: Assembly Program Superoptimization with Large Language Models.” Let’s wrap it up with our crew. Lu, what’s the big picture here?
Lu: The big picture is that we’ve been treating compilers as a black box for decades. This paper shows that an LLM can look inside that black box and find things the compiler missed. That’s a paradigm shift. It means the ceiling for code optimization isn’t fixed anymore.
Meng: And from a practical standpoint, the fact that they got a one point four six times speedup on average, with a seven-billion parameter model that can run on a single GPU, is huge. That’s not a research toy. That’s something you could actually deploy in a build pipeline.
Jane: And the correctness rate of ninety-five percent is the thing that makes it deployable. If the model produced wrong code half the time, nobody would use it. But at ninety-five percent, you can have a human review the output and still save a ton of time.
Tom: And Lalam, you’re the in-house language model. What do you think is the most impactful vision here?
Lalam: The most impactful vision is that this becomes a standard tool in every developer’s toolkit. Not just for squeezing performance out of data center code, but for teaching us about what compilers are missing. Every time the model finds an optimization, that’s a lesson we can feed back into the compiler itself. It’s a feedback loop that makes the whole ecosystem faster over time.
Lu: That’s a great point. The paper even categorizes the types of transformations the model discovers, like loop restructuring and instruction selection. Those categories could be used to improve GCC and LLVM directly.
Meng: And the fact that they open-sourced the benchmark means other researchers can build on this. We’re going to see a wave of papers trying to improve on SuperCoder, and that’s exactly what we need.
Jane: So to summarize: this paper proves that LLMs can be superoptimizers, it gives us a realistic benchmark to measure progress, and it shows that reinforcement learning is the right way to train for this task.
Tom: And it’s not just a one-off result. The authors show that Best-of-N sampling and iterative refinement can push the speedup even higher. So there’s a clear path to making this even better.
Lu: And the limitations are honest too. They rely on test-based validation, which isn’t a proof of correctness. But with ninety-six percent line coverage, it’s a very strong signal.
Tom: Alright, we’ve covered a lot of ground on “SuperCoder.” It’s a paper that’s going to get cited a lot, and I think it marks a real turning point in how we think about code optimization. Thanks to everyone for joining the discussion.
Jane: And to our listeners, if you’re working on compilers, or LLMs, or just care about making software faster, this is a paper you need to read. We’ll be back next time with another exciting paper from arXiv. Until then, keep optimizing.
Tom: Take care, everyone.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization