Efficient Code Embeddings from Code Generation Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Efficient Code Embeddings from Code Generation Models".
Jane: The paper was written by Daria Kryvosheieva, Saba Sturua, Michael Günther and Han Xiao from Massachusetts Institute of Technology and Jina AI GmbH.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, welcome back to the arXiv radio hour. Today we’re looking at a paper called “Efficient Code Embeddings from Code Generation Models,” and honestly, the title alone tells you a lot about where the field is heading.
Jane: It really does, Tom. So normally, when you want to search through code, you use an embedding model — basically a system that turns text into numbers so you can compare how similar two things are. And for a long time, people built these from scratch, or they adapted models that were designed for general text.
Tom: Right, and this paper from Jina AI does something different. They took existing code generation models — the kind that write code when you give them a prompt — and turned those into embedding models. That’s a clever shortcut, because those generation models already understand code deeply.
Jane: Exactly. And the models they used are called Qwen2 point 5-Coder, in two sizes: half a billion parameters and one and a half billion. Those are pretty small by today’s standards, which is part of why the efficiency angle matters.
Lu: I’d love to jump in here, because from a research perspective, this is a really interesting move. Code generation models are trained on massive amounts of code and natural language together. So when you repurpose them for embeddings, you’re inheriting all that understanding without having to train a new model from zero.
Meng: But I’ve got to ask the practical question — how do you actually turn a model that predicts the next token into something that gives you a fixed-size vector for retrieval? That’s not a trivial engineering problem.
Jane: That’s a great question, Meng. The trick they use is called last-token pooling. The model reads the whole input, and then you take the hidden state at the very last token and use that as your embedding. They tested other methods, like averaging all tokens or using a fancy attention-based pooling, but last-token worked best for them.
Tom: And they also add little instruction prefixes before the text, like “find the most relevant code snippet given this query.” That tells the model what kind of task it’s doing, which helps it produce better embeddings.
Lu: The implications here are pretty big. If you can get state-of-the-art retrieval performance from a model that’s a fraction of the size of the big general-purpose ones, that changes the economics of building code search tools.
Meng: Yeah, I mean, smaller models mean faster inference, lower memory usage, and cheaper deployment. For a startup trying to build a code assistant, that’s a real difference.
Jane: And that’s exactly what we’re going to dig into next — how they trained these models and what kind of data they used. Stay with us.
Summary: Tom: So we’ve established that “Efficient Code Embeddings from Code Generation Models” takes code generators and turns them into retrievers. But how did they actually train these things, Jane?
Jane: Good question. They used something called contrastive learning. You take pairs of things that should be similar — like a natural language query and the code that answers it — and you train the model to put those close together in the embedding space, while pushing unrelated pairs apart.
Lu: And the clever part is the data. They didn’t just use the usual docstrings and comments. They pulled from a bunch of sources: benchmark training sets, code search datasets like CoSQA+, and even adapted datasets that were originally built for other purposes.
Meng: I saw they also used GPT-4o to generate synthetic data. That’s interesting — you’re using one AI to teach another.
Jane: Exactly. They used it for areas where real data is scarce, like translating deep learning code between frameworks, or generating programming solutions in multiple languages. They validated the synthetic stuff by manually checking samples.
Tom: And the training itself was pretty quick, right? Like eight hours for the smaller model on four GPUs?
Jane: Yeah, about eight and a half hours for the half-billion parameter model, and twelve hours for the bigger one. That’s remarkably efficient for training an embedding model from scratch.
Lu: It’s efficient because they’re not starting from random weights. They’re starting from a model that already knows how to write code. The contrastive training is just teaching it to organize that knowledge into a retrieval-friendly space.
Meng: So what does the final evaluation look like? Because I’m always skeptical of claims like “state of the art” — what are they actually comparing against?
Jane: They tested on the MTEB-CoIR benchmark, which is a standard set of ten code retrieval tasks, plus some other benchmarks like CoSQA and HumanEval. And their models beat comparable-sized general-purpose models like Qwen3-Embedding-0 point 6B.
Tom: And here’s the kicker — they also beat much bigger models, like jina-embeddings-v4 and even Gemini Embedding on the average score. Their overall average was around seventy-eight to seventy-nine percent, while those big models were in the seventy-four to seventy-seven range.
Lu: That’s the kind of result that makes people sit up and take notice. It suggests that specialization and smart initialization can beat sheer scale for this particular task.
Meng: But I want to know more about what tasks they actually improved on. Because code retrieval isn’t one thing — it’s searching by natural language, searching by code, answering questions, finding completions.
Jane: And that’s exactly what we’re going to break down next — the different task categories and how they tailored the training for each one.
Improvements: Tom: Welcome back. So Jane, we left off with Meng asking about the different kinds of code retrieval tasks. What did the paper actually find there?
Jane: They broke it down into five categories: natural language to code, technical question answering, code to code, code to natural language, and code to completion. For each one, they created a specific instruction prefix.
Lu: That’s a really thoughtful design choice. Instead of one generic instruction, each task gets its own prompt, like “find an equivalent code snippet” for code-to-code, or “find the most relevant answer” for technical QA. That guides the model to produce embeddings that are optimized for that specific kind of search.
Meng: So does that mean you need to know what kind of search you’re doing before you embed your documents? Because that could complicate things in practice.
Jane: Actually, that’s the clever part. They have different prefixes for queries and documents. So when you’re indexing a code snippet, you use the document prefix like “candidate code snippet.” When someone searches, you use the query prefix. That way, the system knows what role each piece of text plays.
Tom: And the results show this works. On code-to-code tasks like CodeChef, they got over ninety-four percent. On natural language to code tasks like CoSQA, they were competitive with much larger models.
Lu: What I find really interesting is the Matryoshka representation learning. That’s a technique where the embedding is trained so you can truncate it — use only the first few hundred dimensions instead of the full vector — and still get decent performance.
Meng: That’s huge for real-world deployment. You can trade off precision for memory and speed depending on your use case. A mobile app might use a shorter embedding, while a server can afford the full one.
Jane: Exactly. And they also did an ablation study on pooling methods, which is where they confirmed that last-token pooling beats mean pooling and latent attention pooling. That kind of rigorous testing is what makes the paper solid.
Tom: So the improvements here aren’t just about the final numbers — it’s about the whole design philosophy. Start with a model that already understands code, give it task-specific guidance, and make the output flexible.
Lu: And that’s a philosophy that could extend beyond code. You could imagine applying this same approach to other specialized domains, like legal documents or medical records, where you have strong generative models but weak retrieval systems.
Meng: I’d love to see how this holds up in production, though. Benchmarks are one thing, but real-world codebases are messy — huge files, weird formatting, legacy code.
Jane: That’s a fair point, and it’s actually a good segue into wrapping up. Let’s bring it all together.
Conclusion: Tom: Alright, let’s wrap up our discussion of “Efficient Code Embeddings from Code Generation Models.” Jane, what’s the big takeaway for our listeners?
Jane: The big takeaway is that you don’t need a massive model to get great code retrieval. By starting with a pre-trained code generation model and fine-tuning it with contrastive learning, these researchers got state-of-the-art results with models that are just half a billion and one and a half billion parameters.
Lu: And they did it efficiently — a few hours of training on four GPUs. That’s accessible to a lot of research groups and companies, not just the big tech giants.
Meng: From my side, the practical impact is clear. Smaller models mean you can deploy code search in more places — on edge devices, in CI pipelines, in IDEs without a huge cloud backend. The Matryoshka truncation adds even more flexibility.
Tom: And they beat models that are several times larger, which is a pretty strong statement about the value of specialization over brute force.
Jane: It really is. And it opens up a question for the future — if this works for code, what else can we apply it to? Scientific papers? Legal documents? Medical records?
Lu: I think that’s the most exciting part. This is a blueprint for building efficient, specialized embedding models in any domain where you have a strong generative model but weak retrieval.
Meng: And the fact that they shared the training recipe — the data sources, the hyperparameters, the ablation studies — means other teams can replicate and build on this.
Tom: So we’ll be watching to see where this goes. For now, we’re saying goodbye to “Efficient Code Embeddings from Code Generation Models” and getting ready to look at the next paper on the arXiv. Thanks for listening, everyone.
Jane: See you next time.
Daria Kryvosheieva, Saba Sturua, Michael Günther, Han Xiao
Massachusetts Institute of Technology · Jina AI GmbH
cs.CL, cs.AI, cs.IR
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: 9 pages, table and evaluations 5-9
Code: https://github.com/ethancaballero/description2code
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 59/100
The gist: The paper introduces jina-code-embeddings, a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify semantically
Key concepts
- Embedding Model
- A system that converts text, such as code, into numerical vectors. These vectors allow users to mathematically compare how similar two pieces of text are, which is fundamental for building effective code search and retrieval tools.
- Last-Token Pooling
- This is the specific engineering trick used to adapt a sequence generation model into an embedding model. Instead of averaging all tokens in the input, the system takes only the hidden state generated at the very last token to create the fixed-size vector representation.
- Contrastive Learning
- A training methodology used to teach similarity. The model is trained using pairs of data—for example, a query and its correct code answer—to ensure that similar pairs are positioned close together in the embedding space while pushing unrelated pairs apart.
Terminology
Summary
The paper introduces jina-code-embeddings, a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify semantically similar code snippets across programming languages. The suite consists of two models: jina-code-embeddings-0.5b and jina-code-embeddings-1.5b, with 494 million and 1.54 billion parameters, respectively.
The authors note that "the rapid adoption of AI-powered development environments like Cursor and Claude Code has transformed software engineering, with code embedding models serving as a critical foundation for retrieval and context engineering of these systems. They argue that while code generation models like Codex can directly synthesize code from natural language prompts,
practical code generation requires contextual understanding of existing codebases, API usage patterns, and integration requirements," positioning code generation systems as retrieval-augmented generation (RAG) architectures where embedding models serve as the critical retrieval component.
A key limitation identified is that "current code embedding models face a fundamental training data limitation. Supervised training typically relies on aligned data such as inline comments, documentation strings, and pedagogical examples from technical documentation—sources that provide insufficient semantic grounding for complex real-world development scenarios. In contrast,
the abundant unaligned code and natural language documentation used to train modern LLMs remains largely underutilized for embedding model development."
The models employ an autoregressive decoder backbone architecture
built on the Qwen2.5-Coder-0.5B and Qwen2.5-Coder-1.5B backbones, which are described as very compact LLMs.
The final hidden layer is transformed into an embedding via last-token pooling. The authors state: We found, after some experimentation, that last-token pooling gave us better performance than mean pooling or latent attention pooling, as documented in Appendix B.
The authors analyzed downstream code embedding tasks and divided them into five categories
: Natural language to code retrieval (NL2Code), technical question answering (TechQA), code-to-code retrieval (Code2Code), code to natural language retrieval (Code2NL), and code to completion retrieval (Code2Completion). For each task, they created English instruction strings that prefix texts passed to the model, with different prefixes for queries and documents:
-
NL2Code: Query prefix
Find the most relevant code snippet given the following query:
; Document prefixCandidate code snippet:
-
TechQA: Query prefix
Find the most relevant answer given the following question:
; Document prefixCandidate answer:
-
Code2Code: Query prefix
Find an equivalent code snippet given the following code snippet:
; Document prefixCandidate code snippet:
-
Code2NL: Query prefix
Find the most relevant comment given the following code snippet:
; Document prefixCandidate comment:
-
Code2Completion: Query prefix
Find the most relevant completion given the following start of code snippet:
; Document prefixCandidate completion:
The models were initialized with weights from the pre-trained Qwen2.5-Coder backbones and trained with a contrastive objective using the InfoNCE loss function.
Pairs of inputs are classed as related or unrelated, and the model learns to embed related items closely together and unrelated items further apart.
The training also uses Matryoshka representation learning to produce truncatable embeddings, so users can make flexible trade-offs between precision and resource usage.
The training data consists of query-document pairs for a variety of code retrieval tasks, largely using docstrings, comments, commit messages, and problem statements as queries, and matching code snippets, diffs, or answers as documents.
Sources include the training splits of MTEB code tasks and the non-MTEB code retrieval dataset CoSQA+,
as well as adapted public datasets originally created for other purposes. Additionally, GPT-4o was used to synthetically generate datasets when available data is scarce,
with synthetic examples validated by manual inspection.
Specific synthetic datasets include SyntheticDLTrans (generated deep learning code translations between frameworks) and a multilingual extension of the CodeChef dataset, generating solutions in eight additional programming languages. This yielded three tasks: CodeChefP2S (problem-to-solution), CodeChefS2S (monolingual solution-to-solution), and CodeChefXLang (crosslingual solution-to-solution).
In each training step, a batch of n query-document pairs is sampled, normalized embeddings are generated, and a similarity matrix is constructed using cosine similarity. The InfoNCE loss is applied with temperature τ = 0.05, batch size n = 512 for the 0.5B model and n = 256 for the 1.5B model, and sequence length 512. Training ran for 1500 steps on four 80GB VRAM A100 GPUs, taking approximately 8.3 hours for the 0.5B model and 12 hours for the 1.5B model.
The models were evaluated on the MTEB-CoIR benchmark (10 tasks spanning text-to-code, code-to-text, code-to-code, and hybrid code retrieval types) and code-related MTEB tasks including CodeSearchNetRetrieval, CodeEditSearchRetrieval, HumanEval, MBPP, DS-1000, WikiSQL, and MLQuestions, as well as CosQA+ and in-house benchmarks.
Results show that Both jina-code-embeddings-0.5b and 1.5b outperform similar-sized general-purpose embedding model Qwen3-Embedding-0.6B and the substantially larger models jina-embeddings-v4 and gemini-embedding-001.
The overall averages were 78.41% for JCE-0.5B and 79.04% for JCE-1.5B, compared to 74.11% for jina-embeddings-v4, 73.49% for Qwen3-Embedding-0.6B, 79.23% for voyage-code-3, and 77.38% for gemini-embedding-001.
An ablation study compared three pooling methods (last-token, mean, and latent-attention) on the 0.5B model with identical training data, hyperparameters, and steps. Last-token pooling achieved the highest overall average (78.41%) compared to mean pooling (77.20%) and latent attention pooling (78.27%), confirming the choice of last-token pooling.
The authors conclude: "By using an autoregressive backbone pre-trained on both text and code, along with task-specific instruction prefixes and last-token pooling, the models excel at a wide variety of tasks and domains related to code retrieval. Despite their smaller size compared to other models, the jina-code-embeddings suite achieves state-of-the-art performance, demonstrating the validity and effectiveness of its unique construction methodology."
Improvements for AI systems
Based on the paper, here are specific improvements I can implement in AI systems:
Improvement: Build a retrieval system using the 0.5B or 1.5B model with the five task-specific instruction prefixes from Table 1.
What it can do:
-
NL2Code: Given a natural language query like
sort an array in descending order,
retrieve the most relevant code snippet from a codebase -
TechQA: Given a programming question like
why does my SQL query return duplicate rows?
, retrieve the most relevant answer from documentation or forums -
Code2Code: Given a Python function, find equivalent implementations in other languages or codebases
-
Code2NL: Given a code snippet, retrieve the most relevant comment or documentation explaining it
-
Code2Completion: Given the first half of a function, retrieve the most relevant completion from a code repository
Abstract
jina-code-embeddings is a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify semantically similar code snippets across programming languages. It makes innovative use of an autoregressive backbone pre-trained on both text and code, generating embeddings via last-token pooling. We outline the training recipe and demonstrate state-of-the-art performance despite the relatively small size of the models, validating this approach to code embedding model construction.
Sources
- MMTEB: Massive Multilingual Text Embedding Benchmark
- Gemini Embedding: Generalizable Embeddings from Gemini
- Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
- jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
- Qwen2.5-Coder Technical Report
- CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
- NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
- GPT-4o System Card
- jina-embeddings-v3: Multilingual Embeddings With Task LoRA
- Representation Learning with Contrastive Predictive Coding
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering