When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
cs.SE, cs.AI, cs.CL, cs.PL
Submitted: 2025-10-19
Updated: 2025-12-09
Comments: Accepted to ICSE 2026 (RECODE workshop)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) with vast context windows offer new avenues for in-context learning (ICL), where providing many examples ("many-shot" prompting) is often assumed to enhance performance.
Terminology
Abstract
Large Language Models (LLMs) with vast context windows offer new avenues for in-context learning (ICL), where providing many examples ("many-shot" prompting) is often assumed to enhance performance. We investigate this assumption for the complex task of code translation. Through a large-scale empirical study of over 90,000 translations, we systematically evaluate the impact of scaling in-context examples from zero-shot to many-shot configurations of up to 625 examples, with prompts spanning from approximately 100,000 to 800,000 tokens. Our findings reveal a "many-shot paradox": while static similarity metrics may modestly improve with more examples, functional correctness consistently peaks with few-shot prompting (5-25 examples). Providing substantially more examples often degrades this crucial functional performance. This study highlights that for code translation, the quality of a few well-chosen examples outweighs sheer quantity, challenging the universal efficacy of "more is better" for ICL and underscoring the task-dependent nature of optimal prompting strategies. Our results have significant implications for effectively leveraging LLMs in software engineering.
Sources
- In-Context Learning with Long-Context Models: An In-Depth Exploration
- Post-Incorporating Code Structural Knowledge into Pretrained Models via ICL for Code Translation
- AlphaTrans: A Neuro-Symbolic Compositional Approach for Repository-Level Code Translation and Validation
- On the generalization of language models from in-context learning and finetuning: a controlled study
- Lost in the Middle: How Language Models Use Long Contexts
- Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures
- Large Language Models Can Be Easily Distracted by Irrelevant Context
- Gemini: A Family of Highly Capable Multimodal Models
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- CodeTransOcean: A Comprehensive Multilingual Benchmark for Code Translation
- More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties