A Taxonomy of Programming Languages for Code Generation
cs.CL
Submitted: 2026-03-31
Updated: 2026-10-01
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: The world's 7,000+ languages vary widely in the availability of resources for NLP, motivating efforts to systematically categorize them by their degree of resourcefulness (Joshi et al., 2020).
Terminology
Abstract
The world's 7,000+ languages vary widely in the availability of resources for NLP, motivating efforts to systematically categorize them by their degree of resourcefulness (Joshi et al., 2020). A similar disparity exists among programming languages (PLs); however, no resource-tier taxonomy has been established for code. As large language models (LLMs) grow increasingly capable of generating code, such a taxonomy becomes essential. To fill this gap, we present the first reproducible PL resource classification, grouping 646 languages into four tiers. We show that only 1.9% of languages (Tier 3, High) account for 74.6% of all tokens in seven major corpora, while 71.7% of languages (Tier 0, Scarce) contribute just 1.0%. Statistical analyses of within-tier inequality, dispersion, and distributional skew confirm that this imbalance is both extreme and systematic. Our results provide a principled framework for dataset curation and tier-aware evaluation of multilingual LLMs.
Sources
- Evaluating Large Language Models Trained on Code
- Large Language Models for Software Engineering: Survey and Open Problems
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
- Multi-lingual Evaluation of Code Generation Models
- The Stack: 3 TB of permissively licensed source code
- StarCoder: may the source be with you!
- Code Llama: Open Foundation Models for Code
- StarCoder 2 and The Stack v2: The Next Generation
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- CodeGen2: Lessons for Training LLMs on Programming and Natural Languages
- On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering