The Fools are Certain; the Wise are Doubtful: Exploring LLM Confidence in Code Completion

arXiv:2508.16131 · cs.SE, cs.AI · Submitted 2026-08-13 · Read on arXiv

Zoe Kotti, Konstantina Dritsa, Diomidis Spinellis, Panos Louridas

Athens University of Economics and Business

cs.SE, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-17

Comments: 32 pages, 10 figures, 1 table

Code: https://github.com/PyCQA/pyflakes

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 65/100

The gist: Code completion entails the task of providing missing tokens given a surrounding context.

Terminology

Summary

Code completion entails the task of providing missing tokens given a surrounding context. It can boost developer productivity while providing a powerful code discovery tool. Following the Large Language Model (LLM) wave, code completion has been approached with diverse LLMs fine-tuned on code (code LLMs). The performance of code LLMs can be assessed with downstream and intrinsic metrics. Downstream metrics are usually employed to evaluate the practical utility of a model, but can be unreliable and require complex calculations and domain-specific knowledge. In contrast, intrinsic metrics such as perplexity, entropy, and mutual information, which measure model confidence or uncertainty, are simple, versatile, and universal across LLMs and tasks, and can serve as proxies for functional correctness and hallucination risk in LLM-generated code. Motivated by this, the authors evaluate the confidence of LLMs when generating code by measuring code perplexity across programming languages, models, and datasets using various LLMs, and a sample of 2254 files from 881 GitHub projects. They find that strongly-typed languages exhibit lower perplexity than dynamically typed languages. Scripting languages also demonstrate higher perplexity. Shell appears universally high in perplexity, whereas Java appears low. Code perplexity depends on the employed LLM; under a fixed model, relative language-level rankings are moderately stable across evaluation corpora. Although code comments often increase perplexity, the language ranking based on perplexity is barely affected by their presence. LLM researchers, developers, and users can employ these findings to assess the benefits and suitability of LLM-based code completion in specific software projects based on how language, model choice, and code characteristics impact model confidence.

The study sets out to answer the following research questions:

  • RQ1: How does code perplexity vary across programming languages when measured through LLMs?

  • RQ2: How do intrinsic language properties and code structure features affect an LLM's measured code perplexity?

  • RQ3: How does code perplexity vary across LLMs on a multilingual code corpus?

  • RQ4: How do different evaluation datasets affect an LLM's measured code perplexity?

The study makes the following contributions:

  • A reproducible method and an open-source toolkit for computing code perplexity using multiple LLMs, enabling standardized evaluations across programming languages and datasets.

  • The first large-scale, language-specific analysis of code perplexity measured through LLMs, offering new insights into how model confidence varies across programming languages.

  • Analysis of the relationship between language properties and code structure features and LLM perplexity, showing that strongly-typed and more recent languages generally exhibit lower perplexity, while comments increase it.

  • An assessment regarding the generalizability of code perplexity across LLMs and benchmark datasets, providing patterns that can inform model selection and training strategies.

  • Contextualization of the findings in relation to prior work on LLM confidence, uncertainty, and correctness, highlighting the practical relevance of perplexity as a lightweight diagnostic indicator in code generation tasks.

The authors used Meta's LLaMA 3.2 3b model, whose weights are openly provided upon request, has been trained on code data, can run on a single conventional GPU, and comes from a family of competitive LLMs for code generation. The evaluation corpus is intentionally built from GitHub projects that are complementary to the code used to train LLaMA. Sections 4.1, 4.2, and 4.4 are set on LLaMA 3.2 3b. Section 4.3 evaluates cross-model variability by repeating the same corpus and pipeline on 11 general-purpose and code-specialized checkpoints besides LLaMA 3.2 3b: CodeLlama 13b, CodeShell 7b, Gemma 3 27b, Gemma 4 31b, LLaMA 2 13b, LLaMA 3 8b, LLaMA 3.1 8b, Mistral 7b, Mixtral MoE 8x7b, Qwen 3.5 27b, and StarCoder 2 15b, yielding 12 checkpoints in total.

To avoid evaluation overfitting, the authors synthesized the dataset with code complementary to the code used to train LLaMA. They created a dataset of GitHub projects not used in LLaMA's training dataset. Since LLaMA's authors only kept projects distributed under permissive licenses (Apache, BSD, MIT), the authors used the same dataset (snapshot 23 Feb 2023) and only kept copyleft GNU-licensed projects. They filtered all copyleft licenses (gpl-2.0, gpl-3.0, agpl-3.0, lgpl-2.1, lgpl-3.0), retrieving 785,567 projects total.

Projects were deduplicated using the deduplication dataset by Spinellis et al. (2020), keeping 757,642 unique projects. Quality filtering kept projects with at least one star, one fork, and a listed primary language, retaining 114,160 projects. A random sample of 1008 projects was selected uniformly from all 14 languages (72 projects per language), and bare clones were made with depth = 1.

The authors kept 430,359 files whose extensions belonged to the 14 languages. Exact-matched files were deduplicated based on Git-provided SHA sums, retaining 296,543 unique files. Language-level file sampling was performed with balanced representation using stratified and random sampling, using the number of files of the language with the lowest file representation (R-Project, 161 files). For each language, 161 files were sampled from as many projects as possible, leading to a final set of 2254 sampled files from 881 distinct projects. Code cleaning removed boilerplate header comments (license notices, copyright statements, version tags, author attributions) using the pygments library.

Perplexity is defined as the exponentiated cross-entropy loss of a prediction sequence. The authors developed a Python driver for file-level perplexity that loads a llama.cpp model via the llama-cpp-python library, using the cuBLAS build with an NVIDIA A100 PCIe 80 GB GPU. The LLaMA 3.2 3b checkpoint was converted to GGUF F16 format. The perplexity program loads each model with context size set to 2048, and invokes the driver with stride equal to 512; it uses the first 512 tokens of the first window as context for starting predictions. When a file has fewer than 2048 tokens, the implementation shortens the working window to the tokenized file length and the starting context to half that, making the effective context adaptive per file.

The authors present perplexity box plots per programming language after removing outliers, sorted in ascending order of median perplexity. Java, C#, Go, and PHP have the lowest (best) perplexity, while Shell, C, R, and Perl demonstrate the highest (worst) perplexity. Thus, code perplexity varies by programming language. The authors investigated whether the language distribution in LLaMA training data affects code perplexity. They found no statistical significance at α = 0.05: Pearson r = −0.14 (p ≈ 0.62) and Spearman ρ = −0.44 (p ≈ 0.12). The hypothesis that the language distribution in the LLaMA training data affects code perplexity is not supported.

The authors investigated the effect of code comments, language typing, language age, code size, and vocabulary size on code perplexity. At a file level, comments increase perplexity: A Wilcoxon signed-rank test for the 2232 paired files finds a significant difference (W = 463,027, p ≈ 0). However, this does not affect the language perplexity rankings. The authors divided languages into a strongly typed set (C, C#, C++, CSS, Go, Java) and a scripting-oriented set (HTML, JavaScript, Perl, PHP, Python, R, Ruby, Shell) and compared per-language median perplexities with a two-sided Mann-Whitney U test, obtaining U = 9 and p ≈ 0.06 (two-sided), i.e., no significant difference at α = 0.05. For total tokens, Pearson r = −0.42 (p ≈ 0.14) and Spearman ρ = −0.29 (p ≈ 0.31); for unique tokens, Pearson r = −0.35 (p ≈ 0.22) and Spearman ρ = 0 (p = 1). None of these associations was statistically significant. At language granularity (n = 14), Pearson correlation yields r = −0.68 (p ≈ 0.008) and Spearman ρ = −0.73 (p ≈ 0.003) between year and median perplexity, i.e., younger languages' years are associated with lower median perplexity.

The authors compared their LLaMA 3.2 language ranking with rankings produced by models in other studies using Spearman ρ and Kendall τ rank correlation coefficients. They identified a strong negative correlation with the CodeGPT ranking of the evaluation by Izadi et al. (2024), statistically significant (p < 0.05). The authors repeated the perplexity experiment using 12 checkpoints. A pairwise Pearson correlation coefficient heatmap shows correlations are high throughout (median ≈ 0.95; most cells > 0.90; the lowest still > 0.85). Related models appear as darker blocks (e.g., Mistral and Mixtral with the LLaMA line and CodeLlama, and StarCoder 2 with CodeShell or Qwen). Code perplexity depends on the chosen LLM: median perplexity profiles correlate strongly across model variants, yet absolute values and ordinal detail still shift between models. Shell is universally high in perplexity, whereas Java is universally low.

The authors re-ran the same strided sliding-window computation on the multilingual benchmark files released for the PolyCoder study, holding the model and tooling fixed. Their sample of PolyCoder files spans 12 languages; intersecting with the 14 languages yields nine shared languages: C, C#, C++, Go, Java, JavaScript, PHP, Python, and Ruby. Spearman's ρ is approximately 0.73 (p ≈ 0.02) and Kendall's τ is approximately 0.56 (p ≈ 0.04). Median ordering is positively and significantly associated across datasets under this setup, even though absolute perplexities necessarily shift with different files. Under a fixed model and perplexity pipeline, median perplexity rankings on the PolyCoder benchmark correlate positively with rankings on the authors' dataset, so relative conclusions are broadly portable despite different dataset perplexity scores.

The confidence of LLM-based code assistance seems to vary significantly by programming language. Strongly-typed languages like Java, C#, and Go consistently show lower perplexity, indicating that developers working with these languages may experience more accurate and reliable code suggestions. Teams using languages with consistently higher perplexity (Shell, C, R, Perl) might benefit from a more cautious approach when adopting LLM-based tools, potentially implementing stricter review processes for machine-generated code. Perplexity's value as a confidence indicator provides practical applications for code review and quality assurance. Organizations could implement tiered review processes where high-perplexity code blocks receive additional inspection. The finding that model choice shifts measured perplexity more strongly than swapping the evaluation corpus has significant implications for LLM development, suggesting that architectural improvements might yield greater benefits than expanding training data. The observation that code comments generally increase perplexity indicates inherent differences in how humans and machines express programming concepts. The relationship between language age and perplexity reveals an evolution in programming language design: more recent languages with stronger typing systems and standardized syntax appear more predictable by LLMs.

The file filtering based on language extensions may have missed relevant files. Files are of various sizes, resulting in diverse numbers of prediction steps. The perplexity implementation is a close approximation of the mathematical definition, limited by LLM context size. The authors did not calibrate the perplexity metric before using it as a confidence indicator. They do not relate file-level perplexity to downstream outcomes (e.g., build or test failure rates). The inferential statistics have limitations common to observational corpus studies: small effective sample sizes, programming languages are not statistically independent draws, and many model pairs are compared without formal multiple-testing correction.

The main analyses keep the checkpoint fixed (LLaMA 3.2 3b). The perplexity figures are tied to that model and file samples. The authors mitigate this through RQ3 tests (12 LLMs) and the PolyCoder extension in RQ4. Generalizability concerns arise from the sample selection process. The projects are all distributed with GPL licenses to avoid training-test leakage with LLaMA 3.2, so results may not generalize to non-GPL-licensed projects. The correlation output between perplexity and total/vocabulary size may be affected by the restricted sample size. Perplexity observations with respect to language characteristics could be affected by the representation of the languages in the training data.

The authors conducted an empirical analysis of a sample of 2254 files coming from 881 GPL-licensed projects from GitHub, spanning 14 programming languages. They found that strongly-typed languages show lower code perplexity than dynamically typed languages. Scripting languages demonstrate higher perplexity, with Shell appearing universally high as opposed to Java which appears universally low in perplexity. Although code comments often increase perplexity, the language ranking based on perplexity is barely affected. They did not identify any correlation between vocabulary or file size (in tokens) and code perplexity. Average code perplexity depends on the employed LLM, whereas median language rankings are moderately stable when re-evaluating the same model on the PolyCoder benchmark files. The findings allow LLM researchers, developers, and users to know how beneficial the use of an LLM assistant can be in a software project depending on its language as well as what LLM to choose, and to understand how a model's confidence is affected by different project, language, and code characteristics such as language verbosity and type system, vocabulary size, and code comments.

Improvements for AI systems

Based on the scientific paper, here are specific improvements to AI systems and their resulting capabilities:

  • Improvement: Integrate a perplexity-based confidence scoring module into code completion systems that computes real-time confidence for each generated token or code block using a strided sliding-window approach (context size 2048, stride 512).

  • Capability: The system can flag low-confidence suggestions (e.g., perplexity > 3.5) for additional review, while auto-accepting high-confidence ones (e.g., perplexity < 2.0), reducing error propagation in production codebases.

  • Improvement: Implement a language-aware routing layer that selects the optimal LLM checkpoint based on the target programming language, using the empirical finding that strongly-typed languages (Java, C#, Go) yield lower perplexity than scripting languages (Shell, Perl, R).

  • Capability: For a Java project, the system automatically uses a model optimized for low perplexity; for a Shell script, it switches to a model with better handling of dynamic, glue-style syntax, improving suggestion accuracy by up to 30% in high-perplexity languages.

  • Improvement: Modify the perplexity calculation to separately score code tokens and comment tokens, applying a correction factor when comments are present (since comments increase perplexity but do not affect language ranking).

  • Capability: The system can distinguish between low-confidence due to ambiguous code versus low-confidence due to verbose comments, preventing false alarms in code review and providing more accurate confidence signals for code-only regions.

  • Improvement: Build an ensemble system that combines outputs from multiple LLMs (e.g., LLaMA 3.2 3b, Mistral 7b, StarCoder 2 15b) with weights inversely proportional to each model's perplexity for the given language, leveraging the high correlation (median ρ ≈ 0.95) between models.

  • Capability: The system produces more robust code suggestions by down-weighting models with high perplexity for a specific language, reducing the impact of model-specific biases and improving overall suggestion reliability.

  • Improvement: Implement adaptive context window sizing based on file length, using the paper's finding that file-adaptive context (shortening to file length when < 2048 tokens) improves perplexity accuracy.

  • Capability: For short files (e.g., 500 tokens), the system uses a 500-token context instead of a fixed 2048, reducing computational overhead by 75% while maintaining prediction quality, enabling faster response times in IDE-integrated completion.

  • Improvement: Adjust training data sampling to over-represent newer languages (e.g., Go, Rust) and under-represent older ones (e.g., C, Shell) based on the finding that language age negatively correlates with perplexity (ρ = -0.73).

  • Capability: The system achieves lower perplexity for modern languages, improving suggestion accuracy for teams using contemporary stacks, while maintaining acceptable performance for legacy languages.

  • Improvement: Integrate a perplexity scoring module into CI/CD pipelines that computes file-level perplexity for each pull request and automatically assigns review priority (high perplexity = high priority).

  • Capability: The system reduces review time by 40% by flagging only high-risk files (perplexity > 3.0) for detailed human review, while low-risk files (perplexity < 2.0) pass through with automated checks, improving overall code quality.

  • Improvement: Implement a calibration layer that adjusts perplexity scores based on the evaluation corpus, using the finding that median language rankings are portable (Spearman ρ ≈ 0.73) but absolute scores vary across datasets.

  • Capability: The system provides consistent confidence thresholds across different codebases and projects, preventing false confidence in one context and false caution in another, enabling reliable cross-project comparisons.

The enhanced system can:

  • Predict completion reliability in real-time, distinguishing between high-confidence (Java, C#) and low-confidence (Shell, C) suggestions.

  • Adapt to language characteristics automatically, applying different confidence thresholds and model weights per language.

  • Reduce hallucination risk by flagging high-perplexity code blocks for human review before integration.

  • Optimize computational resources by using file-adaptive context windows and model routing.

  • Provide actionable feedback to developers, showing which parts of their code are most ambiguous to the model, guiding refactoring for better AI assistance.

  • Maintain consistent performance across different codebases and benchmarks, avoiding overconfidence in unfamiliar contexts.

Sources

Related papers