On the Lexical Superstition of Large Language Models for Code Comprehension: Re-evaluation on Code of Low Lexical Quality
cs.SE, cs.CL, cs.LG
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: 27 pages, 9 figures, 12 tables. Submitted to an ACM journal in September 2025. Preprint; manuscript under review. Corresponding author: Ming Li
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- Evaluating Large Language Models Trained on Code
- Large Language Models of Code Fail at Completing Code with Potential Bugs
- An Extensive Study on Adversarial Attack against Pre-trained Models of Code
- The Llama 3 Herd of Models
- Two Sides of the Same Coin: Exploiting the Impact of Identifiers in Neural Code Comprehension
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates
- CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
- A Survey on Large Language Models for Code Generation
- Code Llama: Open Foundation Models for Code
- LLaMA: Open and Efficient Foundation Language Models
- Qwen2 Technical Report
- Qwen2.5 Technical Report
- Transfer Attacks and Defenses for Large Language Models on Coding Tasks
- DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties