An Empirical Evaluation of Cost-Efficient Large Language Models on Algorithmic Programming Tasks
cs.SE, cs.AI
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 7 pages, 7 figures. Accepted for publication in the IEEE Proceedings of the 2026 International Conference on Advanced Computing Technologies (ICACT)
Code: https://github.com/researchartifacts/ProjectEulerAIEvaluation
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Human vs Machine Mathematical Difficulty on Project Euler: An Experimental Analysis
- MathArena: Evaluating LLMs on Uncontaminated Math Competitions
- Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties