Evaluating Language Models on Cross-Language Code Functional Equivalence
cs.SE, cs.AI, cs.CL
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/dflook/python-min
Terminology
Sources
- Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
- Beyond Resolution Rates: Behavioral Drivers of Coding Agent Success and Failure
- A Multi-agent AI System for Deep Learning Model Migration from TensorFlow to JAX
- CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks
- CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties