InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information
cs.SE, cs.AI, cs.CL
Submitted: 2026-08-30
Updated: 2026-08-30
Code: https://github.com/kmsgk0/InteractBench
Terminology
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Program Synthesis with Large Language Models
- Evaluating Large Language Models Trained on Code
- Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- HardTests: Synthesizing High-Quality Test Cases for LLM Coding
- RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
- Measuring Coding Challenge Competence With APPS
- AgentBench: Evaluating LLMs as Agents
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Can Language Models Solve Olympiad Programming?
- AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
- CodeContests+: High-Quality Test Case Generation for Competitive Programming
- OJBench: A Competition Level Code Benchmark For Large Language Models
- ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests
- LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
- BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
- LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties