DataKernelBench: Can LLMs Optimize Database Queries on GPUs?
cs.CL, cs.AI, cs.DB, cs.LG, cs.PL
Submitted: 2026-08-25
Updated: 2026-08-27
Code: https://github.com/kerneldf/datakernelbench
Project page: https://kerneldf.github.io/datakernelbench
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
- KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
- Evaluating Language Models for Efficient Code Generation
- PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage
- Learned Query Superoptimization
- GPT-4 Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- Devstral: Fine-tuning Language Models for Coding Agent Applications
- Bespoke OLAP: Synthesizing Workload-Specific One-size-fits-one Database Engines
- MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
- TritonRL: Training LLMs to Think and Code Triton Without Cheating
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering