Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
cs.CL
Submitted: 2024-11-15
Updated: 2026-09-20
Comments: Accepted to ICASSP 2026
Code: https://github.com/sustech-nlp/Compound-QA
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- The Llama 3 Herd of Models
- Gemma 3 Technical Report
- Qwen3 Technical Report
- Instruction-Following Evaluation for Large Language Models
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge
- How Easily do Irrelevant Inputs Skew the Responses of Large Language Models?
- Evaluating LLMs with Multiple Problems at once
- LongGenBench: Long-context Generation Benchmark
- The SIFo Benchmark: Investigating the Sequential Instruction Following Ability of Large Language Models
- REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once
- Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
- AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Mistral 7B
- Gemma 2: Improving Open Language Models at a Practical Size
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- InternLM2 Technical Report
- Qwen2.5 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering