A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance
cs.AI, cs.CY
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: Accepted to EMNLP Findings 2026
Code: https://github.com/KimberlyTruong/expert-informed-eval-gen
License: http://creativecommons.org/licenses/by/4.0/
The gist: This paper presents an end-to-end approach for generating context-specific large language model (LLM) benchmark datasets by combining expert input with synthetic data generation.
Terminology
Abstract
This paper presents an end-to-end approach for generating context-specific large language model (LLM) benchmark datasets by combining expert input with synthetic data generation. Existing benchmark construction methods often trade off validity and scalability: datasets designed with domain experts can produce high-quality evaluations but are slow and costly to create, while synthetically generating data may scale efficiently but often results in unrealistic, redundant, or out-of-scope examples. To address this gap, we introduce a schema eliciting key information about the goals, scope, and context of an evaluation task, and use this information to guide synthetic data generation. We further define four criteria grounded in measurement validity for assessing dataset quality: coverage, diversity, content realism, and stylistic realism. Using these criteria, we show how expert-informed scaffolds can guide synthetic data generation toward more valid benchmarks. Through quantitative evaluations and a real-world case study with domain experts, we demonstrate that our approach improves benchmark data quality over existing methods while preserving validity. We additionally analyze how different types of schema information affect different dataset quality criteria, and provide practical guidance on which information to prioritize collecting under resource constraints.
Sources
- Demystifying MMD GANs
- A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
- Stakeholder Participation in AI: Beyond "Add Diverse Stakeholders and Stir"
- BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
- Multi-Sample Prompting and Actor-Critic Prompt Optimization for Diverse Synthetic Data Generation
- Codebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts
- AI and the Everything in the Whole Wide World Benchmark
- Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
- YourBench: Easy Custom Evaluation Sets for Everyone
- Rethinking Model Evaluation as Narrowing the Socio-Technical Gap
- A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data
- BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection