Qworld: Question-Specific Evaluation Criteria for LLMs

arXiv:2603.23522 · cs.CL, cs.AI · Submitted 2026-08-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Qworld: Question-Specific Evaluation Criteria for LLMs".

Jane: The paper was written by Shanghua Gao, Yuchang Su, Pengwei Sui, Curtis Ginder and Marinka Zitnik from Department of Biomedical Informatics, Harvard Medical School and Department of Medicine, Brigham and Women’s Hospital and Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University and Broad Institute of MIT and Harvard and Harvard Data Science Initiative.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: So, we’ve established that "Qworld: Question-Specific Evaluation Criteria for LLMs" fundamentally changes how we think about testing AI. In this segment, we are looking at the paper's summary of its methodology and what that means for those who actually plan to build and deploy these systems.

Jane: The key takeaway from the summary is that by using question-specific worlds, the evaluation process becomes much more granular and less reliant on general textual fluency. It forces the model to demonstrate deep structural understanding relevant only to the specific query at hand.

Lu: What I find particularly useful in this summary is how it formalizes what we mean by "contextual awareness." It’s not enough for a model to just read the context; it has to actively incorporate that context into its reasoning and output.

Meng: From an engineering standpoint, the summary highlights that Qworld allows us to break down complex tasks into verifiable sub-components. Instead of treating a medical query as one big problem, we can test its ability to handle data extraction, cross-referencing, and risk assessment separately.

Lalam: This modularity is huge for development pipelines. It means that if a model fails in one specific area—say, distinguishing between primary and secondary sources—we don't have to throw out the whole model; we just know exactly which module needs reinforcement.

Tom: That speaks volumes about the practical utility of this framework. It gives developers a precise roadmap for improvement, rather than just a vague sense that "it needs more training."

Jane: Exactly. The paper shows that these criteria can be designed to target specific failure modes observed in existing models—for instance, susceptibility to hallucination when dealing with ambiguous inputs.

Lu: And this goes beyond just technical troubleshooting; it helps us define the *ethical* boundaries of the AI. By testing for specific biases or oversimplifications in its responses, we are vetting its alignment with human values.

Meng: The summary also implicitly suggests that these criteria can be used to create guardrails. We can't just hope the model behaves; we have to define the successful behavior for every possible question it might face.

Lalam: It shifts the burden of proof onto the AI itself, forcing it to prove its competency against a defined, rigorous set of standards, which is exactly what users need to see before adopting new technology.

Tom: It really gives structure to what we've been talking about—moving from abstract concepts of 'intelligence' to concrete, measurable criteria.

Jane: And this framework provides the necessary language and tools for the entire academic community to start building these kinds of rigorous, context-aware benchmarks. When we finish this segment, we will look at how Qworld can be improved upon for even greater scalability.

Improvements: Tom: We've seen that "Qworld: Question-Specific Evaluation Criteria for LLMs" is a massive leap forward in rigor and specificity. Now, the paper discusses improvements—how we can make this already strong framework even more robust and applicable across vastly different domains.

Jane: The core message here is that while Qworld is revolutionary, its implementation needs to be optimized for scale and adaptability. The paper suggests methods to streamline the creation of these 'question-specific worlds' without losing their depth.

Lu: I think one crucial improvement suggested is developing meta-criteria—criteria for designing the criteria themselves. Instead of manually building a Qworld for every single use case, we need a system that helps us generate optimal testing parameters efficiently.

Meng: From an engineering standpoint, the scalability challenge is immense. The paper points toward integrating these criteria design principles into automated test suites, meaning we don't have to run human experts through every single test case manually forever.

Lalam: And if we can automate the *design* of the difficult questions—the ones that force the model into its weaknesses—that would be a huge cultural win. It means continuous improvement of our understanding of AI limitations, not just running tests on existing models.

Tom: It sounds like the goal is to make the evaluation system itself self-improving, which is a fantastic concept.

Jane: Exactly. We are moving toward a scientific process where the benchmark improves as rapidly as the technology it is designed to test. The paper discusses how this can be applied across domains, from legal reasoning to creative narrative generation.

Lu: This suggests that Qworld isn't just for technical QA; it can become a fundamental tool for domain experts—like lawyers or doctors—to validate the utility of AI in their specific professional fields.

Meng: And furthermore, addressing the complexity at scale requires developing standardized taxonomies for knowledge gaps. If we know *why* a model fails (e.g., due to temporal reasoning failure vs. ambiguity handling failure), we can target the fix much more precisely.

Lalam: I think this emphasis on process improvement is what will really help adoption. Companies won't just see a list of scores; they will see a clear path laid out for how their AI can be brought up to the required standard using this methodical approach.

Tom: So, these suggested improvements turn Q

Paper discussion segment 3: Tom: We’ve seen Qworld achieve incredible results, but we also know that in real-world deployment, perfection is the enemy of progress, so what does the paper suggest for improvements?

Jane: It's not just about making it better; it's about optimizing how we use this framework. The researchers are looking at ways to make the process scalable so that applying Qworld isn't a huge manual chore for every single test case.

Lu: I think the most exciting improvement is automating the design phase itself, moving beyond just building criteria for existing questions. We could develop meta-criteria—rules for creating better evaluation questions—to test even more thoroughly.

Meng: That’s a massive engineering hurdle, Lu, because designing a question that forces a deep failure mode is much harder than designing one that's just confusing. The system needs to be able to intelligently identify those gaps.

Lalam: And once we have better questions, the cultural implication is profound; AI becomes something we are actively training and guiding toward its highest potential, rather than just accepting whatever output it happens to generate.

Tom: So, we're moving from just testing what to fix, to proactively designing the tests that force a genuine improvement. Jane, how does this help us manage the complexity of different domains like medicine versus creative writing?

Jane: It gives us a way to define context-specific success. Instead of using one giant checklist for everything, we get a tailored rubric for each specific use case, making the evaluation inherently more robust across domains.

Lu: I see it as formalizing human intuition; we are capturing that feeling of 'this response isn't quite right' and turning it into measurable structural requirements.

Meng: From a practical standpoint, if you can automate the generation and refinement of these criteria, the implementation complexity shrinks significantly while retaining that necessary rigor.

Lalam: It ensures that AI is not just fast, but truly capable of supporting nuanced human interactions based on what we actually need in the moment.

Conclusion: Tom: We've seen that Qworld is changing how we test AI by building custom criteria for every single question, which is a huge conceptual shift for everyone working in this space.

Jane: It’s really about moving away from generalized benchmarks and focusing on what the specific intent of a particular query requires, making it much more rigorous.

Lu: I think the most profound implication is that we are defining competence not just by accuracy, but by formal, verifiable adherence to structural demands.

Meng: From a practical standpoint, this means we can finally build scalable and reliable evaluation pipelines that hold up to real-world usage without becoming overly complicated.

Lalam: It allows us to create AI systems that truly understand the context of human needs, not just systems that spit out words that sound convincing.

Tom: You're right, Lalam; it provides a way for us to measure actual capability rather than just a generalized score across the whole dataset.

Jane: And we’ve seen how this works on things like medical queries, where context is everything and failures are dangerous.

Lu: To build on that idea of structure, I'm curious if this framework can be adapted for artistic or creative AI tasks as well, defining "excellence" in a non-factual way.

Meng: That would require defining entirely new categories of criteria, which could be a massive undertaking in terms of engineering effort.

Lalam: But even the goal would be to enhance human expression by providing tools that can interpret and respond to our most nuanced intentions.

Tom: That’s a perfect way to end this discussion, seeing how far this goes from abstract testing into practical, real-world application.

Jane: It gives us so much hope for the future of AI because we are finally moving toward something truly reliable and useful.

Lu: I'm just glad to see the theoretical framework so robust and applicable across different domains of reasoning.

Meng: And I just hope that implementation details will be ready to handle the complexity at scale.

Lalam: It gives us a language to talk about true capability, not just more tokens or more data.

Tom: That's an incredible summary of "Qworld: Question-Specific Evaluation Criteria for LLMs." We really appreciate you all joining us today.

Jane: We’re so excited to see this paper in action, Tom. It marks a significant step forward in ensuring AI is capable of providing the nuanced support we all expect from these tools.

Shanghua Gao, Yuchang Su, Pengwei Sui, Curtis Ginder, Marinka Zitnik

Department of Biomedical Informatics, Harvard Medical School · Department of Medicine, Brigham and Women’s Hospital · Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University · Broad Institute of MIT and Harvard · Harvard Data Science Initiative

cs.CL, cs.AI

Submitted: 2026-08-19

Updated: 2026-08-21

Code: https://github.com/mims-harvard/Qworld

Importance score: 90/100

The gist: The scientific paper introduces One-Question-One-World (Qworld), a method designed to address the difficulty in evaluating large language models (LLMs) on open-ended questions, where "response

Key concepts

Qworld
A framework for evaluating LLMs by creating custom, question-specific criteria. This method moves away from generalized benchmarks, demanding that models prove competency against a defined, rigorous set of standards for every single query.
Contextual Awareness
The ability of an AI model to not just read context but to actively incorporate that specific information into its reasoning and output. The paper formalizes this requirement, making it a measurable structural demand.
Modularity in Testing
A method that allows complex tasks (like medical queries) to be broken down into verifiable sub-components (e.g., data extraction or cross-referencing). This pinpoints exactly which module needs reinforcement when a model fails.
Meta-criteria
Criteria designed for creating the evaluation criteria themselves. Instead of manually building a Qworld for every use case, this system helps generate optimal testing parameters efficiently.

Terminology

Summary

The scientific paper introduces One-Question-One-World (Qworld), a method designed to address the difficulty in evaluating large language models (LLMs) on open-ended questions, where response quality depends on the question’s context. The authors note that Binary scores and static rubrics fail to capture these context-dependent requirements, and existing methods are limited because they define criteria at the dataset level or generate them in a single pass, which limits their ability to explore the evaluation space implied by each question.

Qworld is presented as a solution that generates question-specific evaluation criteria using a recursive expansion tree. The methodology involves taking a given question and allowing Qworld to decompose it into scenarios, perspectives, and fine-grained binary criteria through structured hierarchical and horizontal expansion. These resulting criteria are designed to specify exactly what a high-quality answer must address for that question, forming what the authors refer to as the question’s “world.”

The paper reports significant results regarding the quality of these generated criteria. On HealthBench, Qworld demonstrates strong performance, achieving Coverage of 0.89 and Uniqueness of 0.79. This means Qworld covers 89% of expert-authored criteria and generates 79% novel criteria validated by human experts. Furthermore, human evaluators rated the quality produced by Qworld higher in terms of insight and granularity than those produced by prior methods.

The second major application involves using these customized criteria to benchmark LLMs. When applied to 11 frontier LLMs on HealthBench and Humanity’s Last Exam, Qworld's findings reveal critical distinctions that generic rubrics fail to make. Specifically, Qworld reveals capability differences in dimensions such as long-term impact, equity, error handling, and interdisciplinary reasoning that coarse rubrics do not distinguish.

By shifting the focus from fixed task-level criteria to a dynamic approach—Qworld enables evaluation that adapts to each question rather than relying on fixed task-level criteria. This method allows for a comprehensive analysis of model behavior across various dimensions, providing fine-grained insight into LLM capabilities.

Improvements for AI systems

As a diligent researcher, I have analyzed the Qworld framework and identified specific, high-impact improvements for integrating dynamic, context-aware evaluation into LLM benchmarking pipelines.

The core limitation of current fixed rubrics is that they fail to capture context-dependent requirements—they assume all questions under a task share the same evaluation needs. Qworld solves this by generating question-specific criteria using a Recursive Expansion Tree (RET).

To integrate Qworld into an LLM evaluation pipeline, the following structural and methodological improvements must be implemented:

1. Dynamic Benchmark Generation via RET:

Instead of applying a static rubric C, for every input question Q i in a dataset D, we must dynamically generate a tailored criteria set C i. This requires implementing the three-level hierarchical expansion of the Recursive Expansion Tree (RET):

  • Level 1: Scenario Grounding (=1): Infer the implicit context (e.g., target audience, stakes) of Q i.

  • Level 2: Perspective Elicitation (=2: Derive a set of evaluation perspectives P i that define critical axes for Q i.

  • Level 3: Criteria Instantiation (=3): Generate concrete, measurable criteria C ij under each perspective. This process must employ both Hierarchical Expansion (R h) to refine granularity and Horizontal Expansion (R w) to ensure comprehensive coverage of overlooked dimensions.

2. Multi-Dimensional Scoring System:

The evaluation score S(A, Q i) should not be a single binary metric (e.g., accuracy) but an aggregation of criterion-level scores:

S(A, Q i) = F norm (sum sc(A, Q i))

where each criterion c in C i is evaluated by a dedicated LLM-as-a-Judge, assigning a weighted score alpha c (positive for desirable behavior; negative for harmful/omitted behavior). This allows the the system to track both successes and failures against specific contextual demands.

3. Automated Coverage and Novelty Auditing:

The evaluation pipeline must include automated auditing modules using LLM-based judges to calculate:

  • Coverage: The proportion of expert-authored criteria covered by C i (ensuring alignment with domain standards).

  • Uniqueness: The proportion of generated criteria that are novel, non-obvious dimensions not present in the expert set (identifying gaps in existing knowledge/rubrics).


By adopting this Qworld-based framework, the improved AI evaluation system will achieve capabilities far beyond static benchmarking:

1. Contextual Capability Discrimination:

The system can distinguish between LLMs based on their performance across fine-grained, question-specific dimensions. For example, instead of simply classifying a medical response as Correct, it can separate and quantify distinct failure modes such as:

  • Safety vs. Clarity: Assessing if the model prioritized safety warnings (e.g., risk management) over conversational flow.

  • Technical Rigor vs. Pedagogical Clarity: Determining if a scientific explanation was mathematically rigorous but incomprehensible, or clear but lacking necessary detail.

2. Identification of Systemic Weaknesses (Novel Dimensions):

The system can surface evaluation axes that traditional rubrics overlook, such as:

  • Sustainability & Ethics: Evaluating how well a model addresses long-term impact or ethical considerations in a specific scenario (e.g., recommending an energy-efficient solution for a climate-constrained problem).

  • Equity and Accessibility: Assessing if the model provides advice that is relevant and affordable for diverse user groups, rather than just assuming generalized ideal conditions.

3. Dynamic Benchmark Adaptation:

The system can generate specialized benchmarks for specific use cases. If the input question implies a high-stakes medical scenario, the system will automatically prioritize criteria related to Emergency Recognition and Safety & Risk Management. If the input question is purely abstract reasoning, it will prioritize criteria related to Creativity & Rigor.

4. Mitigation of Benchmark Saturation:

By continually generating novel, highly specific evaluation dimensions (Uniqueness), the system prevents LLMs from gaming a fixed rubric. It ensures that future model improvements are measured against increasingly sophisticated and context-specific standards of excellence, thereby raising the bar for AI performance.

Sources

Related papers