Designing Benchmarks for Knowledge Work

arXiv:2605.23262 · cs.AI · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Designing Benchmarks for Knowledge Work".

Jane: , adhering strictly to its content and maintaining a fastidious, academic tone.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Moving past the general idea, let's look at how "Designing Benchmarks for Knowledge Work" provides specific methodological improvements to our current approach by structuring our vocabulary.

Jane: The authors move beyond using broad, vague labels like "healthcare" and instead introduce a detailed inventory of actual job functions derived from O*NET task statements, which is a massive step up from relying on general categories.

Lu: That's where the O*NET database comes in; it provides this comprehensive view of what people actually do across different industries, giving us a highly grounded vocabulary for the work being done.

Meng: It stops us from grouping tasks too broadly, which I’ve seen causes many systems to fail because they weren't constrained by a specific job role or the complexity of their input materials.

Lalam: By using this granular approach, we are setting the stage for AI to become an integrated partner in professional work, not just some generic chatbot that needs constant retraining.

Tom: The paper really pushes us to think beyond just what is being produced, but how the output can be used in the subsequent steps of a much larger workflow that follows.

Jane: The authors show that by defining the work activity and specifying the setting, we are forcing ourselves to define what the *right* final product should look like for a specific job, rather than just assuming any successful output is fine.

Meng: It’s about ensuring that when we achieve success, we know exactly what criteria were met within an operational framework—like having used the right data source or adhering to a legal standard.

Lu: That context is everything because it isn't just about the outcome; it's about whether that output was possible given the constraints of accessing those specific tools and resources.

Lalam: It allows us to assess *how* we are using AI—whether it is replacing human judgment or assisting a specific type of decision-making process, which is a key cultural shift.

The paper's summary: Tom: Now, let's look at how these ideas are applied in the case studies of "Designing Benchmarks for Knowledge Work," since the authors provide real examples to show us this approach in action.

Jane: They use three benchmarks—GDP VAL, O FFICE QA P RO, and APEX-SWE—to illustrate different types of work claims, proving that specificity is paramount in all tasks.

Lu: It’s fascinating how each benchmark supports a very specific claim; we are seeing a stark contrast to the general claims we usually make when we talk about AI's overall "intelligence."

Meng: I'm especially interested in how they define the *setting* for those cases; it feels very practical and grounded in real-world data, whether that is using actual GitHub issues or specific financial reports.

Lalam: They show us that even when we use complex systems like software engineering, the AI should be judged by its ability to create a verifiable state change within the system itself.

Tom: That’s right; they aren't just asking if the code runs, but whether it creates a specific, verifiable change in the system state, which is far more meaningful than simple function execution.

Jane: The authors clearly show that even when using complex tools and materials—like huge document sets for O FFICE QA P RO—the score is tied to a specific, measurable outcome.

Meng: It’s very telling how they limit the claim based on those defined boundaries; they are making sure we don't assume the full scope of an entire occupation.

Lu: This demonstrates that we are not just measuring general competence, but defining exactly what level of professional responsibility is being tested in that particular benchmark instance.

Lalam: This allows us to move toward a level of trust where AI is used as a highly specific and reliable assistant for every single task it was trained for.

The paper's improvements: Tom: We’ve covered so much ground, moving from the initial critique to the practical application in "Designing Benchmarks for Knowledge Work," showing us how these concepts tie together.

Jane: The overall message that we need to be extremely explicit about what our benchmarks are actually measuring is probably the most important thing to carry away from this discussion.

Lu: It’s a way of enforcing epistemic humility, acknowledging that AI cannot perform all of human knowledge work without specific context and defined tools.

Meng: I think this methodology will lead to much more robust and deployable systems because we know exactly what the scope is and what the success criteria are in any given deployment.

Lalam: It’s about refining our definition of "smart" in a way that ultimately elevates the quality of work and the human experience in it, making us better collaborators.

Tom: So, as we wrap up, let's make sure everyone has a final thought on this incredible paper titled "Designing Benchmarks for Knowledge Work."

Lu: I hope this opens up avenues for more complex simulations that allow us to see how AI interacts across different workflow stages, not just in isolation.

Meng: I’m optimistic about the implementation; we have finally found a structured way to measure reliability rather than just hoping an answer looks right, which is a massive practical win.

Lalam: This ensures that when we build our future vision of AI, it is grounded in verifiable professional capability rather than just a flashy demo that does not work.

Tom: It feels like we are moving toward defining exactly what kind of evidence we need to trust these systems enough to trust them at all, and that is a huge step forward for us.

Conclusion: Tom: Let’s revisit the core idea of "Designing Benchmarks for Knowledge Work" one last time, looking at the implications for how we structure AI evaluation.

Jane: The paper highlights that current evaluation methods are too broad, and scores are often used to support big claims—like a system being good at research synthesis—but these scores rarely actually reflect the full scope of knowledge work.

Lu: I agree with Jane; it’s a fundamental shift in how we view intelligence. We have been treating LLMs as isolated function executors, but the researchers are suggesting that knowledge work is inherently contextual.

Meng: The authors show that this mismatch happens when we only look at visible outputs, which are just "answers" or "patches," instead of the full context of a job requiring resources and judgment.

Lalam: It’s about recognizing that knowledge work requires a specific environment to be effective for the AI to truly help humanity, meaning its ability is tied to the system it is placed in service.

Tom: That leads us straight into the paper's proposed solution, which is summarized in this three-step approach that defines how we measure competence.

Jane: The authors outline a way to tie these benchmark scores back to broader work claims by defining three distinct elements: the activity, the setting, and the product.

Meng: First, they define the *work activity*—what specific job is being done—and second, they specify the *tested setting*, which includes all tools and materials available.

Lu: And third, defining the *scored work product*, which is what’s left behind for review or downstream use; these three elements must be explicit for a coherent claim to be made.

Lalam: We are moving from simple success/failure metrics to understanding the whole ecosystem of how the AI fits into a professional workflow, which is vital for cultural change.

Tom: That sounds like a very rigorous way to ensure we're not overpromising what our models can actually do in real-world settings.

Yining Hua, Hongbin Na, Cyrus Ayubcha, Levi Lian

Harvard University · University of Technology Sydney · Stanford University · Raycaster AI

cs.AI

Submitted: 2026-08-24

Updated: 2026-08-25

Importance score: 88/100

The gist: * I.

Key concepts

O*NET task statements
These are used by the authors to create a detailed inventory of actual job functions across industries. This provides a highly grounded vocabulary for defining work, moving beyond general categories like "healthcare" to specific tasks.
Defining the setting
This involves specifying the environment and all available tools and materials for a task. The authors emphasize that context is important because success depends on whether an output was possible given these specific constraints.
Activity, Setting, and Product
The paper proposes tying benchmark scores to three explicit elements: the work activity (the job done), the tested setting (tools and materials), and the scored work product. This forces clarity in claims about AI competence.

Terminology

Summary

I. Introduction and Problem Statement

The paper addresses the growing body of work on knowledge-work AI, which encompasses labor where knowledge serves as both the source material and the output. While various LLM agents are studied in relation to these tasks—including coding, research, and administrative coordination—the current evaluation methods remain largely tied to the evaluation logic of traditional NLP tasks, such as bounded input-output behavior.

The core problem identified is that benchmark performance scores often fail to reliably indicate a system's capability in real-world deployment settings. This is because the output of knowledge work cannot be understood only through its visible content. A simple final answer or patch may have different meanings depending on the role, materials, setting, and receiving workflow. Consequently, a benchmark that reports only the final output cannot show whether that output can support downstream coordination or continuation (e.g., being checked, revised, filed).

To prevent this mismatch between what a benchmark score measures and the broader knowledge-work capability it is used to claim, this paper proposes a design and reporting approach centered on three critical components: defining the work activity under evaluation, specifying the tested setting, and scoring the appropriate work product.

II. Theoretical Foundations of Knowledge Work

The paper grounds its approach in established theories regarding how professional labor is organized:

  1. Roles and Responsibilities: Expert work is organized through roles, authority, problem areas, and boundaries of responsibility. A visible output does not define the work being evaluated by itself.

  2. Situated Action: Performance depends on the conditions of action—local materials, tools, instructions, and social circumstances.

  3. Boundary Objects and Coordination: Knowledge-work outputs often need to move across actors and systems; they must remain usable in different local contexts while supporting coordination across communities.

The paper defines five key terms:

  • Knowledge work: A form of labor where knowledge is both the source material and the output.

  • Work activity: A recurrent kind of work that appears across many occupations, such as analysis and coordination.

  • Benchmark task: The task given to a system in a benchmark, defined by the prompt, inputs, tools, environment, and expected submission.

  • Tested setting: The benchmark-defined version of the work setting (materials, tools, roles) under which the work activity is evaluated.

  • Work product: The visible content produced by the system (e.g., an answer or code snippet).

  • Work output: The object left by the system for review, filing, execution, or downstream continuation (e.g., a revised document or state change).

  • Supported claim: The strongest claim a score can support, given the work activity, tested setting, and scored work product.

III. Methodology: A Three-Step Approach

The paper develops this framework into three distinct steps for benchmark design and reporting:

  1. Define the Work Activity: This requires selecting a middle level of abstraction—neither too narrow (a component task like retrieval) nor too broad (an entire domain like healthcare). The work activity names the work being claimed while allowing the report to explain how the benchmark task approximates it.

  2. Specify the Tested Setting: The report must explicitly state the conditions under which evaluation occurs, including:

  • Materials: Artifact availability (documents, policies).

  • Tools: System capabilities (search, edit, call APIs).

  • Role and Scope: What the system is allowed to decide or defer.

  • Workflow State: Where the task occurs (drafting, review, routing) and what state should remain after completion. The report must clarify if this is a simplified setting used to represent a broader class of work settings.

  1. Score the Proper Work Product: The benchmark must define the expected work product—the artifact or state left for downstream use. Scoring should move beyond general quality judgments (like helpfulness) and instead check specific components:
  • The rubric must address whether the result used appropriate materials, stayed within the assigned role, and included necessary handoff information.

  • The score provides evidence for the object it checks; if only a final answer is checked, the supported claim is limited to that output.

IV. Implementation: The Work-Activity Inventory

To operationalize this approach, the paper derived a preliminary work-activity reference set from O*NET task statements (a database of job tasks). This involved filtering for knowledge-work occupations, applying a stricter atlas-inclusion screen, and then using techniques like profession-neutral rewriting, UMAP + HDBSCAN clustering, and subsequent expert-panel review. This process resulted in 18 cross-occupation work activities (e.g, analysis, design, record-keeping).

V. Case Analyses

The paper demonstrates the approach through three benchmark case analyses:

  • GDP VAL (Non-code deliverable): The score supports a claim about producing a specified compliance deliverable under a fixed occupational prompt. The gap identified is that the artifact is not linked to an actual grant-review workflow, approval record, or audit trail.

  • ***O FFICE QA P RO (Grounded document analysis):***The case represents record-keeping and analysis by requiring the model to use specific Treasury Bulletin records and apply a calculation. The score supports a claim about grounded numerical analysis over Treasury Bulletin records. The gap is the absence of a reviewable analyst work product (e.g., an analyst memo).

  • ***APEX-SWE (Software engineering):***The score supports a claim about completing a specified integration task in the benchmarked software environment. The gap is that the released tasks do not report on full production deployment, rollback procedures, or long-term maintenance.

VI Conclusion and Future Directions

The paper concludes that benchmark scores only support work-capability claims through the specific work activity, tested setting, and scored product they actually represent. This framing clarifies what kind of evidence a benchmark provides before the system is used in an organization. While deployment evidence remains necessary for real-world claims, benchmarks remain useful as public, repeatable pre-deployment evidence. The final recommendation is that future knowledge-work benchmarks should move toward scoring work products more directly and validate the derived work-activity inventory against other occupational ontologies.

Improvements for AI systems

This research outlines a sophisticated methodology for cross-corpus ontological alignment and divergence analysis in occupational knowledge modeling. Given the high stakes—where misrepresentation of work activities can lead to flawed resource allocation, incorrect job profiling, or systemic operational failures—the methodological findings point toward critical architectural improvements in how AI systems process and fuse heterogeneous knowledge graphs.

Here are the specific improvements I recommend for next-generation AI systems, along with what the improved system can achieve:


Improvement: The current system treats ESCO and O*NET outputs as parallel checks. The improved architecture must integrate a specialized fusion module that doesn't just count occurrences but models the relationship type between source units (tasks to activities vs. competences to activities). This requires moving beyond simple proportional weighting to a relational-type weighting schema.

Improved AI System Capability:

The system can dynamically generate a Divergence Weighting Score for every work activity. Instead of merely stating that ESCO emphasizes 'design' and ONET emphasizes 'troubleshooting,' the system can quantify why the divergence occurs (e.g., "Activity X is primarily defined by discrete actions (ONET) which are necessary precursors to abstract competencies (ESCO).") This allows downstream systems to automatically adjust required training hours, tool requirements, or prerequisite knowledge based on whether the required activity profile is task-heavy or competence-heavy.

Sources

Related papers