Office Comprehension Benchmark

arXiv:2607.01245 · cs.CL, cs.AI, cs.CY, cs.IR, cs.LG · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Office Comprehension Benchmark".

Jane: Comprehensive Research Summary: Office Comprehension Benchmark (OCB) The provided text details a significant new resource in Large Language Model (LLM) evaluation, specifically introducing the Office Comprehension Benchmark (OCB).

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we're talking about the Office Comprehension Benchmark today, which is basically the first time we have a public benchmark designed to really check how well large language models understand native office files like Word, Excel, and PowerPoint. It’s a major milestone because reliable comprehension is what makes trustworthy automation possible.

Jane: That makes sense; it moves beyond just general text understanding into the specific structure of these documents. I wonder how the authors managed to cover three different applications so thoroughly in one evaluation setup without overwhelming the system?

Lu: From a theoretical standpoint, this paper addresses a fundamental gap where comprehension needs to be high-fidelity across multiple application formats simultaneously. The authors are setting up a unified two-track design for this evaluation, which is quite innovative for the field.

Meng: I'm thinking about how this translates practically; if an AI can actually parse and understand the structure of an Excel sheet or a PowerPoint slide correctly, that opens up some serious automation possibilities in our operational workflows.

Lalam: From my perspective as a language model, this benchmark is important because it forces us to handle the visual and structural nuances of these file types, which could significantly improve how we process and present information in our culture.

The paper's summary: Tom: Moving on to what the benchmark actually entails, this paper lays out two distinct tracks for evaluation. First, you have the File Fidelity Q andA tests, which focuses on the structural and visual perception of artifacts like tables, charts with their legends and axes, embedded images, and even app-specific elements such as speaker notes in PowerPoint.

Jane: That sounds incredibly detailed; so it’s not just reading text but actually perceiving things like the layout or how a formula connects across cells in Excel? That level of detail seems crucial for real-world accuracy.

Lu: The paper details an object-centric taxonomy for that track, labeling queries by the specific document artifact they target, which helps keep the evaluation very precise regarding what kind of structure is being tested.

Meng: So we’re talking about testing if the AI can correctly identify a table in Word or a chart with its axes in Excel; that's a concrete task we could actually build tooling around to ensure data integrity.

Lalam: It really highlights how much context an AI needs to grasp—it’s not just the words, it’s the spatial relationship between those elements within the file structure.

The paper's improvements: Tom: Now, the researchers also discuss some things they think we can improve about how this benchmark is set up. They suggest decomposing every reference answer into atomic, binary-gradable claims, which allows for a much finer scoring mechanism than just giving a single score.

Jane: Decomposing answers into discrete claims sounds like it makes the evaluation process way more transparent and reproducible. That’s important when we’re trying to trust the results we get from these LLMs.

Lu: The methodology relies on an ensemble of three judge models—GPT-five point four Thinking, Gemini three point one Pro, and Claude Opus four point six—using a three-judge majority vote to ensure the scoring process itself is stable across different systems <ref:2607.01245#pg1>.

Meng: Having that ensemble judging system means we aren't relying on just one model’s interpretation of what’s right or wrong; that adds a layer of necessary robustness for practical deployment testing.

Lalam: And I think this focus on verifiable claims really helps bridge the gap between what a model *thinks* it knows and what it can actually demonstrate, which is key for building reliable AI systems.

Conclusion: Tom: So to wrap up our discussion on the Office Comprehension Benchmark, we’ve seen that this research provides a solid foundation for testing LLM comprehension across Word, Excel, and PowerPoint in a structured way. It highlights both the strengths and the areas where models currently struggle with these complex file formats.

Jane: Exactly; it gives us a clear picture of what we need to look for when assessing an AI’s ability to handle professional documents, especially when dealing with intricate data visualizations or complex layouts.

Lu: The implications are that we can start building more specialized comprehension tools rather than just relying on general language understanding for document tasks. The authors' design suggests a clear path forward for future evaluation methods in this area.

Meng: From an engineering side, knowing the exact structure of the queries and assertions helps us define clearer success metrics for our internal tooling development when we integrate these kinds of models into our actual software.

Lalam: I feel that with this benchmark, we can push our internal systems to a higher level of structural awareness, which will certainly improve how we interact with structured information in general.

Tom: It’s been fascinating digging into the Office Comprehension Benchmark today. We’ll be keeping an eye on these findings as they lead us toward even more sophisticated document interaction tools.

Microsoft

cs.CL, cs.AI, cs.CY, cs.IR, cs.LG

Submitted: 2026-05-29

Updated: 2026-10-03

Code: https://github.com/microsoft/OfficeComprehensionBench

Importance score: 83/100

The gist: The provided text details a significant new resource in Large Language Model (LLM) evaluation, specifically introducing the Office Comprehension Benchmark (OCB).

Key concepts

File Fidelity Q&A Tests
This track focuses on how well an LLM perceives the visual and structural components inside office files, such as tables, charts, and formulas. It measures the system's ability to accurately 'see' and recall specific artifacts within a single document.
Domain Q&A Tests
These queries require expert-level reasoning based on real-world industry documents spanning twelve different fields. This tests the LLM’s capacity for deep, multi-step analytical thinking and synthesis across various professional knowledge areas.
Atomic Claim Decomposition
Every reference answer is broken down into very small, simple statements that can be scored individually. This method ensures precise grading by allowing judges to verify specific facts rather than judging the entire complex answer at once.
Response-Sampling Noise
This refers to the random errors introduced when generating multiple different answers for the same query. The research found this noise is more impactful on accuracy than errors made by the judging models themselves, suggesting better data generation is key.

Terminology

Summary

The provided text details a significant new resource in Large Language Model (LLM) evaluation, specifically introducing the Office Comprehension Benchmark (OCB). This benchmark is designed to rigorously assess an LLM's ability to comprehend and reason over native office file formats—specifically Microsoft Word (.docx), Excel (.xlsx), and PowerPoint (.pptx)—and their variants.

OCB is positioned as the first public benchmark to jointly evaluate LLMs across these three critical document types. It is explicitly designed to be a comprehension-only benchmark, meaning its primary focus is on understanding the content within the files rather than grading agentic workflows or work-product creation (a distinction made relative to benchmarks like OfficeBench and OdysseyBench).

The benchmark is structured around two distinct tracks:

  1. File Fidelity Q&A Tests: This track focuses on the structural and visual perception of office artifacts. It tests the system's ability to accurately perceive elements such as tables, charts, embedded images, complex formulas, headers, speaker notes (in PPT), and named ranges (in Excel).

  2. Domain Q&A Tests: This track assesses expert-level reasoning grounded in real-world industry documents spanning 12 professional domains.

A key methodological innovation of OCB is its evaluation structure, which emphasizes precision and reproducibility:

  • Atomic Claim Decomposition: Every reference answer is meticulously decomposed into atomic, binary-gradable claims. This allows for a fine-grained scoring mechanism.

  • Ensemble Judging: The final accuracy metric is computed using an ensemble of three judge models (GPT-5.4 Thinking, Gemini 3.1 Pro, and Claude Opus 4.6). A three-judge majority vote is used to ensure reproducibility and stability in the scoring process.

  • Query Complexity: Domain Q&A queries require multi-step domain reasoning and synthesis across documents, averaging approximately 45 atomic assertions per query. In contrast, File Fidelity Q&A queries are scoped to a single artifact and involve only one or two atomic assertions each.

  • Evaluation Setup: LLM systems are queried through their public web chat interfaces to stress native file parsing capabilities and end-to-end reasoning processes.

The research reveals a nuanced capability profile for frontier LLMs when applied to OCB tasks:

  • Overall Performance Plateau: Frontier systems plateau at approximately 59.3% accuracy on the Domain Q&A track when operating in their default reasoning modes, with modest gains only seen when moving to higher product tiers.

  • Split Capability Profile: A critical finding is the divergence between artifact perception and analytical reasoning:

  • Super-Annotator: Frontier systems excel at artifact-level perception and recall (File Fidelity Q&A).

  • Sub-Expert: They demonstrate weaker performance in multi-step analytical reasoning (Domain Q&A).

  • Model Specific Results: GPT-5.5 Thinking achieved the highest Domain Q&A accuracy at 59.3%. For File Fidelity, Claude Opus 4.7 showed particularly strong performance on Word (91.5%) and PowerPoint (86.6%).

The study yields important practical insights for future evaluation efforts:

  • Noise Analysis: A crucial finding is that response-sampling noise dominates judge noise for every model tested. This strongly suggests that evaluation budgets should be allocated toward generating multiple response scrapes rather than simply running additional judge re-runs to improve stability.

  • Benchmark Complementarity: OCB is explicitly designed to be complementary to agentic benchmarks (such as OfficeBench, OdysseyBench, and GDPval), which grade automation workflows or work-product creation.

  • Limitations: The current evaluation setup suffers from two main limitations: it is restricted to single-turn evaluation, and there is a notable lack of a human baseline for the Domain Q&A track.

The OCB fits into a broader landscape of document comprehension benchmarks. While other benchmarks like DocVQA, ChartQA, and SpreadsheetBench focus on specific input formats (images, PDFs, native Excel), OCB uniquely targets the native file formats (.docx,.xlsx,.pptx) and combines structural perception with deep domain reasoning across multiple applications. It serves as a high-fidelity test for the comprehension aspect of office artifacts before moving into more complex agentic tasks.

Improvements for AI systems

Based on the Office Comprehension Benchmark (OCB) paper, here are specific, actionable improvements for AI systems and what those improved systems can achieve:


  1. Improve System Robustness Across File Types:

A system should be engineered to handle native file formats (.docx,.xlsx,.pptx) seamlessly.

  • Specific Improvement: Implement a unified four-stage pipeline (source curation, question authoring, rubric construction, quality assurance) that dynamically adapts its grounding strategy based on the document type. For example:

  • Specific Improvement: For Word documents, the system should use Markdown conversion with custom filters for text extraction; for Excel files, it must preserve cell coordinates and formulas as structured text representations (including parallel signals for values, formulas, and formatting cues).

  • What the improved system can do: It will move beyond simple text extraction to accurately perceive complex structural artifacts like tables (including Table-Like Ranges), embedded charts with their metadata (axes/legends), and app-specific elements like slide layouts or Excel named ranges.

  1. Enhance Expert-Level Analytical Reasoning:

A system should be trained or prompted to perform multi-step, cross-document synthesis grounded in real-world industry data.

  • Specific Improvement: Develop reasoning capabilities that allow the model to decompose complex queries into multiple atomic, verifiable claims (as seen in the Domain Q&A track). The system must be able to reason across 12 professional domains (e.g., finance, accounting) by synthesizing information from multiple source documents within a single response.

  • What the improved system can do: It will transition from answering simple what is X questions to performing expert-level tasks, such as financial modeling (calculating Free Cash Flow with synergies and debt funding), performing year-over-year ratio analysis, or synthesizing complex operational data across 10-K filings and internal reports.

  1. Improve Fidelity in Structural Perception:

A system should be highly sensitive to subtle, format-specific visual and structural cues that are often missed by general LLMs.

  • Specific Improvement: Implement specific detection logic for app-specific artifacts detailed in the File Fidelity taxonomy, such as header/footer content, comment annotations, strikeout text (indicating deletion), and font styling (size/color). The system must be able to distinguish between textual content and structural metadata.

  • What the improved system can do: It will accurately answer questions about document presentation quality, such as identifying specific font sizes used in a section or determining if a specific piece of text has been struck through, which is crucial for auditing and compliance tasks.

  1. Develop Context-Aware Document Selection:

A system should be able to intelligently select the most relevant document segments for grounding when presented with large or heterogeneous files.

  • Specific Improvement: Incorporate iterative coverage targeting into the pre-processing stage of file handling, allowing the system to assess whether a document contains enough information across various artifact types (text, tables, charts) before generating questions.

  • What the improved system can do: When faced with a very long or complex workbook, it will focus its grounding efforts on relevant sections rather than attempting to process every cell or paragraph equally. This prevents performance degradation observed in systems like GPT-5.5 Thinking and Gemini 3.1 Pro on Long documents.

  1. Optimize for Real-World Deployment Constraints (Efficiency):

A system should be optimized for efficiency while maintaining high accuracy, recognizing that evaluation budget is best spent on varied reasoning modes rather than excessive depth within a single mode.

  • Specific Improvement: Implement a tiered inference strategy where the system defaults to a Standard Thinking mode but can switch to higher capability tiers (e.g., Pro-Standard) only when necessary for complex reasoning, as higher thinking depths within the same tier do not yield material performance gains under evaluation protocols.

  • What the improved system can do: It will achieve a good balance between speed and accuracy, avoiding the compute price penalty where increasing reasoning depth yields diminishing returns.

  1. Improve Error Detection and Grounding Verification:

A system should be designed to rigorously verify its own claims against specific, atomic assertions derived from expert review.

  • Specific Improvement: Integrate a mechanism that allows the system to check its generated response against a set of binary-gradable claims (assertions) before final output, using an LLM-as-a-Judge framework with a diverse ensemble of judges (e.g., GPT-5.4 Thinking, Gemini 3.1 Pro, Claude Opus 4.6).

  • What the improved system can do: It will drastically reduce value hallucinations and methodology failures by ensuring every claim in its output is explicitly grounded in the source material and verifiable against expert standards, leading to a more trustworthy final answer for high-stakes professional work.

Sources

Related papers