HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers".
Jane: Understanding chart and table images is essential for applying vision-language models (VLMs) to real-world document understanding, and this work introduces HakushoBench,
Tom: First, who's behind it and why it matters.
Title and authors: Jane: So, focusing on the title and authors of this paper, we see "HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers," which immediately tells us the core focus is on creating a specific test suite for vision-language models. The team listed includes Issa Sugiura, Shuhei Kurita, Yusuke Oda, and Naoaki Okazaki from the Institute of Science Tokyo at NII.
Tom: That’s interesting because when you look at who wrote it, you see a group clearly rooted in Japanese research institutions like NII, which gives them direct access to these kinds of documents and the cultural context needed for this benchmark. It sets a very specific target for the research community.
Lu: The authors are clearly signaling that they’re addressing a real problem: the lack of diverse, non-English chart and table benchmarks that go beyond what is available in English datasets. They aren't just creating another test; they are building a more inclusive resource for VLMs operating in contexts where Japanese documents are common.
Meng: I wonder how much effort went into sourcing thirty-three different governmental white paper series, given the need to ensure the data is high quality and not contaminated by irrelevant images from other sources <ref:2606.01132#pg0>. That kind of careful curation takes serious engineering rigor behind it.
Lalam: The implication here is that for AI development, it pushes us to stop assuming that performance gains in English-centric benchmarks automatically translate to better performance when dealing with documents from other languages or specific cultural backgrounds.
Tom: Precisely; this paper sets a high bar for what a truly robust, real-world VQA dataset needs to look like, and it points directly toward the need for multilingual and culturally aware AI systems.
The paper's summary: Jane: Now that we know the focus is on creating a benchmark, let’s talk about what the paper actually summarizes regarding HakushoBench. Essentially, they describe a three-stage pipeline: first, they use thirty-three Japanese governmental white papers as their data source; second, they filter those images down to five thousand nine hundred three candidates; and third, twenty-one native Japanese-speaking annotators create two thousand fifty-three challenging questions for those images <ref:2606.01132#pg2>.
Tom: That process is what makes it so compelling because they didn't just collect random pictures; they filtered them and designed the questions themselves to demand deep understanding of the chart or table rather than just simple visual recognition. They specifically required questions that couldn't be answered by the text alone, forcing the AI to look at the image for answers.
Lu: The summary really stresses that these two thousand fifty-three pairs are intentionally designed to test complex reasoning dimensions like global context and external knowledge alongside basic visual understanding, which is a significant step up from earlier datasets focused only on local visual cues <ref:2606.01132#pg0>.
Meng: I’m paying attention to the difficulty dimensions they mentioned—global, multi-hop, counting, external knowledge, and visual—because those are exactly the areas where current general AI models often stumble when faced with complex data representation.
Lalam: From my perspective as a model that processes information, this summary confirms that true comprehension requires more than just pattern matching; it demands integrating multiple pieces of information from the chart into a coherent answer, which is vital for cultural understanding.
Tom: It’s clear they are moving away from simple image captioning toward demanding a level of analytical capability that truly mirrors how someone would need to interpret complex policy documents in their daily work.
The paper's improvements: Jane: The paper outlines several ways they improved the process, and one major point is how they structured the QA annotation phase, ensuring each question had to satisfy at least one of those specific difficulty dimensions we discussed earlier. They also used a cross-annotator verification process to make sure the resulting two thousand fifty-three pairs were high quality <ref:2606.01132#pg0>.
Tom: What I find really important about their suggested improvements is that they explicitly acknowledge the bias in existing benchmarks, noting that previous ones are heavily biased toward English conventions and document styles, which sets up this paper as a direct response to that limitation.
Lu: They also mentioned the diversity of image types included—like maps, dashboards, and infographics—which broadens the scope of what a model needs to handle beyond just traditional bar charts or pie charts. That expansion in visual format is key for real-world applicability.
Meng: The authors themselves flagged that their current method, relying on governmental white papers for data sourcing, is limited by its language scope; they note it’s currently Japanese-only, which points to a clear area where future work needs to expand the dataset beyond this initial collection.
Lalam: That limitation is actually an improvement in terms of scalability; because the source material comes from publicly accessible documents, their approach can be naturally extended to other languages and cultural contexts by simply swapping out the source papers, which is a very flexible design.
Tom: So they’re not just fixing one problem; they are providing a scalable framework that addresses language barriers while simultaneously pushing the complexity of the questions themselves to test deeper AI capabilities.
Conclusion: Jane: To wrap up this discussion on HakushoBench, we see that by focusing on governmental white papers and designing highly challenging, multi-dimensional questions, this work successfully creates a necessary resource for testing vision-language models in non-English contexts. It gives us a concrete set of examples to measure how well models truly understand complex visual information drawn from real documents.
Tom: It’s clear the implication here is that we need to develop AI systems that don't just read text or see pictures, but can reason holistically across different data modalities and cultural presentation styles, which is what this paper pushes us toward achieving. It shows us where the current performance gaps are most pronounced.
Lu: I think the big picture here is that by creating such a rigorous benchmark, they are forcing the entire research community to move past superficial understanding and focus on developing models capable of true contextual reasoning in complex visual data. It’s about pushing the limits of what we expect from these systems.
Meng: From a practical standpoint, this paper helps us understand exactly where our current AI deployments might fail when they encounter documents in different languages or formats, allowing us to prioritize where we need to focus our engineering efforts for better robustness.
Lalam: And for me, the impact is that it encourages the development of AI that is genuinely versatile; a model trained on this data will possess a much richer understanding of how information is structured in a real-world governmental setting, which enhances its overall utility across many applications.
Tom: So we’ve covered how HakushoBench was built, what it tests, and why it matters for the future of VLM development. It’s been a really insightful look into the challenges of cross-cultural document understanding. We’ll be keeping an eye on how this benchmark influences the next wave of research in this area.
Issa Sugiura♡, Shuhei Kurita♢♠ Shuhei Kurita♢♠ Yusuke Oda♠ Naoaki Okazaki♡
Institute of Science Tokyo ♢NII ♠NII LLMC
cs.CV
Submitted: 2026-05-31
Updated: 2026-10-02
Comments: Accepted to AACL 2026 (Findings)
Code: https://github.com/stockmarkteam/business-slide-quest
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: Understanding chart and table images is essential for applying vision-language models (VLMs) to real-world document understanding, and this work introduces HakushoBench, a challenging Japanese chart
Key concepts
- HakushoBench
- This is a new benchmark specifically designed to test how well vision-language models can understand Japanese charts and tables. It was built using 33 governmental white papers as the data source, making it useful for testing AI on non-English visual data.
- Dataset Construction Pipeline
- The dataset creation involved three steps: collecting images from 33 Japanese white paper series, filtering out inappropriate ones to get candidate images, and then having native Japanese speakers create complex questions. This process ensured the data was high-quality and diverse.
- Difficulty Dimensions
- Questions in the benchmark are intentionally difficult by satisfying at least one of several dimensions, such as 'Global,' 'Multi-hop,' or 'External knowledge.' This forces models to perform deep reasoning instead of just looking at simple visual cues.
- Performance Gap
- Evaluation showed a significant gap between open-weight and proprietary models. The best open-weight model achieved only 58.6% accuracy, while the best proprietary model performed much better, highlighting the difficulty of complex chart understanding for less powerful models.
Terminology
Summary
Understanding chart and table images is essential for applying vision-language models (VLMs) to real-world document understanding, and this work introduces HakushoBench, a challenging Japanese chart and table VQA benchmark built from 33 governmental white papers. This benchmark matters because it addresses the scarcity of non-English chart and table benchmarks by leveraging publicly accessible governmental white papers as a scalable source for constructing realistic, diverse datasets across various domains.
Dataset Construction
HakushoBench is constructed through a three-stage pipeline to ensure high quality and diversity. First, it involves White papers as a data source,
where images are collected from 33 distinct Japanese governmental white paper series published by agencies like the Japanese Digital Agency. To mitigate contamination risk, the researchers use only the most recent edition from each white paper series.
The initial collection yields 18,539 images, which are then subjected to Image filtering and candidate selection,
where inappropriate images are manually removed to retain 5,903 candidate chart and table images.
QA Annotation Process
The QA annotation phase is performed by 21 native Japanese-speaking annotators hired through a professional annotation agency.
The requirements for the questions are strict: they must (a) be unanswerable from the question text alone, requiring the image to resolve; (b) be natural and well-posed; and (c) admit a single, unambiguous answer expressible in a word, phrase, or short sentence.
To encourage complex reasoning, each accepted question must satisfy at least one of several difficulty dimensions,
including Global,
Multi-hop,
Counting,
External knowledge,
and Visual.
Annotators also assign flags for these difficulty types.
Benchmark Evaluation
The benchmark is designed to assess deep and holistic understanding rather than just local visual cues alone, containing 2,053 unique images spanning over 10 image types. The evaluation involves testing a broad suite of open-weight and proprietary VLMs against existing benchmarks like JGraphQA and ChartQA. Experiments reveal that HakushoBench is more challenging than JGraphQA
for open-weight models, with the best open-weight model achieving only 58.6% accuracy,
highlighting a significant performance gap between open-weight and proprietary models on complex chart and table understanding.
Error Analysis and Findings
Manual error analysis on the best proprietary model, Gemini 3 Pro, shows that even state-of-the-art models exhibit diverse errors including perception, external knowledge, and counting failures.
Representative failure cases include a perception error in reading scatter plot conditions,
a knowledge error in identifying the map location closest to Cape Erimo (襟裳岬),
and a counting error resulting in an off-by-one prediction.
These findings demonstrate that models still struggle with the advanced reasoning and visual understanding required for complex chart and table comprehension.
Limitations and Future Work
The current dataset is limited to Japanese, but the construction approach can be naturally extended to other languages and cultural contexts
because many countries publish analogous governmental reports. The researchers also note a Saturation at the frontier,
suggesting that future work might involve constructing a harder subset by filtering out questions solved by top models or collecting more demanding QA pairs targeting capabilities beyond current frontier models. Ethical considerations confirm that all data is sourced from public white papers, making privacy risks negligible.
extracted final answer: The best open-weight model reaches only 58.6% accuracy on HakushoBench, revealing a 34.9-point accuracy gap between the best proprietary and open-weight models.
[correct answer]: The best open-weight model, Qwen3-VL 8B, reaches only 58.6% accuracy on HakushoBench, and reveals a 34.9-point accuracy gap between the best proprietary and open-weight models.
reasoning: The extracted answer accurately reflects the key finding regarding the performance of open-weight models (58.6%) and the resulting performance gap compared to proprietary models (34.9 points), as stated in Section 1 and Section 6.1 of the paper.
Improvements for AI systems
Here are specific improvements for AI systems based on the findings of HakushoBench, and what those improved systems can achieve:
-
Improve generalization across non-English chart/table understanding by leveraging diverse, domain-specific data sources (like governmental white papers) instead of relying solely on English-centric benchmarks.
-
Develop more robust multimodal reasoning capabilities in Vision-Language Models (VLMs) to handle complex, multi-hop questions that require integrating information across multiple regions of a chart or table.
-
Enhance the performance gap between open-weight and proprietary VLMs for complex visual tasks by focusing training/prompting strategies on improving reasoning and external knowledge integration rather than just local visual cues.
-
Improve model accuracy in specific error modes identified in complex benchmarks, such as counting errors (off-by-one) and knowledge errors (requiring external geographic or domain-specific knowledge).
-
Enable VLMs to perform better on diverse visual formats beyond standard charts, including specialized types like Maps, Infographics, Dashboards, and Table layouts that combine multiple elements.
This improved AI system can:
-
Accurately answer complex questions about data presented in real-world Japanese governmental documents (e.g.,
What is the range of precipitation at this location?
orHow many patrol vessels are deployed?
). -
Demonstrate superior performance on tasks requiring multi-hop reasoning and holistic image understanding, such as calculating ratios from a table or understanding spatial relationships on a map/chart.
-
Exhibit stronger reasoning abilities when using Chain-of-Thought (CoT) prompting, leading to more reliable step-by-step problem solving.
-
Accurately detect and correct common errors in complex visual QA, such as miscounting objects or failing to retrieve necessary external knowledge for geographic context.
-
Be more versatile in interpreting a wide array of document visualizations (Bar, Line, Pie, Area charts, Maps, Tables) encountered across various policy domains (Security, Economy, Society).
Abstract
Understanding chart and table images is essential for applying vision-language models (VLMs) to real-world document understanding. While English benchmarks have advanced rapidly, non-English counterparts remain scarce, leaving it unclear whether this progress generalizes across languages. A key obstacle is the difficulty of collecting realistic and diverse non-English chart and table images at scale. To address this, we leverage governmental white papers as a source for benchmark construction, as they contain naturally occurring charts and tables across diverse formats and domains and are freely accessible in many countries. As a first instantiation, we introduce HakushoBench, a Japanese chart and table VQA benchmark built from 33 governmental white papers. HakushoBench contains 2,053 images spanning over 10 image types, with manually annotated and independently verified QA pairs designed to assess holistic understanding of charts and tables rather than local visual cues alone. Experiments across a broad range of VLMs show that HakushoBench is substantially harder than the existing Japanese benchmark and remains challenging for open-weight models: sub-10B open-weight models reach at most 58.6% accuracy, and even the flagship open-weight model Qwen3.5-397B-A17B trails Gemini 3 Pro by 8.1 points (85.8% vs. 93.9%), highlighting substantial room for improvement in complex chart and table understanding. We release our dataset and code.
Sources
- PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
- A Rosetta Stone for AI Benchmarks
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- DatBench: Discriminative, Faithful, and Efficient VLM Evaluations
- FigureQA: An Annotated Figure Dataset for Visual Reasoning
- Kimi-VL Technical Report
- GPT-4o System Card
- Qwen3-VL Technical Report
- Evaluating Multimodal Large Language Models on Vertically Written Japanese Text
- JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation
- Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- DeepSeek-OCR: Contexts Optical Compression
- POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models