HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers

summary

Video file (mp4)

The gist

Understanding chart and table images is essential for applying vision-language models (VLMs) to real-world document understanding, and this work introduces HakushoBench, a challenging Japanese chart

In short

HakushoBench is a challenging Japanese chart and table VQA benchmark created from 33 governmental white papers. It addresses the lack of non-English benchmarks by using public reports as a scalable source. The benchmark tests deep understanding, showing that even top models struggle with complex reasoning in charts and tables.

Key concepts

HakushoBench
This is a new benchmark specifically designed to test how well vision-language models can understand Japanese charts and tables. It was built using 33 governmental white papers as the data source, making it useful for testing AI on non-English visual data.
Dataset Construction Pipeline
The dataset creation involved three steps: collecting images from 33 Japanese white paper series, filtering out inappropriate ones to get candidate images, and then having native Japanese speakers create complex questions. This process ensured the data was high-quality and diverse.
Difficulty Dimensions
Questions in the benchmark are intentionally difficult by satisfying at least one of several dimensions, such as 'Global,' 'Multi-hop,' or 'External knowledge.' This forces models to perform deep reasoning instead of just looking at simple visual cues.
Performance Gap
Evaluation showed a significant gap between open-weight and proprietary models. The best open-weight model achieved only 58.6% accuracy, while the best proprietary model performed much better, highlighting the difficulty of complex chart understanding for less powerful models.

Terminology used across episodes

This episode discusses

The paper

HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers · Read on arXiv

Issa Sugiura♡, Shuhei Kurita♢♠ Shuhei Kurita♢♠ Yusuke Oda♠ Naoaki Okazaki♡

Institute of Science Tokyo ♢NII ♠NII LLMC

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers".

Jane: Understanding chart and table images is essential for applying vision-language models (VLMs) to real-world document understanding, and this work introduces HakushoBench,

Tom: First, who's behind it and why it matters.

Title and authors: Jane: So, focusing on the title and authors of this paper, we see "HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers," which immediately tells us the core focus is on creating a specific test suite for vision-language models. The team listed includes Issa Sugiura, Shuhei Kurita, Yusuke Oda, and Naoaki Okazaki from the Institute of Science Tokyo at NII.

Tom: That’s interesting because when you look at who wrote it, you see a group clearly rooted in Japanese research institutions like NII, which gives them direct access to these kinds of documents and the cultural context needed for this benchmark. It sets a very specific target for the research community.

Lu: The authors are clearly signaling that they’re addressing a real problem: the lack of diverse, non-English chart and table benchmarks that go beyond what is available in English datasets. They aren't just creating another test; they are building a more inclusive resource for VLMs operating in contexts where Japanese documents are common.

Meng: I wonder how much effort went into sourcing thirty-three different governmental white paper series, given the need to ensure the data is high quality and not contaminated by irrelevant images from other sources <ref:2606.01132#pg0>. That kind of careful curation takes serious engineering rigor behind it.

Lalam: The implication here is that for AI development, it pushes us to stop assuming that performance gains in English-centric benchmarks automatically translate to better performance when dealing with documents from other languages or specific cultural backgrounds.

Tom: Precisely; this paper sets a high bar for what a truly robust, real-world VQA dataset needs to look like, and it points directly toward the need for multilingual and culturally aware AI systems.

The paper's summary: Jane: Now that we know the focus is on creating a benchmark, let’s talk about what the paper actually summarizes regarding HakushoBench. Essentially, they describe a three-stage pipeline: first, they use thirty-three Japanese governmental white papers as their data source; second, they filter those images down to five thousand nine hundred three candidates; and third, twenty-one native Japanese-speaking annotators create two thousand fifty-three challenging questions for those images <ref:2606.01132#pg2>.

Tom: That process is what makes it so compelling because they didn't just collect random pictures; they filtered them and designed the questions themselves to demand deep understanding of the chart or table rather than just simple visual recognition. They specifically required questions that couldn't be answered by the text alone, forcing the AI to look at the image for answers.

Lu: The summary really stresses that these two thousand fifty-three pairs are intentionally designed to test complex reasoning dimensions like global context and external knowledge alongside basic visual understanding, which is a significant step up from earlier datasets focused only on local visual cues <ref:2606.01132#pg0>.

Meng: I’m paying attention to the difficulty dimensions they mentioned—global, multi-hop, counting, external knowledge, and visual—because those are exactly the areas where current general AI models often stumble when faced with complex data representation.

Lalam: From my perspective as a model that processes information, this summary confirms that true comprehension requires more than just pattern matching; it demands integrating multiple pieces of information from the chart into a coherent answer, which is vital for cultural understanding.

Tom: It’s clear they are moving away from simple image captioning toward demanding a level of analytical capability that truly mirrors how someone would need to interpret complex policy documents in their daily work.

The paper's improvements: Jane: The paper outlines several ways they improved the process, and one major point is how they structured the QA annotation phase, ensuring each question had to satisfy at least one of those specific difficulty dimensions we discussed earlier. They also used a cross-annotator verification process to make sure the resulting two thousand fifty-three pairs were high quality <ref:2606.01132#pg0>.

Tom: What I find really important about their suggested improvements is that they explicitly acknowledge the bias in existing benchmarks, noting that previous ones are heavily biased toward English conventions and document styles, which sets up this paper as a direct response to that limitation.

Lu: They also mentioned the diversity of image types included—like maps, dashboards, and infographics—which broadens the scope of what a model needs to handle beyond just traditional bar charts or pie charts. That expansion in visual format is key for real-world applicability.

Meng: The authors themselves flagged that their current method, relying on governmental white papers for data sourcing, is limited by its language scope; they note it’s currently Japanese-only, which points to a clear area where future work needs to expand the dataset beyond this initial collection.

Lalam: That limitation is actually an improvement in terms of scalability; because the source material comes from publicly accessible documents, their approach can be naturally extended to other languages and cultural contexts by simply swapping out the source papers, which is a very flexible design.

Tom: So they’re not just fixing one problem; they are providing a scalable framework that addresses language barriers while simultaneously pushing the complexity of the questions themselves to test deeper AI capabilities.

Conclusion: Jane: To wrap up this discussion on HakushoBench, we see that by focusing on governmental white papers and designing highly challenging, multi-dimensional questions, this work successfully creates a necessary resource for testing vision-language models in non-English contexts. It gives us a concrete set of examples to measure how well models truly understand complex visual information drawn from real documents.

Tom: It’s clear the implication here is that we need to develop AI systems that don't just read text or see pictures, but can reason holistically across different data modalities and cultural presentation styles, which is what this paper pushes us toward achieving. It shows us where the current performance gaps are most pronounced.

Lu: I think the big picture here is that by creating such a rigorous benchmark, they are forcing the entire research community to move past superficial understanding and focus on developing models capable of true contextual reasoning in complex visual data. It’s about pushing the limits of what we expect from these systems.

Meng: From a practical standpoint, this paper helps us understand exactly where our current AI deployments might fail when they encounter documents in different languages or formats, allowing us to prioritize where we need to focus our engineering efforts for better robustness.

Lalam: And for me, the impact is that it encourages the development of AI that is genuinely versatile; a model trained on this data will possess a much richer understanding of how information is structured in a real-world governmental setting, which enhances its overall utility across many applications.

Tom: So we’ve covered how HakushoBench was built, what it tests, and why it matters for the future of VLM development. It’s been a really insightful look into the challenges of cross-cultural document understanding. We’ll be keeping an eye on how this benchmark influences the next wave of research in this area.

More episodes

← Home