WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models

arXiv:2510.22276 · cs.CV, cs.CL · Submitted 2025-10-25 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models".

Tom: Contrastive vision-language models have achieved remarkable progress through largescale pretraining,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, looking at "WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models," the central thesis is testing if global pretraining alone is adequate for culture-specific understanding or if adding natively sourced data can improve performance on those specific cultural benchmarks. They introduce WAON, which they describe as the largest publicly available native Japanese image-text dataset constructed from native Japanese web content in Common Crawl, totaling approximately one hundred fifty-five million examples <ref:2510.22276#pg0,the largest publicly available native Japanese image-text dataset constructed from native>.

Jane: What really matters is their comparison between this native data and other types of data. The paper claims that fine-tuning on WAON consistently achieves stronger performance across all tested baselines when evaluating models on Japanese cultural benchmarks, whereas fine-tuning on English-to-Japanese translated data consistently degrades the model's understanding of Japanese culture.

Lu: That comparison is key because it directly addresses a limitation in prior work where translation was used as a substitute for native context, and this study provides concrete evidence against that approach for cultural tasks.

Meng: It sounds like the authors are making a strong statement about the necessity of cultural specificity over mere linguistic translation when dealing with complex cultural concepts in vision-language models. That's something we need to keep in mind when designing our own systems.

Lalam: For me, this confirms that for capturing subtle visual and contextual cues tied to Japanese culture, the direct source material is what truly matters, not a filtered or translated version of it.

Conclusion: Tom: So, wrapping up this discussion on "WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models," the authors, including Issa Sugiura and Shuhei Kurita, have demonstrated that data origin is a critical factor when it comes to adapting contrastive vision-language models for cultural tasks. Their findings suggest that while global pretraining sets a strong foundation, natively sourced Japanese data is essential for culture-specific fine-tuning because translation alone simply cannot compensate for the lack of original source material.

Jane: In simpler terms, what this means is that if you want an AI to really get the cultural nuances of something specific, like Japanese traditions or scenery, you need to train it on examples that are genuinely from that culture itself rather than relying on translated captions from other languages. The WAON dataset and their benchmark show this effect clearly.

Lu: The implication here is that future work should focus not just on scaling up the total amount of data, but crucially, on ensuring the quality and native sourcing of that data to achieve deep cultural understanding in AI systems.

Meng: From an engineering standpoint, this points toward a need for smarter data pipelines that can effectively filter and prioritize native content over translated content when training models for specialized domains. That's a practical direction we can take.

Lalam: I see the impact as improving the cultural accuracy of applications across various domains, making AI outputs feel much more authentic and less reliant on generalized, often Western-centric, knowledge bases.

Tom: Exactly! The WAON study shows that investing in culturally specific data curation leads to better performance on those very cultural benchmarks. That's a solid foundation for how we should approach building specialized vision models.

Issa Sugiura, Shuhei Kurita, Yusuke Oda, Daisuke Kawahara, Yasuo Okabe, Naoaki Okazaki

Kyoto University · NII LLMC (National Institute of Informatics Laboratory) · Waseda University

cs.CV, cs.CL

Submitted: 2025-10-25

Updated: 2026-10-02

Comments: Accepted to AACL 2026 (Findings)

Code: https://github.com/rom1504/img2dataset

Project page: https://speed1313.github.io/WAON

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: Contrastive vision-language models have achieved remarkable progress through largescale pretraining, but this work investigates whether global pretraining alone is sufficient for culture-specific

Key concepts

WAON
This is the largest publicly available native Japanese image-text dataset, built from web content. It contains about 155 million high-quality image and text pairs sourced directly from Japanese web pages, designed specifically for teaching models about Japanese culture.
WAONBench
A manually curated benchmark created to rigorously test the WAON dataset. It covers eight diverse cultural categories like food, tradition, and scenery. This benchmark helps measure how well a model understands specific aspects of Japanese culture compared to other datasets.
Contrastive Vision-Language Models
These are AI models that learn relationships between images and text by comparing them to find similarities. The study investigates whether using native cultural data improves these models' ability to grasp subtle cultural nuances in Japanese compared to using translated data.
ReLAION (en→ja translation)
This is a dataset created by translating English image-text pairs into Japanese. The study found that fine-tuning on this translated data actually harmed the model's understanding of Japanese culture, proving that translation alone is insufficient for cultural adaptation.

Terminology

Summary

Contrastive vision-language models have achieved remarkable progress through largescale pretraining, but this work investigates whether global pretraining alone is sufficient for culture-specific understanding or if further adaptation with natively sourced data can boost performance. The study introduces WAON, a massive native Japanese image-text dataset, and demonstrates that fine-tuning on this data consistently achieves stronger performance on Japanese cultural benchmarks than fine-tuning on English-to-Japanese translated data.

The gist

Fine-tuning on WAON consistently achieves the best performance across all baselines when evaluating models on Japanese cultural benchmarks, while fine-tuning on ReLAION (en→ja translation) consistently degrades Japanese cultural understanding.

Dataset Construction and Scale

The paper introduces WAON, described as the largest publicly available native Japanese image-text dataset constructed from native Japanese web content in Common Crawl, containing approximately 155 million examples. The construction pipeline involves several cascaded steps to ensure high quality, including:

  1. Extracting Japanese HTML documents from WARC files.

  2. Performing text-based deduplication before downloading images, followed by image-based filtering and deduplication (using pHash).

  3. Applying image quality filtering, such as discarding images whose width or height is below 150 pixels or whose aspect ratio falls outside [0.5, 2.0], and applying NSFW classification to remove unsafe content.

  4. Filtering misaligned pairs using cosine similarity between image and text embeddings computed by SigLIP2-base-patch16-256, setting a threshold of 0.1 for semantic alignment filtering.

The final result after cross-snapshot deduplication yields approximately 155M high-quality image-text pairs.

Benchmark Development

To rigorously test the dataset's utility, the authors introduce WAONBench, a manually curated Japanese cultural benchmark designed to address limitations in prior resources. This benchmark comprises 374 classes spanning eight diverse categories: animal, building, event, everyday, food, nature, scenery, and tradition. It contains 161 classes and 7,654 examples in the Recruit dataset (Honda and Arai, 2024), which was limited to only food categories. WAONBench significantly expands this coverage by providing more than twice as many cultural concepts with 5 images per class for a total of 1,870 images. The visual diversity of WAONBench is quantified, achieving an average pairwise cosine distance of 0.495 compared to the Recruit dataset's 0.425.

Experimental Setup and Results

The experiments fine-tune the siglip2-base-patch16-256 model on three datasets: WAON (155M pairs), ReLAION (ja subset, 120M examples), and ReLAION (en→ja translation, 1,452M pairs). Evaluation is conducted across four benchmarks: WAONBench, Recruit, ImageNet (Western culture), and XM3600. The results show that fine-tuning on WAON consistently achieve[s] the best performance across all baselines. Specifically, WAON achieves a top-1 accuracy of 95.1% on WAONBench and 83.3% on Recruit, outperforming both baseline models and those fine-tuned on ReLAION variants. Conversely, fine-tuning on ReLAION (en→ja translation) consistently degrades Japanese cultural understanding, with performance dropping from 87.8% to 76.8% on WAONBench.

Conclusion and Implications

The study concludes that data origin is a critical factor for cultural adaptation in contrastive vision-language models. The findings suggest that while global pretraining provides a strong foundation, native Japanese data is crucial for culture-specific fine-tuning, as translation alone cannot compensate for the absence of natively sourced data. WAON serves as the largest publicly available natively sourced Japanese image-text dataset, and WAONBench provides a comprehensive benchmark to validate its effectiveness. The authors release WAON and code under the Apache 2.0 License, emphasizing ethical considerations regarding web data usage and content filtering.

Limitations

The study acknowledges several limitations, including taking Japanese as a single case study, fine-tuning only one model architecture (siglip2-base-patch16-256), and lacking a quantitative analysis of each pipeline step's contribution to dataset quality. Ethical considerations address the use of web data from Common Crawl, confirming that no image files are redistributed, and outlining mechanisms for reporting problematic samples. The authors also note that some harmful or biased content may remain despite filtering.

Acknowledgements

The research utilized the “mdx: a platform for building data-empowered society” and ABCI 3.

Improvements for AI systems

Here are specific improvements for AI systems based on the findings presented in the paper:

  1. Improve cross-cultural performance in vision-language models (VLMs) by fine-tuning them with natively sourced, culturally specific data (like WAON) rather than relying solely on globally pretraining or translated data. This allows models to achieve state-of-the-art accuracy on culture-specific benchmarks (e.g., achieving 95.1% top-1 accuracy on WAONBench).

  2. Develop a robust, high-quality dataset construction pipeline for low-resource languages by applying cascaded filtering and deduplication steps across multiple modalities (text, image quality, semantic alignment) to filter out noisy web data effectively.

  3. Enhance the diversity and robustness of cultural understanding in models by creating manually curated benchmarks that significantly increase class coverage beyond existing narrow datasets (like WAONBench expanding from 161 classes to 374 classes across 8 diverse categories).

  4. Improve visual diversity in fine-tuning data for culturally specific tasks by curating benchmark images to ensure wider dispersal in the embedding space, leading to more robust and less category-biased model representations (WAONBench achieving a higher average pairwise cosine distance than the Recruit dataset).

  5. Build specialized Japanese vision-language models that exhibit superior cultural adaptation capabilities by leveraging large-scale global pretraining backbones (like SigLIP2) followed by fine-tuning on natively sourced data, rather than just relying on general multilingual models or translated data.

The improved AI systems can:

  1. Accurately classify images into nuanced Japanese cultural categories (e.g., distinguishing between specific traditional festivals, regional scenery, and everyday objects) with high precision, even when tested on unseen Japanese cultural benchmarks.

  2. Perform better at zero-shot classification tasks involving Japanese cultural concepts because the training data reflects native visual and textual contexts rather than biased translations or Western imagery.

  3. Be more resilient to English-to-Japanese translation artifacts, as their performance will not degrade significantly when fine-tuned on the high-quality WAON dataset, unlike models fine-tuned on translated data.

  4. Generate more diverse and contextually accurate visual representations of Japanese culture by being trained on a benchmark that covers a broad spectrum of cultural domains (animal, building, event, food, etc.), leading to better generalization across different aspects of Japanese identity.

Abstract

Contrastive vision-language models have achieved remarkable progress through large-scale pretraining. Recent work has shown that removing English-only caption filters and pretraining on global data is effective for improving multicultural performance. We study whether such global pretraining is sufficient for culture-specific understanding, or whether further adaptation with natively sourced data can boost performance beyond what global pretraining alone achieves. To enable this investigation, we present WAON, the largest publicly available native Japanese image-text dataset constructed from native Japanese web content in Common Crawl, containing approximately 155 million examples. We also introduce WAON-Bench, a manually curated Japanese cultural benchmark spanning 374 classes. Through comparative fine-tuning experiments on multiple Japanese image-text datasets, we observe that models fine-tuned on WAON consistently achieve stronger performance on Japanese cultural benchmarks than those fine-tuned on English-to-Japanese translated data. Controlled experiments at matched scale, filtering, and training budget across two model families further indicate that native web origin is the primary driver of this gain. We release our dataset, benchmark, model, and code.

Sources

Related papers