WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
summary
The gist
Contrastive vision-language models have achieved remarkable progress through largescale pretraining, but this work investigates whether global pretraining alone is sufficient for culture-specific
In short
Researchers tested if large-scale pretraining alone is enough for vision-language models to understand Japanese culture. They created WAON, a massive native Japanese image-text dataset, and found that fine-tuning on this data significantly outperforms fine-tuning on translated English data when evaluating cultural understanding.
Key concepts
- WAON
- This is the largest publicly available native Japanese image-text dataset, built from web content. It contains about 155 million high-quality image and text pairs sourced directly from Japanese web pages, designed specifically for teaching models about Japanese culture.
- WAONBench
- A manually curated benchmark created to rigorously test the WAON dataset. It covers eight diverse cultural categories like food, tradition, and scenery. This benchmark helps measure how well a model understands specific aspects of Japanese culture compared to other datasets.
- Contrastive Vision-Language Models
- These are AI models that learn relationships between images and text by comparing them to find similarities. The study investigates whether using native cultural data improves these models' ability to grasp subtle cultural nuances in Japanese compared to using translated data.
- ReLAION (en→ja translation)
- This is a dataset created by translating English image-text pairs into Japanese. The study found that fine-tuning on this translated data actually harmed the model's understanding of Japanese culture, proving that translation alone is insufficient for cultural adaptation.
Terminology used across episodes
This episode discusses
- WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models · Paper Radio
- Gradio: Hassle-Free Sharing and Testing of ML Models in the Wild
- Meta CLIP 2: A Worldwide Scaling Recipe
- Gemma 2: Improving Open Language Models at a Practical Size
- Microsoft COCO: Common Objects in Context
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
The paper
WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models · Read on arXiv
Issa Sugiura, Shuhei Kurita, Yusuke Oda, Daisuke Kawahara, Yasuo Okabe, Naoaki Okazaki
Kyoto University · NII LLMC (National Institute of Informatics Laboratory) · Waseda University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models".
Tom: Contrastive vision-language models have achieved remarkable progress through largescale pretraining,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, looking at "WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models," the central thesis is testing if global pretraining alone is adequate for culture-specific understanding or if adding natively sourced data can improve performance on those specific cultural benchmarks. They introduce WAON, which they describe as the largest publicly available native Japanese image-text dataset constructed from native Japanese web content in Common Crawl, totaling approximately one hundred fifty-five million examples <ref:2510.22276#pg0,the largest publicly available native Japanese image-text dataset constructed from native>.
Jane: What really matters is their comparison between this native data and other types of data. The paper claims that fine-tuning on WAON consistently achieves stronger performance across all tested baselines when evaluating models on Japanese cultural benchmarks, whereas fine-tuning on English-to-Japanese translated data consistently degrades the model's understanding of Japanese culture.
Lu: That comparison is key because it directly addresses a limitation in prior work where translation was used as a substitute for native context, and this study provides concrete evidence against that approach for cultural tasks.
Meng: It sounds like the authors are making a strong statement about the necessity of cultural specificity over mere linguistic translation when dealing with complex cultural concepts in vision-language models. That's something we need to keep in mind when designing our own systems.
Lalam: For me, this confirms that for capturing subtle visual and contextual cues tied to Japanese culture, the direct source material is what truly matters, not a filtered or translated version of it.
Conclusion: Tom: So, wrapping up this discussion on "WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models," the authors, including Issa Sugiura and Shuhei Kurita, have demonstrated that data origin is a critical factor when it comes to adapting contrastive vision-language models for cultural tasks. Their findings suggest that while global pretraining sets a strong foundation, natively sourced Japanese data is essential for culture-specific fine-tuning because translation alone simply cannot compensate for the lack of original source material.
Jane: In simpler terms, what this means is that if you want an AI to really get the cultural nuances of something specific, like Japanese traditions or scenery, you need to train it on examples that are genuinely from that culture itself rather than relying on translated captions from other languages. The WAON dataset and their benchmark show this effect clearly.
Lu: The implication here is that future work should focus not just on scaling up the total amount of data, but crucially, on ensuring the quality and native sourcing of that data to achieve deep cultural understanding in AI systems.
Meng: From an engineering standpoint, this points toward a need for smarter data pipelines that can effectively filter and prioritize native content over translated content when training models for specialized domains. That's a practical direction we can take.
Lalam: I see the impact as improving the cultural accuracy of applications across various domains, making AI outputs feel much more authentic and less reliant on generalized, often Western-centric, knowledge bases.
Tom: Exactly! The WAON study shows that investing in culturally specific data curation leads to better performance on those very cultural benchmarks. That's a solid foundation for how we should approach building specialized vision models.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck