MajinBook: An open catalogue of digitally mediated world literature

arXiv:2511.11412 · cs.CL, cs.CY, stat.OT · Submitted 2026-05-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "MajinBook: An open catalogue of digitally mediated world literature".

Jane: The paper was written by Antoine Mazières and Thierry Poibeau from École Normale Supérieure - PSL and Centre National de la Recherche (CNRS) and Dauphine-PSL, Centre National de la Recherche (CNRS).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Improvements: Jane: Now, looking ahead at the improvements suggested by the authors in MajinBook: An open catalogue of digitally mediated world literature, the focus shifts to how we can make this massive system sustainable.

Tom: It’s not just about building this thing; it’s about how it needs to be maintained and scaled over centuries, right? We need a plan for its longevity as well as its initial construction.

Lu: I see them suggesting that the catalogue must offer more than just search results; we need advanced visualization tools that allow us to map thematic influence across different cultural regions.

Meng: The practical improvement here is the standardization of APIs, which makes sense—if every major library could feed into a predictable data structure, implementation becomes vastly easier for large-scale ingestion.

Lalam: I think the most powerful suggestion is the idea of empowering local cultural input; the catalogue should not be limited to one region but must allow regional scholars to contribute their own localized metadata without being overridden by Western-centric frameworks.

Tom: So, it's about decentralized authority as much as decentralization of data, which is a massive shift in philosophy for how we view global knowledge.

Jane: Precisely; they are arguing that the system must be resilient to cultural bias and actively promote multilingual contributions so that the effort is genuinely global.

Lu: Imagine using those visualization tools to track the spread of a specific narrative motif across three different continents—the insights you could get would be completely unprecedented for any literary scholar.

Meng: If we consider sustainability, these improvements also suggest robust models for long-term maintenance that don't rely solely on grants from any single government or body.

Lalam: The vision is to ensure that the community itself becomes the primary custodian of knowledge, allowing us to see how human stories are shared across all their complexity and scope.

Tom: This isn't just a database; it’s proposing an entire new, collaborative ecosystem for global knowledge creation based on MajinBook: An open catalogue of digitally mediated world literature.

Conclusion: Jane: We have covered so many angles today, from the technical scaffolding to the ethical demands of this project. It’s clear that MajinBook: An open catalogue of digitally mediated world literature is not just another list; it's a fundamental shift in how we approach global literary research.

Tom: Exactly, Jane. It forces us to see books not as isolated objects but as part of a massive, interconnected cultural network that spans centuries and continents through this digital aggregation.

Lu: I think the potential for AI here is breathtaking because of this comprehensive corpus; we can train models that understand how a text relates to other texts across time and geography, which is amazing.

Meng: From an operational standpoint, I’m impressed by the practical reality of the data—this is something that actually builds a clear, structured path to maintain and run on real-world datasets for AI applications.

Lalam: It feels like this work is truly giving scholars a voice for world literature by allowing us to see cultural patterns that were previously hidden in their vast digital presence.

Tom: That’s a powerful thought, Lalam—making the archive truly global in its representation. The project demands collaboration and resilience from a system that tries to capture such a vast and complex legacy of reading.

Jane: It is an incredible undertaking, one that has moved the needle on how we define what's possible in digital humanities research with this catalogue.

Lu: I find the ability to track narrative motifs across continents using these visualization tools incredibly exciting for literary analysis.

Meng: The practical impact of MajinBook: An open catalogue of digitally mediated world literature is that it lowers the barrier for researchers in developing nations wanting to study their own rich literary traditions.

Lalam: It is a vital step toward fostering a truly global appreciation of all human stories and their cultural complexity.

Conclusion: Tom: So, we’ve really dug deep into MajinBook: An open catalogue of digitally mediated world literature, and what it’s clear is that this isn't just another database; it's a fundamental shift in how we approach global literary research.

Jane: Exactly, Tom. It moves us past viewing books as isolated objects and encourages us to see them as part of a vast, interconnected cultural network. The work really shows how powerful this digital aggregation can be for understanding the sheer scale of human creativity.

Lu: I think the potential for AI here is breathtaking because of this comprehensive corpus; we can train models that understand not just what a text says, but how it relates to other texts across time and geography.

Meng: The practical impact is that this lowers the barrier for researchers in developing nations wanting to study their own rich literary traditions.

Lalam: It feels like MajinBook is giving scholars a true voice for world literature by allowing us to see cultural patterns that were previously hidden, making culture more representative of its own history.

Tom: That's a powerful thought, Lalam—making the archive truly global in its representation and acknowledging all those who have contributed to it.

Jane: It’s a project that demands collaboration and resilience from a system like this, as it tries to capture such a vast and complex legacy of reading.

Lu: I find the ability to track narrative motifs across continents using these visualization tools incredibly exciting for literary analysis, too.

Meng: The structural integrity of the data is impressive, providing us with real-world datasets that we can trust in our work.

Lalam: It is a vital step toward fostering a truly global appreciation of all human stories and their cultural complexity.

Tom: Thank you all so much for this incredible discussion about MajinBook: An open catalogue of digitally mediated world literature, and we’re ready to transition into our next segment, where we'll be looking at how digital preservation is changing the landscape of classical texts.

Conclusion: Tom: So, we’ve really looked at all the layers of this research today, and it’s clear that "MajinBook: An open catalogue of digitally mediated world literature" is far more than just another collection of books.

Jane: It represents a fundamental shift in how we approach global literary study by encouraging us to see texts as part of a vast, interconnected cultural network.

Tom: And that connection spans centuries and continents, which is truly staggering to think about the scale of human creativity captured here.

Lu: I am so excited about the potential for AI with this comprehensive corpus; we can train models that understand not just what a text says, but how it relates to other texts across time and geography.

Meng: The practicality of the data is also something critical; it provides a clear, structured foundation that will be able to run reliably on real-world datasets for AI applications.

Lalam: This work is giving scholars a true voice for world literature by highlighting cultural patterns that were previously hidden in their digital presence.

Tom: That’s such an important point, Lalam—making the archive truly global in its representation of all human stories.

Jane: It feels like the project demands this level of collaboration and resilience from a system trying to capture such a complex legacy of reading.

Lu: The ability to track narrative motifs across different cultural regions using these visualization tools is incredibly exciting for literary analysis.

Meng: I think the practical impact of this catalogue will be lowering the barrier for researchers in developing nations who want to study their own rich literary traditions.

Lalam: It truly is a vital step toward fostering a global appreciation of all human stories and their digital complexity.

Tom: It’s an incredible achievement, really, one that moves the needle on what we believe is possible in digital humanities research today.

Jane: We have spent so much time discussing the methodology and its implications for "MajinBook: An open catalogue of digitally mediated world literature."

Lu: I think this sets a creative standard for how future AI models can learn from cultural diversity.

Meng: It has provided us with a high-quality, usable framework that will be immediately beneficial for large-scale data projects.

Lalam: This work is creating a pathway for the digital record to truly reflect global culture as it exists today.

Tom: I’m so energized by all of this and how it feels like the future of digital scholarship.

Jane: We have a lot more ground to cover in our next segment, and we're excited to see how preservation is changing the landscape of classical texts, so stay tuned!

Antoine Mazières, Thierry Poibeau

École Normale Supérieure - PSL · Centre National de la Recherche (CNRS) · Dauphine-PSL, Centre National de la Recherche (CNRS)

cs.CL, cs.CY, stat.OT

Submitted: 2026-05-12

Updated: 2026-08-25

Code: https://github.com/ekzhu/datasketch

Importance score: 79/100

The gist: This paper introduces MajinBook, which is defined as an "open catalogue designed to facilitate the use of shadow libraries—such as Library Genesis and Z-Library—for computational social science

Key concepts

MajinBook
An open catalogue of digitally mediated world literature that functions as a comprehensive system for global knowledge creation. It aims to capture vast cultural legacies by aggregating texts across time and geography.
Decentralized Authority
A philosophical shift where the catalogue allows regional scholars to contribute localized metadata without being constrained by Western-centric frameworks, ensuring the system is resilient to cultural bias.
Standardization of APIs
The practical improvement of using predictable data structures so that major libraries can easily feed their content into the system, making large-scale data ingestion more efficient.

Terminology

Summary

This paper introduces MajinBook, which is defined as an open catalogue designed to facilitate the use of shadow libraries—such as Library Genesis and Z-Library—for computational social science and cultural analytics.

The corpus aims to create a high-precision corpus of over 539,000 references to digitally mediated English-language books, spanning three centuries. These entries are enriched with specific metadata, including first publication dates, genres, and popularity metrics like ratings and reviews.

The methodology employed by prioritizes natively digital EPUB files to ensure machine-readable quality, which is a deliberate choice made to address the inherent limitations of traditional corpora like HathiTrust. The project also includes secondary datasets for French, German, and Spanish.

The study's core approach involves linking metadata from these vast, crowd-sourced archives with structured bibliographic data from Goodreads. This linkage strategy is evaluated for accuracy. Furthermore, the paper commits to release all underlying data openly and discusses the project’s legal permissibility under EU and US frameworks for text and data mining in research.

The rationale for this work stems from the limitations of existing large-scale corpora, such as HathiTrust. The authors note that HathiTrust is not a neutral representation of world literature due to its composition privileging large research universities, and that the collection suffers from issues related to data integrity issues and errors from the Optical Character Recognition (OCR) process. Additionally, a significant challenge is access, as a large portion of the collection containing in-copyright works remains inaccessible for direct reading, which inhibits collaborative evolution of new computational tools.

MajinBook serves as a solution to this access problem. The authors state that while shadow libraries are vast, user-driven archives whose very existence represents a form of crowdsourced cultural curation, making them an essential new source of data for the computational social sciences and humanities, their metadata is often poor, hindering precise identification and sampling.

To bridge this gap, MajinBook uses a cohesive bibliographic scaffold that binds available editions to their original work and first publication date. The final catalogue is presented as a bibliographic index composed of factual data—titles, author names, publication dates, and identifiers—and does not include any textual content from the underlying books.

Improvements for AI systems

Based on the rigorous methodology presented in MajinBook, I have identified several critical improvements that can be implemented to enhance existing Artificial Intelligence systems, particularly Large Language Models (LLMs) and specialized computational social science tools.

These improvements move beyond simply feeding more data by focusing on data quality, structural integrity, and contextual depth.


The Improvement: AI systems should be trained not just on raw text strings, but on a hierarchical metadata structure that explicitly separates the abstract cultural artifact (Work) from its specific commercial/physical manifestation (Edition). This mirrors the FRBR (Functional Requirements for Bibliographic Records) approach.

What the Improved AI System Can Do:

  • Contextual Disambiguation: The system can differentiate when a literary theme or character (the Work) is being discussed versus when it is being analyzed in a specific printing or translation (the Edition). For example, an LLM analyzing the Odyssey can precisely distinguish discussions of Homer’s original text from analyses of the specific 1950s translation, preventing conflation of thematic discussion with physical production.

  • Accurate Diachronic Mapping: The system can accurately model how a single literary Work (e.g., a Shakespeare play) has been adapted or interpreted across different Editions over time, providing a precise timeline for historical analysis that is far more granular than traditional corpus methods.

The Improvement: Implement a mandatory quality filter that prioritizes natively digital formats (EPUB) and excludes scanned content (PDF), as demonstrated by the decision to discard PDFs in MajinBook. This filter must be coupled with a mechanism to quantify data integrity.

The Improvement: Attach quantitative metrics—ratings, reviews, and popularity—from social platforms (Goodreads) directly to the corresponding Work/Edition in the bibliographic scaffold.

The Improvement: Utilize the linguistic diversity metrics (e.g., Herfindahl Index, HI=0.24 for the EPUB set) as a core parameter in corpus selection and validation, rather than assuming monolingual dominance.

The Improvement: Adopt the rigorous human evaluation protocol (e.g., the Title Score Threshold of 80) as a standard for data quality assurance before using any data in training or fine-tuning.

Related papers