Scaling Accessible Mathematics on arXiv: HTML Conversion and MathML 4
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Scaling Accessible Mathematics on arXiv".
Tom: We report on ongoing developments in arXiv’s HTML Papers offering, focusing on scaling accessible mathematics through HTML conversion and MathML 4 to better serve STEM readers.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, we're hearing that the core idea here is scaling accessible mathematics on arXiv through HTML conversion and MathML four. That means they are trying to take the LaTeX documents people upload and turn them into HTML so they can be read properly in a web browser.
Jane: Exactly, Tom; it’s about bridging the gap between how mathematicians write their work in LaTeX and how a general audience or an assistive technology can actually interpret those formulas correctly when viewed on the web.
Lu: The title highlights that they are not just converting text to HTML; they are specifically embedding MathML four which is a structured markup language designed for semantic understanding of math.
Meng: I see the challenge in that title; dealing with the "full diversity of STEM documents" mentioned in the paper means their pipeline has to handle a huge variety of mathematical symbols and formatting rules simultaneously.
Lalam: It’s about making sure that when someone reads a paper on physics or computer science, they aren't hitting a wall because the math notation is locked behind an image.
The paper's summary: Tom: The summary points out that their goal is to have a pipeline where once a LaTeX document gets processed successfully, the content stays intact in HTML so it’s usable by the web platform. They are measuring success by hitting ninety percent of arXiv submissions with error-free HTML.
Jane: That ninety percent mark is significant because it suggests that we could make LaTeX documents accessible via HTML at a large scale, which is something that hasn't been achieved before in this way.
Lu: The summary also details their specific technical contributions from two thousand twenty-five and early two thousand twenty-six including community improvements to fidelity and work on corpus-scale conversion aimed at achieving that error-free HTML goal.
Meng: I’m looking closely at the current status; they report that they are currently reaching roughly seventy-five percent success in converting articles without LaTeXML errors, which shows there's still a lot of room for improvement in the conversion accuracy.
Lalam: That seventy-five percent figure tells us that while progress is being made, we’re still working to make this accessible math a standard feature on arXiv rather than an occasional exception.
The paper's improvements: Tom: Beyond just the conversion goal, the paper outlines specific technical improvements they are implementing. They've introduced MathML four Intent annotations to handle novel notations better, especially for speech output in STEM research where standard tools often fail.
Jane: The intent system uses a compact expression syntax and standardizes three vocabularies—core concepts, open concepts, and core properties—to give the AI a much richer understanding of the math than just visual rendering allows.
Lu: They also mention that they are annotating formulae with a ":literal intent property" as their baseline for speech output to make it predictable for non-standard services.
Meng: The paper also mentions an in-progress Rust port of LaTeXML, which is a major move because the goal there is to reduce compute costs and speed up the preview experience on submissions by making the conversion process much faster.
Lalam: That Rust port effort seems really practical; if you can make it run ten to thirty times faster than a Perl reference, that immediately makes it more useful for authors who want instant feedback when they submit their work.
Conclusion: Tom: So, to wrap up, the paper on "Scaling Accessible Mathematics on arXiv: HTML Conversion and MathML four" shows concrete steps toward making LaTeX math accessible via HTML by focusing on fidelity improvements, corpus-scale conversion targets, and introducing MathML Intent for better accessibility.
Jane: It seems the long-term vision is to make born-accessible mathematical communication the default experience on arXiv rather than something authors have to work around.
Lu: The integration of TeX source as an annotation inside the MathML element is a smart way to allow non-standard services to interact directly with the math without needing complex external modifications, which opens up new possibilities for semantic analysis.
Meng: The effort toward a Rust port of LaTeXML is crucial because it addresses the computational cost barrier that currently prevents this from scaling efficiently across the whole corpus.
Lalam: This work on "Scaling Accessible Mathematics on arXiv: HTML Conversion and MathML four" suggests that the future of academic publishing isn't just about how we present information, but ensuring that complex mathematical ideas are inherently structured for broad use.
arXiv · National Institute of Standards and Technology
cs.CL, cs.DL
Submitted: 2026-05-15
Updated: 2026-05-15
Comments: 6 pages, ICMS 2026
Code: https://github.com/arXiv/html_feedback
License: http://creativecommons.org/publicdomain/zero/1.0/
Importance score: 91/100
The gist: We report on ongoing developments in arXiv’s HTML Papers offering, focusing on scaling accessible mathematics through HTML conversion and MathML 4 to better serve STEM readers.
Key concepts
- HTML with MathML
- This method replaces static math images in documents with structured MathML code within HTML. This allows browsers and assistive technologies to understand the mathematical relationships and structure of formulas, making them readable and navigable, unlike simple pictures.
- MathML Intent Annotations
- These are new attributes added to MathML 4 that help standard accessibility tools interpret novel or complex mathematical notations common in STEM research. They categorize formulas into core concepts, open concepts, and properties to provide more reliable speech output.
- LaTeXML Port in Rust
- This is an ongoing project to rewrite the LaTeXML system using the Rust programming language. The goal is to make the conversion process much faster—up to 30 times quicker than previous versions—to reduce computational costs and speed up how arXiv submissions are processed.
Terminology
Summary
We report on ongoing developments in arXiv’s HTML Papers offering, focusing on scaling accessible mathematics through HTML conversion and MathML 4 to better serve STEM readers. The main highlights from 2025 and early 2026 include community improvements to fidelity, corpus-scale conversion work targeting error-free HTML, initial MathML 4 Intent annotations for speech output, and an in-progress Rust port of LaTeXML to reduce compute costs.
The gist
Our HTML output carries MathML rather than formula images.
Background and Goals
arXiv is the world’s largest preprint server, covering mathematics, computer science, physics, and related STEM fields. The dominant distribution format is PDF, which preserves visual fidelity but offers limited structural information for reflow and assistive technologies. HTML with high-quality MathML exposes mathematical structure to the Web Platform to support responsive rendering in a browser, accessible readouts via screen readers, contextual navigation, and machine-readability. The goal of the conversion process is that once a LaTeX document is successfully processed by our system, the authored content survives intact to the final HTML.
Success is measured by reaching 90% of arXiv submissions receive successful HTML conversion,
which would signal that LaTeX-authored documents can be made accessible via HTML at scale.
MathML Intent and Accessibility
The paper utilizes MathML 4 Intent annotations to address the challenges of novel notations in STEM research, where default assistive technology (AT) readouts can be unreliable. The new intent attribute carries a compact expression syntax and standardizes three vocabularies: core concepts, open concepts, and core properties.
The current baseline focuses on syntax-first remediation by annotating formulae with the :literal intent property,
which is trivial to assign systematically and gives a predictable target for speech output.
Alongside this, well-structured MathML Core is emitted from the author’s TeX via LaTeXML’s math grammar. Furthermore, the MathML output carries the original TeX source as an attached annotation inside the element, allowing non-standard services to operate directly on the HTML page without further modification.
Deployment and Pipeline Evolution
The arXiv HTML Papers offering surfaced experimental renditions on every new submission for which LaTeXML produces output. The live pipeline is updated roughly quarterly, folding in kernel improvements from upstream LaTeXML at NIST alongside community enhancements and robustness patches. A key structural change involves an improved kernel for raw LaTeX 3 interpretation, where packages without LaTeXML bindings are loaded raw rather than skipped.
This work also includes sustained improvements in the vectorgraphics and math-rendering paths common in arXiv preprints, leading to broader coverage of TikZ and xy diagrams.
Scaling via Rust Port
To reduce compute costs at arXiv scale and improve the submitter preview experience through faster conversion, an in-progress Rust port of LaTeXML is being developed. This effort has involved roughly 50,000 lines of hand-written Rust covering the TeX-engine emulation and tokenization. Driven by agentic AI assistance (Claude Opus 4.6 and 4.7), the translation effort has advanced significantly, with the pipeline passing all core tests and showing benchmarks running 10–30 times faster than the Perl reference.
The goal is to achieve output parity
in an engineering sense: succeeding on the same inputs while preserving essential document structure for downstream accessibility and rendering.
Impact Signals and Corpus Coverage
The project prioritizes work using two complementary signals:
-
The public arXiv/html feedback issue tracker on GitHub, which captures acute errors reported by readers and authors, with half of approximately 6,000 reports resolved from the start of 2025 through April 2026.
-
Large-scale missing-package statistics gathered by running the LaTeXML pipeline across the historical ar5iv corpus, which covers roughly
90% of arXiv articles that have TeX/LaTeX source.
The meaningful metric for coverage is the fraction of articles that convert without LaTeXML errors and are typically mostly readable, which has slipped to roughly 75%.
The long-term outlook is to make born-accessible mathematical communication the default rather than the exception
on arXiv.
Author Contributions
DG led this work and is the sole developer of the Rust LaTeXML reimplementation. BRM contributes upstream LaTeXML and the v0.9 release path at NIST. BC, arXiv’s technical lead, supervised user-facing architecture and production deployments. JW and JS provide institutional leadership at arXiv, with JW additionally serving as a senior expert on the TeX/LaTeX use of arXiv’s authors.
References
-
arXiv HTML feedback tracker. https://github.com/arXiv/html feedback, accessed April 17, 2026
-
arXiv monthly submissions. https://arxiv.
Improvements for AI systems
Here are specific improvements to AI systems derived from the principles and findings of this paper, categorized by application:
) 1. Enhanced Mathematical Content Ingestion and Rendering Pipeline (For Research Assistants/Knowledge Graphs)
The core improvement lies in transitioning from PDF-centric data ingestion to a structured HTML/MathML representation. The improved system would:
-
Use a LaTeXML-based pipeline (or the proposed Rust port) to automatically convert incoming LaTeX submissions into high-fidelity HTML with embedded MathML 4.
-
Leverage the MathML 4
Intent
annotations (:literal, Core concepts, Open concepts) to enrich the parsed mathematical structure beyond mere visual rendering. -
Use this structured data for downstream AI tasks: an AI could query the
Open concepts
vocabulary to discover novel relationships between mathematical ideas present in disparate papers without relying solely on keyword matching or simple text extraction.
--- 2. Automated Accessibility and Assistive Technology (AT) Support Module
The system would move beyond basic text-to-speech by leveraging the MathML Intent structure:
-
The AI could generate highly contextualized speech output for mathematical expressions, using the defined underscore syntax for literals (
strings meant to be pronounced as-is
). -
For complex STEM documents, the system could dynamically select and apply appropriate MathML Core trees based on the document's
Core concepts
vocabulary (e.g., prioritizing standard notation for secondary education content). -
This allows AI agents to produce accessible summaries of technical papers that are guaranteed to correctly convey mathematical structure to screen readers, moving accessibility from a manual post-processing step to a native output feature.
--- 3. Scalable Error Detection and Corpus Quality Assurance System (For Automated Review/Moderation)
The system can be improved by integrating the signals mentioned in Section 4:
-
Implement a continuous monitoring layer that concurrently tracks two metrics: (a) user-reported conversion failures on individual articles (via GitHub feedback), and (b) large-scale missing package statistics across the historical corpus.
-
The AI would be trained to correlate specific LaTeX packages/macros with failure modes reported by users, allowing it to proactively flag submissions with high risk of accessibility failure before they are fully published.
-
This enables a preventative quality control mechanism that optimizes for both
reported friction
andpervasive errors,
leading to higher overall corpus quality than systems optimized for only one signal.
--- 4. Cost-Effective, High-Velocity Previews (For Submission Feedback Tools)
Utilizing the Rust reimplementation of LaTeXML, the system can dramatically improve developer workflow:
-
Implement a fast conversion service that runs on a Rust backend to generate HTML/MathML previews significantly faster than the Perl reference.
-
This allows arXiv submission systems to provide authors with near-instant feedback on how their LaTeX code will render in HTML, including an initial check for accessibility compliance (e.g.,
This expression is missing a defined literal intent
). -
This reduces compute costs associated with rendering and speeds up the author's iteration cycle, making the process more efficient for high-velocity submissions.
--- 5. Knowledge Graph Construction from TeX Source (For Deep Semantic Analysis)
The system can use the attached original TeX source alongside the MathML output:
-
The AI could build a knowledge graph where nodes represent mathematical concepts, and edges are defined by the structure extracted from both the MathML tree and the underlying TeX syntax.
-
This allows for
non-standard services that rely on TeX syntax
to operate directly on arXiv pages, enabling sophisticated semantic queries that go beyond simple text search (e.g.,Find all theorems where variable X is defined in Section 3 using a custom macro Y
).
Abstract
We report on the ongoing development of arXiv's HTML Papers offering, available on every new TeX/LaTeX submission since its initial release in 2023. The main highlights from 2025 and early 2026 are: (i) community-driven improvements to HTML fidelity and service health, with roughly half of 6,000 user reports resolved; (ii) corpus-scale conversion work aimed at 90% error-free HTML (currently 75%); (iii) initial MathML 4 Intent annotations for accessible speech output; (iv) an in-progress Rust port of LaTeXML, reducing compute costs and enabling faster previews on submission. The arXiv HTML Papers project remains experimental, but is gradually maturing as we better understand the needs of arXiv's readers and the technical opportunities presented by new standards and by advances in programming languages and AI.
Sources
- HTML papers on arXiv -- why it is important, and how we made it happen
- Accessibility for the Working Mathematician
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering