Scaling Accessible Mathematics on arXiv: HTML Conversion and MathML 4

summary

Video file (mp4)

The gist

We report on ongoing developments in arXiv’s HTML Papers offering, focusing on scaling accessible mathematics through HTML conversion and MathML 4 to better serve STEM readers.

In short

arXiv is converting LaTeX documents into accessible HTML using MathML 4 instead of images to improve structure for web platforms and screen readers. Efforts focus on achieving high conversion success rates, handling complex notations with new MathML intent annotations, and scaling the process via a faster Rust port of the LaTeXML system.

Key concepts

HTML with MathML
This method replaces static math images in documents with structured MathML code within HTML. This allows browsers and assistive technologies to understand the mathematical relationships and structure of formulas, making them readable and navigable, unlike simple pictures.
MathML Intent Annotations
These are new attributes added to MathML 4 that help standard accessibility tools interpret novel or complex mathematical notations common in STEM research. They categorize formulas into core concepts, open concepts, and properties to provide more reliable speech output.
LaTeXML Port in Rust
This is an ongoing project to rewrite the LaTeXML system using the Rust programming language. The goal is to make the conversion process much faster—up to 30 times quicker than previous versions—to reduce computational costs and speed up how arXiv submissions are processed.

Terminology used across episodes

This episode discusses

The paper

Scaling Accessible Mathematics on arXiv: HTML Conversion and MathML 4 · Read on arXiv

arXiv · National Institute of Standards and Technology

We report on the ongoing development of arXiv's HTML Papers offering, available on every new TeX/LaTeX submission since its initial release in 2023. The main highlights from 2025 and early 2026 are: (i) community-driven improvements to HTML fidelity and service health, with roughly half of 6,000 user reports resolved; (ii) corpus-scale conversion work aimed at 90% error-free HTML (currently 75%); (iii) initial MathML 4 Intent annotations for accessible speech output; (iv) an in-progress Rust port of LaTeXML, reducing compute costs and enabling faster previews on submission. The arXiv HTML Papers project remains experimental, but is gradually maturing as we better understand the needs of arXiv's readers and the technical opportunities presented by new standards and by advances in programming languages and AI.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Scaling Accessible Mathematics on arXiv".

Tom: We report on ongoing developments in arXiv’s HTML Papers offering, focusing on scaling accessible mathematics through HTML conversion and MathML 4 to better serve STEM readers.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, we're hearing that the core idea here is scaling accessible mathematics on arXiv through HTML conversion and MathML four. That means they are trying to take the LaTeX documents people upload and turn them into HTML so they can be read properly in a web browser.

Jane: Exactly, Tom; it’s about bridging the gap between how mathematicians write their work in LaTeX and how a general audience or an assistive technology can actually interpret those formulas correctly when viewed on the web.

Lu: The title highlights that they are not just converting text to HTML; they are specifically embedding MathML four which is a structured markup language designed for semantic understanding of math.

Meng: I see the challenge in that title; dealing with the "full diversity of STEM documents" mentioned in the paper means their pipeline has to handle a huge variety of mathematical symbols and formatting rules simultaneously.

Lalam: It’s about making sure that when someone reads a paper on physics or computer science, they aren't hitting a wall because the math notation is locked behind an image.

The paper's summary: Tom: The summary points out that their goal is to have a pipeline where once a LaTeX document gets processed successfully, the content stays intact in HTML so it’s usable by the web platform. They are measuring success by hitting ninety percent of arXiv submissions with error-free HTML.

Jane: That ninety percent mark is significant because it suggests that we could make LaTeX documents accessible via HTML at a large scale, which is something that hasn't been achieved before in this way.

Lu: The summary also details their specific technical contributions from two thousand twenty-five and early two thousand twenty-six including community improvements to fidelity and work on corpus-scale conversion aimed at achieving that error-free HTML goal.

Meng: I’m looking closely at the current status; they report that they are currently reaching roughly seventy-five percent success in converting articles without LaTeXML errors, which shows there's still a lot of room for improvement in the conversion accuracy.

Lalam: That seventy-five percent figure tells us that while progress is being made, we’re still working to make this accessible math a standard feature on arXiv rather than an occasional exception.

The paper's improvements: Tom: Beyond just the conversion goal, the paper outlines specific technical improvements they are implementing. They've introduced MathML four Intent annotations to handle novel notations better, especially for speech output in STEM research where standard tools often fail.

Jane: The intent system uses a compact expression syntax and standardizes three vocabularies—core concepts, open concepts, and core properties—to give the AI a much richer understanding of the math than just visual rendering allows.

Lu: They also mention that they are annotating formulae with a ":literal intent property" as their baseline for speech output to make it predictable for non-standard services.

Meng: The paper also mentions an in-progress Rust port of LaTeXML, which is a major move because the goal there is to reduce compute costs and speed up the preview experience on submissions by making the conversion process much faster.

Lalam: That Rust port effort seems really practical; if you can make it run ten to thirty times faster than a Perl reference, that immediately makes it more useful for authors who want instant feedback when they submit their work.

Conclusion: Tom: So, to wrap up, the paper on "Scaling Accessible Mathematics on arXiv: HTML Conversion and MathML four" shows concrete steps toward making LaTeX math accessible via HTML by focusing on fidelity improvements, corpus-scale conversion targets, and introducing MathML Intent for better accessibility.

Jane: It seems the long-term vision is to make born-accessible mathematical communication the default experience on arXiv rather than something authors have to work around.

Lu: The integration of TeX source as an annotation inside the MathML element is a smart way to allow non-standard services to interact directly with the math without needing complex external modifications, which opens up new possibilities for semantic analysis.

Meng: The effort toward a Rust port of LaTeXML is crucial because it addresses the computational cost barrier that currently prevents this from scaling efficiently across the whole corpus.

Lalam: This work on "Scaling Accessible Mathematics on arXiv: HTML Conversion and MathML four" suggests that the future of academic publishing isn't just about how we present information, but ensuring that complex mathematical ideas are inherently structured for broad use.

More episodes

← Home