.tmu: A Low-Entropy Tree-Structured Representation for LLM-Assisted Scientific Writing
cs.CL
Submitted: 2026-03-03
Updated: 2026-09-08
Comments: 24 pages, 8 figures. To be published in EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: As large language models (LLMs) increasingly assist scientific writing, the limitations and token costs of generating TeX become increasingly visible.
Terminology
Abstract
As large language models (LLMs) increasingly assist scientific writing, the limitations and token costs of generating TeX become increasingly visible. This paper analyzes TeX's architectural mismatch with LLM workflows, stemming from its lack of an explicit structural representation, to illustrate its limitations on generated semantics and error localization. As an alternative, we introduce.tmu, a low-entropy tree-structured representation. With its efficient data structure and clear contextual boundaries,.tmu outperforms.tex in the above aspects. Experiments across four LLMs provide evidence for this claim in most evaluated settings. Furthermore, we show that due to its lower information entropy, fine-tuning LLMs on.tmu achieves approximately 43% lower final training loss than on.tex. Our work provides a more scalable and LLM-friendly data representation for LLM-assisted scientific writing.
Sources
- Volume estimates for unions of convex sets, and the Kakeya set conjecture in three dimensions
- DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text
- Towards Accurate and Efficient Document Analytics with Large Language Models
- NeuRaLaTeX: A machine learning library written in pure LaTeX
- MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering