What Should a Large Language Model See? Physical Invariants as a Data Representation for PDE Discovery

arXiv:2608.25189 · cs.LG, cond-mat.soft · Submitted 2026-08-25 · Read on arXiv

cs.LG, cond-mat.soft

Submitted: 2026-08-25

Updated: 2026-08-25

Comments: 6 pages, 1 figure

License: http://creativecommons.org/licenses/by/4.0/

The gist: Understanding how molecular interactions govern macroscopic behaviour is a central challenge in molecular sciences.

Terminology

Abstract

Understanding how molecular interactions govern macroscopic behaviour is a central challenge in molecular sciences. However, conventional theory building cannot keep pace with the vast datasets modern experimentation routinely produces. Large language models offer a promising route to automating theory construction, but a spatiotemporal field cannot be directly placed in a prompt. Existing models generally learn about the data only through a score measuring how well each proposal fits it. Here we introduce data interpretation, a stage that measures the field into the quantities a theorist would consult and supplies them to the model as a direct input. On a benchmark of simulated fields, interpretation nearly triples the accuracy of recovered equations relative to showing the raw data, at negligible computational cost and without any training. By allowing a language model to read field data as a theorist does, data interpretation offers a practical route to automated field theory construction that can coevolve with experimentation.

Sources

Related papers