Subword Segmental BabyLMs: Learning to Tokenise for Sample-Efficient Pretraining

arXiv:2609.01151 · cs.CL · Submitted 2026-09-01 · Read on arXiv

cs.CL

Submitted: 2026-09-01

Updated: 2026-09-01

Code: https://github.com/francois-meyer/subseg-babylms

Terminology

Sources

Related papers