Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?

arXiv:2608.29921 · cs.CL, cs.AI · Submitted 2026-08-30 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-08-30

Updated: 2026-08-30

Project page: https://fractalego.github.io/Sleight-of-Word

Terminology

Sources

Related papers