All You Need Is Non-Commutative Words

arXiv:2608.29314 · cs.CL · Submitted 2026-08-29 · Read on arXiv

cs.CL

Submitted: 2026-08-29

Updated: 2026-08-29

Comments: 11 pages, 3 figures

License: http://creativecommons.org/licenses/by/4.0/

The gist: We represent lexical tokens as unitary matrices and encode each sentence as their ordered product.

Terminology

Abstract

We represent lexical tokens as unitary matrices and encode each sentence as their ordered product. The noncommutativity of matrix product captures word order without positional encodings (PEs). The same algebra yields several capabilities, including antisymmetric self-attention with no query, key, or value projections, and parallel composition of variable-length text chunks at a reduced attention cost. Furthermore, it provides a canonical-coset readout layer that encodes all true unitary degrees of freedom compactly, while supporting continual learning through nested group extensions that enlarge the operator space with each new task preserving prior representations exactly. Across standard text-classification benchmarks, the method matches or exceeds bag-of-words baselines. Achieving higher accuracy on IMDB and comparable performance on AG News. Notably, this is accomplished by replacing the conventional about 30,000-dimensional vocabulary space with a dense, 64-parameter real-valued encoding, highlighting the expressive efficiency of our parameterization.

Related papers