Machine-Interpretable Information: Compiling Documents into Searchable and Readable Protocol States
cs.CL, cs.LG
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 21 pages, 5 figures
Code: https://github.com/gkamradt/LLMTest_NeedleInAHaystack
License: http://creativecommons.org/licenses/by/4.0/
The gist: Long-context language models interface with external knowledge through raw natural language.
Terminology
Abstract
Long-context language models interface with external knowledge through raw natural language. In retrieval-augmented systems, this creates a persistent index-payload schism: dense vectors enable searchable routing, but models must re-ingest lengthy text payloads for reasoning at O(N 2) attention cost. Existing compression methods further produce private states tied to specific architectures. We introduce Machine-Interpretable Information (MII), the first agent-to-agent (A2A) document-to-state protocol. A dual-timescale state-space Writer compiles documents into a canonical, fixed-bandwidth state (56 tokens), and a lightweight Translator maps it into any frozen Reader's embedding space, reducing query-time cost to O(K). The resulting.mii artifact unifies Retrieval (searchable geometry), Reasoning (global memory), and Reconstruction (grounded details) in a single transferable medium. We demonstrate strong cross-model interoperability across heterogeneous LLMs (e.g., Llama, Qwen, Mistral) -- despite the Writer using a legacy GPT-2 vocabulary, forcing genuine semantic translation rather than token-level memorization. Mechanistic probes reveal modular latent structure: entity representations can be causally traced and zero-shot transplanted between unrelated document states while remaining decodable. To address lexical reconstruction under fixed bandwidth, we propose Residual-MII, a cache hierarchy combining compiled global memory with sparse local evidence. On HotpotQA (7,405 queries), Residual-MII exceeds full-context Exact Match at approximately 7% of the attention FLOPs, suggesting a paradigm shift toward compiled, transferable neural document formats.
Sources
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Model-Document Protocol for AI Search
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- PICASO: Permutation-Invariant Context Composition with State Space Models
- MemGPT: Towards LLMs as Operating Systems
- C-Pack: Packed Resources For General Chinese Embeddings
- The Llama 3 Herd of Models
- Extending Context Window of Large Language Models via Positional Interpolation
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
- xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token
- Latent Context Compilation: Distilling Long Context into Compact Portable Memory
- MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
- Longformer: The Long-Document Transformer
- L-Eval: Instituting Standardized Evaluation for Long Context Language Models
- SnapKV: LLM Knows What You are Looking for Before Generation
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering