Text-Preserving Lossy Text Compression: A Study of Strategic Deletion and LLM Reconstruction

arXiv:2605.29000 · cs.CL · Submitted 2026-05-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Text-Preserving Lossy Text Compression".

Jane: Lossy semantic text compression, where an encoder strategically deletes parts of a source text and a large language model (LLM) reconstructs it from the retained skeleton,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, Jane, let's start by getting a handle on this paper called "Text-Preserving Lossy Text Compression: A Study of Strategic Deletion and LLM Reconstruction." Basically, the main idea is that instead of just throwing away text randomly, we strategically delete parts of it so a large language model can piece it back together.

Jane: That sounds really interesting, Tom; so the core thesis here is that we can lose some text while still keep a usable structure for reconstruction by an AI model. It's about finding these artifacts that are useful both before and after the reconstruction process, which is different from just summarizing or rewriting the source material.

Lu: From a creative standpoint, this opens up so many possibilities; imagine we could create compressed versions of documents that retain enough context for downstream tasks without needing to store the full text. This isn't just about saving space; it's about creating smarter inputs for other AI systems.

Meng: I'm interested in how practical this actually is; if we’re talking about pipelines, what does this lossy compression ratio mean in terms of actual data throughput or latency?

Lalam: The LLM perspective here is huge; if we can feed a model a highly structured, compressed skeleton, it means the model doesn't have to process the entire source text every time. This could drastically improve efficiency across many applications.

Tom: Exactly what I mean, Meng; the paper sets up this two-phase encode-decode problem where we get a degraded representation T and then a decoder reconstructs it to get the original text. It claims that with this approach, the effective character-level compression ratio is about one over rkeep, and we only need to store that representation and some lightweight strategy metadata.

Jane: That sounds like a pretty clean framework; so the encoder's job is to strategically delete components from the text based on some intelligent rules, and then a decoder, which is essentially an LLM itself, does the reconstruction from just that skeleton. It’s about finding those right tokens to keep or discard.

Paper summary: Lu: The paper explores a progression of deletion strategies starting from simple things like uniform step deletion up through more complex semantic levels using neural language models to estimate context predictability and delete the lowest-surprisal tokens first. That's a really deep dive into linguistic structure, isn't it?

Meng: It’s cool that they benchmark everything from basic character deletion to hybrid methods combining frequency and entropy signals; I wonder if there are any practical trade-offs in terms of how hard the encoder has to work versus how good the final reconstruction is.

Lalam: The idea of using GPT-two surprisal for entropy-based deletion is fascinating because it lets the deletion strategy become adaptive based on what makes a token contextually surprising, which could lead to much more nuanced compression than just looking at word counts or character positions.

Tom: And the results they found are pretty telling; they tested various retention rates rkeep between zero point one and zero point nine using BERTScore as their main metric against ROUGE-L and CER, which gave them three main findings about the different deletion schemes they tried.

Jane: I read that WordFreq came out as a strong low-cost baseline because it’s fast at the encoder while still being competitive with more expensive semantic methods, even though it only looks at static word frequencies.

Lu: That's an important finding because it suggests that not every complex semantic approach is necessary if you need something quick and efficient; WordFreq serves as a solid starting point for understanding the performance floor of this compression method.

Meng: So, the implication there is that for scenarios where speed matters most, sticking to frequency-based methods might be a sensible choice over immediately jumping to the more computationally intensive neural language model based strategies.

Lalam: And it connects back to how we might use this in culture; if we can have faster, slightly less perfect reconstructions for common text, it could make AI tools much more accessible and usable by a broader audience.

Paper summary: Tom: Moving on to the conclusion of the paper's study, "Text-Preserving Lossy Text Compression: A Study of Strategic Deletion and LLM Reconstruction," what do you guys think about what this work actually means for where we are now?

Jane: I think it really simplifies the idea by focusing on reusable lossy text artifacts rather than just trying to get immediate task performance from a compressed prompt. The goal is to preserve enough meaning, factual anchors, and textual structure so that later recovery is still useful.

Lu: That focus on "reusable lossy text artifacts for later recovery" shifts the goal away from perfect fidelity in every single instance toward creating data that has inherent structural integrity for future use by different AI components.

Meng: From an engineering standpoint, it means we aren't aiming for lossless perfection when bandwidth is tight; instead, we aim for a specific level of meaningful corruption that the decoder can handle reliably.

Lalam: For culture and information dissemination, this suggests a future where complex information can be distributed in smaller packages without losing its essential narrative threads or factual backbone, which could democratize access to detailed content.

Tom: It really boils down to finding a sweet spot where the compression is aggressive enough for bandwidth constraints but not so aggressive that the decoder completely fails to reconstruct anything coherent at all.

Jane: So, in simple terms, this paper shows how we can use an AI model as a sophisticated editor that intelligently prunes text to create a smaller version while ensuring the core message stays intact for later retrieval.

Lu: It’s about designing deletion rules that respect the linguistic structure of the language rather than just treating every character or word the same way.

Meng: That structural awareness is what makes it useful; if we can build better rules, we can build more robust systems for processing large amounts of text efficiently.

Lalam: And as a model, I see this as improving how AI interacts with culture; by making text inherently more resilient to lossy transmission, the underlying knowledge becomes easier to manage and utilize across different platforms.

Conclusion: Tom: So we've been talking about this paper that looks at using strategic deletion to compress text so an AI can rebuild it, and now we're getting to the wrap-up on "Text-Preserving Lossy Text Compression: A Study of Strategic Deletion and LLM Reconstruction."

Jane: That study essentially explores how we can strategically prune parts of a source text, letting a model reconstruct it later without having to store the whole thing. It’s about finding that sweet spot between losing data and keeping enough structure for the AI to function well.

Lu: I think the core idea here is really about designing intelligent pruning rules that respect how language actually works, moving beyond simple character removal to something more nuanced. The authors are mapping out a spectrum of deletion techniques based on linguistic complexity.

Meng: From an engineering standpoint, it’s fascinating how they organized those strategies into these levels—from basic word frequency stuff up to using the model itself to estimate context predictability. It makes the whole process very systematic for building pipelines.

Lalam: For me, I see this as a way to create more resilient information streams; instead of just storing raw data, we could be storing intelligently structured representations that are inherently better at surviving transmission noise or bandwidth limits. That’s a huge cultural shift for how we handle digital knowledge.

Tom: Exactly! And the authors show that whether you use a simple frequency method or something involving the language model's internal understanding of surprise, you get different trade-offs in terms of speed and reconstruction quality.

Jane: It really highlights that there isn't one single best way to compress text; it depends entirely on what you need from the final reconstructed output, whether it’s high fidelity or just enough meaning to be useful.

Lu: That dependence on the specific data domain is something I found particularly interesting during my review of their cross-domain results; they showed that even with a universal framework, the best deletion rule changes depending on whether you're looking at news or Chinese text.

Meng: That points toward a major practical implication for building scalable AI systems; we can’t use one-size-fits-all compression rules if we want it to perform well across different types of content.

Lalam: And that means the future of information processing isn't just about making models bigger, but about making the input data itself smarter and more adaptable to whatever constraints we have on hardware or bandwidth.

Tom: It really boils down to finding a flexible system that adapts its pruning strategy based on the specific text it’s handling, which opens up some really cool avenues for next-generation AI applications.

CUNY Graduate Center

cs.CL

Submitted: 2026-05-27

Updated: 2026-10-01

Comments: Accepted at AACL-IJCNLP 2026 (Main Conference)

Code: https://github.com/hankcs/HanLP

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: Lossy semantic text compression, where an encoder strategically deletes parts of a source text and a large language model (LLM) reconstructs it from the retained skeleton, is studied to find

Key concepts

Encoder-Decoder Framework
The process is split into two parts: an encoder that strategically deletes text, creating a degraded representation, and a decoder (an LLM) that reconstructs the original text from only the deleted skeleton. This allows for compression by storing only the small skeleton and metadata.
Structured Deletion Strategies
Deletion methods are organized into three levels of complexity: Level 1 is simple character deletion; Level 2 focuses on removing predictable words based on frequency statistics; and Level 3 uses a neural model to delete tokens based on contextual predictability (surprisal).
Semantic Reconstruction
The decoder acts as a generative LLM, tasked with taking the sparse, degraded input and generating a full, fluent version of the original text. This reconstruction must preserve meaning without introducing new facts. Fine-tuning smaller models can make this reconstruction highly efficient.
Retention Rate (rkeep)
This parameter defines how much of the original text skeleton is kept during encoding, ranging from 0.1 to 0.9. The study finds that the best compression strategy depends on this rate; lower rates test the decoder's ability to recover meaning and structure.

Terminology

Summary

Lossy semantic text compression, where an encoder strategically deletes parts of a source text and a large language model (LLM) reconstructs it from the retained skeleton, is studied to find practical strategies for context- and bandwidth-limited processing in LLM pipelines. This approach is significant because it addresses the need for textual artifacts that remain useful as text before or after reconstruction, unlike abstractive summarization which rewrites the source.

How it works

The framework is formulated as a two-phase encode-decode problem. The encoder produces a degraded representation T˜ of length rkeep · L by strategically deleting components from T, and the decoder Dϕ reconstructs the original text Tˆ = Dϕ(T˜) from this degraded input alone. The effective character-level compression ratio is 1/rkeep, and only T˜ and lightweight strategy metadata need to be stored or transmitted.

Encoder: Structured Degradation Strategies

The encoder strategies are organized into three levels of linguistic sophistication. Level 1 involves Character-Level Deletion, establishing a baseline with fixed-step character deletion, which is noted as a strict lower-bound benchmark because it is agnostic to linguistic structure. Level 2 focuses on word boundaries, introducing methods like Adaptive Small-Word Removal (WordLen), and frequency statistics, such as the Frequency-Based Static Deletion (WordFreq), which prioritizes removing highly predictable words based on Zipf scores.

Level 3: Semantic-Level Deletion

Higher levels utilize a neural language model to estimate contextual predictability. This includes Entropy-Based Deletion, which computes GPT-2 surprisal and deletes the lowest-surprisal tokens first. More complex methods include Hybrid Frequency–Entropy Deletion, which combines frequency and surprisal signals, such as Hybrid-α that interpolates between WordFreq-like and entropy-like behavior.

Decoder: Semantic Reconstruction

The decoder acts as a generative LLM, reconstructing the skeleton into a fuller version of the source. The primary zero-shot decoder utilized is Gemini 2.0 Flash, prompted to reconstruct a natural, fluent sentence that preserves the original meaning without hallucinating new facts. For compute-efficient alternatives, Strategy-Aware Fine-Tuning (SFT) of models like Llama-3.2-3BInstruct using QLoRA is introduced to make a compact local decoder highly competitive with a stronger zero-shot proprietary decoder under the same reconstruction setting.

Evaluation and Findings

The paper evaluates strategies across various retention rates, including rkeep ∈ [0.1, 0.9], using BERTScore as the primary metric alongside ROUGE-L and CER. Key findings include:

  1. WordFreq is a strong low-cost baseline, remaining competitive with semantic methods while being far faster at the encoder.

  2. Semantic and hybrid methods provide their gains at mild-to-moderate compression, whereas WordFreq is more robust at the lowest retention rates.

  3. Strategy-aware fine-tuning makes a compact local decoder competitive, especially in settings where the skeleton is sparse (low retention).

  4. Cross-domain experiments show that while qualitative trends transfer, the best deletion rule remains dataset-dependent. For instance, Entityrich news text rewards methods that preserve factual anchors.

  5. The framework is most relevant when the target retention rate falls meaningfully below the range where exact recovery is still realistic, shifting the objective to preserving enough meaning, factual anchors, and textual structure for useful downstream recovery.

Limitations and Context

The study acknowledges limitations such as BERTScore being too forgiving at extreme compression, where it can reward genre-consistent generations even when factual anchors have disappeared. Furthermore, the framework is not intended for high-stakes domains like legal or medical settings where exact fidelity is essential. The latency introduced by the LLM decompression step (1–3 s per chunk) makes it unsuitable for real-time access, reinforcing that the lowest-retention regime should be read as a stress test of decoder failure rather than as a generally practical operating point. Lossless codecs are viewed as complementary to this approach, applicable in a cascaded pipeline.

Cross-Domain Generalizability

The framework was tested on English news (BBC News), Wikipedia (Salesforce/wikitext), and Chinese datasets (official news, Zhihu). Results indicate that the overall framework transfers across domains, but specific deletion rules are domain-dependent. For example, in Chinese Wikipedia, Semantic bucket allocation helps most at mild compression, while in Zhihu, Qwen SFT is the strongest method from rkeep = 0.9 through 0.3. This suggests that the effectiveness of a strategy depends on the discourse structure and lexical repetition of the target domain.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on this research, along with what those improved systems could achieve:


AI Systems Improvement Recommendations:

  1. mathbfDevelop Text-Preserving Lossy Compression Pipelines for Contextual Caching and Transmission:

Based on the framework (strategic token deletion + LLM reconstruction), implement a system that compresses long documents or conversation histories into a smaller, text-based skeleton.

  • How: Use the frequency-based deletion (WordFreq) or hybrid methods to create skeletons at specific retention rates (e.g., 50% retention).

  • What it achieves: Enables efficient context caching in agent memory, reduces bandwidth requirements for relaying large texts between services, and allows for lossy compression when exact byte-for-byte fidelity is not required but semantic coherence is paramount.

  1. mathbfImplement Strategy-Aware Local Decoders (QLoRA Fine-Tuning):

Integrate the finding that strategy-aware fine-tuning of compact models (like Llama 3B) makes them competitive with larger zero-shot models (like Gemini 2.0 Flash).

  • How: Fine-tune a small, local decoder model on various degradation patterns derived from different deletion strategies (e.g., training it specifically to recover text corrupted by WordFreq deletions).

  • What it achieves: Creates lightweight, low-latency AI inference endpoints capable of high reconstruction fidelity for specific types of compressed inputs, enabling deployment in resource-constrained environments where a full proprietary API call is too slow or expensive.

  1. mathbfCreate Adaptive Deletion Heuristics for Domain Specificity:

Replace naive deletion with context-aware strategies that adapt based on the text domain and compression level.

  • How: Implement a system that dynamically selects the best deletion rule (e.g., WordFreq for low retention, semantic/hybrid methods for moderate retention) based on preliminary analysis of the input text's characteristics (e.g., entity density vs. conversational register).

  • What it achieves: Maximizes reconstruction quality across diverse datasets like news, Wikipedia, and Reddit by ensuring the chosen deletion strategy aligns with the specific linguistic structures (e.g., preserving technical terms in encyclopedic text vs. conversational flow in Reddit comments).

  1. mathbfEnhance Factual Anchor Preservation for High-Stakes Text:

In critical applications (legal/medical), move beyond BERTScore and integrate named-entity preservation metrics into the compression pipeline's evaluation and selection process.

  • How: Use the NER analysis results to prioritize deletion rules that specifically protect factual anchors, even if they slightly reduce overall semantic similarity scores.

  • What it achieves: Mitigates the risk of factual drift or hallucination in compressed outputs for sensitive domains by ensuring that critical entities (names, dates, numbers) are retained as robust reconstruction anchors.

  1. mathbfDevelop a Cascaded Compression Pipeline:

Design a multi-stage compression workflow leveraging both lossy semantic skeleton creation and lossless byte-level reduction.

  • How: Apply the LLM reconstruction step first to create the semantic skeleton (retaining meaning), and then apply traditional lossless codecs (zlib/LZMA) to that skeleton.

  • What it achieves: Achieves a higher overall compression ratio than purely lossy methods while maintaining semantic coherence, providing a best of both worlds approach for archival storage where downstream consumption is human reading or semantic processing.

  1. mathbfOptimize Encoder Latency vs. Quality Trade-off:

Systematically tune the trade-off between encoder computation (e.g., GPT-2 surprisal inference) and reconstruction quality based on the required operating regime (mild vs. aggressive compression).

  • How: Use the latency data provided in Table 4 to define operational thresholds; for low-latency needs, favor simpler baselines like WordFreq or Step deletion, while reserving complex semantic methods only when higher fidelity is demanded.

  • What it achieves: Provides a computationally efficient path for real-time text processing by allowing engineers to select the simplest viable compression/reconstruction method based on real-time constraints.

Abstract

Traditional lossless text compression preserves every byte, but its gains on natural language are often modest in realistic operating regimes. We study lossy semantic text compression, where the encoder strategically deletes parts of the text and a large language model (LLM) reconstructs the original content from the retained skeleton. We benchmark a progression of deletion strategies, including uniform step deletion, word-length-guided deletion (WordLen), word-frequency-guided deletion (WordFreq), LP-optimized deletion (Opt), entropy-based deletion using GPT-2 surprisal, and hybrid methods that combine frequency and surprisal signals. Evaluation on the BBC News dataset across retention rates keep in [0.1,0.9] shows three main findings. First, WordFreq is a strong low-cost baseline: despite using only a static frequency lookup, it remains competitive with much more expensive semantic methods while being far faster at the encoder. Second, semantic and hybrid methods provide their clearest gains at mild-to-moderate compression, whereas word-frequency deletion is often more robust at the lowest retention rates. Third, QLoRA fine-tuning yields a strong local decoder that is competitive with Gemini 2.0 Flash and is often strongest in decoder-only comparisons. Additional English and Chinese experiments show that the overall framework transfers across domains, while the best deletion rule remains dataset-dependent.

Sources

Related papers