LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
summary
The gist
LensVLM introduces an inference framework and post-training recipe that enables Vision Language Models (VLMs) to scan compressed images and selectively expand only the relevant images to their
In short
LensVLM introduces a framework allowing Vision Language Models to scan compressed images and selectively expand only relevant parts to their uncompressed form using learned tools. This method maintains high accuracy even at extreme compression by treating text as an image and training the model to decide which visual information is necessary for correct reasoning.
Key concepts
- Deterministic Renderer
- This component converts raw text into a sequence of images based on specific settings like font and spacing. It acts as a controlled way to simulate different levels of image compression, allowing the model to test how much detail it can still extract from the visual representation.
- Expand Tool
- A learned tool that the VLM uses during inference to recover selected images from their compressed state into their full, uncompressed form. The model learns when and where to use this tool based on its reasoning needs, such as needing more detail for a specific part of the text.
- Information-Theoretic Tradeoff
- A mathematical analysis quantifying the balance between reading compressed visual data directly and using the Expand tool. It predicts that tool use is most beneficial when direct reading accuracy drops faster than selection accuracy, defining an optimal compression regime for performance.
Terminology used across episodes
This episode discusses
- LensVLM: Selective Context Expansion for Compressed Visual Representation of Text · Paper Radio
- gpt-oss-120b & gpt-oss-20b Model Card
- Glyph: Scaling Context Windows via Visual-Text Compression
- M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
- PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
- ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
- GRIT: Teaching MLLMs to Think with Images
- ColPali: Efficient Document Retrieval with Vision Language Models
- AgentOCR: Reimagining Agent History via Optical Self-Compression · Paper Radio
- ZeroSense:How Vision matters in Long Context Compression
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- RepoQA: Evaluating Long Context Code Understanding
- Language Modelling with Pixels
- HybridFlow: A Flexible and Efficient RLHF Framework
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
- Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
- VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning
- DeepSeek-OCR: Contexts Optical Compression
The paper
LensVLM: Selective Context Expansion for Compressed Visual Representation of Text · Read on arXiv
Apple
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "LensVLM: Selective Context Expansion for Compressed Visual Representation of Text".
Jane: LensVLM introduces an inference framework and post-training recipe that enables Vision Language Models (VLMs) to scan compressed images and selectively expand only the relevant images to their uncompressed form via…
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, to recap what we’ve heard, we’re talking about "LensVLM: Selective Context Expansion for Compressed Visual Representation of Text," which is basically an inference framework designed to help VLMs handle text as compressed images. The main thesis here is that while varying rendering resolution lets you compress the text representation, accuracy usually drops quickly because characters get lost at higher compression rates.
Jane: Exactly, and the authors propose LensVLM as a solution by enabling the model to first scan these compressed images and then use a learned tool to selectively expand only the relevant parts of those images back into their uncompressed form when necessary. This is different from older methods that decide what to keep before they start; this approach makes the expansion decision on demand during inference.
Lu: It’s a really clever idea because it acknowledges the limitation that traditional compression methods have—they lose information irreversibly once a certain threshold is crossed, and LensVLM builds a way around that loss by only spending resources on the parts of the input that matter for the final answer.
Meng: So, essentially, they are giving the VLM an "Expand tool" it can call during its reasoning process to recover high-resolution text or images only when its internal representation indicates those specific details are needed. That’s a concrete mechanism we can actually build with a policy layer.
Lalam: That selective recovery mechanism is what really excites me because it suggests that future AI systems won't just be about processing everything at once; they’ll be about intelligent resource allocation for visual data, which could lead to much more efficient and targeted AI applications in the future.
Tom: And the paper highlights their training process too; they use a three-stage post-training recipe: supervised fine-tuning to teach it from ground truth traces, followed by reinforcement learning to optimize its policy based on maximizing correctness and tool usage. That’s how they instill this selective behavior.
Jane: Right, and the results they report are quite strong; LensVLM maintains accuracy comparable to the full-text upper bound even at a four point three times effective compression rate, which is way better than what previous retrieval-based or text and visual compression baselines could achieve up to ten point one times compression.
Lu: The paper notes that training makes the visual compression robust to rendering choices; they found that it reduced an eighteen-point accuracy spread across different rendering configurations down to under one point, which shows the model is learning a generalized way to handle those input variations.
Meng: That robustness is important for real-world deployment because it means we don't need perfect pre-processing of the input images; the system can handle slight variations in how the text was initially rendered without failing.
Lalam: It really shows that by training this way, we aren't just getting a better compression ratio; we’re getting a more resilient AI that can handle real-world messy inputs without losing critical context.
Conclusion: Tom: So, wrapping up our discussion on "LensVLM: Selective Context Expansion for Compressed Visual Representation of Text," the title itself really captures the essence of what they achieved: selective context expansion for compressed visual representation of text. The authors, Roy Xie and his team, developed a framework that allows VLMs to intelligently manage information access during inference.
Jane: It’s about moving away from a model that blindly processes everything at once and instead giving it a tool to decide precisely which parts of the input need to be fully expanded for an accurate response. In simple terms, it means the AI learns when to spend its processing power on detail versus when it can rely on the compressed representation.
Lu: The implication is that we are moving toward more sophisticated AI systems that don't just process data passively; they actively manage their cognitive load based on what’s needed for the task, which opens up avenues for much deeper levels of contextual understanding.
Meng: From a practical standpoint, this suggests that future AI applications will need to incorporate intelligent resource management into their core design, ensuring the system knows how to dynamically trade off visual detail against computational cost effectively.
Lalam: And for the broader impact, I think this points toward a future where AI can interact with massive amounts of visual information in ways that are both highly efficient and extremely accurate because it’s not trying to process every pixel all the time.
Tom: That’s a big picture idea, Jane; essentially, we are seeing AI develop strategies for intelligent information filtering at the point of use rather than relying on brute force processing of everything. It seems like this work lays a solid foundation for building more nuanced and resource-aware visual reasoning systems.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck