CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
Peng Ling, Yingda Yin, Lingting Zhu, Weikai Chen, Shengju Qian, Zeyu Hu, Xin Wang, Wenming Yang
Shenzhen International Graduate School, Tsinghua University · LIGHTSPEED
cs.CV, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: Accepted to ECCV 2026 as an Oral Presentation
Code: https://github.com/Brucess/CoverPrune
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: CoverPrune is a training-free token pruning framework for 3D Vision-Language Models (3D VLMs) that shifts the pruning objective from maximizing token diversity to preserving visual evidence coverage.
Terminology
Summary
CoverPrune is a training-free token pruning framework for 3D Vision-Language Models (3D VLMs) that shifts the pruning objective from maximizing token diversity to preserving visual evidence coverage. The paper states: We propose a paradigm shift for 3D VLM token pruning: from maximizing diversity to preserving visual evidence coverage.
The method formulates inference-time token pruning as an Optimal Transport (OT) problem, where the retained tokens act as prototypes that distribute their representational mass to the original tokens, and pruning aims to minimize the distortion of this coverage assignment.
To address the challenges of this formulation, CoverPrune introduces three key designs: (1) a Feature-Spatial-Temporal (FST) transport cost that jointly models semantic similarity, spatial proximity, and temporal coherence
; (2) informativeness-aware target capacities to stabilize coverage under aggressive pruning
; and (3) an efficient Spatial-Guided Greedy Selection (SGS) algorithm that approximates the semi-relaxed OT objective. The paper also proposes CoverPrune-Lite, an accelerated variant utilizing spatially structured local matching for minimal overhead,
which uses Morton code ordering and capacity-guided grouping to achieve O(N log N) inference-time complexity.
The paper's main contributions are: a novel pruning paradigm based on coverage rather than diversity; tailored OT solutions including the FST cost, dynamic target capacities, and an efficient optimization algorithm; a lightweight acceleration variant; and state-of-the-art performance across multiple benchmarks.
Experiments were conducted on four benchmarks: ScanQA, SQA3D, Scan2Cap, and VSI-Bench, using GS-Reasoner and VLM-3R as base models. The results show that CoverPrune and CoverPrune-Lite achieve top-tier performance on nearly all reported metrics under matched token budgets
on general 3D tasks. On VSI-Bench, At 20% token retention, CoverPrune preserves 92.4% of full-token performance,
and both variants show much more graceful degradation than other SOTA methods
under aggressive pruning, with substantially higher scores than the strongest competitor at 10% and 5% retention.
Ablation studies show that removing the feature cost causes the largest degradation, while removing geometry or temporal cost incurs smaller but non-negligible drops. Efficiency analysis shows CoverPrune-Lite reduces pruning time to 0.41 seconds compared to 3.47 seconds for DTC, while achieving the best relative accuracy (88.01% vs. 79.85% for DTC) at the same token budget.
Improvements for AI systems
Improvements to AI systems:
-
Coverage-preserving token pruning for 3D VLMs – Replace diversity-maximizing pruning (which drops redundant but semantically critical tokens) with coverage-based pruning that retains tokens acting as prototypes covering the original token distribution. This yields more graceful performance degradation at extreme token budgets (e.g., 5–10% retention) for 3D question answering, captioning, and scene understanding.
-
Optimal transport–based inference-time compression – Formulate token selection as a semi-relaxed OT problem with a Feature-Spatial-Temporal (FST) cost, enabling joint optimization of semantic similarity, 3D spatial proximity, and temporal coherence. This improves robustness for dynamic 3D scenes (e.g., video+LiDAR inputs) where spatial and temporal structure matter.
-
Informativeness-aware dynamic target capacities – Adapt the number of retained tokens per region based on local informativeness (e.g., high-detail objects vs. flat walls), stabilizing coverage under aggressive pruning. This allows the system to allocate more tokens to critical regions, improving accuracy on fine-grained 3D tasks like Scan2Cap.
-
Spatial-Guided Greedy Selection (SGS) algorithm – Approximate the OT objective in near-linear time, reducing inference overhead. This enables real-time or edge-deployment of 3D VLMs without retraining, as pruning happens at inference time.
-
CoverPrune-Lite with Morton code ordering and capacity-guided grouping – Achieve O(N log N) complexity, cutting pruning time from 3.47s to 0.41s (≈8.5× faster) while improving relative accuracy by 8 percentage points over prior DTC method. This makes token pruning feasible for latency-sensitive robotics or AR/VR applications.
-
Task-agnostic plug-in module – Integrate CoverPrune as a training-free layer into any existing 3D VLM (e.g., GS-Reasoner, VLM-3R) to compress input tokens before the LLM backbone, reducing memory and compute without fine-tuning. This improves scalability to larger scenes or higher-resolution 3D inputs.
-
Benchmark-driven robustness – Use the FST cost to generalize across diverse benchmarks (ScanQA, SQA3D, Scan2Cap, VSI-Bench), enabling a single pruning configuration to work across embodied QA, spatial reasoning, and dense captioning tasks, rather than task-specific tuning.
What the improved AI system can do:
-
Process 3D scenes (point clouds, meshes, or Gaussian splats) with up to 95% token reduction while retaining >90% of full-token accuracy on complex reasoning tasks.
-
Operate in real-time on embedded systems (e.g., robots, drones) by pruning tokens in under half a second, enabling on-the-fly scene understanding during navigation or manipulation.
-
Maintain high performance on dynamic 3D data (e.g., sequential LiDAR frames) by leveraging temporal coherence in the pruning cost, avoiding flickering or inconsistent token selection across frames.
-
Automatically focus computational resources on semantically rich regions (e.g., objects, text in scenes) while aggressively compressing uniform areas (walls, floors), improving both speed and accuracy for dense captioning or visual grounding.
-
Serve as a drop-in accelerator for existing 3D VLMs without retraining, reducing GPU memory usage and inference latency, making large 3D models deployable on smaller hardware.
Sources
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Reasoning in Space via Grounding in the World
- 3D Aware Region Prompted Vision Language Model
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
- EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models