MergeOver: Post-Training Token Merging for Recursive Vision Transformers
Junseo Kim, Uraz Odyurt, Amirreza Yousefzadeh
University of Twente
cs.CV, cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: MergeOver is a post-training approach that integrates Token Merging (ToMe) into the recursively weight-shared Sliced Recursive Transformer (SReT), without retraining.
Terminology
Summary
MergeOver is a post-training approach that integrates Token Merging (ToMe) into the recursively weight-shared Sliced Recursive Transformer (SReT), without retraining. The paper states: We propose MergeOver, a post-training approach that integrates Token Merging (ToMe) into the recursively weight-shared Sliced Recursive Transformer (SReT).
The approach resolves three key architectural incompatibilities: (1) token reduction violates the rigid 2D spatial grid required by SReT's hierarchical convolutional pooling layers; (2) unconstrained merging can violate the strict group-size boundaries inherent to the Sliced Group Self-Attention (SGA) of SReT; and (3) SReT's spatial permutations break the tracking necessary for the proportional attention of ToMe.
To address these, MergeOver introduces an Unmerge tracking stack
that restores the original sequence length before each inter-stage convolutional pooling layer by duplicating merged features back into their original spatial coordinates. The paper explains: Before the compressed sequence reaches the inter-stage convolutional pooling layer, the tensor is built back up to its original dimensions by applying the stored Unmerge operations in strict reverse order.
Additionally, the merge rate is constrained by the formula: (N −r) mod g = 0, N −r ≥ g, g = LCM(g1, g2), r ≤ ⌊N/2⌋, ensuring group divisibility and bipartite matching limits. Finally, a token-mass tensor (s) is routed in parallel through the same permutation and inverse permutation functions as the feature sequence to maintain synchronisation.
The paper evaluates four token reduction scheduling strategies: global constant, global linear, stage-wise exponential, and stage-wise single-shot. The stage-wise single-shot schedule performs token reduction only at the first block of each stage and maintains a fixed sequence length throughout subsequent recursive iterations. The paper states: a stage-wise single-shot schedule performs token reduction only at the first block of each stage... This schedule maintains a fixed sequence length throughout the remaining blocks of the stage.
Benchmarked on ImageNet-1K, the selected configuration (rho shot = 0.25) reduces top-1 accuracy by 1.47 percentage points relative to the SReT baseline (from 77.39% to 75.92%). The paper reports: our selected stage-wise single-shot configuration reduces top-1 accuracy by only 1.47 percentage points, which is negligible in many use-cases.
Hardware results are platform- and batch-size-dependent. On an NVIDIA GeForce RTX 4060 Ti GPU, at batch size 16, the selected configuration increases throughput by 21.7% and reduces peak activation memory by 38.4%. At batch size 1, throughput decreases by 21.7% while PAM is reduced by 37.3%. The paper notes: At batch size 1, however, MergeOver degrades GPU throughput for every tested configuration... the GPU requiring sufficient parallel work to amortise the added overhead of MergeOver.
On an Intel Core Ultra 9 285K x86 CPU, the selected configuration reduces latency by 17.9% at batch size 1 and by 30.0% at batch size 16. On a Raspberry Pi 5 (ARM CPU), it reduces latency by 2.4% at batch size 1 and by 17.6% at batch size 16. The paper concludes: These results demonstrate that the hardware benefits of token reduction depend strongly on the processing platform and batch size.
The paper also reports FLOPs reductions: the selected configuration reduces FLOPs from 1.91 G to 1.49 G (a 22.0% reduction). Parameter counts remain unchanged at 4.76 M. The paper notes that a configuration with fewer FLOPs does not necessarily achieve higher throughput,
as the FLOPs estimate does not capture costs such as BSM matching, token-mass tracking, and Unmerge operations.
The paper acknowledges limitations: evaluation is restricted to SReT-Tiny-Distill, the canonical and stage-wise schedules differ in parameterisation and depth indexing, and CPU memory measurements (ΔRSS) are process-level observations not directly comparable to GPU PAM. Future work could combine MergeOver with low-level kernel optimisation and complementary post-training techniques such as quantisation. The paper concludes: "MergeOver provides a post-training baseline for integrating token merging into hierarchical recursive Transformers and establishes a foundation for further low-level optimisation and evaluation, especially on resource-constrained hardware."
Improvements for AI systems
Improvements to AI systems:
-
Adaptive token-merging scheduler for hierarchical transformers – Implement a stage-wise single-shot token reduction schedule (ρ=0.25) that merges tokens only at the first block of each stage, maintaining fixed sequence length thereafter. This reduces FLOPs by 22% (1.91G→1.49G) with only 1.47% top-1 accuracy drop on ImageNet-1K, enabling faster inference on edge devices.
-
Unmerge tracking stack for lossless feature restoration – Integrate a reverse-order unmerge operation that reconstructs original spatial dimensions before inter-stage convolutional pooling. This allows token merging without breaking hierarchical feature maps, preserving spatial coherence for tasks like semantic segmentation or object detection.
-
Group-constrained merge rate with divisibility enforcement – Apply the constraint (N−r) mod g = 0, g = LCM(g1,g2), r ≤ ⌊N/2⌋ to ensure bipartite matching respects Sliced Group Self-Attention boundaries. This prevents attention-group violations, making the method safe for any transformer with grouped attention heads.
-
Parallel token-mass routing for permutation synchronization – Maintain a token-mass tensor (s) that passes through the same permutation and inverse-permutation functions as features. This enables accurate proportional attention tracking during recursive iterations, allowing the model to preserve information density across depth.
-
Platform-aware deployment selector – Use the reported hardware benchmarks to automatically choose between token merging and full inference: enable merging on ARM CPUs (17.6% latency reduction at batch 16) and x86 CPUs (30% at batch 16), but disable it on GPUs at batch size 1 (21.7% throughput degradation) and enable only at batch ≥16 (21.7% throughput gain).
-
Memory-optimized inference pipeline – Leverage the 38.4% peak activation memory reduction (GPU) and 37.3% (batch 1) to fit larger models or longer sequences into limited VRAM, enabling deployment of recursive transformers on 8GB GPUs or mobile NPUs.
What the improved AI system can do:
-
Run hierarchical recursive vision transformers on Raspberry Pi 5 with 17.6% lower latency and 22% fewer FLOPs, making real-time image classification feasible on battery-powered IoT devices.
-
Process batch-16 inference on consumer GPUs (RTX 4060 Ti) with 21.7% higher throughput and 38.4% lower memory, enabling multi-stream video analytics on a single workstation.
-
Maintain near-baseline accuracy (75.92% vs 77.39% top-1) while reducing compute, suitable for edge deployment where accuracy loss under 2% is acceptable.
-
Dynamically switch between merged and full inference based on current hardware load and batch size, maximizing throughput without manual tuning.
-
Extend the same post-training merging to other hierarchical recursive architectures (e.g., audio spectrogram transformers, point-cloud networks) that rely on rigid spatial grids and grouped attention, without retraining.
Abstract
Vision Transformers (ViTs) demonstrate exceptional performance in computer vision but suffer from large parameter counts and quadratic computational complexity, severely limiting their deployment on resource-constrained edge hardware. While recursive weight-sharing reduces parameter counts and token merging mitigates computational and memory bottlenecks, integrating these two paradigms without costly retraining is non-trivial, leaving this intersection largely unexplored. We propose MergeOver, a post-training approach that integrates Token Merging (ToMe) into the recursively weight-shared Sliced Recursive Transformer (SReT). Through an Unmerge tracking stack, constraint-safe merge-rate adjustment, and synchronised token-mass tracking across spatial permutations, MergeOver resolves the spatial and merging constraints of this integration. We further employ a stage-wise single-shot schedule that performs token reduction at the first block of each stage and maintains a fixed sequence length throughout its subsequent recursive iterations. Benchmarked on ImageNet-1K, our selected configuration reduces top-1 accuracy by 1.47 percentage points. On the GPU, it reduces peak activation memory by 37.3% and 38.4% at batch sizes 1 and 16, while throughput decreases by 21.7% at batch size 1 but increases by 21.7% at batch size 16. On a Raspberry Pi 5 (ARM CPU), it reduces latency by 2.4% and 17.6% at batch sizes 1 and 16. These results show that MergeOver can recover a meaningful part of the throughput and memory cost that recursive weight-sharing introduces, without retraining, and provides a baseline for combining token merging with hierarchical recursive transformers.
Sources
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer
- Universal Transformers
- Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations
- Token Merging: Your ViT But Faster
- Towards Joint Quantization and Token Pruning of Vision-Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models