RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models".
Jane: The gist: Our new model family, RT-DETRv4, achieves state-of-the-art results on COCO,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Alright team, we're kicking off our deep dive into this new paper, "RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models". We’re looking at how they tackle the trade-off between making object detectors super fast and keeping them smart.
Jane: That’s right. The core idea here is that when you try to build a very lightweight model for real-time use, you often end up losing the ability to understand complex visual details, which creates this bottleneck that stops further progress and makes it hard to put these models on actual devices.
Lu: Exactly. So this paper proposes a way to fix that by using Vision Foundation Models, or VFMs, which are these massive models trained on huge amounts of data for general vision tasks, to inject those high-level semantic ideas into the smaller detectors without changing the detector's structure when you actually use it.
Meng: So what’s the big claim here? How does this framework actually work to make that transfer stable when you’re dealing with two models that have totally different ways of learning things, like a massive VFM and a small real-time detector?
Tom: Well, the paper introduces a cost-effective and highly adaptable distillation framework. They leverage the rich semantic capacity of VFMs to enhance lightweight object detectors by transferring those semantics during training while keeping the detector architecture exactly the same during inference.
Jane: It sounds like they acknowledge that getting those high-level ideas into smaller models is tricky because there are so many differences in how these two types of models are trained, right?
Lu: That’s it. To handle that disparity, they introduce two main components for stable semantic transfer: first the Deep Semantic Injector, or DSI.
Meng: What exactly does this DSI module do during training? Is it just a fancy way to feed in some extra data?
Lu: The DSI is designed to integrate those high-level representations from the Vision Foundation Model into the deep layers of the detector. It’s a training-only module that gives explicit supervision for feature map F5 by aligning it with semantically rich representations from a VFM through a lightweight feature projector.
Tom: Okay, so we have DSI bringing in the semantics, but how do they make sure this whole process doesn't just break the training? That’s where their second component comes in.
Jane: They have Gradient-guided Adaptive Modulation, or GAM. This strategy dynamically adjusts the intensity of that semantic transfer based on gradient norm ratios during training.
Paper summary: Meng: So instead of just using a fixed loss value for the semantic part, this GAM system is watching how the gradients are behaving and changing how much influence that semantic injection has?
Lu: Precisely. GAM regulates the contribution of a module called AIFI by looking at its gradient norm ratio, rather than just the raw loss magnitude. This ensures a balanced optimization between making good detections and getting that semantic supervision right.
Tom: That sounds like they're trying to keep things stable while pushing for better performance across all those different learning goals simultaneously. So what’s the main training objective they set up?
Jane: They define the total training objective as a combination of the standard detection loss and this new semantic alignment loss. It’s Ltotal = Ldet plus lambdaLDSI, where Ldet is your normal detection loss, and LDSI is that semantic alignment loss.
Meng: And how do they manage that balancing act between the two objectives? Because if you just throw them together, one might totally overpower the other.
Lu: They use GAM to dynamically tune the lambda value based on gradient statistics. They check if the average gradient ratio lies within a target interval, and if it does, they steer AIFI’s contribution toward that boundary to keep things stable near equilibrium.
Tom: That dynamic tuning of lambda is really clever because it lets the system adapt its learning strategy as training progresses. This whole setup is designed to manage those competing objectives effectively.
Jane: So what are the actual results they show from this RT-DETRv4 paper? What are we looking at when we see these models in action?
Tom: The new model family, RT-DETRv4, achieves state-of-the-art results on COCO. They hit AP scores of forty-nine point seven, fifty-three point five, fifty-five point four, and fifty-seven point zero at speeds of two hundred seventy-three one hundred sixty-nine one hundred twenty-four and even a whopping seventy-eight frames per second respectively for the S/M/L/X variants.
Meng: That speed jump is pretty significant when you're talking about real-time use on hardware like a T4 GPU. What does that actually mean for someone who needs to deploy something fast?
Lu: It means even the largest variant, RT-DETRv4-X, gets fifty-seven point zero AP at seventy-eight FPS, which surpasses DEIM-X’s fifty-six point five AP without adding any extra inference overhead whatsoever.
Tom: So we’re seeing better accuracy across the board with models that are significantly faster than previous top contenders like YOLOv13-L and DEIM-L, even if they use less computational budget in some cases.
Paper summary: Jane: It really shows that this approach is practical because it doesn't change the way you run the model once it’s trained. You just get better performance without any extra cost to deploy or run at all.
Lu: They also did some ablation studies to prove the parts are working. They showed that applying DSI helps bring a slight performance improvement, and adding GAM really boosts that gain by zero point five AP on its own.
Meng: I saw those results too, and it confirms that both modules are necessary for the full effect. It’s not just one thing doing the heavy lifting here.
Tom: They also found that using a cosine similarity loss for alignment was better than using a mean squared error loss, which is good to know when you’re tuning those specific parts.
Jane: So, looking at this RT-DETRv4 paper, what does it really change for the person who just listens to our show? It’s about making powerful vision models accessible in real-time applications without needing a massive computing setup.
Lu: This framework offers a flexible way to use whatever Vision Foundation Model you want, like DINOv3 or CLIP, and distill its knowledge into a detector that runs at speed.
Meng: From an engineering standpoint, the fact that it’s agnostic to whether you’re using a CNN or transformer-based detector is huge for scalability in deployment. It can adapt to different setups easily.
Tom: And the training efficiency is good too, because they didn't have to change their entire optimization pipeline just to incorporate these foundation models into the process.
Jane: So, looking at the authors, Zijun Liao, Yian Zhao, Xin Shan and Yu Yan from Peking University and Tsinghua University—they’ve put together a framework that seems really focused on making VFM potential usable in real-time systems.
Lu: It’s a pathway to unlocking the potential of these massive foundation models for efficient visual perception in ways that were previously difficult to achieve.
Meng: It makes sense because if you can get high accuracy with low latency, then applications that need fast vision—like robotics or autonomous systems—become much more viable on edge devices.
Tom: So, the RT-DETRv4 series sets a new benchmark for performance across different scales and speeds on COCO without adding any deployment hurdles. We’re looking at how this distillation method can become a standard way to improve lightweight detectors moving forward.
Conclusion: Tom: So, we’ve seen how RT-DETRv4 uses Vision Foundation Models to boost real-time detection accuracy without adding any extra speed or complexity to the system itself.
Jane: That’s right, and the authors are Zijun Liao, Yian Zhao, Xin Shan, and Yu Yan from Peking University and Tsinghua University.
Lu: They tackled this problem by creating a distillation framework that lets you take knowledge from those massive foundation models and put it into smaller detectors in a very direct way.
Meng: From an engineering standpoint, the big picture here is that we can get much better results on edge devices without having to redesign our entire detection pipeline.
Lalam: This means future vision applications could get way more reliable because the core detection engine gets smarter from a foundation model without needing a massive training overhaul.
Tom: The real implication for us, for anyone listening, is that this isn't just about getting slightly higher accuracy; it’s about making high-level semantic understanding accessible in real-time on devices that don't have huge compute budgets.
Jane: Exactly. It’s a way to unlock the power of those foundation models for practical vision tasks, not just academic experiments.
Lu: I think the most interesting part is how they designed that training objective—how they use gradient ratios to guide the semantic transfer so it stays stable during training.
Meng: Stability is everything when you’re dealing with different learning styles, and managing those competing goals with dynamic tuning sounds like a really solid technical trick.
Tom: It’s a clever way to handle the gap between a huge foundation model and a tiny real-time detector architecture that keeps things balanced during the learning process.
Jane: So, we're looking at how these authors have created something that lets us use the best vision models for speed without sacrificing too much intelligence.
Lalam: It suggests that culture in AI development can move towards integrating these massive general models into specific real-time tasks more smoothly and efficiently than before.
School of Electronic and Computer Engineering, Peking University
cs.CV
Submitted: 2025-10-29
Updated: 2026-10-08
Code: https://github.com/ultralytics/yolov5
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: The gist: Our new model family, RT-DETRv4, achieves state-of-the-art results on COCO, attaining AP scores of 49.7/53.5/55.4/57.0 at corresponding speeds of 273/169/124/78 FPS The core problem
Key concepts
- Vision Foundation Models (VFMs)
- These are very large, powerful models trained on massive datasets that have learned deep, general visual understanding. They can capture complex, high-level semantic information about images. The paper uses these models as a source of rich knowledge to improve the performance of smaller, faster object detectors.
- Deep Semantic Injector (DSI)
- This is a training module designed to integrate the high-level semantic knowledge from VFMs into the deep layers of the lightweight detector. It aligns the detector's internal feature maps with these rich representations using a lightweight feature projector, providing explicit supervision during training.
- Gradient-guided Adaptive Modulation (GAM)
- GAM is a strategy that dynamically adjusts how much semantic information is transferred based on how similar the gradients are between detection and semantic tasks. It regulates the transfer intensity to ensure stable optimization and balanced learning of both detection accuracy and semantic alignment.
Terminology
Summary
The gist: Our new model family, RT-DETRv4, achieves state-of-the-art results on COCO, attaining AP scores of 49.7/53.5/55.4/57.0 at corresponding speeds of 273/169/124/78 FPS
The core problem addressed is the trade-off between lightweight models and feature representation quality.
Real-time object detection often involves a conflict where designing lightweight models to achieve high inference speed inevitably reduces their ability to capture high-level semantics, leading to a semantic bottleneck This limitation hinders further performance improvement and increases the difficulty of practical on-device deployment. The paper proposes a cost-effective and highly adaptable distillation framework that harnesses the capabilities of Vision Foundation Models (VFMs) to enhance lightweight object detectors. By transferring rich semantics from VFMs to real-time detectors during training while keeping the detector architecture unchanged during inference, the method enables significant enhancement without introducing any additional inference or deployment cost.
The framework introduces two key components for stable semantic transfer.
-
Deep Semantic Injector (DSI): This module facilitates the integration of high-level representations from VFMs into the deep layers of the detector. The DSI is a training-only module designed to provide explicit and powerful supervision for the feature map F5. It achieves this by aligning the detector’s feature map F5 with semantically rich representations from a vision foundation model T through a lightweight feature projector P.
-
Gradient-guided Adaptive Modulation (GAM): This strategy dynamically adjusts the intensity of semantic transfer based on gradient norm ratios. It regulates the relative contribution of the AIFI module according to its gradient norm ratio rather than the raw loss magnitude, ensuring balanced optimization between detection and semantic supervision.
The training objective is structured to manage these competing objectives.
The total training objective is defined as Ltotal = Ldet + λLDSI, where Ldet denotes the standard detection loss (e.g., classification and bounding box regression), and LDSI represents the proposed semantic alignment loss. To ensure stable and task-aligned semantic transfer, GAM dynamically tunes λ based on gradient statistics to harmonize the learning of semantic transfer and detection objectives. The update rule for λe+1 is defined based on whether the average gradient ratio r¯e lies within a target interval [ρ−δ, ρ+δ], steering AIFI's effective contribution toward the further boundary of the target range to ensure stable convergence near equilibrium.
The proposed method achieves state-of-the-art results across multiple model scales.
The new model family, RT-DETRv4-(S/M/L/X), achieves 49.7/53.5/55.4/57.0 AP scores on COCO [20] at 273/169/124/78 FPS. Specifically, the RT-DETRv4-L achieves 55.4 AP on COCO at 124 FPS, outperforming YOLOv13-L (53.4 AP) and DEIM-L (54.7 AP) under comparable or even lower computational budgets. The largest variant, RT-DETRv4-X, reaches 57.0 AP, exceeding DEIM-X (56.5 AP) without introducing any inference overhead.
Ablation studies confirm the effectiveness of the proposed modules.
Ablation on DSI and GAM showed that applying DSI can bring a slight performance improvement, while further applying GAM can significantly improve the performance gain (0.5 AP), proving the effectiveness of both. Ablation on semantic injection position indicated that directly applying semantic supervision to backbone features (S3, S4, or S5) individually or jointly yields no improvement, whereas the design aligning only the AIFI output F5 achieves a clear 0.5 AP improvement (54.3 AP). Furthermore, ablation on the alignment loss showed that the cosine similarity loss demonstrates superior performance over MSE Loss. The choice of a linear-based projector was found to yield the best results in Table 4.
The framework offers significant advantages in deployment and scalability.
Deployment Efficiency: Our method introduces zero modification to detector architectures and does not alter the inference pipeline, ensuring that no additional computational cost or latency is incurred. Scalability: The framework is highly general and can be seamlessly applied to detectors with diverse architectures, including CNN-based and transformer-based detectors. It can flexibly incorporate different foundation models, such as DINOv3 [31], MAE [13], or CLIP [27], and even benefit from arbitrarily large models for distilling semantics into real-time detectors. Training Efficiency: Since neither the detector structure nor the optimization pipeline is modified by the incorporation of VFMs, the additional training cost remains minimal.
The framework provides a practical pathway toward unlocking foundation model potential.
Overall, our work provides a practical pathway toward unlocking the potential of foundation models for efficient visual perception. The superiority of GAM in navigating the training dynamics is further illustrated in Figure 5, which plots the validation AP over epochs, consistently remaining above the baseline and all static weight configurations. Our results fully demonstrate the effectiveness and great potential of this approach to enhance real-time detectors without increasing inference or deployment overhead. The framework remains agnostic to both VFM type and scale, introducing no inference or deployment overhead, offering a more flexible and deployment-friendly solution.
In conclusion, the proposed RT-DETRv4 series sets a new state-of-the-art.
The results demonstrate that RT-DETRv4 consistently achieves the best performance across all model scales. The largest variant, RT-DETRv4-X, reaches 57.0 AP, exceeding DEIMv2 (56.5 AP) without introducing any inference overhead. The framework is cost-effective and highly adaptable for enhancing real-time detectors. Our work provides a practical pathway toward unlocking the potential of foundation models for efficient visual perception.
--- Page 1 ---
RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
Zijun Liao1† Yian Zhao1† Xin Shan1 Yu Yan1 Chang Liu2 Lei Lu1 Xiangyang Ji2 Jie Chen1 1 School of Electronic and Computer Engineering, Peking University, Shenzhen, China 2Department of Automation and BNRist, Tsinghua University, Beijing, China zjliao25@stu.pku.edu.cn zhaoyian@stu.pku.edu.cn
Abstract Real-time object detection has achieved substantial progress through meticulously designed architectures and optimization strategies However, the pursuit of high-speed inference via lightweight network designs often leads to degraded feature representation, which hinders further performance improvements and practical on-device deployment. In this paper, we propose a cost-effective and highly adaptable distillation framework that harnesses the rapidly evolving capabilities of Vision Foundation Models (VFMs) to enhance lightweight object detectors. Given the significant architectural and learning objective disparities between VFMs and resource-constrained detectors, achieving stable and taskaligned semantic transfer is challenging. To address this, on one hand, we introduce a Deep Semantic Injector (DSI) module that facilitates the integration of high-level representations from VFMs into the deep layers of the detector. On the other hand, we devise a Gradient-guided Adaptive Modulation (GAM) strategy, which dynamically adjusts the intensity of semantic transfer based on gradient norm ratios. Without increasing deployment and inference overhead, our approach painlessly delivers striking and consistent performance gains across diverse DETR-based models, underscoring its practical utility for real-time detection. Our new model family, RT-DETRv4, achieves state-of-the-art results on COCO, attaining AP scores of 49.7/53.5/55.4/57.0 at corresponding speeds of 273/169/124/78 FPS
--- Page 2 ---
Latency on T4 GPU (ms) 46 C O C O A P v al 49.7 53.5 55.4 57.0 RT-DETRv4 DEIMv2 DEIM D-FINE RT-DETRv3 RT-DETRv2 RT-DETR YOLOv13 YOLOv12 YOLO11 YOLOv10 Figure 1 Compared with existing advanced real-time object detectors on COCO [20] Our RT-DETRv4 models achieve stateof-the art performance.
--- Page 3 ---
Over the past decade, remarkable progress has been driven by increasingly efficient network architectures and end-to-end learning frameworks. In particular, two representative series, YOLO [30] and DETR [3], have profoundly influenced the evolution of object detection paradigms. The YOLO family emphasizes rapid one-stage detection achieving high inference speed and practical deployment efficiency, and the DETR series has reshaped the detection paradigm through its unified modeling of object queries and set-based prediction.
Improvements for AI systems
-
textbfDeep Semantic Injection (DSI) Module Enhancement: Targeted Supervision for F5 Consistency/Improvement on Backbone and AIFI Modules by aligning only F5 with VFM representations, as demonstrated in
configuration (c), considering the pivotal role of the AIFI module within the hybrid encoder, the alignment is conducted on its output F5, which contains the richest semantics.
This design is shown to achieve aclear 0.5 AP improvement (54.3 AP)
over other strategies because it allows gradients to backpropagate,enhancing both modules synergistically.
-
textbfGradient-guided Adaptive Modulation (GAM) for Stable Optimization: Dynamically tune the intensity of semantic transfer based on gradient norm ratios, as described by
dynamically adjusts the intensity of semantic transfer based on gradient norm ratios.
This mechanism prevents issues like insufficient supervision or excessive dominance by ensuringbalanced optimization between detection and semantic supervision,
leading to performance gains over static weight configurations, with GAM achieving the best performance (55.4 AP). -
textbfModel Versatility and Deployment Flexibility: Achieve state-of-the-art results on diverse architectures without increasing overhead, as stated by
Our method introduces zero modification to detector architectures and does not alter the inference pipeline.
This allows for the application of VFMs like DINOv3, MAE, or CLIP to any DETR-based detector (CNN or Transformer), providing amore flexible and deployment-friendly solution
that remains agnostic to VFM type and scale.
Abstract
Real-time object detection has achieved substantial progress through meticulously designed architectures and optimization strategies. However, the pursuit of high-speed inference via lightweight network designs often leads to degraded feature representation, which hinders further performance improvements and practical on-device deployment. In this paper, we propose a cost-effective and highly adaptable distillation framework that harnesses the rapidly evolving capabilities of Vision Foundation Models (VFMs) to enhance lightweight object detectors. Given the significant architectural and learning objective disparities between VFMs and resource-constrained detectors, achieving stable and task-aligned semantic transfer is challenging. To address this, on one hand, we introduce a Deep Semantic Injector (DSI) module that facilitates the integration of high-level representations from VFMs into the deep layers of the detector. On the other hand, we devise a Gradient-guided Adaptive Modulation (GAM) strategy, which dynamically adjusts the intensity of semantic transfer based on gradient norm ratios. Without increasing deployment and inference overhead, our approach painlessly delivers striking and consistent performance gains across diverse DETR-based models, underscoring its practical utility for real-time detection. Our new model family, RT-DETRv4, achieves state-of-the-art results on COCO, attaining AP scores of 49.8/53.7/55.4/57.0 at corresponding speeds of 273/169/124/78 FPS. Code is publicly available at https://github.com/RT-DETRs/RT-DETRv4.
Sources
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Real-Time Object Detection Meets DINOv3
- YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception
- DINOv2: Learning Robust Visual Features without Supervision
- YOLOv3: An Incremental Improvement
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models