ZeBROD: Zero-Retraining Based Recognition and Object Detection Framework

arXiv:2512.04888 · cs.CV · Submitted 2025-12-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "ZeBROD: Zero-Retraining Based Recognition and Object Detection Framework".

Jane: Object detection often suffers from catastrophic forgetting when new products are introduced, necessitating costly and time-consuming model retraining.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we’ve been talking through the ZeBROD framework today, looking at how it tackles the problem of catastrophic forgetting by separating detection from recognition and using a continuous database for classification two <ref:2512.04888#pg1>. It really seems like a neat way to handle the constant influx of new products without needing to retrain everything from scratch.

Jane: That’s right, Tom. We’ve seen how the authors validate this approach with performance metrics showing that ZeBROD achieves seventy-eight point three six mAP in one of their tests while YOLO11n struggled after introducing four new product batches two. It proves that this modular design offers a path forward for handling incremental learning in computer vision systems.

Lu: The authors are essentially arguing that you don't need to completely overhaul the detection mechanism when adding new classes, but rather augment a separate classification layer with reference data two <ref:2512.04888#pg1>. This is a solid argument for decoupling the two major tasks of object localization and recognition.

Meng: From my side, the practical implication is that we could deploy these systems in retail environments where product lines change frequently without needing constant infrastructure rebuilds two <ref:2512.04888#pg1>. The inference speed on edge devices also supports this, which I think is crucial for real-time operation.

Lalam: For the wider culture of AI, ZeBROD demonstrates that AI can be designed to be inherently adaptive and resilient; it doesn't get stuck in outdated knowledge when the environment evolves one. This capability could lead to a much more fluid and responsive interaction between users and intelligent systems.

Tom: That’s the big picture, Jane. So, we see ZeBROD as a framework that provides remarkable accuracy without the need for constant retraining, proving that modularity in AI design can solve real-world deployment headaches two <ref:2512.04888#pg1>. It’s a lot to think about for product teams out there.

Jane: Indeed. The authors have laid out a clear methodology showing how you can achieve zero-shot onboarding of new products just by adding reference embeddings without retraining the core models two <ref:2512.04888#pg1>. It's a very tangible solution to a persistent problem in object detection and recognition systems.

Lu: Moving forward, I think we should focus on exploring how this embedding space itself can be used for more advanced, semantic product relationships rather than just exact matching one. That’s where the real potential for creative AI lies.

Meng: For practical implementation, the next steps will likely involve stress-testing that Qdrant database under heavy load and ensuring the retrieval latency remains consistently low across diverse inventory types two <ref:2512.04888#pg1>. That scalability is what separates a promising idea from a production-ready solution.

Lalam: I’m excited to see how this kind of incremental learning philosophy influences how we design future AI agents; systems that can absorb new experiences without forgetting old ones are key for long-term integration one. This paper gives us a blueprint for that kind of system.

Tom: Well, that wraps up our discussion on ZeBROD, Jane. It’s a study demonstrating how to build robust object detection and recognition systems that adapt to change through thoughtful separation of components two <ref:2512.04888#pg1>. We’ve got some exciting ideas on how this could impact the real world.

Conclusion: Tom: So we’ve just wrapped up our deep dive into ZeBROD, focusing now on what this whole paper actually means for us in practical terms and who came up with this clever idea by putting the framework together.

Jane: It really boils down to how they managed to keep the detection and recognition parts totally separate so you don't have to redo all that heavy training just because a new product shows up.

Lu: The authors, G. I. Parisi and his team, are brilliant for taking this modular approach, especially how they used DeiT features combined with a Qdrant database for the classification step without touching the original detection weights.

Meng: From an engineering standpoint, the most interesting part is that ZeBROD keeps training time constant even after adding new product batches; that means less downtime and simpler maintenance cycles for us at work.

Lalam: This moves AI development toward a much more continuous learning model, which fundamentally changes how we think about updating product catalogs and inventory management systems over the long term.

Tom: Exactly! The title itself, "Zero-Retraining Based Recognition and Object Detection Framework," tells us they’re aiming for a system that adapts on its own without constant manual intervention from our side.

Jane: It’s a way of saying we can onboard new items with just some reference data in the database instead of having to retrain the entire detection model every single time we launch something new.

Lu: I think this separation of concerns—localization versus recognition—is key; it lets us optimize each part independently, which is a really powerful architectural move for complex vision tasks.

Meng: I'm curious about the scalability aspect; if we have thousands of SKUs, how does that Qdrant vector database handle the high-dimensional search and ensure that real-time inference stays snappy on our edge devices?

Lalam: That’s where the impact is huge; imagine an AI system in a store that can instantly recognize a new item just by looking at it, making inventory updates near-instantaneous instead of waiting for massive batch retraining.

Tom: That sounds like a future where product recognition isn't just about reading barcodes, but about seeing and understanding the world continuously.

Jane: It’s about creating an AI that is inherently resilient to change, which makes those retail or manufacturing applications feel much more intuitive for the end-user.

Lu: The implications stretch beyond just detection; this methodology opens up possibilities for building truly lifelong learning systems where knowledge accumulates incrementally over time.

Meng: So we’re looking at a system that can handle rapid product evolution without massive computational overhead, provided the embedding space remains manageable.

Lalam: And on a cultural level, this shows us that AI doesn't have to be brittle; it can be designed to grow and absorb new information gracefully into its existing knowledge base.

Tom: Right! So we’ve seen how this framework uses a clean separation of tasks to achieve recognition without the massive retraining burdens we usually face.

Jane: It really shows that smart design choices in AI architecture can solve persistent problems like catastrophic forgetting in a very elegant way.

Lu: We should definitely keep an eye on how they expand this embedding space; there’s so much creative potential for semantic understanding beyond simple SKU matching here.

Meng: I'm focused on the practical deployment aspect; the paper gives us a solid foundation, and now we need to see how robust it is when we put it through real-world stress testing conditions.

Lalam: This work is a testament to how thoughtful AI design can lead to systems that are not just accurate today, but capable of evolving with the demands of tomorrow’s products.

Priyanto Hidayatullaha, Nurjannah Syakranib, Yudi Widhiyasanac, Muhammad Rizqi Sholahuddind, Refdinal Tubaguse, Zahri Al Adzani Hidayatf, Hanri Fajar Ramadhang, Dafa Alfarizki Pratamah, Farhan Muhammad Yasini

abcdfghi Computer Engineering and Informatics Department, Politeknik Negeri Bandung · Stunning Vision AI

cs.CV

Submitted: 2025-12-04

Updated: 2026-10-01

Comments: This manuscript was first submitted to the Journal of Automation and Intelligence. The preprint version was posted to arXiv afterwards to facilitate open access and community feedback

Code: https://github.com/ultralytics/ultralytics

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 80/100

The gist: Object detection often suffers from catastrophic forgetting when new products are introduced, necessitating costly and time-consuming model retraining.

Key concepts

Object Localization
This is the first step where the system uses a detector (YOLO11n) to find and draw boxes around every product visible in an image. It is trained only on identifying 'product' objects, ensuring it remains stable even when new items appear.
DeIT Feature Extraction
This process takes the cropped image patches from the detected boxes and converts them into a compact numerical representation called an embedding vector (384 dimensions). This embedding captures the unique visual characteristics of each product.
Qdrant Vector Database
This is a high-speed database used to store and search for product embeddings. It uses HNSW graphs to quickly find the closest known product features when trying to identify an unknown item in real-time.

Terminology

Summary

Object detection often suffers from catastrophic forgetting when new products are introduced, necessitating costly and time-consuming model retraining. This study introduces ZeBROD, a Zero-Retraining Based Recognition and Object Detection Framework designed to address this issue by separating object localization from recognition and utilizing a continuous metric-embedding database for classification.

The gist

ZeBROD is a methodology that separates object localization (using YOLO11n) from product identification (using DeIT feature extraction and Qdrant vector database search), allowing new products to be added via few-shot expansion without retraining the model parameters or old weights.

How it works

The framework operates as a modular, two-stage recognition system:

  1. Real-time object detector that localizes all visible products in a checkout image using YOLO11n. The detector is trained only on one class: “product,” and it remains static post-launch to avoid distributional drift. For every recognized box, a picture patch is obtained via cropping and padding to a size of 224 x 224 pixels.

  2. Feature extraction where each detected region feature is mapped into a discriminative embedding space using the Data-efficient Image Transformer (DeIT). The DeiT-Small configuration is used as the backbone, outputting a global representation from the classification token, resulting in an embedding vector of dimension 384. This model is fine-tuned with Proxy Anchor Loss to minimize distance between embeddings of identical products while maximizing distance between different products.

How it works (Continued)

  1. Identification is performed through compound search with reference embeddings stored in Qdrant, a high-performance vector database optimized for high-dimensional data and utilizing Hierarchical Navigable Small World (HNSW) graphs for rapid approximate nearest neighbor (ANN) searches. Reference embeddings are generated by computing the DeIT embeddings of one or more canonical images for each known product SKU, using the centroid of many instances to enhance robustness.

How it works (Continued)

  1. Inference involves finding the embedding vector for each detected patch and querying Qdrant for the top-k nearest neighbors based on cosine similarity:

s i m(ei, elf) = eiT ∙ elf‖ei‖2 ∙ ‖elf‖2

The system returns the SKU that matches the reference vector with the highest score, provided the similarity is higher than a certain threshold (e.g., τ=0.75). If no match is better than τ, the instance is marked as a possible new product for later human operator review. This design enables zero-shot onboarding of new products by only requiring insertion of reference embeddings into Qdrant without retraining the detector or embedding model.

Training and Testing Environment

The study compares the proposed framework (ZeBROD) against classical object detection methods like YOLO11n regarding accuracy, training speed, and inference speed across two environments: a Windows 11 workstation for training and a Raspberry Pi OS edge device for testing. The dataset comprises 140 product categories, initially divided into 100 known goods and 40 new products.

Result and Discussion

In the first stage, using the initial dataset of 100 product categories, ZeBROD achieved a 99.5 mAP detection accuracy on the test dataset with one class. Following DeIT training with Proxy Anchor Loss, it attained an 88.9 mAP, while YOLO11n attained a detection accuracy of 93.1 mAP. After introducing four new product batches, YOLO11n achieved a mean Average Precision (mAP) of 96.7, whereas the ZeBROD framework achieved a mAP of 78.36.

Conclusion

The proposed framework validates that it achieves a remarkable accuracy of 78.36 mAP without retraining the model every time new products are introduced. The average training duration for the classical approach (YOLO11n) increased almost three times following the introduction of four new product batches, while ZeBROD's training length remains constant. Furthermore, it exhibits an average inference time of 580 ms for all products shown in the checkout image on an edge device, making it suitable for real-world conditions and justifying its feasibility in retail applications. The framework eliminates the requirement for barcodes and is rapid enough to be used as a product detector displaying name, price, and total price.

Acknowledgments

This research was supported by Politeknik Negeri Bandung [grant number 108.14/R7/PE.01.03/2025].

References

[1] G. I. Parisi et al., “Continual lifelong learning with neural networks: A review,” Neural Networks, vol. 113, pp.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the ZeBROD framework, and what those improved systems can achieve:


  1. Improve Catastrophic Forgetting Mitigation in Continual Object Detection Systems:

  2. Enable Zero-Retraining Onboarding of New Product SKUs in Retail Environments:

  3. Enhance Model Efficiency for Edge Deployment (Low Latency Inference):

  4. Achieve Robust Recognition Under Real-World Viewpoint and Lighting Variations:

  5. The improved AI system can perform continuous, lifelong learning for object recognition tasks without the prohibitive computational cost or time required by traditional fine-tuning methods. It can adapt to new product introductions (e.g., new SKUs in a retail store) by simply adding their reference embeddings to a vector database (Qdrant), rather than retraining the entire deep neural network.

  6. The system can achieve highly efficient, real-time object detection and identification on resource-constrained edge devices (like Raspberry Pi). It can process multiple products per image with an average inference time of 580 ms, making it suitable for high-throughput Point-of-Sale (POS) systems where rapid price calculation is essential.

  7. The improved system can maintain high classification accuracy even when faced with significant visual variations in product presentation, such as different packaging angles, lighting conditions, and background clutter. This robustness is achieved by decoupling localization (using YOLO11n) from recognition (using DeIT embeddings) and utilizing metric learning for class identification based on feature similarity rather than fixed classification layers.

  8. The system can function as an intelligent inventory management tool that supports zero-shot onboarding of new products instantly upon catalog updates. When a new product is introduced, the system only needs to generate its embedding and insert it into Qdrant; no retraining or model parameter adjustments are required, drastically reducing deployment time and operational expenses for retail businesses.

Abstract

Object detection constitutes the primary task within the domain of computer vision. It is utilized in numerous domains. Nonetheless, object detection continues to encounter the issue of catastrophic forgetting. The model must be retrained whenever new products are introduced, utilizing not only the new products dataset but also the entirety of the previous dataset. The outcome is obvious: increasing model training expenses and significant time consumption. In numerous sectors, particularly retail checkout, the frequent introduction of new products presents a great challenge. This study introduces Zero-Retraining Based Recognition and Object Detection (ZeBROD), a methodology designed to address the issue of catastrophic forgetting by integrating YOLO11n for object localization with DeIT and Proxy Anchor Loss for feature extraction and metric learning. For classification, we utilize cosine similarity between the embedding features of the target product and those in the Qdrant vector database. In a case study conducted in a retail store with 140 products, the experimental results demonstrate that our proposed framework achieves encouraging accuracy, whether for detecting new or existing products. Furthermore, without retraining, the training duration difference is significant. We achieve almost 3 times the training time efficiency compared to classical object detection approaches. This efficiency escalates as additional new products are added to the product database. The average inference time is 580 ms per image containing multiple products, on an edge device, validating the proposed framework's feasibility for practical use.

Sources

Related papers