Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real
summary
In short
The episode discusses a paper showing that using camera geometry beats estimated depth for multi-camera 3D tracking under Sim2Real conditions. The authors found that cross-view geometric consistency is more important than per-image depth accuracy. They conclude that improving the 2D detector's robustness to real-world appearance is the most valuable next step.
Key concepts
- Geometry Beats Estimated Depth
- This refers to using known camera positions and calibrations (geometry) to project 2D detections into world coordinates, rather than estimating depth from each camera independently. The paper shows geometry provides superior consistency across multiple views.
- HOTA
- HOTA is a metric used to balance detection, association, and localization accuracy in multi-object tracking tasks. A higher HOTA score indicates better overall performance in these areas.
- Floor Coherence Diagnostic
- This is an annotation-free tool used to check the consistency of fused depth data. It measures how many points in the fused point cloud land within a certain distance of the known warehouse floor plane, helping identify geometric warping.
Terminology used across episodes
This episode discusses
- Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real · Paper Radio
- Efficiently Reconstructing Dynamic Scenes One D4RT at a Time
The paper
Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real · Read on arXiv
Abdullah Naeem, Ayon Dey, Anav Katwal, Md Tamjidul Hoque, Noman Khan
LSU New Orleans · PinPark, Inc.
The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting; depth is available only for training and validation, so inference is RGB-only. We use two RGB-only routes as a controlled test of one hypothesis: that cross-view geometric consistency, not monocular depth accuracy, governs performance under Sim2Real. The first is a geometry-first pipeline: YOLO11x detection, homography lifting to the world frame, class-level 3D size priors, multi-camera fusion, world-coordinate tracking, and offline tracklet stitching. The second is estimated-depth pseudo-LiDAR: monocular depth (D4RT, Metric3D v2) back-projected into a fused point cloud and passed to a 3D detector (V-DETR), mirroring prior point-cloud winners that used depth. The gap is decisive: geometry-first reaches 13.0 3D HOTA (51.6 LocA), whereas pseudo-LiDAR collapses to 0.12 (9.2 LocA). We trace the collapse to cross-view inconsistency of monocular depth---scale correction is necessary but not sufficient---which domain-adaptation fine-tuning does not repair within budget. Within the geometry pipeline, offline stitching is the only intervention that helps; SAHI detection, appearance Re-ID, learned lifting, RT-DETR ensembling, test-time augmentation, and domain randomization all fail to beat the baseline detector. The bottlenecks are complementary: detection quality bounds the geometry route (DetA), localization consistency bounds pseudo-LiDAR (LocA). We release a complete, reproducible RGB-only pipeline and ablation.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real".
Jane: The paper was written by Abdullah Naeem, Ayon Dey, Anav Katwal, Md Tamjidul Hoque and Noman Khan from LSU New Orleans and PinPark, Inc..
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. Tom here, and I've got Jane with me. We're looking at a paper that just hit arXiv, and the title alone tells you where this is going: "Geometry Beats Estimated Depth: RGB-Only Multi-Camera three dee Tracking under Sim2Real."
Jane: Tom, I love this title because it's basically a thesis statement. The authors are saying, look, everyone's been trying to estimate depth from a single camera to do three dee tracking, and they're arguing that old-school geometry — knowing where the cameras are and how they're calibrated — beats that approach hands down.
Tom: And the numbers back it up hard. We're talking thirteen point zero four HOTA for the geometry approach versus zero point one two for the depth-estimation approach. That's not a small gap, that's two orders of magnitude.
Jane: Right, and HOTA is the metric that balances detection, association, and localization accuracy in multi-object tracking. So a score of thirteen versus basically zero means the depth-based approach just completely fell apart.
Tom: Now, this is for the AI City Challenge two thousand twenty-six Track one. The setup is warehouse cameras, multiple synchronized views, tracking people and forklifts and robots. And the twist is that the training data is synthetic, but the test data includes real scenes. That's the Sim2Real part.
Jane: So the models train on simulated warehouses and then have to work on real ones. And depth maps are available during training, but at inference time, you only get RGB images. That's why the authors had to choose between reconstructing depth from the images or using the camera geometry directly.
Tom: And their choice was geometry. They detect objects in 2D, use the known camera calibration to project those detections onto the warehouse floor in world coordinates, and then track in that three dee world space.
Jane: It's almost like the old-school approach, right? Before deep learning, people did multi-camera tracking with homographies and ground-plane assumptions. The authors are saying that foundation is still more reliable than the fancy monocular depth models.
Tom: And that's the provocative part. The deep learning community has spent years pushing monocular depth estimation forward, and this paper says, for this specific task, under domain shift, the calibrated geometry wins.
Jane: I think the key insight is that cross-view consistency matters more than per-image accuracy. When you estimate depth from each camera independently, the depths don't agree across cameras. But when you use the calibration, every camera projects onto the same ground plane, so they have to agree.
Tom: That's the hypothesis they're testing, and we're going to dig into how they actually ran that experiment. But first, Jane, what do you think the impact of this is beyond the challenge itself?
Jane: I think it's a reality check for the field. It says, don't throw away your camera calibration just because you have a neural network. For industrial applications, warehouses, security, robotics — where cameras are fixed and calibrated — geometry is cheap, reliable, and doesn't need massive training data.
Tom: And that's exactly where we're headed next. We're going to look at the summary of the paper and how they structured this comparison. Stay with us.
Summary: Jane: Back with the paper "Geometry Beats Estimated Depth: RGB-Only Multi-Camera three dee Tracking under Sim2Real." Tom, we've set the stage — let's talk about what the authors actually built.
Tom: So they built two complete pipelines. The first is the geometry-first approach: YOLO11x for 2D detection, then homography lifting to project the bottom-center of each bounding box onto the warehouse floor plane. They use class-level size priors for the three dee box dimensions, fuse detections across cameras, and track in world coordinates.
Jane: And the second pipeline is the pseudo-LiDAR approach, which is what previous winners of this challenge used when they had depth maps provided. You estimate depth from each camera, back-project into a point cloud, fuse across cameras, and run a three dee detector on that point cloud.
Tom: The authors used two different monocular depth models — D4RT and Metricthree dee v2 — and a transformer-based three dee detector called V-DETR. And they even fine-tuned the detector on estimated-depth clouds to reduce the domain gap.
Jane: And the result was a complete collapse. zero point one two HOTA versus thirteen point zero four for the geometry approach. The localization accuracy dropped from fifty-one point six to nine point two. That's the LocA component, which measures how well the predicted three dee boxes overlap the ground truth.
Tom: So the boxes were just in the wrong places. And the authors diagnosed why: the monocular depth estimates were cross-view inconsistent. Each camera reconstructed the scene slightly differently, so when you fused the point clouds, you got warped, fragmented geometry.
Jane: They even built a diagnostic for this. They call it floor coherence. Since the warehouse floor is flat and at a known height, they measured how many points in the fused cloud landed within thirty centimeters of the floor, and how many ended up below it.
Tom: For the provided depth maps, thirty-nine percent of points were in that band and zero percent were below the floor. For the estimated depth from D4RT, only twenty percent were in the band and twenty percent were below the floor. Metricthree dee didn't reconstruct the floor at all.
Jane: That's the smoking gun. If the floor is warped, then everything resting on the floor — people, forklifts, robots — is misplaced. The geometry approach never has this problem because the homography forces everything onto the calibrated ground plane.
Tom: And here's the kicker: the authors say scale correction wasn't enough. D4RT needed a four point three times correction to reach metric scale, but even after that, the cross-view disagreement remained. So it's not just about getting the depth magnitude right, it's about consistency across cameras.
Jane: I think that's the most important scientific contribution here. The field has been focused on per-image depth accuracy, but this paper shows that cross-view consistency is the property that actually matters for multi-camera three dee perception.
Tom: And that reframes the whole problem. If you're building a multi-camera system, you should be optimizing for agreement between cameras, not for matching some ground-truth depth map per image.
Jane: Exactly. And that's a design principle that could change how people approach this. But the paper also has a lot to say about what didn't work within the geometry pipeline itself. That's coming up next.
Improvements: Tom: We're back with "Geometry Beats Estimated Depth: RGB-Only Multi-Camera three dee Tracking under Sim2Real." Jane, the authors didn't just compare two pipelines — they ran a whole ablation study on the geometry approach.
Jane: And the results are almost as interesting as the main comparison. They tried a bunch of interventions, and most of them made things worse. Only one thing actually helped.
Tom: That one thing is offline tracklet stitching. The online tracker fragments identities when an object disappears for a while and reappears. Stitching relinks those fragments after the fact, using position and timing. It raised the association score from fourteen point five three to sixteen point seven one and HOTA from twelve point four nine to thirteen point zero four.
Jane: And it's a pure association fix. It doesn't touch detection or localization, it just relabels identities. The authors say it's the only intervention that helped.
Tom: Meanwhile, everything else failed. SAHI sliced detection, which tiles the image to find small objects, more than doubled the detection count but lowered HOTA. The problem was false positives — it traded precision for recall and flooded the system with noise.
Jane: Appearance-based Re-ID also failed. The idea was to use object crops to match identities across time, but in real warehouse scenes, the crops are low-resolution and self-similar. Everyone's wearing similar clothes, the lighting is bad. Geometry was more reliable than appearance.
Tom: They also tried replacing the calibrated homography with a learned MLP that maps bounding boxes to three dee coordinates. That was catastrophic — zero point nine four HOTA. The learned lift just doesn't have the metric grounding that calibration provides.
Jane: And they tried ensembling YOLO with RT-DETR, test-time augmentation, domain-randomized training with YOLO26 — all of it either didn't help or made things worse.
Tom: So the pattern is clear. The geometry route is bounded by detection quality. The authors say DetA — detection accuracy — is the ceiling. Association is not the bottleneck, and localization is relatively stable. It's the detections that are limiting everything.
Jane: And that's a really clean diagnosis. If you want to improve this system, you don't work on tracking or fusion or depth. You work on making the 2D detector better at finding small, distant, occluded objects in real warehouse scenes.
Tom: The validation numbers support that. Recall for PalletTruck is only zero point two three nine. Forklift is zero point six zero one. NovaCarter is zero point eight zero six. So the classes with low recall — the low-profile, self-similar vehicles — are dragging down the whole system.
Jane: And the synthetic-to-real gap is the culprit. On the training split, recall was above zero point eight seven for all classes. On validation, it drops to zero point six zero three overall. The detector learns the synthetic appearance but doesn't transfer to real scenes.
Tom: So the highest-value next step, according to the authors, is closing that Sim2Real detection gap. Not building better depth estimators, not fancier trackers — just making the detector more robust to real-world appearance.
Jane: And that's a useful message for the community. It's easy to reach for the newest three dee perception technique, but sometimes the bottleneck is the boring 2D detector.
Tom: Speaking of boring 2D detectors, we should talk about the first page of the paper, where they lay out the problem and the hypothesis. That's next.
First Page: Jane: We're still on "Geometry Beats Estimated Depth: RGB-Only Multi-Camera three dee Tracking under Sim2Real." Tom, let's go back to the very beginning of the paper, because the framing is important.
Tom: The first page sets up the challenge: multi-camera three dee perception in large indoor warehouses. You have synchronized cameras, you need to detect and track people, forklifts, mobile robots, humanoids, transporters, and pallet trucks. The output is a single file with world-coordinate three dee boxes and identities for every frame.
Jane: And the difficulty is that training data is synthetic, while the hidden test set includes real videos with visual stressors. Plus, depth maps are only available for training and validation. At inference, you're RGB-only.
Tom: The authors state their central hypothesis right there: cross-view geometric consistency matters more than per-image depth accuracy. The geometry-first lift maximizes consistency because all cameras share one calibrated ground plane. Estimated-depth pseudo-LiDAR maximizes per-image detail but sacrifices consistency.
Jane: And they're very explicit about what they're testing. Both routes address the same task, same data, same metric. The only difference is whether you preserve cross-view consistency or sacrifice it for per-image depth accuracy.
Tom: I like how they frame the contributions. They're not just presenting a system — they're presenting scientific findings. The geometry-versus-depth comparison, the floor-coherence diagnostic, the ablation principle, and a reproducible baseline.
Jane: The floor-coherence diagnostic is clever because it's annotation-free. You don't need ground-truth depth to measure it. You just need to know where the floor is in the world coordinate system. That's a practical tool anyone can use to check whether their multi-camera depth fusion is working.
Tom: And it separates scale error from consistency error. You can have the right scale but still have inconsistent geometry across views. The diagnostic lets you see which problem you have.
Jane: The authors also mention that they release the full pipeline and ablation. That's valuable for the community — a reference point for Sim2Real three dee perception that others can build on.
Tom: One thing that struck me on the first page is the list of classes. Person, Forklift, NovaCarter, Transporter, FourierGR1T2, AgilityDigit, PalletTruck. These are real industrial vehicles and robots. This is applied research for actual warehouses.
Jane: And that's the impact. If this works, you can deploy it in real distribution centers, manufacturing plants, logistics hubs. You can track workers and vehicles for safety, efficiency, automation.
Tom: But the paper is honest about the current state. thirteen HOTA is not a solved problem. The authors are clear that detection quality is the bottleneck and that more work is needed.
Jane: Still, the scientific contribution stands. They've shown that geometry beats estimated depth for this task, and they've given the community a diagnostic to understand why. That's a solid foundation.
Tom: And we're about to wrap up with our final thoughts. Lu and Meng have been listening, and I think they have some perspectives to add.
Conclusion: Tom: Alright, we're closing out our discussion of "Geometry Beats Estimated Depth: RGB-Only Multi-Camera three dee Tracking under Sim2Real." Jane, let's bring in Lu and Meng for final thoughts.
Jane: Lu, you've been quiet. What's your take on the big picture here?
Lu: I think the most exciting implication is that this reframes the research agenda for multi-camera three dee perception. Instead of chasing better monocular depth models, we should be building systems that explicitly enforce cross-view consistency. The floor-coherence metric is a step in that direction, but there's room for much more sophisticated consistency constraints.
Meng: From an engineering standpoint, I appreciate that the geometry approach is computationally cheap. You're running a 2D detector and doing matrix math. No heavy depth estimation, no point cloud processing. That means it can run in real time on modest hardware, which matters for actual deployment in warehouses.
Tom: And that's a real advantage. The pseudo-LiDAR route requires running a depth model per camera, back-projecting, fusing, and then running a three dee detector. That's a lot of compute. The geometry route is much lighter.
Jane: But Meng, you're an engineer — what would you want to see before deploying this in a real warehouse?
Meng: I'd want to see the detection quality improved. The paper is clear that DetA is the bottleneck. I'd invest in collecting real warehouse data, fine-tuning the detector, maybe using semi-supervised learning to leverage unlabeled real footage. That's where the return on investment is.
Lu: And I'd add that the cross-view consistency idea could be pushed further. Instead of just using homographies, you could train a network to predict geometry that is explicitly consistent across views, using the calibration as a hard constraint. That's a promising direction that this paper opens up.
Tom: So the future work is clear: better detection, and consistency-aware geometry learning. And the paper gives us the tools to measure progress.
Jane: I think the lasting contribution is the controlled comparison. The authors didn't just show that one approach works better — they explained why, with a diagnostic that isolates the cause. That's how science should be done.
Tom: And for anyone working on multi-camera three dee tracking, especially under domain shift, this paper is a must-read. It might save you from going down the pseudo-LiDAR path and wasting months.
Jane: Agreed. We'll be watching for follow-up work on the detection side. Thanks to Lu and Meng for joining us, and to our listeners for tuning in.
Tom: That's it for "Geometry Beats Estimated Depth: RGB-Only Multi-Camera three dee Tracking under Sim2Real." Next up, we've got a paper on efficient video transformers that I'm really excited about. See you then.
Jane: Take care, everyone.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language