Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real

arXiv:2608.07579 · cs.CV, cs.AI · Submitted 2026-08-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real".

Jane: The paper was written by Abdullah Naeem, Ayon Dey, Anav Katwal, Md Tamjidul Hoque and Noman Khan from LSU New Orleans and PinPark, Inc..

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. Tom here, and I've got Jane with me. We're looking at a paper that just hit arXiv, and the title alone tells you where this is going: "Geometry Beats Estimated Depth: RGB-Only Multi-Camera three dee Tracking under Sim2Real."

Jane: Tom, I love this title because it's basically a thesis statement. The authors are saying, look, everyone's been trying to estimate depth from a single camera to do three dee tracking, and they're arguing that old-school geometry — knowing where the cameras are and how they're calibrated — beats that approach hands down.

Tom: And the numbers back it up hard. We're talking thirteen point zero four HOTA for the geometry approach versus zero point one two for the depth-estimation approach. That's not a small gap, that's two orders of magnitude.

Jane: Right, and HOTA is the metric that balances detection, association, and localization accuracy in multi-object tracking. So a score of thirteen versus basically zero means the depth-based approach just completely fell apart.

Tom: Now, this is for the AI City Challenge two thousand twenty-six Track one. The setup is warehouse cameras, multiple synchronized views, tracking people and forklifts and robots. And the twist is that the training data is synthetic, but the test data includes real scenes. That's the Sim2Real part.

Jane: So the models train on simulated warehouses and then have to work on real ones. And depth maps are available during training, but at inference time, you only get RGB images. That's why the authors had to choose between reconstructing depth from the images or using the camera geometry directly.

Tom: And their choice was geometry. They detect objects in 2D, use the known camera calibration to project those detections onto the warehouse floor in world coordinates, and then track in that three dee world space.

Jane: It's almost like the old-school approach, right? Before deep learning, people did multi-camera tracking with homographies and ground-plane assumptions. The authors are saying that foundation is still more reliable than the fancy monocular depth models.

Tom: And that's the provocative part. The deep learning community has spent years pushing monocular depth estimation forward, and this paper says, for this specific task, under domain shift, the calibrated geometry wins.

Jane: I think the key insight is that cross-view consistency matters more than per-image accuracy. When you estimate depth from each camera independently, the depths don't agree across cameras. But when you use the calibration, every camera projects onto the same ground plane, so they have to agree.

Tom: That's the hypothesis they're testing, and we're going to dig into how they actually ran that experiment. But first, Jane, what do you think the impact of this is beyond the challenge itself?

Jane: I think it's a reality check for the field. It says, don't throw away your camera calibration just because you have a neural network. For industrial applications, warehouses, security, robotics — where cameras are fixed and calibrated — geometry is cheap, reliable, and doesn't need massive training data.

Tom: And that's exactly where we're headed next. We're going to look at the summary of the paper and how they structured this comparison. Stay with us.

Summary: Jane: Back with the paper "Geometry Beats Estimated Depth: RGB-Only Multi-Camera three dee Tracking under Sim2Real." Tom, we've set the stage — let's talk about what the authors actually built.

Tom: So they built two complete pipelines. The first is the geometry-first approach: YOLO11x for 2D detection, then homography lifting to project the bottom-center of each bounding box onto the warehouse floor plane. They use class-level size priors for the three dee box dimensions, fuse detections across cameras, and track in world coordinates.

Jane: And the second pipeline is the pseudo-LiDAR approach, which is what previous winners of this challenge used when they had depth maps provided. You estimate depth from each camera, back-project into a point cloud, fuse across cameras, and run a three dee detector on that point cloud.

Tom: The authors used two different monocular depth models — D4RT and Metricthree dee v2 — and a transformer-based three dee detector called V-DETR. And they even fine-tuned the detector on estimated-depth clouds to reduce the domain gap.

Jane: And the result was a complete collapse. zero point one two HOTA versus thirteen point zero four for the geometry approach. The localization accuracy dropped from fifty-one point six to nine point two. That's the LocA component, which measures how well the predicted three dee boxes overlap the ground truth.

Tom: So the boxes were just in the wrong places. And the authors diagnosed why: the monocular depth estimates were cross-view inconsistent. Each camera reconstructed the scene slightly differently, so when you fused the point clouds, you got warped, fragmented geometry.

Jane: They even built a diagnostic for this. They call it floor coherence. Since the warehouse floor is flat and at a known height, they measured how many points in the fused cloud landed within thirty centimeters of the floor, and how many ended up below it.

Tom: For the provided depth maps, thirty-nine percent of points were in that band and zero percent were below the floor. For the estimated depth from D4RT, only twenty percent were in the band and twenty percent were below the floor. Metricthree dee didn't reconstruct the floor at all.

Jane: That's the smoking gun. If the floor is warped, then everything resting on the floor — people, forklifts, robots — is misplaced. The geometry approach never has this problem because the homography forces everything onto the calibrated ground plane.

Tom: And here's the kicker: the authors say scale correction wasn't enough. D4RT needed a four point three times correction to reach metric scale, but even after that, the cross-view disagreement remained. So it's not just about getting the depth magnitude right, it's about consistency across cameras.

Jane: I think that's the most important scientific contribution here. The field has been focused on per-image depth accuracy, but this paper shows that cross-view consistency is the property that actually matters for multi-camera three dee perception.

Tom: And that reframes the whole problem. If you're building a multi-camera system, you should be optimizing for agreement between cameras, not for matching some ground-truth depth map per image.

Jane: Exactly. And that's a design principle that could change how people approach this. But the paper also has a lot to say about what didn't work within the geometry pipeline itself. That's coming up next.

Improvements: Tom: We're back with "Geometry Beats Estimated Depth: RGB-Only Multi-Camera three dee Tracking under Sim2Real." Jane, the authors didn't just compare two pipelines — they ran a whole ablation study on the geometry approach.

Jane: And the results are almost as interesting as the main comparison. They tried a bunch of interventions, and most of them made things worse. Only one thing actually helped.

Tom: That one thing is offline tracklet stitching. The online tracker fragments identities when an object disappears for a while and reappears. Stitching relinks those fragments after the fact, using position and timing. It raised the association score from fourteen point five three to sixteen point seven one and HOTA from twelve point four nine to thirteen point zero four.

Jane: And it's a pure association fix. It doesn't touch detection or localization, it just relabels identities. The authors say it's the only intervention that helped.

Tom: Meanwhile, everything else failed. SAHI sliced detection, which tiles the image to find small objects, more than doubled the detection count but lowered HOTA. The problem was false positives — it traded precision for recall and flooded the system with noise.

Jane: Appearance-based Re-ID also failed. The idea was to use object crops to match identities across time, but in real warehouse scenes, the crops are low-resolution and self-similar. Everyone's wearing similar clothes, the lighting is bad. Geometry was more reliable than appearance.

Tom: They also tried replacing the calibrated homography with a learned MLP that maps bounding boxes to three dee coordinates. That was catastrophic — zero point nine four HOTA. The learned lift just doesn't have the metric grounding that calibration provides.

Jane: And they tried ensembling YOLO with RT-DETR, test-time augmentation, domain-randomized training with YOLO26 — all of it either didn't help or made things worse.

Tom: So the pattern is clear. The geometry route is bounded by detection quality. The authors say DetA — detection accuracy — is the ceiling. Association is not the bottleneck, and localization is relatively stable. It's the detections that are limiting everything.

Jane: And that's a really clean diagnosis. If you want to improve this system, you don't work on tracking or fusion or depth. You work on making the 2D detector better at finding small, distant, occluded objects in real warehouse scenes.

Tom: The validation numbers support that. Recall for PalletTruck is only zero point two three nine. Forklift is zero point six zero one. NovaCarter is zero point eight zero six. So the classes with low recall — the low-profile, self-similar vehicles — are dragging down the whole system.

Jane: And the synthetic-to-real gap is the culprit. On the training split, recall was above zero point eight seven for all classes. On validation, it drops to zero point six zero three overall. The detector learns the synthetic appearance but doesn't transfer to real scenes.

Tom: So the highest-value next step, according to the authors, is closing that Sim2Real detection gap. Not building better depth estimators, not fancier trackers — just making the detector more robust to real-world appearance.

Jane: And that's a useful message for the community. It's easy to reach for the newest three dee perception technique, but sometimes the bottleneck is the boring 2D detector.

Tom: Speaking of boring 2D detectors, we should talk about the first page of the paper, where they lay out the problem and the hypothesis. That's next.

First Page: Jane: We're still on "Geometry Beats Estimated Depth: RGB-Only Multi-Camera three dee Tracking under Sim2Real." Tom, let's go back to the very beginning of the paper, because the framing is important.

Tom: The first page sets up the challenge: multi-camera three dee perception in large indoor warehouses. You have synchronized cameras, you need to detect and track people, forklifts, mobile robots, humanoids, transporters, and pallet trucks. The output is a single file with world-coordinate three dee boxes and identities for every frame.

Jane: And the difficulty is that training data is synthetic, while the hidden test set includes real videos with visual stressors. Plus, depth maps are only available for training and validation. At inference, you're RGB-only.

Tom: The authors state their central hypothesis right there: cross-view geometric consistency matters more than per-image depth accuracy. The geometry-first lift maximizes consistency because all cameras share one calibrated ground plane. Estimated-depth pseudo-LiDAR maximizes per-image detail but sacrifices consistency.

Jane: And they're very explicit about what they're testing. Both routes address the same task, same data, same metric. The only difference is whether you preserve cross-view consistency or sacrifice it for per-image depth accuracy.

Tom: I like how they frame the contributions. They're not just presenting a system — they're presenting scientific findings. The geometry-versus-depth comparison, the floor-coherence diagnostic, the ablation principle, and a reproducible baseline.

Jane: The floor-coherence diagnostic is clever because it's annotation-free. You don't need ground-truth depth to measure it. You just need to know where the floor is in the world coordinate system. That's a practical tool anyone can use to check whether their multi-camera depth fusion is working.

Tom: And it separates scale error from consistency error. You can have the right scale but still have inconsistent geometry across views. The diagnostic lets you see which problem you have.

Jane: The authors also mention that they release the full pipeline and ablation. That's valuable for the community — a reference point for Sim2Real three dee perception that others can build on.

Tom: One thing that struck me on the first page is the list of classes. Person, Forklift, NovaCarter, Transporter, FourierGR1T2, AgilityDigit, PalletTruck. These are real industrial vehicles and robots. This is applied research for actual warehouses.

Jane: And that's the impact. If this works, you can deploy it in real distribution centers, manufacturing plants, logistics hubs. You can track workers and vehicles for safety, efficiency, automation.

Tom: But the paper is honest about the current state. thirteen HOTA is not a solved problem. The authors are clear that detection quality is the bottleneck and that more work is needed.

Jane: Still, the scientific contribution stands. They've shown that geometry beats estimated depth for this task, and they've given the community a diagnostic to understand why. That's a solid foundation.

Tom: And we're about to wrap up with our final thoughts. Lu and Meng have been listening, and I think they have some perspectives to add.

Conclusion: Tom: Alright, we're closing out our discussion of "Geometry Beats Estimated Depth: RGB-Only Multi-Camera three dee Tracking under Sim2Real." Jane, let's bring in Lu and Meng for final thoughts.

Jane: Lu, you've been quiet. What's your take on the big picture here?

Lu: I think the most exciting implication is that this reframes the research agenda for multi-camera three dee perception. Instead of chasing better monocular depth models, we should be building systems that explicitly enforce cross-view consistency. The floor-coherence metric is a step in that direction, but there's room for much more sophisticated consistency constraints.

Meng: From an engineering standpoint, I appreciate that the geometry approach is computationally cheap. You're running a 2D detector and doing matrix math. No heavy depth estimation, no point cloud processing. That means it can run in real time on modest hardware, which matters for actual deployment in warehouses.

Tom: And that's a real advantage. The pseudo-LiDAR route requires running a depth model per camera, back-projecting, fusing, and then running a three dee detector. That's a lot of compute. The geometry route is much lighter.

Jane: But Meng, you're an engineer — what would you want to see before deploying this in a real warehouse?

Meng: I'd want to see the detection quality improved. The paper is clear that DetA is the bottleneck. I'd invest in collecting real warehouse data, fine-tuning the detector, maybe using semi-supervised learning to leverage unlabeled real footage. That's where the return on investment is.

Lu: And I'd add that the cross-view consistency idea could be pushed further. Instead of just using homographies, you could train a network to predict geometry that is explicitly consistent across views, using the calibration as a hard constraint. That's a promising direction that this paper opens up.

Tom: So the future work is clear: better detection, and consistency-aware geometry learning. And the paper gives us the tools to measure progress.

Jane: I think the lasting contribution is the controlled comparison. The authors didn't just show that one approach works better — they explained why, with a diagnostic that isolates the cause. That's how science should be done.

Tom: And for anyone working on multi-camera three dee tracking, especially under domain shift, this paper is a must-read. It might save you from going down the pseudo-LiDAR path and wasting months.

Jane: Agreed. We'll be watching for follow-up work on the detection side. Thanks to Lu and Meng for joining us, and to our listeners for tuning in.

Tom: That's it for "Geometry Beats Estimated Depth: RGB-Only Multi-Camera three dee Tracking under Sim2Real." Next up, we've got a paper on efficient video transformers that I'm really excited about. See you then.

Jane: Take care, everyone.

Abdullah Naeem, Ayon Dey, Anav Katwal, Md Tamjidul Hoque, Noman Khan

LSU New Orleans · PinPark, Inc.

cs.CV, cs.AI

Submitted: 2026-08-04

Updated: 2026-08-11

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 63/100

Key concepts

Geometry Beats Estimated Depth
This refers to using known camera positions and calibrations (geometry) to project 2D detections into world coordinates, rather than estimating depth from each camera independently. The paper shows geometry provides superior consistency across multiple views.
HOTA
HOTA is a metric used to balance detection, association, and localization accuracy in multi-object tracking tasks. A higher HOTA score indicates better overall performance in these areas.
Floor Coherence Diagnostic
This is an annotation-free tool used to check the consistency of fused depth data. It measures how many points in the fused point cloud land within a certain distance of the known warehouse floor plane, helping identify geometric warping.

Terminology

Summary

Summary

This paper, Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real, addresses the AI City Challenge 2026 Track 1, which evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting. The challenge requires detecting and tracking people and mobile objects—including forklifts, mobile robots, humanoids, transporters, and pallet trucks—across synchronized warehouse cameras, outputting world-coordinate 3D bounding boxes and object identities per frame. A key constraint is that depth is available only for training and validation, so inference is RGB-only.

The authors frame their work as a controlled test of one central hypothesis: that cross-view geometric consistency, not monocular depth accuracy, governs performance under Sim2Real. They compare two RGB-only routes. The first is a geometry-first pipeline: YOLO11x detection, homography lifting to the world frame, class-level 3D size priors, multi-camera fusion, world-coordinate tracking, and offline tracklet stitching. The second is estimated-depth pseudo-LiDAR: monocular depth (D4RT, Metric3D v2) back-projected into a fused point cloud and passed to a 3D detector (V-DETR), mirroring prior point-cloud winners that used provided depth.

The main result is decisive: geometry-first reaches 13.0 3D HOTA (51.6 LocA), whereas pseudo-LiDAR collapses to 0.12 (9.2 LocA). The authors trace the collapse to cross-view inconsistency of monocular depth—scale correction is necessary but not sufficient—which domain-adaptation fine-tuning does not repair within budget.

Within the geometry pipeline, the authors report that offline stitching is the only intervention that helps; SAHI detection, appearance Re-ID, learned lifting, RT-DETR ensembling, test-time augmentation, and domain randomization all fail to beat the baseline detector. They identify complementary bottlenecks: detection quality bounds the geometry route (DetA), localization consistency bounds pseudo-LiDAR (LocA).

The paper's contributions are: (1) a systematic geometry-vs-depth comparison showing explicit cross-view geometry outperforms learned monocular depth by roughly two orders of magnitude (13.0 vs 0.12 HOTA); (2) a geometric-consistency diagnostic—an annotation-free floor-coherence metric that quantifies cross-view geometric consistency, separates scale error from consistency error, and predicts the pseudo-LiDAR collapse before any detector is run; (3) a design principle from ablation grouping interventions into detector-side, geometry-side, and association-side categories, concluding "the geometry route is bounded by detection quality, so only association-side relinking (offline stitching) helps, while adding detections (detector-side) or replacing calibrated geometry with learned substitutes (geometry-side) hurts"; and (4) a reproducible baseline pipeline with full ablation.

The method details are as follows. For 2D detection, the authors trained an Ultralytics YOLO11x detector at 1280 input resolution, using a low confidence threshold during inference to retain recall. For world-frame lifting, each 2D detection's bottom-center point is projected into the ground/world plane using camera calibration and homography, with the configuration invert-homography –footpoint-y 1.00 –world-plane xy producing a median validation lift error of approximately 0.633 m. Class-level priors assign vertical coordinate and 3D box dimensions, with yaw estimated heuristically. Multi-camera fusion merges detections within a class-specific spatial radius per frame and class. World-coordinate tracking is performed separately per class, associating fused detections to active tracks using world-coordinate distance and motion prediction, with optional coasting to bridge short gaps. Offline tracklet stitching links same-class fragments when one begins shortly after another ends within a gap budget and within a distance threshold of the velocity-extrapolated endpoint; this raises AssA from 14.53 to 16.71 at essentially unchanged DetA and LocA.

Detector validation metrics show high precision (0.907 overall) but limited recall (0.603 overall), with PalletTruck particularly weak (recall 0.239, mAP50 0.309). On the training split, class-wise recall was above 0.87 for all classes and above 0.98 for PalletTruck, indicating the model can learn categories but generalizes poorly to validation/test appearance due to domain shift, small or distant objects, camera artifacts, illumination changes, and low object-background contrast.

Leaderboard results show the best configuration is YOLO11x + geometry baseline with offline tracklet stitching: HOTA 13.0413, DetA 10.7897, AssA 16.7105, LocA 51.5785. Rejected interventions include a Light Re-ID variant (HOTA 10.5120), low-confidence recall bump (12.0249), learned lift MLP (0.9418), YOLO+RT-DETR ensemble (10.8819), YOLO11x TTA (11.1725), SAHI sliced detection (11.4943), and YOLO26 with stress augmentation (10).

Qualitative analysis using video overlays shows that on synthetic scenes the world boxes over-detect with mis-placed ground-plane bases (lift error), while on real scenes the 2D detector finds people but detections fail to propagate into world tracks, causing under-coverage. The score decomposition shows LocA is relatively stable near 51, AssA is higher than DetA, and the main limitation is the availability and quality of scene-level detections. On a validation frame, predictions far outnumber ground truth (Pred=120 vs GT=67; 74 false positives against 46 true positives), bounding DetA. The synthetic-to-real gap is evident: detection recall is markedly lower on real scenes, contributing disproportionately to missed detections.

The pseudo-LiDAR negative result is detailed. Monocular depth is not metric out of the box: D4RT required a near-constant 4.3× correction, while Metric3D v2 is metric by construction but places clouds inconsistently with its floor floating above the ground plane. Even after scale correction, monocular depth disagrees across views, yielding warped, non-planar geometry. The floor-coherence proxy measures the share of points within ±0.3 m of the ground plane and below it: provided-depth clouds place 39% of points in-band with 0% below, scale-corrected D4RT places only 20% in-band with 20% below, and Metric3D leaves the floor essentially unreconstructed. The provided-depth-trained V-DETR applied to estimated-depth clouds scored 0.12 HOTA with 9.2 LocA, and domain-adaptation fine-tuning did not yield usable detections within budget. The authors interpret this as a failure of geometric consistency, not detection: LocA collapses from 51.6 to 9.2 because fused monocular depth does not agree across cameras.

The conclusion states: We conclude that, absent inference-time depth, explicit geometry is the more reliable foundation, and that closing the Sim2Real detection-quality gap—not learned monocular depth—is the highest-value next step.

Improvements for AI systems

Based on the paper’s findings, here are the specific improvements I can implement in an AI system for RGB-only multi-camera 3D tracking under Sim2Real conditions:

Improvement: Replace any learned monocular-depth or learned-lifting modules with a hard-coded, calibration-based homography lift. The system uses YOLO11x 2D detections → bottom-center footpoint projection → class-prior 3D dimensions → multi-camera fusion in world coordinates.

What it can do: Achieves 13.04 HOTA (51.58 LocA) versus 0.12 HOTA (9.23 LocA) for pseudo-LiDAR. It maintains cross-view metric consistency by construction—all cameras share one calibrated ground plane—eliminating the dominant failure mode.


Summary of what the improved system can do: It reliably produces valid leaderboard submissions (13.04 HOTA) with stable localization (51.58 LocA), correctly identifies detection quality as the binding constraint, avoids all known failure modes (pseudo-LiDAR, learned lifting, Re-ID, SAHI, ensembling), and provides a diagnostic to reject depth-based routes before expensive 3D detection runs.

Abstract

The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting; depth is available only for training and validation, so inference is RGB-only. We use two RGB-only routes as a controlled test of one hypothesis: that cross-view geometric consistency, not monocular depth accuracy, governs performance under Sim2Real. The first is a geometry-first pipeline: YOLO11x detection, homography lifting to the world frame, class-level 3D size priors, multi-camera fusion, world-coordinate tracking, and offline tracklet stitching. The second is estimated-depth pseudo-LiDAR: monocular depth (D4RT, Metric3D v2) back-projected into a fused point cloud and passed to a 3D detector (V-DETR), mirroring prior point-cloud winners that used depth. The gap is decisive: geometry-first reaches 13.0 3D HOTA (51.6 LocA), whereas pseudo-LiDAR collapses to 0.12 (9.2 LocA). We trace the collapse to cross-view inconsistency of monocular depth---scale correction is necessary but not sufficient---which domain-adaptation fine-tuning does not repair within budget. Within the geometry pipeline, offline stitching is the only intervention that helps; SAHI detection, appearance Re-ID, learned lifting, RT-DETR ensembling, test-time augmentation, and domain randomization all fail to beat the baseline detector. The bottlenecks are complementary: detection quality bounds the geometry route (DetA), localization consistency bounds pseudo-LiDAR (LocA). We release a complete, reproducible RGB-only pipeline and ablation.

Sources

Related papers