A DPT-style head fuses multi-scale features into a dense depth map, distilled from DepthPro so every pixel gets supervision, not just sparse LiDAR points.
Qualitative Results
KITTI predictions: 2D overlays with IoU scores, and LiDAR point clouds with ground truth (green) and predictions (red). Drag to orbit, scroll to zoom.
Loading 3D Scene…
Overview
Monocular 3D detection predicts object position, size and orientation from a single image, which is inherently ambiguous in depth. 2D foundation models like MOBIUS excel at open-vocabulary recognition but lack the geometric reasoning 3D requires.
With Google Research and ETH CVG (Prof. Marc Pollefeys), I built MOBIUS-3D, which lifts MOBIUS into 3D with a lightweight depth predictor and a depth-guided transformer decoder. The key is dense depth distillation from the DepthPro foundation model, enabling accurate 3D detection without dense LiDAR supervision.
The Challenge
2D models train on millions of images across thousands of categories. 3D datasets are small, expensive to annotate and cover a handful of classes, so 3D detectors are heavy, specialized models trained from scratch that miss the semantic knowledge of 2D foundation models.
The goal: lift a lightweight 2D model into 3D while keeping its efficiency, and close the annotation gap by distilling geometry from a depth foundation model.
Method
MOBIUS-3D adds three components to the frozen MOBIUS backbone: a depth predictor, a depth encoder, and a depth-guided transformer decoder.
Deformable attention turns features into depth-aware tokens, sampling a few points per query instead of the full map: O(N·K) instead of O(N²).
Object queries attend to 3D positional embeddings, built by projecting pixels into metric space with the predicted depth and camera intrinsics.
Monocular detectors usually supervise depth with sparse LiDAR (< 5 % of pixels). We distill from a frozen DepthPro teacher instead, aligned to metric scale with a least-squares scale + shift fit to LiDAR. This single change gives the largest gain in the ablation.
Results
KITTI 3D (Car, IoU = 0.7). MOBIUS-3D has the best APBEV, the main 3D localization metric, at every difficulty level.
| Method | AP3D (%) | APBEV (%) | ||||
|---|---|---|---|---|---|---|
| Easy | Mod. | Hard | Easy | Mod. | Hard | |
| MonoDETR | 28.05 | 20.76 | 16.76 | 37.60 | 27.28 | 23.38 |
| MonoDETRNext-A | 32.95 | 25.01 | 21.92 | - | - | - |
| MOBIUS-3D (Ours) | 31.92 | 24.22 | 22.58 | 41.83 | 31.64 | 30.04 |
Ablation: Depth Supervision Strategy
Depth supervision is the most impactful design choice, and dense distillation wins across the board.
| Depth Supervision | AP3D (%) | APBEV (%) | ||||
|---|---|---|---|---|---|---|
| Easy | Mod. | Hard | Easy | Mod. | Hard | |
| None | 19.39 | 12.16 | 10.31 | 31.79 | 22.12 | 19.30 |
| Object-Level Only | 17.35 | 11.10 | 10.25 | 27.52 | 18.70 | 17.86 |
| Sparse LiDAR | 23.15 | 16.69 | 15.54 | 31.42 | 23.93 | 22.15 |
| Dense Distillation (Ours) | 31.92 | 24.22 | 22.58 | 41.83 | 31.64 | 30.04 |