ETH Zürich · Computer Vision & Geometry Group · Fall 2025

MOBIUS goes 3D

Efficient monocular 3D object detection via depth-guided transformers and dense depth distillation.

ContextSemester Project, Fall 2025
PartnersGoogle Research × ETH CVG
SupervisionProf. Dr. Marc Pollefeys
BenchmarkKITTI 3D Object Detection
Python PyTorch Transformers 3D Vision
01

Qualitative Results

KITTI predictions: 2D overlays with IoU scores, and LiDAR point clouds with ground truth (green) and predictions (red). Drag to orbit, scroll to zoom.

KITTI 2D detection panel

Loading 3D Scene…

Toggles
02

Overview

Monocular 3D detection predicts object position, size and orientation from a single image, which is inherently ambiguous in depth. 2D foundation models like MOBIUS excel at open-vocabulary recognition but lack the geometric reasoning 3D requires.

With Google Research and ETH CVG (Prof. Marc Pollefeys), I built MOBIUS-3D, which lifts MOBIUS into 3D with a lightweight depth predictor and a depth-guided transformer decoder. The key is dense depth distillation from the DepthPro foundation model, enabling accurate 3D detection without dense LiDAR supervision.

The Challenge

2D models train on millions of images across thousands of categories. 3D datasets are small, expensive to annotate and cover a handful of classes, so 3D detectors are heavy, specialized models trained from scratch that miss the semantic knowledge of 2D foundation models.

The goal: lift a lightweight 2D model into 3D while keeping its efficiency, and close the annotation gap by distilling geometry from a depth foundation model.

03

Method

MOBIUS-3D adds three components to the frozen MOBIUS backbone: a depth predictor, a depth encoder, and a depth-guided transformer decoder.

1 · Depth Predictor

A DPT-style head fuses multi-scale features into a dense depth map, distilled from DepthPro so every pixel gets supervision, not just sparse LiDAR points.

2 · Depth Encoder

Deformable attention turns features into depth-aware tokens, sampling a few points per query instead of the full map: O(N·K) instead of O(N²).

3 · Depth-Guided Decoder

Object queries attend to 3D positional embeddings, built by projecting pixels into metric space with the predicted depth and camera intrinsics.

Key Insight · Dense Depth Distillation

Monocular detectors usually supervise depth with sparse LiDAR (< 5 % of pixels). We distill from a frozen DepthPro teacher instead, aligned to metric scale with a least-squares scale + shift fit to LiDAR. This single change gives the largest gain in the ablation.

04

Results

KITTI 3D (Car, IoU = 0.7). MOBIUS-3D has the best APBEV, the main 3D localization metric, at every difficulty level.

Method AP3D (%) APBEV (%)
Easy Mod. Hard Easy Mod. Hard
MonoDETR 28.05 20.76 16.76 37.60 27.28 23.38
MonoDETRNext-A 32.95 25.01 21.92 - - -
MOBIUS-3D (Ours) 31.92 24.22 22.58 41.83 31.64 30.04

Ablation: Depth Supervision Strategy

Depth supervision is the most impactful design choice, and dense distillation wins across the board.

Depth Supervision AP3D (%) APBEV (%)
Easy Mod. Hard Easy Mod. Hard
None 19.39 12.16 10.31 31.79 22.12 19.30
Object-Level Only 17.35 11.10 10.25 27.52 18.70 17.86
Sparse LiDAR 23.15 16.69 15.54 31.42 23.93 22.15
Dense Distillation (Ours) 31.92 24.22 22.58 41.83 31.64 30.04