Zoom - Scroll mouse wheel or pinch
Rotate - Left drag
Pan - Right drag
Pick a point - Left click
Self-Guided Sparse 3D Refinement (SSR). We lift the coarse point map from a base model onto a sparse voxel shell and refine the depth via lightweight sparse 3D convolutions. Operating in 3D space rather than the image plane avoids feature mixing across discontinuities between depth layers, recovering fine-detailed geometry that 2D decoders typically over-smooth.
Through iterations, each pass re-voxelizes based on the current estimate and refines itself, progressively recovering cleaner and finer thin structures.
Zoom - Scroll mouse wheel or pinch
Rotate - Left drag
Pan - Right drag
Refine step - Drag the slider below
Zoom - Scroll mouse wheel or pinch
Rotate - Left drag
Pan - Right drag
Switch baselines from the dropdown above
Zero-shot averages across multiple benchmarks. Lower Rel is better; higher δ and F1 are better.
State-of-the-art fidelity. Across 9 zero-shot benchmarks, SSR sets the best global and local accuracy, with the largest gains on strict fine-detail metrics (δ0.01). It matches Depth Pro's boundary sharpness at much lower resolution, while yielding far cleaner 3D structure.
| Method | Local | Boundary | Global | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Depth | Point | F1↑ | Relative Depth | Relative Point | Metric Depth | ||||||
| Reld↓ | δd0.01↑ | Relp↓ | δp0.01↑ | Reld↓ | δd1↑ | Relp↓ | δp1↑ | Reld↓ | δd1↑ | ||
| Depth Pro | 4.15 | 46.9 | 3.01 | 45.2 | 16.3 | 8.10 | 91.3 | 11.1 | 87.6 | 29.0 | 54.0 |
| UniDepth V2 | 5.60 | 45.3 | 4.60 | 42.6 | 14.5 | 6.98 | 92.8 | 10.2 | 88.6 | 23.7 | 71.2 |
| Depth Anything 3 | 10.5 | 41.8 | 7.39 | 44.7 | 6.61 | 7.89 | 91.4 | 10.6 | 87.5 | 17.4 | 70.1 |
| UniK3D | 5.52 | 46.2 | 4.37 | 45.0 | 15.1 | 6.95 | 92.8 | 9.63 | 89.6 | 17.9 | 77.8 |
| Pixel-Perfect Depth | 5.24 | 39.8 | 4.39 | 39.1 | 12.7 | 7.17 | 92.5 | 20.3 | 69.8 | – | – |
| InfiniDepth | 5.82 | 43.2 | 4.29 | 43.7 | 19.3 | 8.15 | 91.1 | 11.6 | 87.6 | – | – |
| MoGe-2 | 3.81 | 47.1 | 3.19 | 46.6 | 15.6 | 5.98 | 94.1 | 8.73 | 90.4 | 15.6 | 77.3 |
| Ours (ViT-L without refinement) | 3.60 | 48.89 | 2.93 | 50.26 | 15.62 | 5.56 | 94.67 | 7.94 | 91.97 | 14.74 | 80.97 |
| Ours (ViT-L) | 3.76 | 52.5 | 2.79 | 55.9 | 16.0 | 5.52 | 94.7 | 7.91 | 92.0 | 15.0 | 82.7 |
| Ours (ViT-G) | 3.66 | 54.1 | 2.80 | 56.8 | 17.3 | 4.80 | 95.6 | 7.43 | 92.4 | 15.8 | 82.7 |
2D image-space decoders mix features across depth discontinuities, blurring boundaries and thin structures. SSR instead refines geometry directly in 3D, preserving surface separation and recovering finer detail.
Our Self-Guided Sparse 3D Refiner (SSR) discretizes the coarse point map into a thin sparse voxel shell and refines it with sparse 3D convolutions. Because occluding and occluded surfaces fall into separate voxels, cross-boundary feature bleeding is avoided — something a parameter-matched 2D U-Net cannot achieve.
A factorized representation (u, v, log Z) keeps the image-axis coordinates fixed and lets the refiner predict only log-depth residuals. Each iteration re-voxelizes the updated geometry, progressively sharpening surface boundaries in a self-guided loop.
With K=0 the model matches the MoGe-2 base runtime; the 3-step refiner adds only moderate overhead (121 ms). Trained with K=3, it generalizes to K=7 at test time without degradation, and scales further with a ViT-G backbone.
We believe the SSR module can serve as a general-purpose dense geometry prediction head and a modular add-on for broader geometry foundation models.
@misc{kong2026finedetailmonoculargeometryestimation,
title={Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement},
author={Lingyu Kong and Ruicheng Li and Ruicheng Wang and Sicheng Xu and Chengtang Yao and Jianfeng Xiang and Jiaolong Yang},
year={2026},
eprint={2607.17967},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2607.17967},
}