Try it for yourself: https://foxglove.dev/examples 🔥
We had a lot of fun training the DROID dataset using DepthAnythingV2 models, achieving fine-grained details and enhanced depth accuracy, all visualized through Foxglove.
DepthAnythingV2, trained on 595K synthetic labeled images and 62M+ real unlabeled images, is a cutting-edge monocular depth estimation (MDE) model. It surpasses V1 by producing more robust depth predictions through three improvements: utilizing synthetic images, scaling the teacher model, and leveraging large-scale pseudo-labeled real images.
Compared to models built on Stable Diffusion, the DepthAnythingV2 models are over 10x faster and more accurate, with parameter scales ranging from 25M to 1.3B, ensuring broad applicability. Fine-tuned with metric depth labels, these models demonstrate strong generalization capabilities, as shown in this visualization example.
Additionally, we integrated time-series data for torques and velocities from the 7 DoF Franka Robotics Panda arm used in the dataset.
The DepthAnythingV2 project was made possible by contributions from Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao.
Check out the project and 🤗 space for more information.
Hugging Face space: https://buff.ly/3Y8bUev
GitHub project page: https://buff.ly/47VCMCU