Pumpire: Unified Benchmark for Metric Distance Estimation
Abstract
We present Pumpire, a unified benchmark for evaluating metric point-pair distance estimation capability of both image- and video-level 3D foundation models, with or without depth priors. In contrast to previous approaches that normally evaluate depth and camera intrinsics separately or evaluate point-clouds with geometric similarity metrics, which cannot directly reflect models' point-to-point distance estimation capability, Pumpire directly assesses point-to-point distances from the reconstructed geometry. To this end, we collect a large-scale and diverse dataset (pumpire-6k) comprising 100 real-world scenes, each annotated with physically measured point-pair distances and containing 64 frames, for a total of 6,400 frames. Building on this dataset, we establish a holistic evaluation protocol that covers both image- and video-level 3D foundation models and enables direct assessment of point-pair distance errors and cross-setting comparison. We conduct extensive experiments across 29 baseline configurations of representative 3D foundation models and provide a comprehensive analysis of the results. By offering this benchmark, we target the more fundamental ability to perceive and estimate physical scale in the reconstructed 3D space, which prior evaluation protocols have largely overlooked. The project page can be found at https://pumpire.github.io/
Community
We present Pumpire, a unified benchmark for evaluating metric point-pair distance estimation in image- and video-level 3D foundation models, with or without depth priors. We first manually collect a large-scale dataset comprising 6,400 frames from 100 real-world scenes, annotated with physically measured point-pair distances. We then extensively evaluate 29 baseline configurations, providing a comprehensive analysis of their metric accuracy and cross-view consistency.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- SpatialCrafter: Single Image World Modeling with Generative 3D Proxies (2026)
- FounRef: Robust, Structure-Preserving, and Fast Metric Refinement of Frozen Monocular Foundation Priors with Sparse Anchors (2026)
- TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views (2026)
- PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation (2026)
- Geometry-Grounded Unified 3D Perception for Autonomous Driving (2026)
- VisTa3D: A Dataset and Benchmark for Thin Object Reconstruction from Vision, Tactile, and 3D Point Clouds (2026)
- GeoWeaver: Accurate Long-Sequence 3D Reconstruction via Hierarchical Geometric Assembly (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper