百科.dev
全部条目AI 编程趋势榜开源项目技术资讯提交条目
登录
< 返回工具列表
P

pixel-perfect-sfm

> 编程语言
开源

利用特征度量细化的像素完美结构从运动法 (ICCV 2021 最佳学生论文奖)

1.5K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

利用特征度量细化的像素完美结构从运动法 (ICCV 2021 最佳学生论文奖)

# Pixel-Perfect Structure-from-Motion ### Best student paper award @ [ICCV 2021](http://iccv2021.thecvf.com/) We introduce a framework that **improves the accuracy of Structure-from-Motion (SfM) and visual localization** by refining keypoints, camera poses, and 3D points using the direct alignment of deep features. It is presented in our paper: - [Pixel-Perfect Structure-from-Motion with Featuremetric Refinement](https://arxiv.org/abs/2108.08291) - Authors: [Philipp Lindenberger](https://scholar.google.com/citations?user=FMVAi2YAAAAJ&hl=en)\*, [Paul-Edouard Sarlin](https://psarlin.com/)\*, [Viktor Larsson](http://people.inf.ethz.ch/vlarsson/), and [Marc Pollefeys](http://people.inf.ethz.ch/pomarc/) - Website: [psarlin.com/pixsfm](https://psarlin.com/pixsfm/) (videos, slides, poster) Here we provide `pixsfm`, a Python package that can be readily used with [COLMAP](https://colmap.github.io/) and [our toolbox hloc](https://github.com/cvg/Hierarchical-Localization/). This makes it easy to **refine an existing COLMAP model or reconstruct a new dataset with state-of-the-art image matching**. Our framework also improves visual localization in challenging conditions.

The refinement is composed of 2 steps: 1. **Keypoint adjustment:** before SfM, jointly refine all 2D keypoints that are matched together. 2. **Bundle adjustment:** after SfM, refine 3D points and camera poses. In each step, we optimize the consistency of dense deep features over multiple views by minimizing a featuremetric cost. These features are extracted beforehand from the images using a pre-trained CNN. With `pixsfm`, you can: - reconstruct and refine a scene using hloc, from scratch or with given camera poses - localize and refine new query images using hloc - run the keypoint or bundle adjustments on a COLMAP database or 3D model - evaluate the refinement with new dense or sparse features on the ETH3D dataset Our implementation scales to large scenes by carefully managing the memory and leveraging parallelism and SIMD vectorization when possible. ## Installation `pixsfm` requires Python >=3.6, GCC >=6.1, and COLMAP 3.8 [installed from source](https://colmap.github.io/install.html#build-from-source). The core optimization is implemented in C++ with [Ceres >= 2.1](https://github.com/ceres-solver/ceres-solver/) but we provide Python bindings with high granularity. The code is written for UNIX and has not been tested on Windows. The remaining dependencies are listed in `requirements.txt` and include [PyTorch](https://pytorch.org/) >=1.7 and [pycolmap](https://github.com/colmap/pycolmap) + [pyceres](https://github.com/cvg/pyceres) built from source: ```bash # install COLMAP following colmap.github.io/install.html#build-from-source, tag 3.8 sudo apt-get install libhdf5-dev git clone https://github.com/cvg/pixel-perfect-sfm --recursive cd pixel-perfect-sfm pip install -r requirements.txt ``` To use other local features besides SIFT via COLMAP, we also require [hloc](https://github.com/cvg/Hierarchical-Localization/): ```bash git clone --recursive https://github.com/cvg/Hierarchical-Localization/ cd Hierarchical-Localization/ pip install -e . ``` Finally build and install the `pixsfm` package: ```bash pip install -e . # install pixsfm in develop mode ``` We highly recommend to use `pixsfm` with a working GPU for the dense feature extraction. All other steps can only run on the CPU. Having issues with compilation errors or runtime crashes? Want to use the codebase as a C++ library? Check our [FAQ](./doc/FAQ.md). ## Tutorial The Jupyter notebook [`demo.ipynb`](./demo.ipynb) demonstrates a minimal usage example. It shows how to run Structure-from-Motion and the refinement, how to align and compare different 3D models, and how to localize and refine additional query images.


Visualizing mapping and localization results in the demo.

## Structure-from-Motion ### End-to-end SfM with hloc Given keypoints and matches computed with hloc and stored in HDF5 files, we can run Pixel-Perfect SfM from a Python script: ```python from pixsfm.refine_hloc import PixSfM refiner = PixSfM() model, debug_outputs = refiner.reconstruction( path_to_working_directory, path_to_image_dir, path_to_list_of_image_pairs, path_to_keypoints.h5, path_to_matches.h5, ) # model is a pycolmap.Reconstruction 3D model ``` or from the command line: ```bash python -m pixsfm.refine_hloc reconstructor \ --sfm_dir path_to_working_directory \ --image_dir path_to_image_dir \ --pairs_path path_to_list_of_image_pairs \ --features_path path_to_keypoints.h5 \ --matches_path path_to_matches.h5 ``` Note that: - The final refined 3D model is written to `path_to_working_directory` in either case. - Dense features are automatically extracted (on GPU when available) using a pre-trained CNN, [S2DNet](https://github.com/germain-hug/S2DNet-Minimal) by default. - The result `debug_outputs` contains the dense features and optimization statistics. ### Configurations We have fine-grained control over all hyperparameters via [OmegaConf](https://omegaconf.readthedocs.io/) configurations, which have sensible default values defined in `PixSfM.default_conf`. See [Detailed configuration](#detailed-configuration) for a description of the main configuration entries and their defaults. [Click to see some examples] For example, dense features are stored in memory by default. If we reconstruct a large scene or have limited RAM, we should instead write them to a cache file that is loaded on-demand. With the Python API, we can pass a configuration update: ```python refiner = PixSfM(conf={"dense_features": {"use_cache": True}}) ``` or equivalently with the command line [using a dotlist](https://omegaconf.readthedocs.io/en/2.1_branch/usage.html#from-command-line-arguments): ```bash python -m pixsfm.refine_hloc reconstructor [...] dense_features.use_cache=true ``` We also provide ready-to-use configuration templates in [`pixsfm/configs/`](./pixsfm/configs/) covering the main use cases. For example, [`pixsfm/configs/low_memory.yaml`](./pixsfm/configs/low_memory.yaml) reduces the memory consumption to scale to large scene and can be used as follow: ```python refiner = PixSfM(conf="low_memory") # or python -m pixsfm.refine_hloc reconstructor [...] --config low_memory ``` ### Triangulation from known camera poses [Click to expand] If camera poses are available, we can simply triangulate a 3D point cloud from an existing reference COLMAP model with: ```python model, _ = refiner.triangulation(..., path_to_reference_model, ...) ``` or ```bash python -m pixsfm.refine_hloc triangulator [...] \ --reference_sfm_model path_to_reference_model ``` By default, camera poses and intrinsics are optimized by the bundle adjustment. To keep them fixed, we can simply overwrite the corresponding options as: ```python conf = {"BA": {"optimizer": { "refine_focal_length": False, "refine_extra_params": False, # distortion parameters "refine_extrinsics": False, # camera poses }}} refiner = PixSfM(conf=conf) refiner.triangulation(...) ``` or equivalently ```bash python -m pixsfm.refine_hloc triangulator [...] \ 'BA.optimizer={refine_focal_length: false, refine_extra_params: false, refine_extrinsics: false}' ``` ### Keypoint adjustment The first step of the refinement is the keypoint adjustment (KA). It refines the keypoints from tentative matches only, before SfM. Here we show how to run this step separately. [Click to expand] To refine keypoints stored in an hloc HDF5 feature file: ```python from pixsfm.refine_hloc import PixSfM refiner = PixSfM() keypoints, _, _ = refiner.refine_keypoints( path_to_output_keypoints.h5, path_to_input_keypoints.h5, path_to_list_of_image_pairs, path_to_matches.h5, path_to_image_dir, ) ``` To refine keypoints stored in a COLMAP database: ```python from pixsfm.refine_colmap import PixSfM refiner = PixSfM() keypoints, _, _ = refiner.refine_keypoints_from_db( path_to_output_database, # pass path_to_input_database for in-place refinement path_to_input_database, path_to_image_dir, ) ``` In either case, there is an equivalent command line interface. ### Bundle adjustment The second contribution of the refinement is the bundle adjustment (BA). Here we show how to run it separately to refine an existing COLMAP 3D model. [Click to expand] To refine a 3D model stored on file: ```python from pixsfm.refine_colmap import PixSfM refiner = PixSfM() model, _, _, = refiner.refine_reconstruction( path_to_input_model, path_to_output_model, path_to_image_dir, ) ``` Using the command line interface: ```bash python -m pixsfm.refine_colmap bundle_adjuster \ --input_path path_to_input_model \ --output_path path_to_output_model \ --image_dir path_to_image_dir ``` ## Visual localization When estimating the camera pose of a single image, we can also run the keypoint and bundle adjustments before and after PnP+RANSAC. This requires reference features attached to each observation of the reference model. They can be computed in several ways. [Click to learn how to localize a single image] 1. To recompute the references from scratch, pass the path to the reference images: ``` … ``` The default localization configuration can be accessed with `QueryLocalizer.default_conf`. 2. Alternatively, if dense reference features have already been computed during the pixel-perfect SfM, it is more efficient to reuse them: ```python refiner = PixSfM() model, outputs = refiner.reconstruction(...) features = outputs["feature_manager"] # or load the features manually features = pixsfm.extract.load_features_from_cache( refiner.resolve_cache_path(output_dir=path_to_output_sfm) ) localizer = QueryLocalizer( reference_model, # pycolmap.Reconstruction 3D model dense_features=features, ) ``` We can also batch-localize multiple queries equivalently to [`hloc.localize_sfm`](https://github.com/cvg/Hierarchical-Localization/blob/master/hloc/localize_sfm.py): ```python pixsfm.localize.main( dense_features, # FeatureManager or path to cache file reference_model, # pycolmap.Reconstruction 3D model path_to_query_list, path_to_image_dir, path_to_image_pairs, path_to_keypoints, path_to_matches, path_to_output_results, config=config, # optional dict ) ``` ## Example: mapping and localization We now show how to run the featuremetric pipeline on the Aachen Day-Night v1.1 dataset. First, download the dataset by following [the instructions described here](https://github.com/cvg/Hierarchical-Localization/tree/master/hloc/pipelines/Aachen_v1_1#installation). Then run `python examples/sfm+loc_aachen.py`, which will perform mapping and localization with SuperPoint+SuperGlue. As the scene is large, with over 7k images, we cache the dense feature patches and therefore require about 350GB of free disk space. Expect the sparse feature matching to take a few hours on a recent GPU. We also show in [`examples/refine_sift_aachen.py`](examples/refine_sift_aachen.py) how to start from an existing COLMAP database. ## Evaluation We can evaluate the accuracy of the pixel-perfect SfM and of camera pose estimation on the ETH3D dataset. Refer to the paper for more details. First, we download the dataset with `python -m pixsfm.eval.eth3d.download`, by default to `./datasets/ETH3D/`. ### 3D triangulation [Click to expand] We first need to install the [ETH3D multi-view evaluation tool](https://github.com/ETH3D/multi-view-evaluation): ```bash sudo apt install libpcl-dev

Issues· 48 开放

查看全部 Issues在 GitHub 打开

暂无开放 Issues,或尚未同步最近议题。

> 标签

C++3d-visiondeep-learningfeature-matchingstructure-from-motion

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月17日
分类编程语言
定价开源

> 相关工具

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言