Upgrade to Pro — share decks privately, control downloads, hide ads and more …

HOLO: Homography-Guided Pose Estimator Network ...

Sponsored · Your Podcast. Everywhere. Effortlessly. Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
Avatar for Aki Teshima Aki Teshima
August 28, 2026

HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps

Paper from CVPR2026 introduced in 67th Kanto Computer Vision reading group

Avatar for Aki Teshima

Aki Teshima

August 28, 2026

More Decks by Aki Teshima

Other Decks in Science

Transcript

  1. HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on

    SD Maps Xuchang Zhong, Xu Cao, Jinke Feng, Hao Fang tomoaki_teshima tomoaki0705 tomoaki_teshima tomoaki0705
  2. Introduction • Goal : Localization • Input : 6 image

    • Output : Location + angle 2026/Aug/29 第67回 コンピュータビジョン勉強会@関東(後編) 3
  3. Localization issue • GPS : Noisy • Visual Localization •

    Matching base • Regression base • Cross modal 2026/Aug/29 第67回 コンピュータビジョン勉強会@関東(後編) 4
  4. Matching based visual localization 11 18 28 19 2026/Aug/29 第67回

    コンピュータビジョン勉強会@関東(後編) 5
  5. Regression based visual localization Feature extraction -> cross modal fusion

    -> pose decoding 7 6 2026/Aug/29 第67回 コンピュータビジョン勉強会@関東(後編) 6
  6. Cross-modal Homography Estimation 29 RGB/NIR/depth/SD map 4 22 2026/Aug/29 第67回

    コンピュータビジョン勉強会@関東(後編) 7
  7. Related works • Visual localization • Matching based method •

    Pose retrieval な問題に帰結 • 学習位置が離散的で、学習位置にひっぱられがち • Regression based method • 位置情報を入れないため、位置精度おちがち • Cross-modal Homography estimation • アノテーション済のHomographyと位置情報が必要 2026/Aug/29 第67回 コンピュータビジョン勉強会@関東(後編) 8
  8. BEV Perception Module Camera Encoder : 画像特徴量抽出で遠さを推 定 Neural View

    Transformation : 特徴量を俯瞰画 像に変換 BEV Encoder : 俯瞰画像特徴量 Semantic feature : Road / Building 2026/Aug/29 第67回 コンピュータビジョン勉強会@関東(後編) 10
  9. Map Process Module 2026/Aug/29 Map Rasterization : Open Street Map

    はベク タ形式で保存されてるのでラスタ画像形式に変 換 Map Encoder : BEV 特徴量マップ 第67回 コンピュータビジョン勉強会@関東(後編) 11
  10. Homography Guided pose estimation Feature Fusion : 初期位置合わせ Iterative Pose

    Refinement : 反復的な調整 最終的にHomographyが計算されたら、BEV特 徴マップの中心と向きをSDマップ空間に投影 できる=>現在地GET 2026/Aug/29 第67回 コンピュータビジョン勉強会@関東(後編) 12
  11. Experiment • Dataset : nuScenes (BostonとSingapore) • 性能評価 • Recall@Xm

    • Recall@X° • Absolute Position Error • Absolute Orientation Error 2026/Aug/29 第67回 コンピュータビジョン勉強会@関東(後編) 13
  12. Ablation study (Homography Estimation) • タイトルにもあるHomographyを使って良くなったの? • 良くなってる • Cross

    Modal の方がCross Attention より良い 2026/Aug/29 第67回 コンピュータビジョン勉強会@関東(後編) 16
  13. Ablation study (Iterative Strategy) • 反復的に改善して効果あり • Iterationが増えると計算時間が増える • 2

    iterations で 25.6FPS • 6 iterations で 19.9FPS 2026/Aug/29 第67回 コンピュータビジョン勉強会@関東(後編) 18
  14. Summary • 画像と地図データからHomographyを計算して現在地推定 • Homographyの教師データが不要 • 学習データにない位置でも精度よく位置推定できる • Matching based

    methodよりよい精度が出る • 直接位置推定するのではなく、Homographyを求める 2026/Aug/29 第67回 コンピュータビジョン勉強会@関東(後編) 19
  15. 余談 • Homographyを計算して自車位置推定をする修士論文(フレーム間) • 本論文も車載カメラを使った自車位置推定 • BEV便利! • BEV便利! •

    Homographyは画像のマッチングで計算 • 時代はDNNでHomography推定 • 単眼カメラ5FPS • 6台カメラで25FPS • 気分はチャオズ 2026/Aug/29 第67回 コンピュータビジョン勉強会@関東(後編) 20
  16. Reference 4: Si-Yuan Cao, Jianxin Hu, Zehua Sheng, and Hui-Liang

    Shen. Iterative deep homography estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1879–1888, 2022 6: Shuai Chen, Tommaso Cavallari, Victor Adrian Prisacariu, and Eric Brachmann. Map-relative pose regression for visual re-localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20665–20674, 2024 7: Siyan Dong, Shuzhe Wang, Shaohui Liu, Lulu Cai, Qingnan Fan, Juho Kannala, and Yanchao Yang. Reloc3r: Largescale training of relative camera pose regression for generalizable, fast, and accurate visual localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16739–16752, 2025. 11: Ted Lentsch, Zimin Xia, Holger Caesar, and Julian FP Kooij. Slicematch: Geometry-guided aggregation for cross-view pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17225–17234, 2023. 18: Paul-Edouard Sarlin, Daniel DeTone, Tsun-Yi Yang, Armen Avetisyan, Julian Straub, Tomasz Malisiewicz, Samuel Rota Bulo, Richard Newcombe, Peter Kontschieder, and Vasileios Balntas. Orienternet: Visual localization in 2d public maps with neural matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21632–21642, 2023. 19: Paul-Edouard Sarlin, Eduard Trulls, Marc Pollefeys, Jan Hosang, and Simon Lynen. Snap: Self-supervised neural maps for visual positioning and semantic understanding. Advances in Neural Information Processing Systems, 36:7697–7729, 2023. 22: Sanghyeob Song, Jaihyun Lew, Hyemi Jang, and Sungroh Yoon. Unsupervised homography estimation on multimodal image pair via alternating optimization. Advances in Neural Information Processing Systems, 37:61306–61327, 2024 28: Zimin Xia, Olaf Booij, and Julian FP Kooij. Convolutional cross-view pose estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(5):3813–3831, 2023 29: Junchen Yu, Si-Yuan Cao, Runmin Zhang, Chenghao Zhang, Zhu Yu, Shujie Chen, Bailin Yang, and Hui-Liang Shen. Sshnet: Unsupervised cross-modal homography estimation via problem reformulation and split optimization. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 16685–16694, 2025. 2026/Aug/29 第67回 コンピュータビジョン勉強会@関東(後編) 21