Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
論文解説 Mask2Former
Search
koharite
June 15, 2022
Research
5.8k
11
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
論文解説 Mask2Former
Presentation for explaining the paper Mask2Former presented at CVPR2022.
koharite
June 15, 2022
More Decks by koharite
See All by koharite
論文解説 DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
koharite
0
310
論文解説 DTPP: Differentiable Joint Conditional Prediction and Cost Evaluationfor Tree Policy Planning in Autonomous Driving
koharite
0
190
論文解説 Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
koharite
0
310
論文解説 DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models
koharite
0
360
論文解説 EfficientViT: Memory Efficient Vision Transformer with Cascaded Group Attention
koharite
0
740
論文解説 CoCa: Contrastive Captioners are Image-Text Foundation Models
koharite
0
1.3k
論文解説 LoRA : Low Rank Adaptation of Large Language Models
koharite
3
2.6k
論文解説 ControlNet
koharite
1
6.7k
論文解説 InstructGPT : Training language models to follow instructions with human feedback
koharite
4
3.9k
Other Decks in Research
See All in Research
MM-OVSeg: Multimodal Optical–SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing
satai
3
110
IA for theory
gpeyre
1
360
Sequences of Logits Reveal the Low Rank Structure of Language Models
sansantech
PRO
1
310
LA-Bench 2025:実験指示から実行可能手順を生成するためのデータセット/LA-Bench 2025: A Dataset for Generating Executable Experimental Procedures from Experimental Instructions
stktu
0
130
2026年度 生成AI を活用した論文執筆ガイド/ワークショップ / 2026 Academic Year Guide to Writing Papers Using Generative AI - Workshop
ks91
PRO
0
210
敵対生成プロンプト同時探索による内省型プロンプト最適化
kinoue_smarthr
0
370
nlp2026 In-Context Learningに基づく経路案内のための地理的知識の活用方法に関する検討
takashiinui
0
130
MIRU2026 チュートリアル講演2:三次元データ処理の動向
nnchiba
6
4.3k
[最先端NLP勉強会2026] Checklists Are Better Than Reward Models For Aligning Language Models
nzw0301
1
130
2025年度秋葉原ウォーカブルプロジェクト調査報告 「アキバらしいウォーカブル」とは何か
izumiyama_lab
1
180
[Fishers] DIVER OSINT CTF 2026 特化AIエージェントハーネスで挑戦するOSINT CTF
analokmaus
0
500
Sleuthcon Keynote - How Cybercriminals (ab)use AI
fr0gger
0
290
Featured
See All Featured
コードの90%をAIが書く世界で何が待っているのか / What awaits us in a world where 90% of the code is written by AI
rkaga
63
45k
Evolving SEO for Evolving Search Engines
ryanjones
0
260
So, you think you're a good person
axbom
PRO
2
2.1k
The Cost Of JavaScript in 2023
addyosmani
55
10k
Groundhog Day: Seeking Process in Gaming for Health
codingconduct
0
290
Done Done
chrislema
186
16k
Sharpening the Axe: The Primacy of Toolmaking
bcantrill
46
2.9k
Leo the Paperboy
mayatellez
8
2.2k
Why You Should Never Use an ORM
jnunemaker
PRO
61
10k
The browser strikes back
jonoalderson
0
1.5k
Chrome DevTools: State of the Union 2024 - Debugging React & Beyond
addyosmani
10
1.3k
Balancing Empowerment & Direction
lara
6
1.2k
Transcript
論⽂解説 Masked-attention Mask Transformer for Universal Image Segmentation Takehiro Matsuda
2 論⽂情報 • タイトル:Masked-attention Mask Transformer for Universal Image Segmentation
• 論⽂: https://arxiv.org/abs/2112.01527 • コード: https://github.com/facebookresearch/Mask2Former • 投稿学会: CVPR2022 • 著者: Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, Rohit Girdhar • 所属:Facebook AI Research (FAIR), University of Illinois at Urbana-Champaign (UIUC) 選んだ理由: • Transformerを使ったユニバーサルなアーキテクチャを提案し、セグメンテーション タスクについてSemantic, Instance, Panopticの違いによらず使える • Semantic, Instance, PanopticそれぞれでこれまでのSOTAを超える性能を達成した。
3 論⽂概要 Panoptic Instance Semantic Transformer DecoderにMasked Attentionを導⼊する Transformer decoderをMulti-scaleにする。
学習で得られたMask領域におけるMasked Attentionにより 局所的な特徴を精度良く捉える。 Panoptic: COCO Panopnic val2017 Instance: COCO val2017 Semantic: ADE20K SOTAを達成 Ground Truth Prediction Ground Truth Prediction
4 Segmentationの違い Pixel毎にクラスを認識 指定したクラス の存在する場所 を認識、同じク ラスでも別個体 は分ける (空などを対象ク ラスにしなけれ
ば識別されない Pixelがある) Thingはinstanceと して認識、 Stuff(空や道路)も 認識
5 関連論⽂ DETR( Detection Transformer) : Object DetectionでTransformerを導⼊ MaskFormer: SegmentationでTransformerによるMaskを作り出し、推定する
FAIR (Facebook AI Research)が出しているTransformerを使った画像認識に 関する⼀連の論⽂の流れ DETR MaskFormer TransformerでGlobalな特徴や関係を抽出できる が、⼩さい物体の認識は若⼲苦⼿だったことや ⼤きな計算リソースが必要だった点を改良する。
6 Transformer概説 https://www.slideshare.net/SSII_Slides/ssii2022-ts1-transformer (⽜久⽒資料より)
7 Transformer概説
8 Transformer概説
9 Transformer概説
10 Transformer概説
11 Transformer概説
12 Transformer概説
13 Transformer概説
14 DETR Anchorの設定やNMS(Non Maximum Suppression)を必要としない。
15 DETR ⾼解像度の近傍Pixel(領域) の特徴はCNNネットワーク でエンコードして取得(W, H は1/32, Cは2048) CNNから取り出された画像の特徴量からAttentionを⽤い て各物体の位置や種類の情報に変換
事前に決められた個数Nの物体を予測する 他の予測内容を考慮して⾃⾝の予測するEncoder-Decoder ネットワーク Transformerの出⼒を物体の位置座 標・クラスラベルにデコードする ネットワーク
16 MaskFormer TransformerでSemantic SegmentationとPanoptic Segmentationを⾏う Ground Truth Prediction Ground Truth
Prediction
17 MaskFormer Per-Pixel Classification is Not All You Need for
Semantic Segmentation Binary mask predictionsを取得する transformer decoderでN個のclass predictionsと mask embeddingsを取得 Binary MaskにたいしてPixelごとのmask lossを算出 Maskごとにクラス推定のlossを算出 Segmentation TaskをMask classificationとして、 (1) 画像からN個のbinary mask 領域を作成 (2) 各マスク領域をK個の認識 カテゴリそれぞれに所属 する確率をだす
18 Mask2Former MaskFormerの弱点を改良 • ⼩さな対象の精度が悪い • ⼤きなコンピュータリソース • ⻑い学習時間 panoptic
segmentation (57.8 PQ on COCO) instance segmentation (50.1 AP on COCO) Semantic segmentation (57.7 mIoU on ADE20K). SOTAを達成
19 Masked Attention Masked attention 画像全体から学習されるcross-attentionに変わり、 オブジェクトクエリの予測に基づいて⽣成され たマスクを使って特定領域内でAttentionをとる。 通常のcross attention
Masked attention ⼩物体や物体境界などの細部の認識が改善さ れるのではないか。 We hypothesize that local features are enough to update query features and context information can be gathered through self-attention.
20 Multi-scale high-resolution features Pixel Decoderで元画像の1/32, 1/16, 1/8の Feature Pyramidを作り、Transformer
Decoder もそれぞれに対応する Transformer Decoder 3 x L layers 画像系ではよく使われる解像度のPyramid構造を採⽤ ⼩さなオブジェクトの認識性能を上げる
21 Optimization improvements 通常のTransformer Decoder layerはquery featuresを⽣み出すのにself-attention module, cross- attention,
feed-forward networkを順に送るが、 SelfとMasked(Cross) -attentionの順番を 変え、query featuresを学習可能にした。 Dropoutをなくした。 (これまではresidual connectionsと attention mapsに適応していた)
22 Computer resource reduction MaskFormerでは1つの画像で32GメモリのGPUが必要だった。 PointRendやImplicit PointRendから着想を得て、mask lossを計算するのに、mask全体でなく、 K(=12544=112 x112)個のランダムサンプルされた点で計算する。
推論とground truthとのfinal lossはimportance samplingで別にとったK個のサンプルされた点で⾏う。 最終的に、Mask2Formerでは1つの画像で必要なメモリが18GBから6GBまで削減された。 ⾼解像のMask predictionのため
23 PQ Metrics Average IoU 正しく認識できたものの 割合(F1 scoreに似たもの) IoU >=0.5でTP
Panoptic Segmentationの性能評価指標
24 Experiment – Panoptic Segmentation COCO panoptic val 2017 with
133 categories
25 Panoptic Segmentation Visualization GT GT predict predict
26 Experiment – Instance Segmentation COCO val 2017 with 80
categories
27 Instance Segmentation Visualization GT GT predict predict
28 Experiment – Semantic Segmentation ADE20K val with 150 categories
Single scale Multi scale
29 Semantic Segmentation Visualization GT GT predict predict
30 参考資料 DETR https://arxiv.org/abs/2005.12872 https://github.com/facebookresearch/detr MaskFormer https://arxiv.org/abs/2107.06278 https://github.com/facebookresearch/MaskFormer Panoptic Segmentation
https://arxiv.org/abs/1801.00868 Transformerの最前線 (オムロンサイニックエックス ⽜久⽒) https://www.slideshare.net/SSII_Slides/ssii2022-ts1-transformer