Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
[Journal club] MOKA: Open-Vocabulary Robotic Ma...
Search
Sponsored
·
SiteGround - Reliable hosting with speed, security, and support you can count on.
→
Semantic Machine Intelligence Lab., Keio Univ.
PRO
November 15, 2024
Technology
440
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
[Journal club] MOKA: Open-Vocabulary Robotic Manipulation through Mark-Based Visual Prompting
Semantic Machine Intelligence Lab., Keio Univ.
PRO
November 15, 2024
More Decks by Semantic Machine Intelligence Lab., Keio Univ.
See All by Semantic Machine Intelligence Lab., Keio Univ.
[RSJ26] Building a VLA Model Based on Self-Distilled Classification
keio_smilab
PRO
0
200
[RSJ26] AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
keio_smilab
PRO
0
160
[RSJ26] NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields
keio_smilab
PRO
0
220
[RSJ26] Flow as Flow: Modeling Robot Velocity Fields as Probability Velocity Fields
keio_smilab
PRO
1
220
[RSJ26] Hierarchy-Aware Multimodal Retrieval-Augmented Generation for Embodied Question Answering
keio_smilab
PRO
1
190
[MIRU26] Open-Vocabulary Intention-Guided Object Detection in Diverse Scenes
keio_smilab
PRO
0
220
[Journal club] Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
keio_smilab
PRO
0
52
[MIRU26] To What Extent Does MLLM-as-a-Judge Exhibit Cross-Model Preference Bias?
keio_smilab
PRO
0
310
[Journal club] FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
keio_smilab
PRO
0
47
Other Decks in Technology
See All in Technology
顧客の成果創出とプロダクトの成長を 両立するためのFDE
sansantech
PRO
0
390
地方移住と都心キャリアの両立は「金・時間・人」のリソースをフル活用すれば実現できる!〜Snowflake女子会 vol.8
snowwmn0824
0
130
ボードゲームの遊び相手をFoundation Modelsで作る / iOSDC Japan 2026
genda
0
220
その Lambda、8分で 管理者権限まで奪われます
k1nakayama
7
3.8k
「今盗んで、後で解く」に備える ― AWSのポスト量子暗号入門
yama3133
2
270
spanner-autoscalerに学ぶ CRD設計パターン 〜自動化と緊急時対応を両立する Kubernetesコントローラーの作り方〜
tkuchiki
0
200
行動するAIのためのオントロジー | DevRev — Encraft #26.pdf
dvrv_tknrszk
1
580
人にやさしく、AIにやさしく、書き手を選ばないIaCのガードレール再考 / Rethinking IaC Guardrails for Humans and AI Alike
kohbis
5
1.7k
AI駆動開発で仕様はどこまで書くべきか? ― 人とAIの責務境界から考える開発プロセスの実践
takahiromatsui
1
210
CI/CDではもう遅い - 人とAIが迂回しないDevSecOps Verify基盤の再設計 -
kintotechdev
1
400
ログラスのマルチプロダクトを 支える認証基盤 〜テナントごとに異なる統制とどう向き合うか〜
dada4386
2
160
Azure Copilot Resiliency Agentをいろいろ試してみる
tomokusaba
0
130
Featured
See All Featured
"I'm Feeling Lucky" - Building Great Search Experiences for Today's Users (#IAC19)
danielanewman
230
23k
Art, The Web, and Tiny UX
lynnandtonic
304
22k
Google's AI Overviews - The New Search
badams
0
1.6k
Practical Orchestrator
shlominoach
192
12k
Marketing Yourself as an Engineer | Alaka | Gurzu
gurzu
0
310
Git: the NoSQL Database
bkeepers
PRO
432
67k
GraphQLとの向き合い方2022年版
quramy
50
15k
Java REST API Framework Comparison - PWX 2021
mraible
34
9.7k
WENDY [Excerpt]
tessaabrams
14
39k
jQuery: Nuts, Bolts and Bling
dougneiner
66
8.6k
How to Ace a Technical Interview
jacobian
280
24k
Chasing Engaging Ingredients in Design
codingconduct
0
340
Transcript
慶應義塾大学 杉浦孔明研究室 名字氏名 MOKA: Open-Vocabulary Robotic Manipulation through Mark-Based Visual
Prompting Kuan Fang, Fangchen Liu, Pieter Abbeel, Sergey Levine (UC Berkeley) RSS 2024 慶應義塾大学 杉浦孔明研究室 是方諒介 Fang, K., Liu, F., Abbeel, P., Levine, S. "MOKA: Open-Vocabulary Robotic Manipulation through Mark-Based Visual Prompting.“ RSS 2024.
概要 背景 ✓ open-vocabularyな指示文に基づく物体操作タスク ✓ 基盤モデルの常識的な知識への期待 提案 ✓ VLMによるhigh/low-levelな2段階のreasoning ✓
VQAに帰着したkeypoint予測に基づくaffordance検出 結果 ✓ 実機において階層的な物体操作タスクを実施し,既存手法を上回る成功率 ✓ ロボティクス基盤モデルによる拡張性を示唆 2
背景:open-vocabularyな指示文に基づく物体操作 ◼ 課題 ◼ 指示文の曖昧さ,複雑性,階層性 ◼ 多様かつ未知の物体/環境への汎化 → 常識的な知識を持つ基盤モデルに期待
LLMは視覚情報が欠落し,3D空間の認知に弱い ☺ VLMにより,視覚と軌道生成との中間的な affordance表現をkeypointとして獲得 3 "Insert the pink roses into the vase." "Put the scissors in the hand."
関連研究:VLMによるkeypoint予測を扱う手法は少ない 4 手法 概要 Code as Policies [Liang+, ICRA23] LLMにより,指示文を実行可能なコードに変換
VLMを用いておらず,視覚的な接地が不十分 VoxPoser [Huang+, CoRL23] voxel value mapを構築し,LLM / VLMを用いてプランニング 性能がvoxel mapの解像度に依存 ViLa [Hu+, 23] GPT-4Vを用いたプランニング low-levelなスキルを事前に定義する必要がある Code as Policies VoxPoser ViLa
提案手法:Marking Open-vocabulary Keypoint Affordances (MOKA) ◼ VLM (GPT-4V) によるhigh /
low-levelな2段階のreasoning ◼ affordance検出を,keypoint / waypoint選択に関するVQAに帰着 ◼ 対象物体の候補点/全体をgrid状に分割した候補領域を観測画像に重畳 5
high-level reasoning:階層的な指示文をサブタスクに分解 ◼ サブタスクごとに把持物体,干渉物体,操作方向を特定 ◼ GroundedSAM [Ren+, 24] により対象物体のセグメンテーションマスクを取得 6
Grounding DINO [Liu+, 23] + SAM [Kirillov+, ICCV23] :プロンプト :指示文 :初期の観測画像
low-level reasoning (1/2):マーキングによる視覚的なプロンプト ◼ VLMは座標を直接予測するより候補から選択する方が正確 (cf. SoM [Yang+, 23]) ◼
keypoint候補:PointNet [Qi+, CVPR17] による輪郭上の 点 + 物体の中心1点 ◼ waypoint候補:観測画像全体をgrid状に分割 → そこから一様に1点をサンプリング 7 SoM
low-level reasoning (2/2):VLMの「選択」によるkeypoint / waypoint予測 ◼ サブタスクごとに把持,作用,干渉keypoint,および動作waypointを選択 8 :プロンプト, :サブタスク,
:現在の観測画像, :マーキング処理
成功例に基づく改良:in-context learning, policy distillation ◼ in-context learning ◼ 3つの成功例(画像,出力)をVLMのプロンプトに追加 ◼
policy distillation ◼ ロボティクス基盤モデル Octo [Ghosh+, 23] ◼ RT-X [Vuong+, CoRL23] データセットの800Kの軌道でpre-trained ◼ 本タスクにおいて,50の軌道でfine-tuning 9 Octo RT-X
定量的結果:既存手法を上回るタスク成功率 [%] ◼ それぞれ2つのサブタスクから成る,合計4タスクを各々10回試行 ◼ 考察 ✓ すべてのサブタスクにおいて,既存手法と同等または上回った ✓ 蒸留の寄与より,data
generatorとしての応用可能性を示唆 10
定性的結果 (1/2):階層的なタスクを正確に実施 ◼ Table Wiping ◼ Laptop Packing 11 "Unplug
the charge cable and close the lid of the laptop." "Move the eyeglasses onto the yellow cloth and use the brush to sweep the snack package to the right side of the table."
定性的結果 (2/2):異なる指示文,配置,形容に対して頑健 ◼ 同じタスクに関して,多様な条件で評価 12
まとめ 背景 ✓ open-vocabularyな指示文に基づく物体操作タスク ✓ 基盤モデルの常識的な知識への期待 提案 ✓ VLMによるhigh/low-levelな2段階のreasoning ✓
VQAに帰着したkeypoint予測に基づくaffordance検出 結果 ✓ 実機において階層的な物体操作タスクを実施し,既存手法を上回る成功率 ✓ ロボティクス基盤モデルによる拡張性を示唆 13
Appendix:疑似コード 14
Appendix:high-level reasoningに用いるプロンプト 15
Appendix:low-level reasoningに用いるプロンプト (1/2) 16 入力に関する説明 keypoint / waypointの定義
Appendix:low-level reasoningに用いるプロンプト (2/2) 17 出力に関する説明
Appendix:その他のタスク 18 ◼ Watch Cleaning ◼ Gift Preparation
Appendix:Ablation Study 19
Appendix:エラー分析 20