Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
[RSJ26] AnoleVLA: Lightweight Vision-Language-A...
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
Semantic Machine Intelligence Lab., Keio Univ.
PRO
August 29, 2026
Technology
140
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
[RSJ26] AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
Semantic Machine Intelligence Lab., Keio Univ.
PRO
August 29, 2026
More Decks by Semantic Machine Intelligence Lab., Keio Univ.
See All by Semantic Machine Intelligence Lab., Keio Univ.
[RSJ26] Building a VLA Model Based on Self-Distilled Classification
keio_smilab
PRO
0
200
[RSJ26] NarrativeFlow: Flow-Based Vision-Language-Action Model Using Robot Velocity Fields
keio_smilab
PRO
0
170
[RSJ26] Flow as Flow: Modeling Robot Velocity Fields as Probability Velocity Fields
keio_smilab
PRO
1
220
[RSJ26] Hierarchy-Aware Multimodal Retrieval-Augmented Generation for Embodied Question Answering
keio_smilab
PRO
1
190
[MIRU26] Open-Vocabulary Intention-Guided Object Detection in Diverse Scenes
keio_smilab
PRO
0
210
[Journal club] Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
keio_smilab
PRO
0
49
[MIRU26] To What Extent Does MLLM-as-a-Judge Exhibit Cross-Model Preference Bias?
keio_smilab
PRO
0
300
[Journal club] FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
keio_smilab
PRO
0
38
[Journal club] DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
keio_smilab
PRO
1
78
Other Decks in Technology
See All in Technology
例外の正しい扱い方 そのエラー try-catchして大丈夫?
jinwatanabe
3
510
安心して変更できるWebフロントエンドの作り方
pirosikick
5
2.6k
Issue 駆動でスペシャリストの意図を届ける、AI 実装のアクセシビリティ向上
thkt
0
120
新機種発売前に見直そう!端末移行で再ログインが要るアプリ・要らないアプリは何が違うのか 〜シームレスに再開できる設計と実装〜
zozotech
PRO
0
200
【技術的負債conf】事業成長に伴う技術的負債の説明責任とAIによるモニタリング、認知的負債について
i35_267
3
1.7k
その Lambda、8分で 管理者権限まで奪われます
k1nakayama
5
2.7k
Deployment の 先にある AI Agent 基盤 - kagent vNext、Agent Substrate、Hermes から読み解く Agent Runtime の現在地 / k8s-matsuri-2-ai-agent-platform-amsy810
masayaaoyama
4
630
AI de Idea
kawaguti
PRO
2
120
積み重なった技術負債への挑戦 〜初手としての全社ゴト化〜
techtekt
PRO
0
1.3k
Azure Serverless 2026:Production-ready な AI エージェント基盤 / Azure Serverless 2026: Production-Ready AI Agent Platform
miyake
1
150
LLMに渡さなかった仕事
nanaism
0
430
今話題のAI「Jev」って何? 宇宙最速で学ぶ会
minorun365
PRO
26
15k
Featured
See All Featured
世界の人気アプリ100個を分析して見えたペイウォール設計の心得
akihiro_kokubo
PRO
74
42k
BBQ
matthewcrist
89
10k
Rebuilding a faster, lazier Slack
samanthasiow
85
9.6k
Building AI with AI
inesmontani
PRO
1
1.2k
Mobile First: as difficult as doing things right
swwweet
225
10k
個人開発の失敗を避けるイケてる考え方 / tips for indie hackers
panda_program
123
22k
jQuery: Nuts, Bolts and Bling
dougneiner
66
8.6k
The Limits of Empathy - UXLibs8
cassininazir
1
670
Ecommerce SEO: The Keys for Success Now & Beyond - #SERPConf2024
aleyda
1
2.2k
Distributed Sagas: A Protocol for Coordinating Microservices
caitiem20
333
23k
What’s in a name? Adding method to the madness
productmarketing
PRO
24
4.2k
技術選定の審美眼(2025年版) / Understanding the Spiral of Technologies 2025 edition
twada
PRO
120
120k
Transcript
深層状態空間モデルに基づく 軽量VLAによる物体操作 慶應義塾大学 髙木裕輔 神原元就 八島大地 妹尾幸樹 戸倉健登 杉浦孔明
Motivation: VLAのメモリ消費量・推論時間を削減したい 背景: VLAを実機で動作させる際、推論時の計算コストが問題 メモリ消費量が大きく高性能な計算機が必要 軌道生成の研究 [Kambara+, RA-L26], [Kaichi+,
IROS26],... 推論時間が長くアームの動作がjerkyに 本研究: 軽量かつ高速なVLAを提案 “Pick up the cube and put it into the basket.” 推論時間 [ms] OpenVLA jerkyな動作 GPUメモリ消費 17GB以上 𝜋0 𝜋0.5 提案手法 𝜋0.5 推論時の動作 [Jiaming+, 26] GPUメモリ消費量 [GB] 2
関連研究: VLAの計算コストを抑制する研究 手法 特徴 SmolVLA 0.45Bの軽量VLA ・action chunkingを活用 Transformerに基づくバックボーン・長系列の処理が困難
深層状態空間モデルに基づくバックボーン 画像上の接触点 / 姿勢予測のみ・軌道を生成しない [Shukor+, 25] RoboMamba [Liu+, NeurIPS24] SmolVLA RoboMamba 3
提案手法: 深層状態空間モデルに基づく軽量なVLA 4
提案手法: 深層状態空間モデルに基づく軽量なVLA 入力を埋め込みトークン系列を生成 ロボット状態 ロボット状態の時間差分 画像 指示文 5
提案手法: 深層状態空間モデルに基づく軽量なVLA バックボーンにて軌道生成に必要な情報を集約 入力系列 バックボーンLLM 出力系列 6
提案手法: 深層状態空間モデルに基づく軽量なVLA 最終トークンを利用しチャンク長の軌道を生成 軌道 最終トークンに 情報が集約 7
高速かつメモリ消費量の少ないバックボーンLLM ▪ Mamba [Gu+, COLM24] による系列処理 隠れ状態 入力 約6Mの学習可能パラメータ 出力
ブロックの多層化 ☺ 𝒪(𝑁)での系列処理 cf. Transformerの計算量は𝒪(𝑁 2 ) ☺ 370Mパラメータの軽量モデル cf. RoboMamba: 2.8B 8
実験設定: シミュレーション・実機ロボットにおける実験 ▪ シミュレーション実験: Meta-World [Yu+, CoRL19] ▪ 実機実験: モバイルマニピュレーションタスク(HSRを使用)
▪ リーダ・フォロワシステムを用い、データを収集 タスク数 エピソード数 試行回数/タスク シミュレーション 50 2,500 10 実機 5 250 10 Meta-Worldのタスク例 実機実験のデータ収集 9
定量的結果: シミュレーションにてベースラインを上回った 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 提案手法 10
定量的結果: シミュレーションにてベースラインを上回った 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 平均成功率 提案手法 +22pt 11
定性的結果: シミュレーションで様々なタスクに成功 "Grasp a stick and pull a box with
the stick." "Sweep a puck off the table." 提案手法 ベースライン手法 提案手法 ベースライン手法 ☺ 目標位置に物体を移動 不適切な位置で停止 ☺ 物体を把持し移動 物体を把持できず失敗 12
実機実験: モバイルマニピュレーションタスクは困難 ベースの回転・移動に伴う視点変化があり困難 ”Insert the lemon into the cup.
” 提案手法 事前学習データセットに モバイルタスクが少ない 2.4% 𝜋0.5 x2 その他 97.6% 𝜋0.5 の事前学習データセット 13
定量的結果: モバイルマニピュレーションにてベースラインを上回る 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 提案手法 100 100 Mobile
pick Mobile move Mobile open Mobile push 100 100 Mobile close 14
定量的結果: モバイルマニピュレーションにてベースラインを上回る 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 提案手法 100 100平均成功率 100
100 +11pt Mobile pick Mobile move Mobile open Mobile push Mobile close 15
定量的結果: モバイルマニピュレーションにてベースラインを上回る 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 提案手法 100 100平均成功率 100
100 +21pt Mobile pick Mobile move Mobile open Mobile push Mobile close 16
まとめ: 深層状態空間モデルに基づく軽量なVLA ▪ 背景:VLAにおいて推論時のメモリ消費量・推論時間が課題 ▪ 新規性:深層状態空間モデルに基づく軽量なバックボーン ▪ 結果:シミュレーション・実機のタスク成功率でベースライン手法を上回った LLMの事前学習知識を活用するため蒸留を用いた軽量VLA ◼
9/4 13:44~ @大会議室A “Self-Distilled Classificationに基づくVLAの構築” 17
Appendix
定量的結果:高速かつ推論時のメモリ消費量が少ない 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 提案手法 19
定量的結果:高速かつ推論時のメモリ消費量が少ない 𝜋0.5 VLA-Adapter TinyVLA SmolVLA 提案手法 1 × 10 1
× 3 20
Ablation Study: 加速度損失・DeepSSMが性能向上に寄与 w/o 加速度損失 DeepSSM → Transformer 提案手法 21
加速度に対する損失関数を導入した2段階訓練 ▪ 軌道の時間差分を効率的に捉えるため2段階訓練を導入 ▪ Stage1: エンドエフェクタの速度に対するL1損失 ▪ Stage2: 加速度に対するL1損失. .
を追加 22
エラー分析: 対象物体の位置の認識が困難 ▪ シミュレーションの失敗例20例に対してエラー分析を実施 ▪ 物体位置の認識エラー: 空間的に誤った位置にアームを移動させる ▪ 把持点推定の失敗: 適切な位置で把持できず物体を落とす
▪ 動作の未完了: 動作途中でアームが停止する エラーカテゴリ エラー数 物体位置の認識エラー 10 把持点推定の失敗 6 動作の未完了 4 合計 20 23
実機実験:提案手法の失敗例 早くグリッパを閉じる アームでボトルを倒す 24