Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
生成AIなんでも展示会vol6 LT登壇資料 NexteraBERT
Search
Rikka Botan
September 23, 2026
Research
75
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
生成AIなんでも展示会vol6 LT登壇資料 NexteraBERT
生成AIなんでも展示会 vol.6での登壇資料です。
Rikka Botan
September 23, 2026
More Decks by Rikka Botan
See All by Rikka Botan
HackSick vol.7 LT資料【LLMアーキテクチャ入門・事前学習時の躓き所解説】 スパースなAttention・状態空間モデル
rikkabotan7
0
180
【ローカルAIに向き合う展示会vol.2】液体時間定数型モジュールを用いた オリジナルの双方向エンコーダーモデルNexteraBERT 推論速度向上検討並びにダウンストリーム評価
rikkabotan7
0
200
【生成AIなんでも展示会vol.5 LT登壇】NexteraBERT発表資料
rikkabotan7
1
200
SSE: Stable Static Embedding
rikkabotan7
0
44
【ローカルAI LT大会】SSE: Stable Static Embedding ー速度低下を伴わず 静的埋め込みモデルの潜在能力を引き出す Dynamic Tanh手法の提案
rikkabotan7
0
120
SEA Model series Op.1: Saint Lupinus pre-release
rikkabotan7
0
160
Other Decks in Research
See All in Research
秋葉原ウォーカブル基礎調査報告書
izumiyama_lab
1
120
Cross-Media Information Spaces and Architectures
signer
PRO
0
370
2026年度 生成AI を活用した論文執筆ガイド/ワークショップ / 2026 Academic Year Guide to Writing Papers Using Generative AI - Workshop
ks91
PRO
0
230
EIRによる不正端末のブロッキング 5G時代におけるデバイス識別と不正対策の進化
stellarcraft
0
140
全国町字単位空き家率推定データver1.0データ仕様
microbaseinc
0
260
学術バーQ AI研究最前線:自己教師あり学習による画像モデルの事前学習
naok615
0
110
CVPR2026論文紹介_VLMにとって良いvision encoderとは何か?Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
kobayashi31
1
230
Vector Map as Language: Toward Unified Remote Sensing Vector Mapping
satai
3
280
Visual SLAM未来予測 / Future Prediction in Visual SLAM
koide3
1
1k
「AIとWhyを深堀る」をAIと深堀る
iflection
0
650
ros2-perf-multihost: 分散システムにおける客観的なアーキテクチャ評価フレームワーク
takasehideki
0
270
Spatial Active Noise Control Based onSound Field Interpolation Incorporating Physical Constraints
skoyamalab
0
190
Featured
See All Featured
How to Grow Your eCommerce with AI & Automation
katarinadahlin
PRO
2
280
Helping Users Find Their Own Way: Creating Modern Search Experiences
danielanewman
31
3.4k
Discover your Explorer Soul
emna__ayadi
2
1.3k
Chrome DevTools: State of the Union 2024 - Debugging React & Beyond
addyosmani
10
1.3k
A Soul's Torment
seathinner
8
3.6k
How To Speak Unicorn (iThemes Webinar)
marktimemedia
1
580
The World Runs on Bad Software
bkeepers
PRO
72
12k
Testing 201, or: Great Expectations
jmmastey
46
8.3k
Let's Do A Bunch of Simple Stuff to Make Websites Faster
chriscoyier
508
140k
Conquering PDFs: document understanding beyond plain text
inesmontani
PRO
4
3.1k
Deep Space Network (abreviated)
tonyrice
0
320
How to Create Impact in a Changing Tech Landscape [PerfNow 2023]
tammyeverts
56
3.5k
Transcript
わたしのこだわりは、知の架橋と知の転換です
NexteraBERT: Input-Dependent Gating Liquid Mixer and Length-Adaptive Attention for Fast,
Long-Context Bidirectional Encoders 入力依存型Liquid Mixerと長さ適応型Attentionによる 高速な長文対応双方向エンコーダ
Vol.6 生成AIなんでも展示会
自己紹介 / About us り っ か 六花 ぼ た
ん 牡丹 Rikka Botan 独立研究者(機械学習 / 代数学 / 数理論理学) Independent researcher (machine learning / algebra / mathematical logic) ◆趣味 お菓子作り・紅茶・クラシック鑑賞・お洋服 ◆最近の活動 Silver Award: Liquid AI Hackathon Series | Tokyo 記事執筆(Mamba, LFM2 (LTCs) 関連) SSE Modelシリーズの公開 X(Twitter) Portfolio
目録 / Contents 1 主な貢献 / Main Contributions 2 知の架橋
/ Bridging Knowledge 3 知の転換 / A Shift in Thinking 4 経験のユニーク性 / The Uniqueness of Experience
主な貢献 / Main Contributions 独自のアーキテクチャ(NexteraBERT) からなるモデルをフルスクラッチで構築 下記の点で世界最高クラスの性能を示しました。 ・65k tokensの推論でModernBERT-baseより5.22倍 高速
(LFM2.5 Encoder 230Mより2.01倍高速) ・15倍少ないtokensでの学習で、GLUEでは ModernBERT-baseと同等、MTEB v2では凌駕 (OptiBERT比では5.27倍) ・学習長の8倍でもほぼ劣化しない外挿性 (損失 +3.8%のみ増加 (他のモデルは+83%増加)) ・Code & Long Retrieval(CodeSearchNet, MultiLongDocRetrieval)でModernBERT-baseを凌駕
主な貢献 / Main Contributions 技術詳細は発表の大筋からされるため割愛します(詳細は論文参照) 論文
知の架橋 / Bridging Knowledge Self Attentionの非効率性の打開のために、 近年の言語モデルではSelf Attention以外のMixerを 併用するのがデファクトスタンダードになっています。 計算神経科学
(NexteraBERTでもLiquid Time-Constantsを使用) 制御工学 入力に応じた Mamba, Gated Delta Networks, Kimi Delta Attention, 少ない変数で システムの適応的変化 Liquid Time-Constantsなどは 状態の遷移を表現 言語モデルの文脈から現れたものではなく、 制御工学や計算神経科学(神経生物学) Mamba, GDN LTCs から端を発したものです。 言語モデルという、「枠」に囚われたままでは Transformer 辿り着けない領域であり、 他の学術分野をも活用するという視点が 自然言語処理 新たな効率性への扉を開いています。 言葉の構造と意味を 数値処理で扱う
知の転換 / A Shift in Thinking エンコーダーモデルの学習において、 多くのモデルがマスク率が 一定で行われてきました。 →本当に一定が適切?
→マスク率を線形に下げていくと 収束性が高まる。 →より少ない学習で高い性能に到達 前提を疑い、模索することが重要
経験のユニーク性 / The Uniqueness of Experience 遠回りだと思っていた経験が、 成果につながりました。 知を架橋するには、二つ以上の分野を知っている必要があります。 前提を疑うには、その前提の外側を知っている必要があります。
(例えば私は以前、制御工学を学んでいてその視点を少し持っていました。)
経験のユニーク性 / The Uniqueness of Experience あなたの経験の組み合わせは あなたにしかないものです。 あなただけの見方で作った 面白い世界を見せてください。
12