Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
Direct Preference Optimization
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
Henry Cui
February 24, 2024
Science
490
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Direct Preference Optimization
Henry Cui
February 24, 2024
More Decks by Henry Cui
See All by Henry Cui
プロダクション言語モデルの情報を盗む攻撃 / Stealing Part of a Production Language Model
zchenry
1
260
Diffusion Model with Perceptual Loss
zchenry
0
540
レンズの下のLLM / LLM under the Lens
zchenry
0
250
Go with the Prompt Flow
zchenry
0
260
Mojo Dojo
zchenry
0
280
ことのはの力で画像の異常検知 / Anomaly Detection by Language
zchenry
0
760
驚愕の事実!LangChainが抱える問題 / Problems of LangChain
zchenry
0
350
MLOps初心者がMLflowを触る / MLflow Brief Introduction
zchenry
0
230
{{guidance}}のガイダンス / Guidance of guidance
zchenry
0
210
Other Decks in Science
See All in Science
第67回コンピュータビジョン勉強会論文紹介「RoboWheel: A Data Engine from Real-World Human Demonstrations for Cross-Embodiment Robotic Learning」
x_ttyszk
0
230
presen_司法書士学員会.pdf
tagtag
PRO
1
110
サンプル対応のない複数遺伝子発現プロファイルに対するテンソル分解型統合解析の要約
tagtag
PRO
0
260
科学で迫る勝敗の法則-スポーツデータ分析の最前線 (刈谷市連携講座.2026年7月) / The principle of victory discovered by science. at Kariya City, 2027.07
konakalab
0
190
機械学習 - 決定木からはじめる機械学習
trycycle
PRO
0
1.7k
データベース06: SQL (3/3) 副問い合わせ
trycycle
PRO
1
1.2k
ssmonline #51 ヤマサキ春のサメ祭り 2026 / ssmjp Yamasaki Spring JAWS Festival 2026
naospon
1
150
Inside the Mind of an LLM
baggiponte
0
350
HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps
tomoaki0705
0
160
データベース01: データベースを使わない世界
trycycle
PRO
1
1.5k
[第67回 CV勉強会@関東] CV × Scientific Figures / kantoCV 67th CVPR 2026
lychee1223
0
240
因果探索の発展と展望
sshimizu2006
2
1.1k
Featured
See All Featured
Abbi's Birthday
coloredviolet
4
10k
個人開発の失敗を避けるイケてる考え方 / tips for indie hackers
panda_program
123
22k
Navigating the Design Leadership Dip - Product Design Week Design Leaders+ Conference 2024
apolaine
2
450
RailsConf & Balkan Ruby 2019: The Past, Present, and Future of Rails at GitHub
eileencodes
141
35k
jQuery: Nuts, Bolts and Bling
dougneiner
66
8.6k
A Modern Web Designer's Workflow
chriscoyier
699
190k
Public Speaking Without Barfing On Your Shoes - THAT 2023
reverentgeek
1
580
AI: The stuff that nobody shows you
jnunemaker
PRO
10
1.1k
Future Trends and Review - Lecture 12 - Web Technologies (1019888BNR)
signer
PRO
0
3.8k
The Success of Rails: Ensuring Growth for the Next 100 Years
eileencodes
47
8.4k
職位にかかわらず全員がリーダーシップを発揮するチーム作り / Building a team where everyone can demonstrate leadership regardless of position
madoxten
69
66k
SERP Conf. Vienna - Web Accessibility: Optimizing for Inclusivity and SEO
sarafernandez
2
1.6k
Transcript
Direct Preference Optimization 機械学習の社会実装勉強会第32回 Henry 2024/2/24
内容 ▪ NeurIPS 2023 Outstanding Main Track Runner-Ups 受賞 ▪
著者に有名な先生が多い 2
モチベーション ▪ 大量テキストで学習した言語モデルを望ましい挙動に微調整 する必要(Alignment) • 大量コードの平均能力でなく、少量存在の優れたコードに • 一般大衆のもつ誤認識でなく、それを修正すべき ▪ Alignmentを達成するために、現状2段階の複雑な強化学習
手法を使うので、それと理論上等価なシンプルな手法を提案 3
RLHFアプローチの3ステップ ▪ SFT: Supervised fine-tuning ▪ Rewardモデルを学習する • RewardモデルがBradley-Terry (BT)に従う想定
• BTの仮定で導出する損失関数 ▪ RL Fine-tune • Rewardモデルを使って、下記損失関数でfine-tune ▪ 提案法はRewardとRL Fine-tuneをまとめて、rewardモデルを 使わずに学習 4
提案法DPO ▪ RL Fine-tuneの損失関数の最適解 ▪ 上記最適解をrewardモデルを取り出すよう書き換える • Your Language Model
Is Secretly a Reward Model ▪ Rewardモデルを学習する損失関数に代入する • BTモデルのお陰で、Zが消える • Directに言語モデルを最適化できるようになる 5
実験 ▪ 3つのタスクで評価 • controlled sentiment generation • summarization •
single-turn dialogue ▪ 複数スケールのデータセットでRHLFと同等またはそれ以上の 性能を確認 ▪ 多数のオープンソース言語モデルに実装 6