Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Direct Preference Optimization
Search
Henry Cui
February 24, 2024
Science
480
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Direct Preference Optimization
Henry Cui
February 24, 2024
More Decks by Henry Cui
See All by Henry Cui
プロダクション言語モデルの情報を盗む攻撃 / Stealing Part of a Production Language Model
zchenry
1
260
Diffusion Model with Perceptual Loss
zchenry
0
530
レンズの下のLLM / LLM under the Lens
zchenry
0
240
Go with the Prompt Flow
zchenry
0
250
Mojo Dojo
zchenry
0
280
ことのはの力で画像の異常検知 / Anomaly Detection by Language
zchenry
0
750
驚愕の事実!LangChainが抱える問題 / Problems of LangChain
zchenry
0
330
MLOps初心者がMLflowを触る / MLflow Brief Introduction
zchenry
0
230
{{guidance}}のガイダンス / Guidance of guidance
zchenry
0
210
Other Decks in Science
See All in Science
機械学習 - DBSCAN
trycycle
PRO
0
2k
不動産業界における業界特化のデータ整備とAI活用 ─Vertical DataとVertical AI─
estie
1
940
人生を変えた一冊「独学大全」のはなし / Self-study ENCYCLOPEDIA: The Book Which Change My Life #独学大全 #EM推し本
expajp
0
200
因果探索の発展と展望
sshimizu2006
2
1k
データベース01: データベースを使わない世界
trycycle
PRO
1
1.5k
Utiliser Bitcoin sans Internet
rlifchitz
0
360
1. CPC理論の展開と集合的知能モデル(JSAI2026 KS-27 集合的予測符号化と新たな知性の時代)
hayashiyus884
1
350
20260820_アウトカムが二値のデータに対するCausal Impact@LINEヤフー Data Science Share #2 / Causal Impact for Binary Outcomes
brainpadpr
3
1.3k
20260722【JAWS-UG東京 ランチタイムLT会 #37④】AWS Well-Architectedフレームワークに沿った回答をするAIエージェントを作ってみた
nozakijcom
1
140
Conwayの法則を"ちゃんと"使うために — 原典でConwayは何を言っていたのか
bonotake
10
6.8k
Massey Ratings for Match Outcome Prediction in Table Tennis: Evidence of Greater Stability than the ITTF World Ranking
konakalab
0
140
ハミルトン・ヤコビ方程式の解の性質と物理的意味
enakai00
0
890
Featured
See All Featured
Organizational Design Perspectives: An Ontology of Organizational Design Elements
kimpetersen
PRO
1
810
Stop Working from a Prison Cell
hatefulcrawdad
274
21k
The World Runs on Bad Software
bkeepers
PRO
72
12k
Tell your own story through comics
letsgokoyo
1
1.1k
Future Trends and Review - Lecture 12 - Web Technologies (1019888BNR)
signer
PRO
0
3.7k
Designing for humans not robots
tammielis
254
26k
Build your cross-platform service in a week with App Engine
jlugia
234
19k
How GitHub (no longer) Works
holman
316
150k
Design of three-dimensional binary manipulators for pick-and-place task avoiding obstacles (IECON2024)
konakalab
0
570
Embracing the Ebb and Flow
colly
88
5.2k
Understanding Cognitive Biases in Performance Measurement
bluesmoon
32
3k
Keith and Marios Guide to Fast Websites
keithpitt
413
23k
Transcript
Direct Preference Optimization 機械学習の社会実装勉強会第32回 Henry 2024/2/24
内容 ▪ NeurIPS 2023 Outstanding Main Track Runner-Ups 受賞 ▪
著者に有名な先生が多い 2
モチベーション ▪ 大量テキストで学習した言語モデルを望ましい挙動に微調整 する必要(Alignment) • 大量コードの平均能力でなく、少量存在の優れたコードに • 一般大衆のもつ誤認識でなく、それを修正すべき ▪ Alignmentを達成するために、現状2段階の複雑な強化学習
手法を使うので、それと理論上等価なシンプルな手法を提案 3
RLHFアプローチの3ステップ ▪ SFT: Supervised fine-tuning ▪ Rewardモデルを学習する • RewardモデルがBradley-Terry (BT)に従う想定
• BTの仮定で導出する損失関数 ▪ RL Fine-tune • Rewardモデルを使って、下記損失関数でfine-tune ▪ 提案法はRewardとRL Fine-tuneをまとめて、rewardモデルを 使わずに学習 4
提案法DPO ▪ RL Fine-tuneの損失関数の最適解 ▪ 上記最適解をrewardモデルを取り出すよう書き換える • Your Language Model
Is Secretly a Reward Model ▪ Rewardモデルを学習する損失関数に代入する • BTモデルのお陰で、Zが消える • Directに言語モデルを最適化できるようになる 5
実験 ▪ 3つのタスクで評価 • controlled sentiment generation • summarization •
single-turn dialogue ▪ 複数スケールのデータセットでRHLFと同等またはそれ以上の 性能を確認 ▪ 多数のオープンソース言語モデルに実装 6