Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
論文紹介_Learning Dynamic Contextualised Word Embed...
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
ShitoRyo
January 10, 2024
Research
160
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
論文紹介_Learning Dynamic Contextualised Word Embeddings via Template-based Temporal Adptation
ShitoRyo
January 10, 2024
More Decks by ShitoRyo
See All by ShitoRyo
論文紹介_LSC-Eval: A General Framework to Evaluate Methods for Assessing Dimensions of Lexical Semantic Change Using LLM-Generated Synthetic Data
lexusd
0
38
Tutorial of Coding Environment for Research by Docker
lexusd
0
54
Computational Approaches for Diachronic Semantic Change Detection_2024_8
lexusd
0
64
論文紹介_Are Embedded Potatoes Still Vegetables_ On the Limitation of WordNet Embeddings for Lexical Semantics
lexusd
0
170
論文紹介_Interpretable Word Sense Representations via Definition Generation_ The Case of Semantic Change Analysis
lexusd
0
140
論文紹介_Twitter Topic Classification
lexusd
0
120
論文紹介_What is Done is Done_ an Incremental Approach to Semantic Shift Detection
lexusd
0
140
Demoの作り方_研究会チュートリアル
lexusd
0
180
論文紹介_Ruddit_Norms of Offensiveness for English Readdit Comments
lexusd
0
77
Other Decks in Research
See All in Research
超効率化への挑戦:1bit LLMの現状と展望
yumaichikawa
0
650
データサイエンティストの就労意識~2015 → 2026 一般(個人)会員アンケートより
datascientistsociety
PRO
0
750
論文紹介: Understanding Epistemic Language with a Language-augmented Bayesian Theory of Mind
hisaokatsumi
0
150
Research Engineerという仕事 / Research Engineering: Bridging Research and Business
chck
1
290
Model Discovery and Graph Simulation: A Lightweight Gateway to Chaos Engineering
anatolykr
0
290
LINEヤフー データサイエンス Meetup「三井物産コモディティ予測チャレンジ」の舞台裏-AlpacaTechパート
gamella
1
650
Vector Map as Language: Toward Unified Remote Sensing Vector Mapping
satai
3
260
typst の使い方:言語学を研究する学生のために
gitomochang
0
570
XDPerf: A High-Performance Traffic Generator Built with WASM and eBPF
takehaya
1
290
Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance
satai
3
130
進学校の生徒にはア行の苗字が多いのか
ozekinote
0
560
SAKURAONE:An Open Ethernet-based AI HPC System And Its Observed Workload Dynamicsin a Single-Tenant LLM Development Environment
yuukit
1
580
Featured
See All Featured
Why You Should Never Use an ORM
jnunemaker
PRO
61
10k
Build The Right Thing And Hit Your Dates
maggiecrowley
39
3.4k
So, you think you're a good person
axbom
PRO
2
2.1k
Making the Leap to Tech Lead
cromwellryan
135
10k
Design of three-dimensional binary manipulators for pick-and-place task avoiding obstacles (IECON2024)
konakalab
0
570
Getting science done with accelerated Python computing platforms
jacobtomlinson
2
470
Claude Code のすすめ
schroneko
67
230k
How to make the Groovebox
asonas
2
2.4k
Build your cross-platform service in a week with App Engine
jlugia
234
19k
svc-hook: hooking system calls on ARM64 by binary rewriting
retrage
2
550
Put a Button on it: Removing Barriers to Going Fast.
kastner
60
4.6k
B2B Lead Gen: Tactics, Traps & Triumph
marketingsoph
0
230
Transcript
ACL 2023 2023.11.29 2024.1.10 M2 凌 志棟 1
概要 この論文何やった: 時期や社会環境の違いによる語義変化に言語モデルを適応させるために、 Promptを用いたDynamic Contextualized Word Embedding (DCWE)の学習 貢献 •
Promptを利用してMLMを時間適応するための方法を提案 • 先行研究の手法より性能がよい+効率が良い 2 時間や社会などの言語外要素 に対応する表現 文脈を考慮した単語表現
関連研究 • Dynamic Word Embedding (DWE) ◦ Word2VecやLSTMを学習するときに、言語外の情報(時間・社会)を Encode [Welch
et al.2020] ◦ 社会要因より時間が語義に多く影響を与える [Hoffman et al. 2021] • Dynamic Contextualized Word Embedding (DCWE) ◦ DCWEs: 時間・社会情報をType-based表現にEncodeし、Token-based表現に変換[Hoffman et al. 2021] ◦ ↑以前相田さんが紹介した ◦ TempoBERT:訓練テキストに時期 Tokenを加え、それをMaskしてBERTに当てさせる[Rosin et al.2022] 今回の提案手法はContextualized Word Embedding を時間適応することを目的 3
Prompt-based Time Adaptation Main idea:2つの時期に意味変化が起きた頻出単語を使ってPromptを作る。 異なる時期T1 T2のコーパスC1 C2に対して Pivot単語w 、
Anchor単語u,v wはC1 C2に頻出な単語、u,v はC1,C2においてwと関連する単語 このような(w,u,v)をTupleといい、これによってpromptを作成 4
Tuple Selection Methods Frequency-based Diversity-based Context-based 5
Tuple Selection Methods Frequency-based • Pivot単語w: すべての単語wのScoreを降順して上位k個を選ぶ • Anchors単語u,v: •
Frequency-based Tuple集合が 6
Tuple Selection Methods Diversity-based • 意味変化した単語w = uとvの集合 が違うものが欲しい ◦ 式(1)を計算したスコアのTopにある単語を選択し、
DiversityスコアでRe-rankしTop-kを取る Diversity-based Tuple集合が 7
Tuple Selection Methods Context-based • PMIでAnchorを探すのは2つの問題がある ◦ コーパス内の低頻度語に対応しにくい ◦ PMIは一回2つの単語しか扱えない、他の文脈語に対応できない
• 単語xの平均ベクトル: • Tuple(w,u,v)に対して、C1の単語をw1,u1,v1、C2の単語をw2,u2,v2 g(a,b)はaとbのCos類似度 このように得られたTupleは 8
Prompts Generation Prompts from manual templates Prompts from automatic templates
9
Prompts Generation Prompts from manual templates • 人手で書いたテンプレートに穴埋め:e.g. <w> is
associated with <u> in <T1>, whereas it is associated with <v> in <T2> <〇>にTuple(w,u,v)とu,vの時期T1 T2を入れる 10
Prompts Generation Prompts from automatic templates • Tuple(w,u,v)用いて、T5でPromptを自動生成 • 変換ルール に従って生成
uの用例S 1 とvの用例S 2 をそれぞれC1,C2から抽出 最後にBeam searchで多様なPromptを獲得する 11
Examples of Prompts 12 人手で書いたPromptは長い 長文のLikelihoodが低くなるので、生成されたPromptsは短い傾向
Time Adaptation By Fine-Tuning のTuplesを使って穴埋め・生成したPromptsでMLMをFine-Tuning • ランダムにPromptの一個Tokenをマスクしてモデルに当てさせる • PromptのAnchor単語だけをマスクする方法も試したが、結果に差はなかった 13
Expriments Datasets • Yelp: 2010 & 2020 • Reddit: 2019.9~2020.4
• ArXiv: 2001 & 2020 • Ciao: 2000 & 2011 Evaluation Metric T2でのPerplexity : lower the better 14 Baselines: • BERT-base-uncased • BERT(T1):T1でFine-tuning • BERT(T2):T2でFine-tuning • FT(model,template):提案手法 Hyperparameters • weight decay=0.01 • batch size=4 • learning rate=3x10-8 • k={500,1000,2000,5000,10000} • Epoch=20
Results 15
Results 16
Results 17
Results 18
Conclusion まとめ • 本研究は人手作成と自動生成したPromptでMLMを時間適応する手法を提案 • 複数のデータセットで先行研究より低いperplexityを得られた 今後の課題 • 多言語モデルに適用できるように提案手法を拡張 19