Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
【論文紹介】SPECTER: Document-level Representation Le...
Search
Kaito Sugimoto
November 02, 2020
Research
620
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
【論文紹介】SPECTER: Document-level Representation Learning using Citation-informed Transformers
研究室の日本語輪読会で発表したスライドです。
内容に問題や不備がある場合は、お手数ですが hellorusk1998 [at] gmail.com までご連絡お願いいたします。
Kaito Sugimoto
November 02, 2020
More Decks by Kaito Sugimoto
See All by Kaito Sugimoto
ChatGPTを活用した病院検索体験の改善 〜病院探しをもっと楽しく〜
hellorusk
0
170
【論文紹介】Word Acquisition in Neural Language Models
hellorusk
0
380
【論文紹介】Toward Interpretable Semantic Textual Similarity via Optimal Transport-based Contrastive Sentence Learning
hellorusk
0
320
【論文紹介】Unified Interpretation of Softmax Cross-Entropy and Negative Sampling: With Case Study for Knowledge Graph Embedding
hellorusk
0
590
【論文紹介】Modeling Mathematical Notation Semantics in Academic Papers
hellorusk
0
380
【論文紹介】Detecting Causal Language Use in Science Findings / Measuring Correlation-to-Causation Exaggeration in Press Releases
hellorusk
0
210
【論文紹介】Efficient Domain Adaptation of Language Models via Adaptive Tokenization
hellorusk
0
550
【論文紹介】SimCSE: Simple Contrastive Learning of Sentence Embeddings
hellorusk
0
1.2k
【論文紹介】Automated Concatenation of Embeddings for Structured Prediction
hellorusk
0
340
Other Decks in Research
See All in Research
セマンティック通信勉強会 6Gに向けたデバイス間効率的な通信の技術紹介・課題・今後展望
satai
3
300
「AIとWhyを深堀る」をAIと深堀る
iflection
0
590
LA-Bench 2025:実験指示から実行可能手順を生成するためのデータセット/LA-Bench 2025: A Dataset for Generating Executable Experimental Procedures from Experimental Instructions
stktu
0
140
Karkada さんの論文 × 2 の紹介: (1) Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models, (2) Symmetry in language statistics shapes the geometry of model representations
eumesy
PRO
1
370
2026年度 生成AI を活用した論文執筆ガイド/ワークショップ / 2026 Academic Year Guide to Writing Papers Using Generative AI - Workshop
ks91
PRO
0
220
Vector Map as Language: Toward Unified Remote Sensing Vector Mapping
satai
3
210
RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent
satai
3
500
SOTAのさらに先へ:厳しい推論制約下での高性能モデルのPost-Training
analokmaus
0
1.5k
Sleuthcon Keynote - How Cybercriminals (ab)use AI
fr0gger
0
300
最先端NLP 2026 論文紹介: Wait, Wait, Wait... Why Do Reasoning Models Loop? / SNLP Paper Review: Wait, Wait, Wait... Why Do Reasoning Models Loop?
tkng
0
140
Apache Gravitinoで実現する Icebergカタログ統合とアクセスの一元化
matsumooon
0
470
医療LLMの現在地〜最新研究から社会実装までを考える〜
kento1109
1
1.7k
Featured
See All Featured
Digital Ethics as a Driver of Design Innovation
axbom
PRO
1
390
Marketing to machines
jonoalderson
1
5.7k
Helping Users Find Their Own Way: Creating Modern Search Experiences
danielanewman
31
3.3k
Product Roadmaps are Hard
iamctodd
55
12k
A Modern Web Designer's Workflow
chriscoyier
698
190k
Stop Working from a Prison Cell
hatefulcrawdad
274
21k
技術選定の審美眼(2025年版) / Understanding the Spiral of Technologies 2025 edition
twada
PRO
120
120k
svc-hook: hooking system calls on ARM64 by binary rewriting
retrage
2
530
Why Mistakes Are the Best Teachers: Turning Failure into a Pathway for Growth
auna
0
220
[Rails World 2023 - Day 1 Closing Keynote] - The Magic of Rails
eileencodes
38
3k
Visualizing Your Data: Incorporating Mongo into Loggly Infrastructure
mongodb
49
10k
Unlocking the hidden potential of vector embeddings in international SEO
frankvandijk
0
900
Transcript
SPECTER: Document-level Representation Learning using Citation-informed Transformers Cohan et al.,
ACL 2020 杉本 海人 Aizawa Lab. B4 2020/11/02 1 / 17
読んだ論文 ACL 2020 https://www.aclweb.org/anthology/2020.acl-main.207.pdf 2 / 17
どんな論文? • 文書間の関係の情報(引用ネットワークなど)を BERT に取り入 れて, document representation を生成する方法を新たに提案 •
論文の分類や推薦などの多くの downstream task で有効性を確認 なぜ読んだか: • context-aware citation recommendation という, 論文の特定の位置 からその箇所に対応づけるべき論文を選ぶタスクに興味を持っ ているが, BERT はまだ殆ど使われていない. 論文の埋め込みは BERT でどのように行うのが良いのかに興味 があった 3 / 17
論文の背景 • BERT のような pre-trained のニューラル言語モデルが word や sentence 単位の埋め込みにおいて有用であることは広く研究さ
れてきたが, document 全体の埋め込みに関しては相対的に研究が 少ない • 特に scientific paper analysis において, 引用ネットワークの埋め込 み自体は Graph Convolutional Network など研究されてきたが, そ れを BERT の 学習時に活かせていなかった 4 / 17
関連研究 hyperdoc2vec: Distributed Representations of Hypertext Documents (ACL 2018): (BERT
ではないが引用の情報を意識して学習) 5 / 17
関連研究 A Context-Aware Citation Recommendation Model with BERT and Graph
Convolutional Networks (2019) (BERT と GCN を組み合わせて論文推薦, ここでは引用の情報は BERT とは別物になっている) 6 / 17
提案手法 SPECTER: Scientific Paper Embeddings using Citation-informed TransformERs • 論文の埋め込みを
Transformer ベースで得る新たな手法 • Transformer を SciBERT(Semantic Scholar の論文で pre-trained さ れた BERT)で初期化する • SciBERT はすでに論文の中身に関する言語情報を獲得している と考えられるが, 論文間の関係情報は一切考慮していない。これ を考慮できるようにさらに学習する 7 / 17
提案手法: Training クエリの論文 PQ だけでなく, positive paper P+, negative paper
P− も加 えた 3 つ組 を入力として使う。 8 / 17
提案手法: Training P+: PQ が引用した論文 P−: 2 種類の選び方がある。1 つは, PQ
が引用していない論文からラ ンダムに 1 つ選ぶ。 もう 1 つは, P+ が引用しているにもかかわらず PQ が引用していない 論文からランダムに 1 つ選ぶ(hard neagtives) 。もし全くクエリに関 係ない論文なら, そのクエリが引用した論文とも全く関係ないのは自 明であるが, hard negatives では自明でない例を学習するということに なる。 9 / 17
提案手法: Training PQ, P+, P− それぞれの論文を入力として, Transformer モデルに入れ, [CLS] トークンの出力から埋め込みを得る.
入力形式は, 基本的には「論文のタイトル + abstract」としている。後 の実験で, abstract を使わないタイトルのみ場合や author(著者), venue(会議名)のメタ情報を入力に加えた場合とも比較している。 これら 3 つの埋め込みについて, 以下のような TripletMarginLoss を計 算し, back propagation する. 10 / 17
提案手法: Evaluation • クエリ論文 P を学習した Transformer モデルに入れ, [CLS] トーク
ンの出力から埋め込みを得る • 推論時には引用ネットワーク情報が不要というのがポイント 11 / 17
実験: pre-trained model の作成 • Semantic Scholar から 146K のクエリ論文を訓練用に,
32K の論文 を validation 用に抽出した。クエリ論文 1 つに対し, 最大 5 つの PQ, P+, P− の 3 つ組を作成。 5 つの 3 つ組のうち 2 つは hard negatives, 3 つは easy negatives となっている。 累計 684K の 3 つ組を訓練用に, 145K を validation 用に用意した • https://github.com/allenai/specter 12 / 17
実験: タスク・データセット scientific paper embeddings を包括的に評価するための新たなフレーム ワークである SCIDOCS を用意した(この論文のもう 1
つのポイ ント) 。 SCIDOCS では論文に関する 7 つのタスクで評価する。 • MeSH Classification • Paper Topic Classification • Citation Prediction (Direct Citations) • Citation Prediction (Co-Citations) • User Activity (Co-Views) • User Activity (Co-Reads) • Recommendation 13 / 17
結果 14 / 17
分析: Ablation Study 15 / 17
分析: Visualization 16 / 17
分析: Comparison with Task Specific Fine-Tuning 17 / 17