Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
テキストメディア特論 「会社名」の抽出
Search
Lamron
October 01, 2023
Research
140
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
テキストメディア特論 「会社名」の抽出
Lamron
October 01, 2023
More Decks by Lamron
See All by Lamron
テキストメディア特論 類似した「名前」の同一性の判定
lamrongol
0
94
Blueskyでは何が話し合われているか。「情報技術は民主主義を生み、今は殺そうとしている」
lamrongol
0
7.7k
要約: Formal Approaches in Categorization: Chapter.5 Semantics without categorization
lamrongol
0
4.6k
Blueskyの「今」がわかる!Bot
lamrongol
0
2k
Other Decks in Research
See All in Research
Anthropic が提案する LLM の内部状態を自然言語で説明可能にした Natural Language Autoencoders / Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
shunk031
0
190
typst の使い方:言語学を研究する学生のために
gitomochang
0
570
重要だけど測れていないもの:高齢者ケアの見えない課題
theoriatec2024
0
490
Claude Code × autoresearch 実践
mathbullet
0
260
Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance
satai
3
110
MIRU2026 チュートリアル講演2:三次元データ処理の動向
nnchiba
6
4.6k
[最先端NLP勉強会2026] Checklists Are Better Than Reward Models For Aligning Language Models
nzw0301
1
300
Karkada さんの論文 × 2 の紹介: (1) Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models, (2) Symmetry in language statistics shapes the geometry of model representations
eumesy
PRO
1
590
260624_NLP-colloquium: Hubness
de9uch1
2
210
Source Code Diff Revolution
tsantalis
0
120
HackSick vol.7 LT資料【LLMアーキテクチャ入門・事前学習時の躓き所解説】 スパースなAttention・状態空間モデル
rikkabotan7
0
170
Research Engineerという仕事 / Research Engineering: Bridging Research and Business
chck
1
290
Featured
See All Featured
How People are Using Generative and Agentic AI to Supercharge Their Products, Projects, Services and Value Streams Today
helenjbeal
1
300
The Anti-SEO Checklist Checklist. Pubcon Cyber Week
ryanjones
0
220
Exploring anti-patterns in Rails
aemeredith
3
480
No one is an island. Learnings from fostering a developers community.
thoeni
21
3.8k
How to make the Groovebox
asonas
2
2.4k
Chrome DevTools: State of the Union 2024 - Debugging React & Beyond
addyosmani
10
1.3k
Building Adaptive Systems
keathley
44
3.2k
Groundhog Day: Seeking Process in Gaming for Health
codingconduct
0
310
Kristin Tynski - Automating Marketing Tasks With AI
techseoconnect
PRO
0
490
So, you think you're a good person
axbom
PRO
2
2.1k
The Art of Programming - Codeland 2020
erikaheidi
57
14k
Navigating Weather and Climate Data
rabernat
0
500
Transcript
「会社名」の抽出 @lamrongol
「~社」などの表現から会社名を判断する方法には限界 がある 切れ目の判断が難しい(「・」は切れ目か否か、など) 「オラクル」のように「~社」の形になってないものは社名と判 断できない 「東電」などの略称もある
あらかじめどのような会社名があるか登録しておけばよ い
Wikipedia の利用 Wikipediaの特徴 各項目には多くの場合「千葉県の会社」などカテゴリが 付与されている 一定の規則に基づいた文書が大量にある
人手による更新・訂正が行われるので正確性がある程 度保証されている 大量の「会社名」データを手に入れることができる (Wikipediaのデータベース・ダンプを利用)
略称の取得 略称と正式名称の関連も取得できる 例)「日立」というリンクから「日立製作所」につな がっている場合 「日立」=「日立製作所」と関連付けられる
Wikipedia以外からの取得 Web上にはWikipedia以外の文書も大量にある しかし、それらはWikipediaのように「企業」であることが 明記されてるわけではない だが、量は圧倒的に多いのでなんとか活用したい 周りの文章から「会社名」であることを判断できな
いか? 「〇〇は東証一部に上場した~」 「〇〇は1997年に創業した~」
構造化されてない文章からの会社名の取得 まず、Wikipediaなど構造化されているデータを「訓 練データ」として用いる 前後の単語から、会社名を判断する確率モデルを作 る 構造化されてないデータ(ブログの文章等)に対して これを適用し、会社名を取り出す
P(会社名|創業)= N(会社名∧創業) N(創業)
関連研究の応用 Support Vector Machineを用いた日本語固有表 現抽出[山田 et al] 前後の単語の素性(単語自体だけでなく、品詞の
種類なども含む)ベクトルの集合に対してSVMを行 い、学習させる