Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
TTS Skins: Speaker Conversion via ASR
Search
peisuke
November 20, 2020
Technology
460
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
TTS Skins: Speaker Conversion via ASR
Interspeech2020音声読み会発表資料
peisuke
November 20, 2020
More Decks by peisuke
See All by peisuke
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
peisuke
0
260
VGGT: Visual Geometry Grounded Transformer
peisuke
1
1.9k
AI for Kids:小学生に画像認識を教えてみた話
peisuke
1
110
LangGraphで始めるマルチエージェントシステム
peisuke
14
5.1k
Self-RAG: Learning to Retrieve, Generate and Critique through Self-Reflections
peisuke
9
1.6k
Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields
peisuke
0
14k
LangChain Toolsの運用と改善
peisuke
5
3k
GNeRF: GAN-based Neural Radiance Field without Posed Camera
peisuke
1
850
A Quantum Computational Approach to Correspondence Problems on Point Sets
peisuke
0
780
Other Decks in Technology
See All in Technology
「守りたい体験」を渡すだけで E2E を生成させられるようになった話
hinac0
2
910
インシデント事例と パッケージの全量解析に学ぶ ソフトウェアサプライチェーンの守り方 / supply-chain-attack-defense
flatt_security
0
570
副作用のある Lambda でも Lambda Power Tuning は使えるのか / lambda-power-tuning-side-effects
koukihosaka
1
130
しくみを学んで使いこなそう GitHub Copilot app
torumakabe
2
300
『モデル + ハーネス』で読み解く AIエージェント入門
oracle4engineer
PRO
1
110
非定型なドキュメントを効率よくリファクタする 〜えぇ!?仕様書27本の移行が1日で終わったって!?〜
subroh0508
2
590
AIが実装を自走する時代の認知負債との戦い
lycorptech_jp
PRO
2
990
Multicaで30個のミニプロジェクトをAIエージェント運用して見えてきたこと
eiei114
1
560
AI Native なプロダクト組織の立ち上げ方 : 生産性 100 倍への挑戦
mikesorae
0
770
公式ドキュメントの歩き方etc
coco_se
1
120
Playwright × AI Agent でE2Eテストはどう変わるか AI駆動テストの可能性と実用検証の結果
taiga7543
2
690
10年目を迎えた「ABEMA」がどのように AI 活用を推進して、AI 駆動開発にシフトしているのか / How ABEMA, entering its 10th year, is promoting the use of AI and shifting toward AI-driven development
miyukki
0
320
Featured
See All Featured
Navigating Team Friction
lara
192
16k
The agentic SEO stack - context over prompts
schlessera
0
850
世界の人気アプリ100個を分析して見えたペイウォール設計の心得
akihiro_kokubo
PRO
72
40k
We Have a Design System, Now What?
morganepeng
55
8.2k
The Director’s Chair: Orchestrating AI for Truly Effective Learning
tmiket
1
220
Digital Projects Gone Horribly Wrong (And the UX Pros Who Still Save the Day) - Dean Schuster
uxyall
1
2.1k
Tips & Tricks on How to Get Your First Job In Tech
honzajavorek
1
610
How GitHub (no longer) Works
holman
316
150k
RailsConf & Balkan Ruby 2019: The Past, Present, and Future of Rails at GitHub
eileencodes
141
35k
Lessons Learnt from Crawling 1000+ Websites
charlesmeaden
PRO
1
1.4k
Music & Morning Musume
bryan
47
7.3k
SEO Brein meetup: CTRL+C is not how to scale international SEO
lindahogenes
1
2.8k
Transcript
TTS Skins: Speaker Conversion via ASR Authors: A. Polyak, L.
Wolf, Y. Taigman presenter: @peisuke
2016 ABEJA 2016 Twitter @peisuke Github https://github.com/peisuke Qiita https://qiita.com/peisuke SlideShare
https://www.slideshare.net/FujimotoKeisuke
• • TTS Skins: Speaker Conversion via ASR • •
• ASR WaveNet • • ASR
• Text-to-Speech 100 • • Text-to-Speech • • TTS
• ASR F0
• • Jasper: An End-to-End Convolutional Neural Acoustic Model •
https://github.com/NVIDIA/OpenSeq2Seq • • 1DConv-BN-ReL • Skip-Connection • Pre-trained •
• WaveNet • condition • https://github.com/NVIDIA/nv-WaveNet • • • •
F0 •
• • Look up table pytorch Embedding • • •
F0 • • fine tuning
• • LibriTTS VCTK • • Many-to-many seen unseen •
TTS • • • MOS • Mel cepstral distortion • Speaker classification • • WaveNet AutoEncoder • PPG
Seen • Seen-to-seen • A B • • Identification F0
LibriTTS VCTK MOS MCD Identification MOS MCD Identification Full method 3.78±0.83 96.12 4.08±0.75 8.76±1.72 98.97 w/o F0 3.61±0.83 96.96 3.59±0.96 8.99±1.5 96.89 AE baseline 2.89±0.88 29.19 3.46±1.07 9.45±1.63 69.26 PPG 2.82±0.91 94.01 2.67±0.93 9.19±1.50 98.77 PPG2 2.87±1.00 95.77 3.03±1.06 9.18±1.52 96.24
Uneen • Uneen-to-seen • A B • LibriTTS VCTK MOS
MCD Identification MOS MCD Identification Full method 3.70±0.80 97.10 4.05±0.74 8.94±1.53 98.33 w/o F0 3.67±0.82 97.15 3.62±0.99 9.25±1.62 95.69 AE baseline 3.02±0.89 32.55 3.83±0.91 9.65±1.51 66.20 PPG 2.79±0.93 94.05 2.89±0.93 9.45±1.45 97.45 PPG2 2.71±0.93 95.43 3.19±1.04 9.79±1.86 97.25
TTS • TTS • TTS LibriTTS VCTK MOS MCD Identification
MOS MCD Identification Original TTS 4.25±0.77 10.12±1.27 4.37±0.80 14.52±2.40 Full method 3.67±0.81 8.13±0.95 96.06 4.17±0.88 12.68±2.17 99.25 w/o F0 3.47±0.76 8.43±0.97 96.66 3.75±1.07 13.06±2.26 96.36 AE baseline 3.02±0.84 9.38±1.09 60.26 3.85±1.05 13.81±2.29 75.56 PPG 2.91±0.94 8.52±0.93 96.63 3.50±0.83 12.45±1.92 98.36 PPG2 2.85±0.87 8.76±1.06 95.08 3.66±1.03 12.57±2.10 97.62
• The voice conversion challenge 2018 • 1 81 4-5
• Hub Spoke • Hub Spoke MOS Similarity MOS Similarity Ours 3.84±0.85 2.87±1.14 4.00±0.55 3.14±0.97 N10 3.92±0.75 2.83±1.20 3.98±0.52 3.13±0.97 N17 3.27±0.95 2.77±1.17 3.40±0.88 3.05±0.96
• • TTS • ASR F0 Conditional WaveNet •