Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
TTS Skins: Speaker Conversion via ASR
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
peisuke
November 20, 2020
Technology
460
0
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
TTS Skins: Speaker Conversion via ASR
Interspeech2020音声読み会発表資料
peisuke
November 20, 2020
More Decks by peisuke
See All by peisuke
EgoX: Egocentric Video Generation from a Single Exocentric Video
peisuke
0
190
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
peisuke
0
280
VGGT: Visual Geometry Grounded Transformer
peisuke
1
2k
AI for Kids:小学生に画像認識を教えてみた話
peisuke
1
120
LangGraphで始めるマルチエージェントシステム
peisuke
14
5.1k
Self-RAG: Learning to Retrieve, Generate and Critique through Self-Reflections
peisuke
9
1.7k
Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields
peisuke
0
14k
LangChain Toolsの運用と改善
peisuke
5
3k
GNeRF: GAN-based Neural Radiance Field without Posed Camera
peisuke
1
870
Other Decks in Technology
See All in Technology
Claude Codeを「使うほど育つ」AI秘書にするノウハウ
minorun365
PRO
33
31k
AI時代の「技術的負債」の変質ー概念の終焉と再解釈、エージェントと共に向かう先
nwiizo
1
3.2k
Issue 駆動でスペシャリストの意図を届ける、AI 実装のアクセシビリティ向上
thkt
0
130
品質と信頼性を地続きにする
grimoh
2
950
Claude in Chrome 入門 / Introduction to Claude in Chrome
cielo1985
0
890
Slack上でインフラをトラブルシュートする! Agentic Platform Engineeringの第一歩
teru0x1
5
1.9k
[Kiro Meetup #7] Kiro Crew Dive Deep
konippi
0
260
30座EKS, 180次升級淬煉的EKS Upgrade Skill 的歷程
eric8230
0
210
研究開発部の紹介 / Sansan R&D Profile
sansan33
PRO
5
25k
Railsのように考える: See through the Master
snoozer05
PRO
5
1.3k
Deployment の 先にある AI Agent 基盤 - kagent vNext、Agent Substrate、Hermes から読み解く Agent Runtime の現在地 / k8s-matsuri-2-ai-agent-platform-amsy810
masayaaoyama
4
710
Antigravity SDK for the Java Developer
glaforge
0
180
Featured
See All Featured
Designing for Timeless Needs
cassininazir
1
480
Digital Projects Gone Horribly Wrong (And the UX Pros Who Still Save the Day) - Dean Schuster
uxyall
1
3k
Art, The Web, and Tiny UX
lynnandtonic
304
22k
Reflections from 52 weeks, 52 projects
jeffersonlam
356
21k
Joys of Absence: A Defence of Solitary Play
codingconduct
1
530
Efficient Content Optimization with Google Search Console & Apps Script
katarinadahlin
PRO
1
880
Effective software design: The role of men in debugging patriarchy in IT @ Voxxed Days AMS
baasie
1
530
Kristin Tynski - Automating Marketing Tasks With AI
techseoconnect
PRO
0
520
コードの90%をAIが書く世界で何が待っているのか / What awaits us in a world where 90% of the code is written by AI
rkaga
63
46k
Testing 201, or: Great Expectations
jmmastey
46
8.3k
Being A Developer After 40
akosma
91
590k
Exploring the relationship between traditional SERPs and Gen AI search
raygrieselhuber
PRO
3
4.3k
Transcript
TTS Skins: Speaker Conversion via ASR Authors: A. Polyak, L.
Wolf, Y. Taigman presenter: @peisuke
2016 ABEJA 2016 Twitter @peisuke Github https://github.com/peisuke Qiita https://qiita.com/peisuke SlideShare
https://www.slideshare.net/FujimotoKeisuke
• • TTS Skins: Speaker Conversion via ASR • •
• ASR WaveNet • • ASR
• Text-to-Speech 100 • • Text-to-Speech • • TTS
• ASR F0
• • Jasper: An End-to-End Convolutional Neural Acoustic Model •
https://github.com/NVIDIA/OpenSeq2Seq • • 1DConv-BN-ReL • Skip-Connection • Pre-trained •
• WaveNet • condition • https://github.com/NVIDIA/nv-WaveNet • • • •
F0 •
• • Look up table pytorch Embedding • • •
F0 • • fine tuning
• • LibriTTS VCTK • • Many-to-many seen unseen •
TTS • • • MOS • Mel cepstral distortion • Speaker classification • • WaveNet AutoEncoder • PPG
Seen • Seen-to-seen • A B • • Identification F0
LibriTTS VCTK MOS MCD Identification MOS MCD Identification Full method 3.78±0.83 96.12 4.08±0.75 8.76±1.72 98.97 w/o F0 3.61±0.83 96.96 3.59±0.96 8.99±1.5 96.89 AE baseline 2.89±0.88 29.19 3.46±1.07 9.45±1.63 69.26 PPG 2.82±0.91 94.01 2.67±0.93 9.19±1.50 98.77 PPG2 2.87±1.00 95.77 3.03±1.06 9.18±1.52 96.24
Uneen • Uneen-to-seen • A B • LibriTTS VCTK MOS
MCD Identification MOS MCD Identification Full method 3.70±0.80 97.10 4.05±0.74 8.94±1.53 98.33 w/o F0 3.67±0.82 97.15 3.62±0.99 9.25±1.62 95.69 AE baseline 3.02±0.89 32.55 3.83±0.91 9.65±1.51 66.20 PPG 2.79±0.93 94.05 2.89±0.93 9.45±1.45 97.45 PPG2 2.71±0.93 95.43 3.19±1.04 9.79±1.86 97.25
TTS • TTS • TTS LibriTTS VCTK MOS MCD Identification
MOS MCD Identification Original TTS 4.25±0.77 10.12±1.27 4.37±0.80 14.52±2.40 Full method 3.67±0.81 8.13±0.95 96.06 4.17±0.88 12.68±2.17 99.25 w/o F0 3.47±0.76 8.43±0.97 96.66 3.75±1.07 13.06±2.26 96.36 AE baseline 3.02±0.84 9.38±1.09 60.26 3.85±1.05 13.81±2.29 75.56 PPG 2.91±0.94 8.52±0.93 96.63 3.50±0.83 12.45±1.92 98.36 PPG2 2.85±0.87 8.76±1.06 95.08 3.66±1.03 12.57±2.10 97.62
• The voice conversion challenge 2018 • 1 81 4-5
• Hub Spoke • Hub Spoke MOS Similarity MOS Similarity Ours 3.84±0.85 2.87±1.14 4.00±0.55 3.14±0.97 N10 3.92±0.75 2.83±1.20 3.98±0.52 3.13±0.97 N17 3.27±0.95 2.77±1.17 3.40±0.88 3.05±0.96
• • TTS • ASR F0 Conditional WaveNet •