Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
Automating Machine Learning
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
Andreas Mueller
July 15, 2016
Science
1.2k
4
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Automating Machine Learning
Andreas Mueller
July 15, 2016
More Decks by Andreas Mueller
See All by Andreas Mueller
PyCon India - Commodity Machine Learning; past, present and future
amueller
0
2.8k
Engineering Scikit-Learn V2
amueller
0
330
Advanced Machine Learning with Scikit-Learn for Pycon Amsterdam
amueller
0
320
Scikit-learn: New project features in 0.17
amueller
0
150
Bootstrapping machine learning
amueller
0
160
PyData Berlin 2014 Keynote: Commodity machine learnin
amueller
0
210
Advanced Machine Learning with Scikit-Learn
amueller
1
780
Machine Learning With Scikit-Learn ODSC SF 2015
amueller
4
1.8k
Machine Learning With Scikit-Learn - Pydata Strata NYC 2015
amueller
1
3k
Other Decks in Science
See All in Science
生成AI・プレプリント時代における 研究成果公開の再設計 ― トップカンファレンス文化はどこへ向かうのか / Redesigning the Dissemination of Research Outputs in the Age of Generative AI and Preprints — Where Is the Top-Conference Culture Heading?
ykiyota
0
30k
科学で迫る勝敗の法則-スポーツデータ分析の最前線 (刈谷市連携講座.2026年7月) / The principle of victory discovered by science. at Kariya City, 2027.07
konakalab
0
160
GKE上でオセロの強化学習やってみた
akasan
0
110
Massey Ratings for Match Outcome Prediction in Table Tennis: Evidence of Greater Stability than the ITTF World Ranking
konakalab
0
150
コーヒー豆様核 (Coffee-bean nuclei) における形態学的サブタイピングと精選・焙煎特性の同定
jagupath
PRO
0
170
機械学習 - SVM
trycycle
PRO
2
1.2k
データベース15: ビッグデータ時代のデータベース
trycycle
PRO
1
570
AlgorithAlgorihms for Decision Making
mickey_kubo
0
140
Visual Linear Algebra - Lecture at Shosen Grande
hiranabe
0
500
Inside the Mind of an LLM
baggiponte
0
330
20260820_アウトカムが二値のデータに対するCausal Impact@LINEヤフー Data Science Share #2 / Causal Impact for Binary Outcomes
brainpadpr
3
1.4k
Build your own LLM, Live, with MicroGPT
ianozsvald
0
150
Featured
See All Featured
Ten Tips & Tricks for a 🌱 transition
stuffmc
1
230
SEO in 2025: How to Prepare for the Future of Search
ipullrank
3
3.8k
Building a A Zero-Code AI SEO Workflow
portentint
PRO
0
720
Typedesign – Prime Four
hannesfritz
42
3.2k
The SEO Collaboration Effect
kristinabergwall1
1
560
Documentation Writing (for coders)
carmenintech
77
5.5k
Visualization
eitanlees
152
17k
Design and Strategy: How to Deal with People Who Don’t "Get" Design
morganepeng
133
19k
Reality Check: Gamification 10 Years Later
codingconduct
0
2.3k
Why Our Code Smells
bkeepers
PRO
340
58k
How to Think Like a Performance Engineer
csswizardry
28
2.8k
Noah Learner - AI + Me: how we built a GSC Bulk Export data pipeline
techseoconnect
PRO
0
430
Transcript
Andreas Mueller (NYU Center for Data Science, scikit-learn) Automatic Machine
Learning?
Why?
Issues with current tools (scikit-learn)
Flow chart / selecting model
Selecting Hyper-Parameters
Scikit-learn: Explicit is better than implicit make_pipeline( OneHotEncoder(), Imputer(), StandardScaler(),
SVC())
What? from automl import AutoClassifier clf = AutoClassifier().fit(X_train, y_train) >
Current Accuracy: 70% (AUC .65) LinearSVC(C=1), 10sec > Current Accuracy: 76% (AUC .71) RandomForest(n_estimators=20) 30sec > Current Accuracy: 80% (AUC .74) RandomForest(n_estimators=500) 30sec
Step 1: Automate Parameter Selection
Step 2: Automate Model Selection
Step 3: Automate Pipeline Selection
How?
Formalizing the Search Space Discrete and Continuous Parameters Conditional Parameters
Fixed pipeline vs flexible pipeline
Formalizing the Search Space Discrete and Continuous Parameters Conditional Parameters
Fixed pipeline vs flexible pipeline
Search Methods
Exhaustive Search (Grid Search)
Randomized Search
Bayesian Optimization (SMBO)
None
None
None
Gaussian Processes
Random Forest Based (SMAC)
Non-parametric (TPE)
None
None
Warm-starting and Meta-learning
Meta-Learning optimization Algorithm + Parameters Dataset 1
Meta-Learning optimization Algorithm + Parameters Dataset 3 optimization Algorithm +
Parameters Dataset 2 optimization Algorithm + Parameters Dataset 1
Meta-Learning Meta-Features 1 optimization Algorithm + Parameters Dataset 3 optimization
Algorithm + Parameters Dataset 2 optimization Algorithm + Parameters Dataset 1 Meta-Features 2 Meta-Features 3 ML model
Meta-Learning Meta-Features 1 optimization Algorithm + Parameters Dataset 3 optimization
Algorithm + Parameters Dataset 2 optimization Algorithm + Parameters Dataset 1 Meta-Features 2 Meta-Features 3 ML model New Dataset ML model Algorithm + Parameters
Meta-Features
Existing Approaches
auto-sklearn (Hutter, Feurer, Eggensperger) http://automl.github.io/auto-sklearn/stable/
Autoweka
Hyperopt-sklearn
TPot
Spearmint https://github.com/HIPS/Spearmint
Scikit-optimize
Within Scikit-learn • GridSearchCV • RandomizedSearchCV • BayesianSearchCV (coming) •
Searching over Pipelines (coming) • Built-in parameter ranges (coming)
TODO Clean separation of: • Model Search Space • Pipeline
Search Space • Optimization Method • Meta-Learning • Exploit prior knowledge better! • Usability • Runtime consideration
TODO Clean separation of: • Model Search Space • Pipeline
Search Space • Optimization Method • Meta-Learning • Exploit prior knowledge better! • Usability • Runtime consideration • Data subsampling
Criticism
Randomized Search works well
Do we need 100 Classifiers? Do we need Complex pipelines?
I don’t want a black-box!
46 http://oreilly.com/pub/get/scipy
47 Material • Random Search for Hyper-Parameter Optimization (Bergstra, Bengio)
• Efficient and Robust Automated Machine Learning (Feurer et al) [autosklearn] • http://automl.github.io/auto-sklearn/stable/ • Efficient Hyperparameter Optimization and Infinitely Many Armed Bandits (Lie et. al) [hyperband] https://arxiv.org/abs/1603.06560 • Scalable Bayesian Optimization Using Deep Neural Networks [Snoek et al]
48 @amuellerml @amueller
[email protected]
http://amueller.io Thank you.