Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
Automating Machine Learning
Search
Sponsored
·
Your Podcast. Everywhere. Effortlessly.
Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
→
Andreas Mueller
July 15, 2016
Science
1.2k
4
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Automating Machine Learning
Andreas Mueller
July 15, 2016
More Decks by Andreas Mueller
See All by Andreas Mueller
PyCon India - Commodity Machine Learning; past, present and future
amueller
0
2.8k
Engineering Scikit-Learn V2
amueller
0
330
Advanced Machine Learning with Scikit-Learn for Pycon Amsterdam
amueller
0
320
Scikit-learn: New project features in 0.17
amueller
0
150
Bootstrapping machine learning
amueller
0
160
PyData Berlin 2014 Keynote: Commodity machine learnin
amueller
0
210
Advanced Machine Learning with Scikit-Learn
amueller
1
780
Machine Learning With Scikit-Learn ODSC SF 2015
amueller
4
1.8k
Machine Learning With Scikit-Learn - Pydata Strata NYC 2015
amueller
1
3k
Other Decks in Science
See All in Science
白金鉱業Meetup Vol.25 【初学者向け発表枠】「で、この施策って効いてるの」に答える効果検証の基礎 ~ATE / ATT / CATE / LATEを現場の問いに翻訳する~
brainpadpr
0
160
Physical AIを支えるWeights & Biases
olachinkei
1
600
[NLP2026 参加報告会] AI for Science まとめ / NLP2026
lychee1223
0
2k
HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps
tomoaki0705
0
150
明治薬科大学講義_ビッグデータ解析を支えるデータベース技術とクラウドコンピューティング
ktatsuya
1
180
機械学習 - DBSCAN
trycycle
PRO
0
2.1k
機械学習 - K-means & 階層的クラスタリング
trycycle
PRO
0
2.1k
Controller Design for Symbolic Input-Output Systems Using Feedforward Neural Networks
konakalab
0
110
「念のためのログ保存」を組織全体でやめるためのポリシーと仕組み作り
i2tsuki
4
380
東北地方における過去20年間の降水量の変化
naokimuroki
1
500
Understanding CVP Waveforms: Interpretation and Clinical Implications in Anesthesiology
taka88
0
880
presen_司法書士学員会.pdf
tagtag
PRO
1
110
Featured
See All Featured
Learning to Love Humans: Emotional Interface Design
aarron
275
41k
SEO Brein meetup: CTRL+C is not how to scale international SEO
lindahogenes
2
2.9k
Agile that works and the tools we love
rasmusluckow
331
22k
Art, The Web, and Tiny UX
lynnandtonic
304
22k
Winning Ecommerce Organic Search in an AI Era - #searchnstuff2025
aleyda
2
2.1k
Rebuilding a faster, lazier Slack
samanthasiow
85
9.6k
Faster Mobile Websites
deanohume
310
32k
What's in a price? How to price your products and services
michaelherold
247
13k
Navigating the moral maze — ethical principles for Al-driven product design
skipperchong
2
540
Odyssey Design
rkendrick25
PRO
2
810
How To Speak Unicorn (iThemes Webinar)
marktimemedia
1
580
The Curious Case for Waylosing
cassininazir
1
510
Transcript
Andreas Mueller (NYU Center for Data Science, scikit-learn) Automatic Machine
Learning?
Why?
Issues with current tools (scikit-learn)
Flow chart / selecting model
Selecting Hyper-Parameters
Scikit-learn: Explicit is better than implicit make_pipeline( OneHotEncoder(), Imputer(), StandardScaler(),
SVC())
What? from automl import AutoClassifier clf = AutoClassifier().fit(X_train, y_train) >
Current Accuracy: 70% (AUC .65) LinearSVC(C=1), 10sec > Current Accuracy: 76% (AUC .71) RandomForest(n_estimators=20) 30sec > Current Accuracy: 80% (AUC .74) RandomForest(n_estimators=500) 30sec
Step 1: Automate Parameter Selection
Step 2: Automate Model Selection
Step 3: Automate Pipeline Selection
How?
Formalizing the Search Space Discrete and Continuous Parameters Conditional Parameters
Fixed pipeline vs flexible pipeline
Formalizing the Search Space Discrete and Continuous Parameters Conditional Parameters
Fixed pipeline vs flexible pipeline
Search Methods
Exhaustive Search (Grid Search)
Randomized Search
Bayesian Optimization (SMBO)
None
None
None
Gaussian Processes
Random Forest Based (SMAC)
Non-parametric (TPE)
None
None
Warm-starting and Meta-learning
Meta-Learning optimization Algorithm + Parameters Dataset 1
Meta-Learning optimization Algorithm + Parameters Dataset 3 optimization Algorithm +
Parameters Dataset 2 optimization Algorithm + Parameters Dataset 1
Meta-Learning Meta-Features 1 optimization Algorithm + Parameters Dataset 3 optimization
Algorithm + Parameters Dataset 2 optimization Algorithm + Parameters Dataset 1 Meta-Features 2 Meta-Features 3 ML model
Meta-Learning Meta-Features 1 optimization Algorithm + Parameters Dataset 3 optimization
Algorithm + Parameters Dataset 2 optimization Algorithm + Parameters Dataset 1 Meta-Features 2 Meta-Features 3 ML model New Dataset ML model Algorithm + Parameters
Meta-Features
Existing Approaches
auto-sklearn (Hutter, Feurer, Eggensperger) http://automl.github.io/auto-sklearn/stable/
Autoweka
Hyperopt-sklearn
TPot
Spearmint https://github.com/HIPS/Spearmint
Scikit-optimize
Within Scikit-learn • GridSearchCV • RandomizedSearchCV • BayesianSearchCV (coming) •
Searching over Pipelines (coming) • Built-in parameter ranges (coming)
TODO Clean separation of: • Model Search Space • Pipeline
Search Space • Optimization Method • Meta-Learning • Exploit prior knowledge better! • Usability • Runtime consideration
TODO Clean separation of: • Model Search Space • Pipeline
Search Space • Optimization Method • Meta-Learning • Exploit prior knowledge better! • Usability • Runtime consideration • Data subsampling
Criticism
Randomized Search works well
Do we need 100 Classifiers? Do we need Complex pipelines?
I don’t want a black-box!
46 http://oreilly.com/pub/get/scipy
47 Material • Random Search for Hyper-Parameter Optimization (Bergstra, Bengio)
• Efficient and Robust Automated Machine Learning (Feurer et al) [autosklearn] • http://automl.github.io/auto-sklearn/stable/ • Efficient Hyperparameter Optimization and Infinitely Many Armed Bandits (Lie et. al) [hyperband] https://arxiv.org/abs/1603.06560 • Scalable Bayesian Optimization Using Deep Neural Networks [Snoek et al]
48 @amuellerml @amueller
[email protected]
http://amueller.io Thank you.