Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
Spark Machine Learning 101 @HadoopCon
Search
Chu-Yu Hsu
September 19, 2015
Technology
420
1
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Spark Machine Learning 101 @HadoopCon
Chu-Yu Hsu
September 19, 2015
Other Decks in Technology
See All in Technology
新たなDBアーキテクチャ「LTAP」にDeep Dive!!
inoutk
0
150
それでも、技術なブログを書く理由 #kichijojipm / Why I Still Write Tech Blogs Even Now
shinkufencer
0
1.2k
ソフトウェアアーキテクチャ研修【MIXI 26新卒技術研修】
mixi_engineers
PRO
2
990
【5分でわかる】セーフィー エンジニア向け会社紹介
safie_recruit
0
53k
ウォーターフォール開発案件のPMとしてAI活用を模索している話
hatahata021
2
200
QAと開発の両側から進める AI活用 -QAプロセスAI支援ツールキットと Inner Loop / Outer Loopの取り組み-
legalontechnologies
PRO
2
360
脱Jenkins、インターン生が挑んだCIツールGitHubActions移行
mixi_engineers
PRO
1
260
plamo-3-translateの開発
pfn
PRO
0
150
BigQuery を検索ソースとした AI Agent の作り方って 〇〇 通りあんねん
satohjohn
0
140
AI Agent を本番環境へ―― Microsoft Foundry × Azure Serverless で作る Enterprise-Ready な基盤
shibayan
PRO
1
830
文字起こし基盤の信頼性
abnoumaru
0
150
Claude Mythos、Fable...フロンティアAIの最新動向と企業のセキュリティ対策
flatt_security
0
170
Featured
See All Featured
Leading Effective Engineering Teams in the AI Era
addyosmani
9
2.2k
AI Search: Where Are We & What Can We Do About It?
aleyda
0
7.7k
A Tale of Four Properties
chriscoyier
163
24k
WENDY [Excerpt]
tessaabrams
11
39k
No one is an island. Learnings from fostering a developers community.
thoeni
21
3.8k
Raft: Consensus for Rubyists
vanstee
141
7.6k
Color Theory Basics | Prateek | Gurzu
gurzu
0
400
Chasing Engaging Ingredients in Design
codingconduct
0
240
JavaScript: Past, Present, and Future - NDC Porto 2020
reverentgeek
52
6k
Marketing to machines
jonoalderson
1
5.6k
Optimising Largest Contentful Paint
csswizardry
37
3.8k
Let's Do A Bunch of Simple Stuff to Make Websites Faster
chriscoyier
508
140k
Transcript
Spark Machine Learning 101 Chu-Yu Hsu @ HadoopCon 2015
About Me Chu-Yu Hsu, 許儲⽻羽 • Software Engineer • Machine
Learning Practicer • Used Spark ML and Python in daily work and Kaggle competition • http://blog.chuyuhsu.ml
Outline • Introduction to Spark ML • Alternative Least Squares
(ALS) • Hands-on example
None
Apache Spark MLlib • To Make practical machine learning easy
and scalable • spark.mllib - the primary API • spark.ml - a higher-level API for constructing ML workflows Apache Spark spark.mllib spark.ml
What’s in MLlib Utilities Data types Basic statistics Classification and
regression SVM Logistic regression Linear regression Naive Bayes Decision trees Ensembles of trees Isotonic regression Collaborative filtering Alternating least squares (ALS) Clustering K-means Gaussian mixture Power iteration clustering Latent Dirichlet allocation Streaming k-means Dimensionality reduction SVD PCA Frequent pattern mining FP-growth Optimization Stochastic gradient descent Limited-memory BFGS https://spark.apache.org/docs/latest/mllib-guide.html
ML Workflow can be VERY complex
Types of Recommenders • Editorial and hand curated • Simple
aggregates • Tailored to individual users
Who Uses Recommenders
Approaches • Content based method • Item based method •
Model based method
Collaborative Filtering • One of mostly known “Recommendation Algorithm” •
Widely used in E-commerce application • The data size can be enormous • Need to be delivered as soon as possible
Collaborative Filtering Main idea: Find set N of other users
whose ratings are “similar” to X’s ratings
Users Preferences • This is a baby example • Users:
> 2M • Items: > 30M • Sparsity: > 2%
Low Rank Assumption • Matrix can be reduced to the
product of low rank matrixes • That is also understood as “latent factors” • We assume that the low factor can represent the hidden factors we do not know Action Romance Thriller
Low Rank Assumption Action Romance Thriller Action Romance Thriller
Matrix Factorization
• Our goal is to find P and Q such
that (Sum of Square Error): • Root Mean Square Error (RMSE)
Alternative Least Squares • Because p and q are both
unknown, the object function is not convex • If fix one of the unknowns > can be solved as a least squares problem
Amazon Reviews Dataset 35 million ratings, 6.6 million users, 2.4
million products on 16-node (m3.2xlarge) https://github.com/apache/spark/pull/3720
Resources
Resources
And More Resources • Source code examples https://github.com/apache/spark/tree/master/ examples •
Apache Spark JIRA https://issues.apache.org/jira/browse/spark
Dataset • MovieLens Dataset http://grouplens.org/datasets/movielens/ • “ratings.dat” UserID::MovieID::Rating::Timestamp • “movies.dat”
MovieID::Title::Genres
Conclusion • Spark MLlib grows fast, but still need some
time • Spark MLlib is a strong tool, if you use it right • Sharpening ML skills is first priority
Q&A Visit me on: http://blog.chuyuhsu.ml Github: http://github.com/ChuyuHsu Thanks
References • https://spark.apache.org/docs/latest/mllib-guide.html • http://www.slideshare.net/jeykottalam/mllib • http://www.slideshare.net/PetrZapletal1/mllib-and-machine-learning-on-spark • https://databricks.com/blog/2014/07/23/scalable-collaborative-filtering-with- spark-mllib.html
• https://github.com/apache/spark/pull/3720 • https://www.hakkalabs.co/articles/spark-mllib-making-practical-machine- learning-easy-and-scalable • http://www.slideshare.net/databricks/practical-machine-learning-pipelines- with-mllib