Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Features
Speaker Deck
PRO
Sign in
Sign up for free
Search
Search
etcd - mission-critical key-value store - CoreO...
Search
Brandon Philips
May 09, 2016
Programming
420
3
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
etcd - mission-critical key-value store - CoreOS Fest 2016
Brandon Philips
May 09, 2016
More Decks by Brandon Philips
See All by Brandon Philips
Node.js Workflow with Minikube and Skaffold
philips
0
300
Manage the App on Kubernetes
philips
0
370
Production Backbone Monitoring Containerized Apps
philips
0
230
KubeCon EU 2017: Dancing on the Edge of a Volcano
philips
1
870
rkt - KubeCon EU keynote - 2017
philips
1
310
FOSDEM_Keynote_2017-_.pdf
philips
0
170
Tectonic Summit Day 2 Keynote
philips
0
420
Kubernetes: Simple to Manage Anywhere (self-hosted, Tectonic upgrade demo)
philips
0
450
KubeCon Keynote 2016- Distributed Systems Simplified on Kubernetes
philips
2
590
Other Decks in Programming
See All in Programming
壊れたパーサから始める関数型設計と構成的なパーサ #fp_matsuri
raiga0310
2
390
AIが無かった頃の素敵な出会いの話
codmoninc
1
210
変わらないものが、変わるものを決める — 意図駆動開発 × イベントソーシング × イミュータブル | What Doesn't Change Decides What Can — IDD × Event Sourcing × Immutability
tomohisa
0
210
5分で問診!Composer セキュリティ健康診断
codmoninc
0
580
Claude Opus 4.6以後の受託開発エンジニアの変化(Claude Code開発ノウハウ大公開スペシャルbyクラスメソッド)
iidatakuma
1
850
AI時代、エンジニアはどう育つのか -未経験エンジニアの成長を間近で見て考えたこと-
thasu0123
0
140
琵琶湖の水は止められてもNet--HTTPのリトライは止められない / You might be able to stop the water flow of Lake Biwa but you can't stop Net::HTTP retries
luccafort
PRO
0
430
自作OSでスライド発表する
uyuki234
1
3.9k
全PRの83%がAIレビューだけでマージできるようになった開発組織はその後どうなったか
athug
0
310
エンジニアにデザインハーネスを 〜デザインプロセスを規定するためのハーネス〜 / Design harness from an engineer's perspective
rkaga
2
1.7k
act1-costs.pdf
sumedhbala
0
250
AI駆動開発を妨げる技術的負債の解消アプローチ / ai-refactoring-approach
minodriven
17
9.3k
Featured
See All Featured
ラッコキーワード サービス紹介資料
rakko
1
4M
AI: The stuff that nobody shows you
jnunemaker
PRO
9
840
Sam Torres - BigQuery for SEOs
techseoconnect
PRO
0
440
How To Stay Up To Date on Web Technology
chriscoyier
790
250k
Winning Ecommerce Organic Search in an AI Era - #searchnstuff2025
aleyda
1
2.1k
The agentic SEO stack - context over prompts
schlessera
0
850
Taking LLMs out of the black box: A practical guide to human-in-the-loop distillation
inesmontani
PRO
3
2.3k
Understanding Cognitive Biases in Performance Measurement
bluesmoon
32
3k
Neural Spatial Audio Processing for Sound Field Analysis and Control
skoyamalab
0
380
Collaborative Software Design: How to facilitate domain modelling decisions
baasie
1
270
AI Search: Where Are We & What Can We Do About It?
aleyda
0
7.7k
Fashionably flexible responsive web design (full day workshop)
malarkey
408
67k
Transcript
Brandon Philips @BrandonPhilips |
[email protected]
etcd - mission-critical key-value store
Uncoordinated Upgrades
... ... ... ... ... ... Unavailable Uncoordinated Upgrades
Motivation CoreOS cluster reboot lock - Decrement a semaphore key
atomically - Reboot and wait... - After reboot increment the semaphore key
3 CoreOS updates coordination
CoreOS updates coordination 3
... CoreOS updates coordination 2
... ... ... CoreOS updates coordination 0
... ... ... CoreOS updates coordination 0
... ... CoreOS updates coordination 0
... ... CoreOS updates coordination 0
... ... CoreOS updates coordination 0
... ... ... CoreOS updates coordination 0
CoreOS updates coordination
Store Application Configuration config
config Start / Restart Start / Restart Store Application Configuration
config Update Store Application Configuration
config Unavailable Store Application Configuration
Requirements Strong Consistency - mutual exclusive at any time for
locking purpose Highly Available - resilient to single points of failure & network partitions Watchable - push configuration updates to application
Requirements CAP - We want CP - We want something
like Paxos
Common problem GFS Paxos Big Table Spanner CFS Chubby Google
- “All” infrastructure relies on Paxos
Common problem Amazon - Replicated log powers ec2 Microsoft -
Boxwood powers storage infrastructure Hadoop - ZooKeeper is the heart of the ecosystem
COMMON PROBLEM #GIFEE and Cloud Native Solution
10,000 Stars on Github 250 contributors Google, Red Hat, EMC,
Cisco, Huawei, Baidu, Alibaba...
THE HEART OF CLOUD NATIVE Kubernetes, Cloud Foundry Diego, Project
Calico, many others
ETCD KEY VALUE STORE Fully Replicated, Highly Available, Consistent
PUT(foo, bar), GET(foo), DELETE(foo) Watch(foo) CAS(foo, bar, bar1) Key-value Operations
DEMO play.etcd.io
Runtime Reconfiguration Point-in-time Backup Extensive Metrics etcd Operationality
ETCD v3 Successor of etcd v2
ETCD v3 Better Performance
ETCD v3 More Efficient APIs
Multi-Version Put(foo, bar) Put(foo, bar1) Put(foo, bar2) Get(foo) -> bar2
Multi-Version Put(foo, bar) Put(foo, bar1) Put(foo, bar2) Get(foo, 1) ->
bar
Tx.If( Compare(Value("foo"), ">", "bar"), Compare(Version("foo"), "=", 2), ... ).Then( Put("ok","true")...
).Else( Put("ok","false")... ).Commit() Mini-Transactions
l = CreateLease(15 * second) Put(foo, bar, l) l.KeepAlive() l.Revoke()
Leases
w = Watch(foo) for { r = w.Recv() print(r.Event) //
PUT print(r.KV) // foo,bar } Streaming Watch
Synchronization LoC
ETCD v2 machine coordination -> O(10k)
ETCD v3 app/container coordination -> O(1M)
Performance 1K keys
Performance Snapshot caused performance degradation etcd2 - 600K keys
Performance etcd2 - 600K keys Snapshot triggered elections
ZooKeeper Performance Non-blocking full snapshot Efficient memory management
Performance ZooKeeper default
Performance Snapshot triggered election ZooKeeper default
Performance Snapshot ZooKeeper default
Performance GC ZooKeeper snapshot disabled
Reliable Performance - Similar to ZooKeeper with snapshot disabled -
Incremental snapshot - No Garbage Collection Pauses - Off-heap storage
Performance etcd3 /ZooKeeper snapshot disabled
Performance etcd3 /ZooKeeper snapshot disabled
Memory 10GB 2.4GB 0.8GB 512MB data - 2M 256B keys
Reliability 99% at small scale is easy - Failure is
infrequent and human manageable 99% at large scale is not enough - Not manageable by humans 99.99% at large scale - Reliable systems at bottom layer
HOW DO WE ACHIEVE RELIABILITY WAL, Snapshots, Testing
Write Ahead Log Append only - Simple is good Rolling
CRC protected - Storage & OSes can be unreliable
Snapshots Torturing DBs for Fun and Profit (OSDI2014) - The
simpler database is safer - LMDB was the winner Boltdb an append only B+Tree - A simpler LMDB written in Go
Testing Clusters Failure Inject failures into running clusters White box
runtime checking - Hash state of the system - Progress of the system
Testing Cluster Health with Failures Issue lock operations across cluster
Ensure the correctness of client library
TESTING CLUSTER dash.etcd.io
etcd/raft Reliability Designed for testability and flexibility Used by large
scale db systems and others - Cockroachdb, TiKV, Dgraph
etcd vs others Do one thing
etcd vs others Only do the One Thing
etcd vs others Do it Really Well
etcd Reliability Do it Really Well
ETCD v3.0 BETA Efficient and Scalable
BETA AVAILABLE TODAY github.com/coreos/etcd
FUTURE WORK Proxy, Caching, Watch Coalescing, Secondary Index
GET INVOLVED github.com/coreos/etcd
Brandon Philips @BrandonPhilips |
[email protected]
etcd - mission-critical key-value store
Thank you!