Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
Incident Response / infra study 3
Search
tjun
June 16, 2020
Technology
3.8k
3
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
Incident Response / infra study 3
Infra Study Meetup #3の発表資料です。
https://forkwell.connpass.com/event/176885/
tjun
June 16, 2020
More Decks by tjun
See All by tjun
SREとしてスタッフエンジニアを目指す / SRE Kaigi 2025
tjun
17
19k
CloudNative環境におけるトラブルシューティングガイド / CloudNative Days Tokyo 2023
tjun
6
2.7k
2023-12-07 SRE Talk クラウドと長く付き合う
tjun
0
270
インシデント対応を改善しよう/2024 TechFeed Experts Night 17
tjun
1
580
メルペイにおけるマイクロサービス運用の苦労と改善 / CloudNative Days Tokyo2020
tjun
16
4.7k
絶え間なく変化するメルカリ・メルペイにおけるSREの組織と成長 / SRE Next 2020
tjun
6
21k
メルペイのマイクロサービスとCloud Native / CloudNative Days Kansai2019
tjun
22
23k
メルペイを支えるGKEとCloud Spanner / 2019 Google Cloud Architect Night 1
tjun
1
2.6k
メルペイのマイクロサービスの構築と運用 / CloudNative Days Tokyo2019
tjun
26
15k
Other Decks in Technology
See All in Technology
生成AIを使って「人が」考える技術 ― AI時代の人機共想と実践ノウハウ|UNITT AC2026
ishiirikie
0
550
事業活動を AI Ready にする攻めと守りのデータエンジニアリング / data-engineering-for-ai-ready-business
pei0804
1
210
Incremental HTTP
kazuho
5
2k
Kernel testing frameworks
ennael
PRO
0
110
AI時代のAPI開発を加速する品質ガードレール / API Quality Guardrails in the AI Era
yokawasa
0
140
あなたの知らないAmazon VPC Route Server/Amazon VPC Route Server you don't know about
masakiokuda
2
220
オブザーバビリティを高める AI エージェント体験を考える / Designing AI Agent Experiences That Enhance Observability
aoto
PRO
2
310
私の推しは「聞いてから進む」AIです -AI-DLCに一人でアプリを作らせた話
yama3133
1
170
More Freedom on the Same Shared GPU Cluster: A Small Team’s Experience with vCluster
nttcom
0
130
全人類(ほぼ)AWS Organizations の上でAWSを利用している、その世界を知る話
htan
0
110
AgentCoreで実践するハーネスエンジニアリング
yakumo
1
280
ミイダス株式会社 テックチームのご紹介 / MIIDAS Tech Team
miidas
0
150
Featured
See All Featured
The Director’s Chair: Orchestrating AI for Truly Effective Learning
tmiket
1
310
From π to Pie charts
rasagy
1
390
"I'm Feeling Lucky" - Building Great Search Experiences for Today's Users (#IAC19)
danielanewman
230
23k
Chasing Engaging Ingredients in Design
codingconduct
0
340
Navigating the moral maze — ethical principles for Al-driven product design
skipperchong
2
600
Breaking role norms: Why Content Design is so much more than writing copy - Taylor Woolridge
uxyall
1
440
Groundhog Day: Seeking Process in Gaming for Health
codingconduct
0
390
30 Presentation Tips
portentint
PRO
1
420
Building a Modern Day E-commerce SEO Strategy
aleyda
45
9.2k
Reality Check: Gamification 10 Years Later
codingconduct
0
2.3k
Jamie Indigo - Trashchat’s Guide to Black Boxes: Technical SEO Tactics for LLMs
techseoconnect
PRO
0
690
Efficient Content Optimization with Google Search Console & Apps Script
katarinadahlin
PRO
1
910
Transcript
Incident Response Infra Study Meetup #3 LT Merpay SRE @tjun
Junichiro Takagi https://speakerdeck.com/tjun/infra-study-3
「インシデント対応やってますか?」
今日のテーマ Incident Response • できればやりたくない • でもSREをやるなら避けられない • どうすれば、より健全なIncident Responseができるか
今日の話は https://response.pagerduty.com/ の超ざっくりしたまとめ なので、詳しくは読んでほしい
はじめに Incident とは 予期せず提供しているサービスが利用できない状態になったり、 期待している機能が提供できない状態
はじめに Incident とは 予期せず提供しているサービスが利用できない状態になったり、 期待している機能が提供できない状態 Incident Response とは Incidentを解決・管理するための組織的なしくみ。 問題を解決するだけでなく、被害を減らしたり解決までの時間やコストを減らす
取り組みも含まれる。 エンジニアだけじゃなく、Customer Support、PM、PRなども関わる。
Incident 前に やること • 心構え: Incidentは必ず起きる…! • Incident, Severity を定義する
• Trigger を用意する • 役割を決める(Incident Commander等) • コミュニケーションの仕組みを 用意する
Incident 中に やること • 心構え: 慌てない • 必要なメンバーを招集する • 役割ごとに必要な対応を行う
◦ Incident Commander 関係者に連絡しSlackで指示を出す ◦ エンジニア 問題を調査し解決方法を提案・実行する
Incident 後に やること • 心構え: Blameless ( 人を責めない ) •
Post-mortem(振り返り) を行う ◦ What Happened? ◦ Impact ◦ Resolution ◦ Timeline ◦ うまくできたこと、だめだったこと ◦ Action Items
Incident Response をはじめよう 1. インシデントを定義する 2. コミュニケーションの仕組みを作る ◦ アラート設定、Slackで集まるChannel、などを用意 3.
インシデント対応の役割を決める ◦ Incident Commanderを決める 4. Post-mortemのテンプレを作る ◦ https://landing.google.com/sre/sre-book/chapters/postmortem/ などが参考になる 5. 練習する 6. 実際のインシデントで実行する
まとめ • Incident Response はSREだけのものではない、組織的な 仕組みづくりが必要。できるところから始めよう • 適切な準備をして、健全な運用を作りましょう