menu
{ "item_title" : "On-Call In Action", "item_author" : [" Quan Huynh "], "item_description" : "In today's always-on world, downtime is not an option. Your users expect seamless service, 24/7. Your business depends on it. But how do you guarantee that reliability when complex systems inevitably encounter turbulence? The answer lies in a world-class on-call capability.On-Call In Action is your practical playbook for building just that. This isn't just another theoretical tome; it's a hands-on guide to navigating the high-stakes reality of modern on-call. We'll equip you with the SRE principles, incident management lifecycles, and effective alerting strategies (leveraging the Versus Incident project as our real-world example) that form the backbone of resilient operations.This book, On-Call In Action, is your friendly guide to making on-call work better. We'll show you: Why being on-call is so important.What to do when a problem (we call it an incident) happens.How to set up good alerts so you only get called for big problems. We'll even show you how with a free tool called Versus Incident.How to check if your services are running well (using simple goals).How to learn from mistakes without blaming anyone, so things get better.How to make good on-call schedules so people don't get too tired.How to create a supportive team for on-call work.Stop just reacting to problems and start engineering reliability. Whether you're a tech person who is on-call, a manager, or just curious, this book will give you clear advice and real examples. We want to help you build an on-call system that keeps your services running and your team feeling good. This book contains 11 chapters: Chapter 1 Foundations: Why On-Call Matters & SRE PrinciplesChapter 2 Anatomy of an Incident: The Management LifecycleChapter 3 Effective Alerting: Strategy and Routing Use Versus IncidentChapter 4: Integrating Monitoring Sources and Escalation Policies: A Case StudyChapter 5: Measuring Reliability: SLIs, SLOs, and Error BudgetsChapter 6: Putting It All Together: Practical Examples of Unified Alerting & TemplatingChapter 7: Learning from Failure: Blameless PostmortemsChapter 8: Sustainable On-Call: Scheduling and Managing BurnoutChapter 9: Effective IncidentChapter 10: The On-Call Ecosystem: Tooling and Future TrendsChapter 11: On-Call in Action: Digital Customer Onboarding in Banking", "item_img_path" : "https://covers3.booksamillion.com/covers/bam/9/79/828/355/9798283556314_b.jpg", "price_data" : { "retail_price" : "19.99", "online_price" : "19.99", "our_price" : "19.99", "club_price" : "19.99", "savings_pct" : "0", "savings_amt" : "0.00", "club_savings_pct" : "0", "club_savings_amt" : "0.00", "discount_pct" : "10", "store_price" : "" } }
On-Call In Action|Quan Huynh

On-Call In Action : Site Reliability Engineering Best Practices for Building Resilient Systems

local_shippingShip to Me
In Stock.
FREE Shipping for Club Members help

Overview

In today's "always-on" world, downtime is not an option. Your users expect seamless service, 24/7. Your business depends on it. But how do you guarantee that reliability when complex systems inevitably encounter turbulence? The answer lies in a world-class on-call capability.

"On-Call In Action" is your practical playbook for building just that. This isn't just another theoretical tome; it's a hands-on guide to navigating the high-stakes reality of modern on-call. We'll equip you with the SRE principles, incident management lifecycles, and effective alerting strategies (leveraging the Versus Incident project as our real-world example) that form the backbone of resilient operations.

This book, "On-Call In Action," is your friendly guide to making on-call work better. We'll show you:

  • Why being on-call is so important.
  • What to do when a problem (we call it an "incident") happens.
  • How to set up good alerts so you only get called for big problems. We'll even show you how with a free tool called "Versus Incident."
  • How to check if your services are running well (using simple goals).
  • How to learn from mistakes without blaming anyone, so things get better.
  • How to make good on-call schedules so people don't get too tired.
  • How to create a supportive team for on-call work.
Stop just reacting to problems and start engineering reliability. Whether you're a tech person who is on-call, a manager, or just curious, this book will give you clear advice and real examples. We want to help you build an on-call system that keeps your services running and your team feeling good.

This book contains 11 chapters:

  • Chapter 1 Foundations: Why On-Call Matters & SRE Principles
  • Chapter 2 Anatomy of an Incident: The Management Lifecycle
  • Chapter 3 Effective Alerting: Strategy and Routing Use Versus Incident
  • Chapter 4: Integrating Monitoring Sources and Escalation Policies: A Case Study
  • Chapter 5: Measuring Reliability: SLIs, SLOs, and Error Budgets
  • Chapter 6: Putting It All Together: Practical Examples of Unified Alerting & Templating
  • Chapter 7: Learning from Failure: Blameless Postmortems
  • Chapter 8: Sustainable On-Call: Scheduling and Managing Burnout
  • Chapter 9: Effective Incident
  • Chapter 10: The On-Call Ecosystem: Tooling and Future Trends
  • Chapter 11: On-Call in Action: Digital Customer Onboarding in Banking

This item is Non-Returnable

Details

  • ISBN-13: 9798283556314
  • ISBN-10: 9798283556314
  • Publisher: Independently Published
  • Publish Date: May 2025
  • Dimensions: 11 x 8.5 x 0.39 inches
  • Shipping Weight: 0.96 pounds
  • Page Count: 182

Related Categories

You May Also Like...

    1

BAM Customer Reviews