Agent Reliability Engineering : Failure, Recovery, and the Discipline of Running Autonomous AI in Production
Overview
Your agent worked in the demo. It worked in the pilot. Then one morning you did not learn it had failed from a dashboard. You learned from a customer, an invoice that would not reconcile, a number wrong in a way no alert caught. You opened its logs and every signal was green. It had reported doing the work, and until that morning you had not known that "did the work" and "reported doing the work" were two different sentences.
Almost anyone can build an agent now. Almost nobody can run one. Roughly one in ten reaches production, and the ones that do fail in ways classic reliability tools were never built to catch, because an agent's behavior is sampled, not specified, and it can be up and wrong at the same time.
Agent Reliability Engineering names the discipline of running autonomous AI in production and gives it a body of practice, the way Site Reliability Engineering did for infrastructure. It stands on one idea with teeth: reliability, not capability, is the binding constraint on autonomy.
Inside: the five ways agent systems fail and the containment move for each; the compounding math that tells you how long an agent can run before it will fail; how to make an agent recover from any crash; SLOs, eval-gated deploys, incident response, and on-call for systems that do not stop when you sleep; a production-readiness gate; and a maturity model that sequences the work. Every claim carries a receipt. Every chapter ends with an artifact you can put to work the same week.
Volume 5 of The AI-Native Builder Canon.
This item is Non-Returnable
Customers Also Bought
Details
- ISBN-13: 9798186228363
- ISBN-10: 9798186228363
- Publisher: Independently Published
- Publish Date: July 2026
- Dimensions: 9 x 6 x 0.83 inches
- Shipping Weight: 1.19 pounds
- Page Count: 404
Related Categories
