Book Summary
Book 2 of AWS Industrial Cloud. The Amazon listing names infrastructure-as-code state, zero-downtime delivery, EKS with Karpenter, serverless resilience, SLO burn-rate alerting, chaos drills, and production post-mortems. Kindle and Kindle Unlimited. ISBN-13 978-9360121020.

The book

AWS Industrial Cloud: Volume 2 is Book 2 of 3 in AWS Industrial Cloud by Vatsal Shah. The Kindle edition was published on 11 September 2026. Language: English. ISBN-13 978-9360121020. ASIN B0HJGQQLNY.

The full listing title is: AWS Industrial Cloud — Volume 2: DevOps, Production Incident Engineering, EKS Fleets, Chaos Engineering & SRE Systems.

The Amazon description says architectures that look fine on slides break in production, and that this volume is the operating manual for that break.

Open the Kindle edition on Amazon.

What is inside

The description on the Amazon page names these blocks:

  1. Infrastructure as code and state — S3 and DynamoDB remote state locking, cross-account OIDC for GitHub Actions, drift detection, and GitOps reconciliation.
  2. Zero-downtime delivery — blue/green on ECS with CodeDeploy, canary and linear shifts, and multi-region rollback alarms.
  3. EKS fleets and Karpenter — node provisioning, VPC CNI address space, and Cilium observability.
  4. Serverless resilience — Lambda SnapStart, SQS visibility timeouts, dead-letter redrive, and Step Functions Distributed Map.
  5. Observability and error budgets — ADOT, X-Ray trace context, and multi-window burn-rate alerts.
  6. Chaos drills — Fault Injection Simulator, multi-AZ partition tests, and abort triggers.
  7. Production post-mortems — the listing says ten failure write-ups, including DNS stampedes, state-file corruption, and split-brain writes.
  8. DOP-C02 drills — the listing says ten scenario questions with rationales.

Who it helps

Beginner. Start with the pipeline chapter. Learn what remote state is, and what a rollback alarm is for, before you own an on-call week.

Builder. Use Karpenter, SQS, and Step Functions when the system has to provision, retry, and stop.

SRE. Use the burn-rate and chaos chapters when the question is how fast the error budget is going, and how you abort a drill.

CxO and business owner. Use the post-mortem set. The listing frames them as production failures with a timeline and a remediation, which is the language a review meeting already uses. This page does not restate a dollar figure from the listing as a measured result.

How to use it

  1. Lock the state. Say where Terraform state lives and who can write it.
  2. Name the release path. Blue/green or canary, and the alarm that rolls it back.
  3. Set one burn-rate alert. One service, one window, one owner.
  4. Read one post-mortem from the ten the listing includes, and write the matching abort rule.

What you leave with

  • A state-lock rule.
  • A release path with a rollback alarm.
  • One burn-rate alert with an owner.
  • One failure pattern, taken from the listing's post-mortems, that you will not rediscover on a Friday.

The series on this site

Questions about using this in a team workflow: contact me.

Vatsal Shah

Vatsal Shah

AI Leader · Solution Architect · TPM

I design autonomous AI systems, enterprise architectures, and publish deep technical content for global organisations.