Scenarios
Every level is a production incident: a story, a traffic storm, and an SLA to survive. Build the architecture on the canvas, and a deterministic physics engine grades the wreckage.
Spontaneous Combustion
SEV-3CachingYou just posted your app, and the internet took you at your word. Big mistake, or the best one, it hit the front page. A thousand people click. Then five thousand. Then 50,000 curious engineers arrive in the same instant, straight into a database you sized for "me and three friends." The waves keep coming after the peak, and when the dust settles you still have steady traffic. Congratulations. Now survive it.
DoNotTouch.java
SEV-2Async & BackpressureIt's Black Friday, and your inventory service is in the line of fire. Normally you'd just rewrite it, but this is the sacred legacy service the CEO authored himself. Refactoring it is seppuku for your career. It's synchronous, it's untouchable, and the write storm is already here.
Schrödinger's Avatar
SEV-2Consistency & ReplicationA user just updated their profile picture, but on refresh they still see their old face. Maybe it's your platform's way of saying they looked better before. Meanwhile Finance just saw the AWS bill for your "highly available" read replicas.
Friendly Fire
SEV-1Resilience"No one hires juniors anymore." Well, your company just did. Meet Tim. Tasked with fixing latency on the payment service, Tim made it "more robust" with a strict timeout and a 3x retry, so now every slightly-slow query times out and multiplies into three more. He's accidentally launched a DDoS against your own database, and you can't undo his timeout, only work around it.
Gaslighting as a Service
SEV-3Consistency & ReplicationYour CEO came back from a tech talk having "learned" that Postgres is legacy, so your user profiles now live in a multi-node cluster with read and write consistency both dialed to ONE for the benchmarks. Bob blocks his troll, refreshes, and the troll is still there. He refreshes again and they vanish. The data can't decide what's true.
Time To Leave (TTL)
SEV-2CachingTo speed up logins, the security team set the RBAC cache to an infinite TTL, and the Identity Database has never been happier. Then HR fired Ron over Zoom and IT revoked his credentials. The database knows he's gone. The cache does not. Ron is at the Starbucks across the street, still pulling company data at 2ms latency.
Ephemeral Commerce
SEV-2Load BalancingBecause databases are too slow, Ron (former employee) decided to keep every shopping cart in a single instance's memory. Traffic is spiking, that instance is about to fall over, and the autoscaler that saves you also cycles instances out from under live sessions.
First Come, First Severed
SEV-3Rate LimitingBob's product API serves four customers a second, and four was always enough. Then a competitor's scraper found the endpoint: 500 requests a second, none of them buying anything. Postgres serves everyone in the order they show up, so your paying customers are now stuck behind a bot photocopying the price list, and the database is buckling.
Cache Me If You Can
SEV-2CachingA storefront cached its 10 core product pages for 100 seconds each. Mathematically tidy, architecturally lethal. Pages that warm up together expire together, so every 100 seconds the cache blinks out at once and normal traffic turns into a synchronized mob charging the database.
us-east-1
SEV-1ResilienceOctober 2025: a DNS automation script in Virginia decides the DynamoDB endpoint no longer sparks joy and quietly deletes the record. The database is fine, nobody can find it. Every lookup hangs and every SDK retries like a pigeon pecking a dead feeder. Also, your worker fleet is enthusiastically DDoSing the void. AWS says the fix is "in progress."
Kamikaze Query
SEV-0ResilienceNovember 2025: a query change at Cloudflare doubles a global rule file, blowing past the memory limit of every edge-proxy parser at once. The result is a synchronized, worldwide reboot loop, and because the poisoned file regenerates every few minutes, the network keeps lurching back to life just long enough for you to close the incident channel before it explodes again.
Reply All
SEV-1ResilienceLast month, shard 3 went dark for a while. Since a search isn't finished until all shards answer, thousands of users saw empty results. The on-call "fixed" it by configuring the gateway to retry the entire query thrice whenever a shard hiccuped. It worked. Everyone moved on.
Multiverse of Madness
SEV-1ResilienceTim pruned a "spare" VLAN from a trunk port as he figured it was unused. Little did he know it was carrying the node heartbeats, and he accidentally split the link between the east and west zones clean in half. At T = 200ms the zones stop seeing each other, and both are completely fine, which is the problem. The standby has to decide whether the primary is dead or just delayed. Guess wrong and every order starts existing in two states at once. It gets worse, at T = 600ms the primary actually dies. At T = 900ms it comes back, considering itself the primary.
Déjà Vu
SEV-2Async & BackpressureFriday's deploy gave the billing worker a memory leak and is now stuck in a reboot loop. It eats RAM until it falls over, restarts and gets back to work. The queue in front of it guarantees delivery, so the crashes cost nothing anyone could see.
This is Fine 🔥
SEV-1Load BalancingFour database shards, users spread evenly by key. Tuesday, 3 AM: a brand-new account posts a meme, and by breakfast it is everywhere. Millions are loading one profile, and everyone who laughs smashes follow.
Not My Type
SEV-2Async & BackpressureThe CEO wanted AI in the company by next sprint. It landed on Tim's desk, and he delivered. His vibecoded ops agent turns user requests into JSON and drops them on the jobs queue. It worked in the demo.
ENOSPC
SEV-2Async & BackpressureA partner team just discovered their bulk export button and weighed it down with a coffee mug. 500 requests a second, 10 MB of uncompressed analytics payloads each.
Silent Night
SEV-1ResilienceEvery alert your paging pipeline fires gets routed, queued, delivered, logged. Nobody thinks about it, which is the highest compliment infrastructure can get. The problem is that when it dies, it doesn't page you. Instead it just goes quiet until the damage is done.
Musical Chairs
SEV-1ResilienceThe reliability team armed the chaos rig against your checkout pool. It kills one instance at a time, and holds each one down for the rest of the exercise. The drill is scheduled over the night's big sale. Most of the crowd is just browsing seat maps, a minority is mid-checkout with money.