Scenarios
Every level is a production incident: a story, a traffic storm, and an SLA to survive. Build the architecture on the canvas, and a deterministic physics engine grades the wreckage.
Spontaneous Combustion
SEV-3CachingYou just posted your app, and the internet took you at your word. Big mistake, or the best one, it hit the front page. A thousand people click. Then five thousand. Then 50,000 curious engineers arrive in the same instant, straight into a database you sized for "me and three friends." The waves keep coming after the peak, and when the dust settles you still have steady traffic. Congratulations. Now survive it.
DoNotTouch.java
SEV-2Async & BackpressureIt's Black Friday, and your inventory service is in the line of fire. Normally you'd just rewrite it, but this is the sacred legacy service the CEO authored himself. Refactoring it is seppuku for your career. It's synchronous, it's untouchable, and the write storm is already here.
Schrödinger's Avatar
SEV-2Consistency & ReplicationA user just updated their profile picture, but on refresh they still see their old face. Maybe it's your platform's way of saying they looked better before. Meanwhile Finance just saw the AWS bill for your "highly available" read replicas.
Friendly Fire
SEV-1Resilience"No one hires juniors anymore." Well, your company just did. Meet Tim. Tasked with fixing latency on the payment service, Tim made it "more robust" with a strict timeout and a 3x retry, so now every slightly-slow query times out and multiplies into three more. He's accidentally launched a DDoS against your own database, and you can't undo his timeout, only work around it.
Gaslighting as a Service
SEV-3Consistency & ReplicationYour CEO came back from a tech talk having "learned" that Postgres is legacy, so your user profiles now live in a multi-node cluster with read and write consistency both dialed to ONE for the benchmarks. Bob blocks his troll, refreshes, and the troll is still there. He refreshes again and they vanish. The data can't decide what's true.
Time To Leave (TTL)
SEV-2CachingTo speed up logins, the security team set the RBAC cache to an infinite TTL, and the Identity Database has never been happier. Then HR fired Ron over Zoom and IT revoked his credentials. The database knows he's gone. The cache does not. Ron is at the Starbucks across the street, still pulling company data at 2ms latency.
Ephemeral Commerce
SEV-2Load BalancingBecause databases are too slow, Ron (former employee) decided to keep every shopping cart in a single instance's memory. Traffic is spiking, that instance is about to fall over, and the autoscaler that saves you also cycles instances out from under live sessions.
First Come, First Severed
SEV-3Rate LimitingBob's product API serves four customers a second, and four was always enough. Then a competitor's scraper found the endpoint: 500 requests a second, none of them buying anything. Postgres serves everyone in the order they show up, so your paying customers are now stuck behind a bot photocopying the price list, and the database is buckling.
Cache Me If You Can
SEV-2CachingA storefront cached its 10 core product pages for 100 seconds each. Mathematically tidy, architecturally lethal. Pages that warm up together expire together, so every 100 seconds the cache blinks out at once and normal traffic turns into a synchronized mob charging the database.
us-east-1
SEV-1ResilienceOctober 2025: a DNS automation script in Virginia decides the DynamoDB endpoint no longer sparks joy and quietly deletes the record. The database is fine, nobody can find it. Every lookup hangs and every SDK retries like a pigeon pecking a dead feeder. Also, your worker fleet is enthusiastically DDoSing the void. AWS says the fix is "in progress."
Kamikaze Query
SEV-0ResilienceNovember 2025: a query change at Cloudflare doubles a global rule file, blowing past the memory limit of every edge-proxy parser at once. The result is a synchronized, worldwide reboot loop, and because the poisoned file regenerates every few minutes, the network keeps lurching back to life just long enough for you to close the incident channel before it explodes again.
Reply All
SEV-1ResilienceLast month, shard 3 went dark for a while. Since a search isn't finished until all shards answer, thousands of users saw empty results. The on-call "fixed" it by configuring the gateway to retry the entire query thrice whenever a shard hiccuped. It worked. Everyone moved on.
Multiverse of Madness
SEV-1ResilienceTim pruned a "spare" VLAN from a trunk port as he figured it was unused. Little did he know it was carrying the node heartbeats, and he accidentally split the link between the east and west zones clean in half. At T = 200ms the zones stop seeing each other, and both are completely fine, which is the problem. The standby has to decide whether the primary is dead or just delayed. Guess wrong and every order starts existing in two states at once. It gets worse, at T = 600ms the primary actually dies. At T = 900ms it comes back, considering itself the primary.