HavocLabAll scenarios →
SEV-1Resilience

Reply All

Last month, shard 3 went dark for a while. Since a search isn't finished until all shards answer, thousands of users saw empty results. The on-call "fixed" it by configuring the gateway to retry the entire query thrice whenever a shard hiccuped. It worked. Everyone moved on. Tonight, it's failing again. Every failed query is re-sent to all three, twice more, including the two shards that already answered it. Nothing errors, but everything is late. The Job: Get latency back under the SLA before the gateway drowns the cluster. The queries it swallows still have to come back. Somebody in this diagram is asking too many people.

Free to play, no account needed. A deterministic engine grades your design against this scenario's SLA - error rate, latency, and cost.