A team deploys a buggy build. 10,000 users' events get corrupted. With Kafka, they reset the offset and reprocess. With SQS, those messages are gone forever. Same problem. Completely different recovery story.
1/8
Your rate limiter probably isn't a rate limiter. Put a second server behind a load balancer and a 100-requests-per-minute limit quietly becomes 100 x N — each instance counts alone, no shared state, no errors, no alerts. Just 5x the traffic you meant to block.
1/8
A misconfigured route table is usually why "the service can't reach the internet" eats 40 minutes of debugging — and the actual fix takes 30 seconds.
1/8
You added an index. Your query is still slow. This isn't a mystery — there are exactly 7 reasons indexes get silently ignored in production. And most of them are your fault.
1/8
You add Redis. Response times drop from 200ms to 5ms. You celebrate. Three weeks later, a customer sees a price that was updated an hour ago. You didn't have a caching problem before. Now you have two.