A client's automation ran perfectly for six months, then their volume jumped tenfold overnight and it fell apart in three places at once. This is the post-mortem, unedited.
What broke
- Rate limits: parallel branches hammered a downstream API until it started returning 429s.
- Race conditions: two events for the same record ran at once and overwrote each other.
- No queue: with nothing to absorb the spike, failures cascaded instead of retrying.
The fix we should have shipped on day one: a queue in front of anything that talks to a rate-limited API, plus a lock per record. Both are cheap before you need them and painful after.
Nothing here was exotic. It was all the stuff you skip when volume is low and regret when it isn't.
Written by Kaan
Founder of AI Builders Stack. Builds automations for clients by day, writes up what broke by night.