Dnevni recovery posao kao zaštitna mreža automatizacije
Kompletan tekst beleške je trenutno na engleskom. Srpski urednički prevod je deo sledeće content faze.
Event-driven workflows are appealing because they're fast: something happens, and the next step fires immediately instead of waiting for the next scheduled poll. The catch is that webhooks, API calls, and third-party integrations occasionally fail silently. A missed event with no backup means a stalled record that nobody notices until someone asks why nothing happened.
The instinct is often to make the live path more defensive: add retries, add fallback triggers, add more error handling directly in the event handler. In my experience that usually makes the live path harder to reason about without actually closing the gap, because the failure modes you're protecting against (a webhook that never arrives at all) can't be caught by retry logic inside a handler that never ran.
A cleaner fix is a separate, deliberately small recovery job that runs on a fixed schedule, once a day is often enough, and does one narrow thing: look for records that should have moved forward but didn't, and nudge a small, bounded number of them.
The word 'bounded' matters. A recovery job that tries to fix everything it finds can turn one missed event into a burst of duplicate work, especially if the live path recovers on its own moments later. Capping the recovery job to a small number of records per run, and having it check whether something is already in progress before touching anything, keeps it firmly in a supporting role instead of a competing one.
This also makes the recovery job easy to reason about in isolation. It doesn't need to know why the live path might have failed. It only needs to know: what does 'stuck' look like for this record type, and what is the single next step to unstick it. That's a much smaller problem than trying to build a fully fault-tolerant event pipeline.
The other benefit is visibility. A scheduled job that runs once a day and logs what it recovered gives you a passive health signal for the live path. If the recovery job is quietly picking up two or three records every night, that's useful information about how often the live path is actually failing, even if nothing was ever technically broken enough to page anyone.
None of this replaces fixing the root cause of missed events where you can. But for the cases you can't fully eliminate, a small, bounded, once-a-day recovery job is often a better trade than trying to make the fast path perfect.
