← All notes

The cancellation arrived before the payment

4 min read

A workflow that reacts to events is really a workflow that reads a story: the order was placed, then paid, then shipped. Build it that way and it works for months, because most of the time the story does arrive in order.

Stripe's webhook documentation says plainly that it does not guarantee delivery of events in the order they were generated. Its own example is a subscription being created, which produces a small burst of events — the subscription created, an invoice created, an invoice paid, a charge created — with no promise about the sequence in which your endpoint sees them.

The same page notes two related facts: the same event can be delivered to you more than once, and occasionally two distinct event objects are generated for the same underlying thing.

None of this is a malfunction. It is the stated behaviour of a system that delivers over the internet and prefers delivering twice to delivering never.

What it looks like in a small business

The damage is rarely dramatic, which is why it survives so long.

The record created twice. The "customer created" event is still in flight when the "customer updated" event arrives. Your workflow looks for a customer, does not find one, and creates a second record — the everyday route to two copies of the same person.

The status that goes backwards. Shipped, then paid, then placed. If each step writes the current status, the last one to arrive wins, and the order in your system ends up as "placed" three days after it went out of the door.

The email sent about a thing that no longer applies. The cancellation was processed before the confirmation event arrived, so the customer gets a cheerful confirmation of an order they just cancelled. Nothing failed. The workflow acted on a fact that was true when it was generated and stale when it arrived — the mechanism behind a whole category of automations telling customers something untrue.

The double action. Two deliveries of one event, two runs, two invoices. This is the ordinary duplicate-run problem covered in what happens when an automation runs the same job twice, with the wrinkle that here the platform is behaving exactly as documented.

The design change that fixes most of it

Treat an event as a notification that something happened, not as a description of the current state.

In practice: when an event arrives, ask the source system what the situation is now, and act on the answer. The event tells you which order to look at. The order tells you what to do. Out-of-order delivery stops mattering, because you never trusted the sequence in the first place — you looked.

This is one line of extra work in most tools and it removes an entire class of problem. If your platform makes the lookup awkward, that is worth knowing before you build the workflow rather than after.

The rest of the guardrails

Deduplicate on the event identifier. Keep the identifiers you have already processed and ignore repeats. The vendor recommends exactly this, and it is the difference between "may be delivered more than once" being a footnote and being a second invoice.

Make repeated runs harmless. If running the same step twice would send two emails or take two payments, the step needs a guard of its own — check whether the thing has already been done, immediately before doing it.

Do not order by timestamp. Timestamps tell you when something happened at the source, not what has happened since. Current state comes from asking.

Give status changes a direction. If your workflow writes statuses, let it move an order forward but not backward. A shipped order should not be talked back into "placed" by a late arrival.

Delay the irreversible ones. Anything that spends money or speaks to a customer benefits from a short pause and a re-check. A few minutes between event and action absorbs almost all of the reordering you will ever see.

Why this bites small businesses specifically

Large teams hit this early, because they have volume: with thousands of events a day, the rare ordering quirk shows up in the first week and gets designed out.

A small business gets months of clean behaviour first. The workflow is built, trusted, and quietly relied on — and then one busy afternoon produces a burst of activity on the same order, the events race, and the failure looks like a one-off. It is filed as a glitch. It recurs at the next peak, which is inevitably the week you can least afford it.

The fix is not more monitoring. It is the assumption change: read the state, do not infer it.

If you want your existing workflows checked for the places where they infer state from the order of arrival, that is part of a process audit: $299, three business days.

Related: what happens when an automation runs the same job twice is the duplicate half of this problem, and worth reading first if your workflows move money.