← All notes

Your automation switched itself off

4 min read

A workflow that fails every time is not useful to anybody, so platforms eventually stop running it. Zapier states the threshold: "Your Zap will automatically turn off if 95% of its runs result in errors in the last 7 days."

As engineering, that is correct. Retrying a hopeless job forever wastes your allowance and the platform's capacity.

As operations, it produces a specific and expensive situation. Something has been broken for a week. Then the last remaining evidence that it exists — the error messages — stops, because the workflow is no longer running. The business gets quieter, and quiet reads as fine.

The week you did not notice

Work backwards from the rule and the timeline is uncomfortable.

For a workflow to hit 95% errors over seven days, it must have been failing nearly every run for most of a week. Whatever it was doing — creating invoices, adding customers, notifying the warehouse — has not happened since.

The most common cause is not a broken workflow at all. Connections expire: "You must reconnect your app connection if your password changed recently or if your organization enforced additional security controls." Somebody changed a password on Monday, and by the following Monday the workflow is off. Nothing was edited, nothing was deployed, and there is no obvious moment to point at.

Zapier does describe advance warning: "If your account is on a Team or Enterprise plan, Zapier will send you an email notification to the account owner and provide a grace period before turning off your Zap" — 72 hours on Enterprise, 24 on Team. The page says nothing either way about smaller plans, which is itself a good reason to find out what your own plan does rather than assuming a warning is coming.

And even on the plans that warn, the warning goes to the account owner, which in a small business is often whoever signed up years ago, possibly using an address nobody reads. That is the same weak point as not knowing whose accounts the automation runs on, showing up at the worst moment.

Why a disabled workflow is worse than a failing one

A failing workflow generates evidence. Every run leaves an error, and errors accumulate somewhere a person could look.

A disabled workflow generates nothing. It has no history, no errors, no runs. If your monitoring is built on noticing failures, it is now watching something that has stopped producing failures and will report that everything is well. This is the more advanced version of the problem in why automations fail silently: there, the failure does not speak up. Here, there is nothing left to speak.

There is also a restart trap. When the workflow is switched back on, it resumes from now. The week of orders that never reached the accounting system does not flow through on its own — they are simply missing, and finding out exactly what is missing means comparing two systems by hand for the period.

Making it noticeable

Watch for absence, not just errors. One check per important workflow: has it run in the last day (or hour) as expected? A count of zero is the alarm. This catches disabling, expiry, deletion, and someone switching it off by accident, all with one rule.

Watch the business outcome as well. Invoices created today, orders synced today, messages sent today. Those numbers are what you actually care about, and they fall to zero for every possible cause including the ones nobody predicted.

Put the platform's notifications somewhere two people see. Account owner address to a shared mailbox, not an individual. Then check that a real notification actually lands there, rather than assuming it will.

Treat repeated errors as an incident on day one. The 95% rule means a workflow that is mostly failing is on a countdown. Errors that "have been appearing for a few days" are not background noise, they are a deadline.

Write down the restart procedure while things are calm. Reconnect the account, test, switch on, then — the part everyone forgets — work out what fell through the gap and replay it. Knowing in advance which report tells you what is missing turns a bad day into an hour.

Know your plan's behaviour. Whether you get a warning, how long the grace period is, and where the message goes. Three facts, findable in ten minutes, and worth more than any amount of general vigilance.

The uncomfortable question underneath

If a workflow can be off for a week and the only symptom is that things feel a bit quiet, then nobody was watching the process — they were watching the tool.

That is worth sitting with, because it applies beyond this one rule. Every automated process needs an answer to "how would I know if this stopped?" that does not depend on the automation itself being alive to tell you. Usually the answer is a number somebody looks at daily, and usually it does not exist yet.

Working out which of your automations could stop without anyone noticing, and what the missing number is for each, is part of a process audit: $299, three business days.

Related: why automations fail silently, and how to make them speak up — start there if you have no way of telling whether your workflows ran today.