Zapier has been the plumbing between small business software for years. Form submitted, add a row to a spreadsheet, send a notification. Rigid, reliable, and limited to things you could describe as a rule.
The AI steps change what fits in that sentence. You can now put a step in the middle of a workflow that reads something and forms a judgement about it. That is a meaningful expansion, and it is also where automations start going wrong in new ways.
A rule is deterministic. If the form field says "emergency", route it to the on-call number. It behaves identically every time, which means you can test it once and trust it.
A judgement is not. Ask a model to read an incoming enquiry and decide whether it is urgent, and it will be right most of the time and occasionally wrong in ways you did not anticipate. It might classify a politely-worded burst pipe as routine.
Both belong in your automations, but they need different treatment. Rules can run unattended. Judgements need either a low cost of being wrong, or a person somewhere in the path. Confusing the two is the source of most automation failures we are asked to fix.
Start where a wrong answer is cheap and the manual version is tedious.
Notice what is absent: nothing here sends a message to a customer without a human seeing it. That restraint is deliberate, and it should last until you have watched a workflow behave correctly for a couple of months.
The mechanics are simpler than people expect. You add a step, choose an AI action, and write instructions describing what should come out. The data from earlier steps is available to reference.
Three things separate a workflow that holds up from one that breaks quietly:
Do not ask for a summary. Ask for a specific structure: a category from a fixed list, a one-line summary, and an urgency rating of high, medium, or low. Free-form output breaks whatever step comes next.
Include an explicit "unclear" option and route those to a person. Without one, the model will pick something rather than admit uncertainty, and you will never see the cases it found difficult.
Everyone tests with a clean example. Test instead with the enquiry that was three words long, the one written in a hurry with no punctuation, the one that mentioned two different problems. Those arrive weekly and they are where the workflow fails.
Automations fail silently. A rule that breaks throws an error you notice; a judgement that degrades just makes worse decisions quietly. Build in a way to see what it is doing. Logging every AI decision to a spreadsheet costs one extra step and means you can review a week's output in five minutes.
Watch the cost as well. AI steps consume more of your task allowance than simple ones, and a workflow triggered by every incoming email will burn through a plan faster than expected. Filter before the AI step so the model only sees what needs judging.
Set a review date. Six weeks after building an automation, read fifty of its decisions and check how many you would have made differently. If it is more than a handful, tighten the instructions or move the task back to a person.
If a task happens twice a month, automating it will cost more time than it saves, including the maintenance. If the task requires knowledge that lives only in your head, no instruction set will capture it. If the process is not yet stable, automating it locks in the current version, which is a problem if the current version is wrong.
And if a mistake reaches a customer directly, keep a person in the loop. The efficiency argument for removing that person is always smaller than it looks, and the cost of the failure is always larger. A workflow that saves you fifteen minutes a day is a good outcome. A workflow that sends an inappropriate message to a client at 2am is not offset by any amount of saved time.
No. The AI steps are configured by writing plain instructions describing what you want returned. The skill required is not programming but precision: describing the desired output specifically enough that later steps can rely on it. Vague instructions produce inconsistent output, which is the usual reason a first attempt disappoints.
Something internal, frequent, and low-stakes. Routing incoming enquiries, extracting data from emailed invoices, or summarising the week's activity are good starting points. Avoid anything that sends messages to customers until you have watched a simpler workflow behave correctly for several weeks.
Reliable enough for classification and extraction, not reliable enough to be unsupervised on anything consequential. They will be correct most of the time and occasionally wrong without warning. Design for that: give the model an explicit uncertain option, route those cases to a person, and log decisions so you can audit them.
Zapier connects to more services than the alternatives, which usually settles it for a small business using ordinary software. Competing platforms offer more control over complex logic and can be cheaper at high volume. If your tools are all mainstream and your workflows are straightforward, breadth of integration matters more.
Want this handled for you?
Get a free audit of your website, Google reviews, and local SEO — we’ll show you exactly where you’re losing customers. Delivered in 24 hours, no sales call.
Get my free audit → or book a 15-min callWe help local businesses in Stamford, Greenwich, Norwalk, and Fairfield County implement AI marketing that generates real results.
Get Your Free AI Marketing Audit →