Almost every guide to A/B testing was written for companies with a hundred thousand visitors a month. If your site gets four hundred, that advice will not just waste your time, it will hand you confident-sounding conclusions that are pure noise. AI changes part of this picture. It does not change the part that matters most, and being clear about which is which will save you a lot of wasted effort.
You show version A to half your visitors and version B to the other half, then count which produced more of the thing you care about: phone calls, form submissions, bookings. The statistics behind it exist to answer one question. Is the difference you are looking at real, or is it the same wobble you would get flipping a coin two hundred times?
Answering that requires volume. The smaller the improvement you are hunting, the more visitors you need before it becomes visible. A change that lifts enquiries by half will show up quickly. A change that lifts them by a few percent may never be provable on a local business website, no matter how patient you are. That is not a failure of the tool. It is arithmetic.
Historically the bottleneck was producing alternatives. Someone had to sit down and write eight different headlines for the emergency plumbing page. A language model will give you fifteen in under a minute, and perhaps four will be worth putting in front of real people. Your job shifts from writing to judging, which is faster and, frankly, more suited to someone who knows the trade.
A classic A/B test holds the split at 50/50 until it ends. A multi-armed bandit gradually sends more traffic to whichever version is winning while the test is still running. For a low-traffic site this is usually the better shape: you lose less revenue to the losing version, and you do not have to guess a stopping point in advance. The trade-off is that you learn less cleanly about why one version won.
The most common testing mistake is stopping the moment a result looks good. Modern tools will tell you when a lead is within the range of random chance, and some will refuse to declare a winner at all. Take that refusal seriously rather than overriding it.
This is the unglamorous truth. Testing is for refining something that already works. If the fundamentals are broken, you are optimising the paint on a car with no wheels. In our audit of 622 local business websites, 51.0% had no click-to-call link at all, and 30.4% were missing a meta description. Adding a tappable phone number to a site where half your visitors are on a phone will do more for you than any amount of headline testing.
The honest sequence is: fix what is plainly missing, then measure a baseline for a couple of months, then test. If you are unsure which stage you are at, a straightforward website audit will tell you faster than a testing tool will.
Test one thing at a time. Multivariate testing, where several elements change at once, needs far more traffic than a local business will ever have.
Pick a tool that integrates with whatever built your site rather than one that requires a developer. Define the conversion before you launch, in writing, so you cannot move the goalposts later. Run for a minimum of two full weeks so you capture both weekday and weekend behaviour, and never end a test early because Tuesday looked promising. Keep a simple log of every test you have run, including the failures. After a year that log is more valuable than any individual result, because it tells you what your particular customers respond to.
If a test has run for six weeks with no clear winner, the two versions are probably equivalent and your attention belongs elsewhere. If your traffic is seasonal and you are mid-season, pause rather than pollute the data. And if testing has become a way of avoiding a harder decision, such as whether your pricing page is honest about cost, stop testing and go make that decision.
There is no universal threshold, but under a few hundred conversions per month you will struggle to prove anything but very large differences. If you get five enquiries a week, spend your effort on fixing obvious gaps and improving your offer instead. Testing becomes genuinely useful once you have enough volume that a two-week test produces a meaningful count on both versions.
Partly. Automated tools can generate variations, allocate traffic and call results without you watching. What they cannot do is decide what is worth testing, because that depends on your margins, your capacity and what you actually want more of. Automation handles the mechanics well; the strategy still needs a human who understands the business.
Not by default. Generated copy tends toward a generic, slightly overcooked register. Give the model examples of how you already speak to customers, then edit what comes back. Treat the output as a first draft from a keen but new employee: useful raw material, never publishable untouched.
Stopping early. A version pulls ahead in the first few days, it looks decisive, and the test gets called. Early leads reverse constantly. Decide your end date and your minimum sample before launching, write them down, and hold to them even when the interim numbers are tempting.
Want this handled for you?
Get a free audit of your website, Google reviews, and local SEO — we’ll show you exactly where you’re losing customers. Delivered in 24 hours, no sales call.
Get my free audit → or book a 15-min callWe help local businesses in Stamford, Greenwich, Norwalk, and Fairfield County implement AI marketing that generates real results.
Get Your Free AI Marketing Audit →