How to run an A/B test when you do not have the traffic
Most business sites cannot detect the effects they are testing for. How to know when a test is pointless, and what to do instead.

Part of Measuring what a website does
A marketing lead at a B2B company with about nine hundred monthly visitors to its pricing page wants to test a shorter headline against the current one. The tool they use makes this one click: create variant, split traffic, wait for a winner. Three weeks later the dashboard shows variant B ahead by four points, not yet "significant" by the tool's own badge, and the team has to decide whether to ship it, extend the test, or start over. Nobody in that meeting asks the question that would have made the meeting unnecessary: given nine hundred visitors a month and the size of headline effects that actually occur, was this test ever able to produce an answer.
It almost never was, and the tool did not say so, because saying so is not in the tool's interest. Every A/B testing product is sold on the promise of certainty, and certainty scales with traffic in a way the sales page does not mention. This is not a case against testing. It is a case against testing before checking whether the traffic can support one, which takes five minutes and would have saved this team three weeks.
Run the calculation before the test, not after
The two-proportion sample size formula is not exotic — it is the same one behind the table in our piece on measuring what a website does, and it is worth walking through with a second, more typical set of inputs, because a 2% baseline is not everyone's situation.
Take a B2B contact page converting at 4%, a plausible rate for a page whose traffic already has some intent behind it. The team wants to detect a 25% relative lift — 4% to 5%, a change most people would call meaningful rather than marginal. At the conventional thresholds (95% confidence, 80% power), that requires roughly 6,700 visitors per variant, about 13,500 in total. A more modest and arguably more realistic target, a 10% relative lift from 4% to 4.4%, needs close to 39,000 per variant — nearly 80,000 visitors in total, for a change nobody would call dramatic if they saw it happen.
| Change you want to detect | Visitors needed per variant | Both variants |
|---|---|---|
| 4.0% → 5.0% (+25% relative) | ~6,700 | ~13,500 |
| 4.0% → 4.4% (+10% relative) | ~39,000 | ~78,000 |
Against 800 visitors a month to the page, the first test takes roughly seventeen months. The second does not finish inside a normal person's tenure at the company. This is not a pessimistic scenario — 800 monthly visitors to a contact or pricing page is a comfortable number for a mid-sized B2B site, well above what most companies in that bracket actually see. The arithmetic does not care how good the headline is. It only cares how many people saw it, and halving the effect size you are trying to detect roughly quadruples the sample required, which is the detail that makes "just run it longer" such a bad answer — longer does not scale the way people assume it does.
Testing a headline tweak at this volume is a category error, not bad luck
The instinct, once the numbers above land, is to blame execution — wrong tool, wrong significance threshold, wrong duration. It is worth being precise about why that is not the problem. A test compares two things that are already close to each other and tries to detect a small difference against natural variation. The smaller the difference you are testing for, the more data it takes to see it over the noise, which is exactly the relationship the table shows.
A headline rewrite, a button color, a slightly shorter form — these produce small effects by design, because the page around them did not change. At low traffic, testing them is not a weaker version of testing a redesign; it is attempting to measure something the available data physically cannot resolve, the same way a kitchen scale cannot report the weight of a letter to the milligram. The fix is not more patience. It is choosing a different kind of change or a different kind of evidence, which is the rest of this piece.
The one case where small-traffic testing works is when the effect itself is large, not subtle. A broken form that starts working, a phone number that becomes clickable, a nine-second load that becomes two — these can double or triple an outcome, and a doubling is visible in a fraction of the sample a 20% lift would need. If you are not confident the change you are testing could plausibly double the outcome, the honest question is not whether to test longer, it is whether to test at all.
Peeking at the dashboard is not free, even when the effect is real
Almost nobody runs a test by calculating the sample size, waiting for it, and looking once. Almost everybody checks the dashboard every day or two and stops the first time it turns green. This feels like discipline — watching closely, catching the result as soon as it arrives — and it is the single most common way an A/B test lies to the person running it.
The 5% false-positive rate printed on the significance badge is calculated for one look at the data, at a sample size decided in advance. Every additional look is another chance for the gap between two identical variants to cross the significance line purely by chance, the same way flipping a coin ten more times gives ten more chances to hit an unlikely streak. Checking daily and stopping on the first "significant" result does not calibrate to 5% — it degrades, and it degrades badly, because the procedure being run is no longer the one the p-value was computed for. A nominal 5% risk under continuous peeking with an early stop can end up several times higher than the number on the screen, which means a meaningful share of "winning" variants declared this way are not actually better than what they replaced — they are noise that got lucky before the sample size was reached.
The fix costs nothing and is entirely a matter of discipline: decide the sample size before the test starts, using the calculation above, and do not look at the result — not the direction, not the magnitude — until that number is reached. If the organisation cannot commit to not looking, it should not run the test; it should pick one of the alternatives below instead, honestly, rather than run a test it will not let finish on its own terms.
Before-and-after can be defensible, but only when you can name what else changed
Given all of the above, the temptation is to skip the split test and just compare the month before a change to the month after. This is not automatically wrong, but it borrows a different kind of risk in exchange for the sample-size problem: instead of noise inside the test, you get confounding outside it.
It is defensible when the change is large enough to plausibly explain the shift by itself, the before and after windows are close enough in time that nothing seasonal moved between them, and nothing else launched at the same time — no new campaign, no press mention, no pricing change, no competitor event. It is not defensible across a quarter boundary, around a seasonal business, alongside a paid campaign that also changed, or whenever the "after" period is short enough that a single large client's procurement visit could be sitting inside it. The tell that before-and-after has failed is the same tell as everywhere else in measurement: if you cannot say in one sentence why this change and not something else caused the shift, you have a correlation with a launch date attached to it, not a result.
Reading the before-and-after number against a stable baseline of what your conversion rate already looks like for pages like yours is the cheapest sanity check available — a shift that lands inside the range the page has moved on its own in other months is not a result no matter how neatly it lines up with the launch date.
The alternatives that produce evidence a test cannot, at this volume
None of this means giving up on evidence. It means switching to evidence that does not need hundreds of conversions to be trustworthy.
Watch five to ten people use the actual page, live or recorded, and you will usually find the same friction point independently in the first three sessions — that convergence is the signal, not a percentage. Read the actual text people type into a "how did you hear about us" or an open support ticket field; a dozen people independently describing the same confusion about pricing is more actionable than a test that could not have detected a 10% shift anyway. Ask the sales team what prospects say on the first call, since they hear the objections the analytics never captures. None of this produces a p-value, and none of it needs one, because the question it answers is not "did this specific change move the number" but "what is actually wrong", which is the question a small site should be asking in the first place. It also tends to surface changes large enough to test properly later — a page that gets rebuilt around a genuine confusion point is a big enough change that the exceptions above start to apply, and a real A/B test becomes possible again once there is something worth testing.
There is a second-order benefit worth naming plainly, since it rarely gets said out loud: a site that loads slowly or breaks on mobile will sink any test run on it regardless of what the variants say, so it is worth checking what Core Web Vitals are actually worth before spending months on a headline experiment the underlying page was never going to let succeed.
The order that actually works, for most sites below a few thousand monthly conversions on the page in question, is this: calculate the sample size first. If it fits inside a quarter, run the test properly and do not peek. If it does not, stop calling it a test — do the qualitative work, make the bigger change the qualitative work points to, and use before-and-after with its eyes open about what else moved. A test that cannot finish is not a more rigorous decision process than a considered launch. It is a slower one wearing rigour's clothes.
Questions people ask
- How long should I run an A/B test before deciding it failed?
- You should know the answer before you start, not while you are running it. Calculate the visitors per variant a real effect would need, divide by your monthly traffic to the page, and that is your answer — if it comes out at eight months, the test failed at the planning stage, not the reporting stage.
- Is it ever fine to just launch a redesign without testing it?
- Yes, and for most sites below a few thousand monthly conversions it is the more honest option. A test that cannot reach significance in a reasonable time is not more rigorous than a considered launch, it is just slower and dressed up as science.
- What sample size does an A/B test actually need?
- It depends on your baseline rate and how large an effect you are trying to detect, and the relationship is not linear — halving the effect size roughly quadruples the sample needed. Run the two-proportion calculation with your own numbers before committing to a test; do not borrow someone else's rule of thumb.
- Can I trust a test that reached significance after three days?
- Only if you calculated the required sample beforehand and it arrived in three days. If you were watching the dashboard and stopped because a green number appeared, the result is not trustworthy regardless of what the p-value says, because checking repeatedly and stopping on the first win is a different procedure from the one the p-value was calculated for.