TheProduct Playbook

A/B Test

Two versions of one thing, shown to two comparable groups at the same time, where only a single element differs. You measure which version drives the behavior you care about, and you ship the winner. That's it.

Every other experiment in this catalog runs before the thing exists. The A/B test is the odd one out. It needs a live product with real traffic to run against, because it's not asking whether anyone wants the thing. It's asking which version of a thing they already use gets more of them to act on it.

The question it answers

An A/B test answers one narrow, sharp question. Between two real options on live traffic, which one converts better?

Not "will anyone buy at all." You already have the yes. You have a shipped product and people moving through it. The question now is optimization. Which headline, which button, which line of copy gets more of the people who already showed up to act? You're not validating the idea anymore. You're tuning how often it converts.

The discipline that makes it honest is the control. You change exactly one thing between the two versions, so the difference in results can only be that one thing. Same audience, same traffic source, same timing, one variable. Break that and you've got two pages that differ in five ways and a result that tells you nothing about which of the five mattered.

The way these experiments are organized underneath this page, and the catalog it sits in, comes from Testing Business Ideas by David Bland and Alex Osterwalder. The A/B test is one of the 44 experiments they map. Credit where it's earned.

How to run it cheaply

The minimum version is one page, one change, and a clean random split.

  1. Pick one element you believe matters. A headline, a button color, a price frame, one line of copy. One. The whole method dies if you change two things at once.
  2. Build the variant. Clone the live version (call it the control), change the one element, leave everything else identical.
  3. Split traffic evenly and at random. Which version a person lands on is a coin flip. Most analytics and page tools do this split for you.
  4. Write the success number and the end date down before you turn it on. Which behavior counts as a win (a click, a signup, a purchase), and how big a gap you need before you'll call it real and not noise.
  5. Run it until you have enough visitors that the gap can't be chance, then check which version got the higher conversion rate and ship that one. "Enough visitors that it can't be chance" has a real name, statistical significance, and your analytics or A/B tool calculates it for you. You don't eyeball it.

The piece people skip is that last one. An A/B test needs real traffic to mean anything. A 6% versus 4% gap across 40 visitors is just noise. The same gap across 4,000 visitors is a decision you can ship.

A worked example

Picture a checkout page that converts at 4%, and a hunch that a clearer money-back guarantee would lift it. The company, the page, and every number here are invented to make the mechanics concrete. The method is real.

You don't redesign the page. That would change ten things and teach you nothing about which one moved the dial. You run the current page as the control against an identical page with one change: a single money-back guarantee line, placed right above the buy button.

Traffic splits evenly and at random. Before launch you write the bar down. The guarantee version has to beat control by a margin big enough that it can't be chance, across enough visitors to trust it. It lands at 4.9% against 4.0%, with enough traffic behind it that the gap isn't luck.

One change, one clean read. You ship the guarantee line, and that 0.9 points is real money on a page that was always going to get that traffic anyway. Then you test the next single thing. Change five things at once and conversion might climb, but you'll never know which one did it, so you can't ship the good part. One variable. Always.

When to reach for it, and when not

Reach for an A/B test when you have a live product with enough traffic to split, and a specific question of the form "does this version beat that one?" Both halves have to be true. Live traffic, and a one-variable question. Miss either and this isn't your tool.

Don't reach for it to find out whether anyone wants the thing at all. That's the wrong end of the catalog. If you don't have a product live yet, or you're still asking "will anyone reach for this," you want the demand tests up front, the landing-page test and the fake-door test. An A/B test optimizes a yes you already have. It can't manufacture one you don't.

Don't confuse it with its neighbors either. The Wizard of Oz test fakes a feature that doesn't exist to see if people use it. The A/B test pits two real, shipped versions against each other to see which performs. One asks "does this work at all, faked behind a curtain." The other asks "which of these two real things works better." Don't run one thinking you ran the other.

And don't run it on a page that doesn't get the traffic. This is the one experiment here that can't be run cheaply if the numbers aren't there, because the whole read depends on a sample big enough to trust. Low traffic, no result. When that's you, go back down the ladder to a test that learns from five people instead of five thousand.

The rule that ties this catalog together still holds: find the riskiest assumption, name the question under it, run the cheapest experiment that answers it. The A/B test is the top rung, the one you reach for last, once a real product with real traffic has earned the right to be tuned. The full catalog has the other twelve recipes in this Toolbox.