Marketing

What is A/B testing? Answering "which is better" with data, from ads to email

What is A/B testing and how do you do it? Hypothesis, the single-variable rule, sample/duration, statistical significance, limits, and tying the result to real revenue via your CRM.

Rocketly · 2026-07-10

"Is this headline better, or that one?" "Should the button say 'Buy Now' or 'Try It Now'?" These debates end the same way in most teams: the loudest or most senior person's opinion wins. Yet in marketing there's a world of difference between "I think" and "it's been proven." A/B testing is exactly the method that ends this debate with data, not opinion: it shows two versions at the same time and tells you with numbers which actually works better.

But when set up wrong, A/B testing can look data-driven while actually producing misleading results — tests stopped early, done with insufficient samples, or changing ten things at once produce guesses dressed up as "proof." In this guide we cover what A/B testing is, its correct process, why statistical significance matters, its limits, and how to tie the result to real revenue.

What is A/B testing?

A/B testing is preparing two versions of something (A and B), splitting your audience in two, showing each group one version, and measuring which performs better on a metric you define (clicks, form fills, purchases). Because the two versions compete under the same conditions in the same time window, the difference between them can be attributed to the element you changed rather than to chance. It's a way to answer "which is better?" with a controlled experiment.

At its core, this is the scientific method applied to marketing: you form a hypothesis, run a controlled experiment, and observe the result. The difference is that you work with your real customers instead of a lab.

What do you test?

The range of things you can A/B test is wide: a landing page's headline, the call-to-action button's text and color, an email's subject line, an ad image, page layout, form length, even pricing presentation. As a rule, any prominent element that can influence a visitor's decision can be tested. But rather than trying to test everything, starting with elements that can have a large impact (headline, offer, main button) is how you use limited traffic most efficiently.

The A/B testing process

1Hypothesis2Two Variants (A/B)3Split Traffic4Measure5Apply the Winner
A healthy test doesn't start randomly but with a hypothesis, and ends by applying the winner.

A healthy A/B test has five steps. First you form a hypothesis ("if I make the headline benefit-focused, form fills will rise"). Then you prepare two variants that test this hypothesis (the existing one and the new one). You split the traffic in two (half the visitors see A, half see B). You measure the metric you defined. And once enough data accumulates, you apply the winner. This loop turns a guessing "I think" into a repeatable learning system.

One variable at a time

The golden rule of A/B testing is to change only a single element in one test. If you change the headline, the button, and the image at the same time and B wins, you can't know what caused the win — maybe the headline worked but the new image hurt, and the net result is misleading. The single-variable rule lets you attribute the result confidently to that one element you changed. If you want to test multiple elements at once, that's no longer A/B but multivariate testing, and it requires far more traffic.

Sample and duration: a matter of patience

The sneakiest trap of A/B testing is jumping to a conclusion before enough data accumulates. Saying "B won" because B is ahead in the first 20 visitors is like flipping a coin twice and saying "tails comes up more often." A reliable result requires a sufficient number of visitors and conversions; that number depends on your current conversion rate and the size of the difference you want to detect. Run the test to a predetermined duration or sample size, and don't "keep peeking" before reaching that threshold.

Statistical significance: noise or real?

Statistical significanceLow confidenceHigh confidence
Significance indicates that the probability the difference is a coincidence is low enough.

Seeing a difference between two variants isn't enough; you need to be reasonably sure that difference isn't chance. Statistical significance measures exactly this: how low is the probability that the difference you observed arose by coincidence? In small samples, differences are often coincidental — A ahead today, B tomorrow; this fluctuation is "noise," not "signal." Most testing tools calculate significance for you; your job is not to decide before the tool says "confident enough."

The "keep peeking and decide early" problem comes in here: if the winner changes every time you look at the test, you don't yet have enough data to decide. Being patient isn't a virtue in A/B testing — it's a necessity.

Where do you run A/B tests?

A/B testing is most valuable in three areas. On landing pages, headline, offer, and button tests directly affect conversion. In email marketing, subject-line and content tests raise open and click rates. In ads, image, copy, and targeting tests improve budget efficiency — in this last area, analyzing ad creatives before publishing provides a pre-screen that complements the test. Most of these platforms offer built-in A/B testing tools, so a separate piece of software is often unnecessary.

The limits of A/B testing

A/B testing is powerful but not a cure-all. Its biggest limit is traffic: on a page that gets a few visitors a day, seeing a statistically significant difference between two variants can take months — during which the result often stays "inconclusive." In low-traffic situations there are alternatives: trying variants sequentially rather than simultaneously, or turning to qualitative feedback (user observation, surveys). A/B testing shines when you have enough volume; forcing it without volume leads to mistaking noise for reality.

Tying the result to real revenue

The classic measurement of A/B testing is clicks or form fills — but the real question is whether the winning variant brought more revenue, not more clicks. Sometimes a more-clicked headline brings lower-quality leads and converts to fewer sales. To catch this trap, you need to track the leads from your forms in your CRM and measure how many sales and how much revenue each variant actually turned into. That way you answer "which variant won" with the real result reflected in your bank, not surface-level clicks.

Judge the winning variant by its revenue

Rocketly ties the leads from your forms and campaigns to real revenue, so you pick your A/B test's winner by closed sales, not just clicks.

Start Free

Common mistakes

  • Stopping the test early: Declaring "it won" before enough data is mistaking noise for signal.
  • Changing many things at once: If a winner emerges, you can't know what caused the win.
  • Ignoring significance: Mistaking a coincidental difference for a real one leads to a wrong decision.
  • Testing trivial things: Small details like a button color's shade waste limited traffic.
  • Measuring only clicks: A more-clicked variant can bring lower-quality leads; you miss it if you don't measure revenue.
  • Forcing it on low traffic: Without enough volume, A/B testing produces uncertainty; alternative methods are healthier.

Getting-started checklist

  • 1. Write a clear hypothesis. "If I change this, that metric will rise."
  • 2. Set a single variable. Change only one element in a test.
  • 3. Set the sample/duration threshold in advance. Don't decide before that threshold.
  • 4. Wait for significance. Don't declare a winner before the tool says "confident enough."
  • 5. Apply the winner, record the learning. Every test feeds the next hypothesis.
  • 6. Measure revenue in the CRM. Count closed sales as the winner, not clicks.

Frequently asked questions

How much traffic do you need for an A/B test?

There's no exact number; it depends on your current conversion rate and the size of the difference you want to detect. Roughly, the lower the conversion rate and the smaller the difference you're seeking, the more traffic you need. With very little traffic, reaching a meaningful result can take months; in that case consider other methods.

How long should I run a test?

Until you reach the sample size or duration you set in advance. At least one full week is usually recommended (to cover behavioral differences between days of the week), but the real determinant is reaching a sufficient number of conversions and statistical significance.

What's the difference between A/B testing and multivariate testing?

A/B testing compares a single element with two versions. Multivariate testing tries different combinations of multiple elements at once. Multivariate testing gives more information but requires far more traffic; for most SMBs simple A/B testing is more practical and reliable.

Is the winning variant always permanent?

No. A test's result is for that audience, that period, and that context; when season, campaign, or audience change, the result can change too. A/B testing isn't a one-time victory but an ongoing culture of learning — you apply the winner, then continue with a new hypothesis.

A/B testing is the most powerful tool for turning an "I think" culture into an "it's been proven" culture in marketing — but only when set up correctly. Start with a clear hypothesis, change one variable, be patient for enough data and statistical significance, and measure the result not by clicks but by real revenue in your CRM. Once you make this a habit, every campaign gets a little better than the last — because you're no longer guessing, you're learning.