What Is Experiment and How A/B Tests Prove What Works
What is experiment? Learn how A/B tests work, from hypothesis to significance, with CRO examples and pitfalls to avoid.

An experiment is a controlled test that compares a change with a control through random assignment to establish causality, rather than trying something and watching what happens. In UK government guidance, this means splitting users randomly and equally between an original version and a variant, then comparing their outcomes.
You may be facing this question after changing a landing-page headline, moving a checkout button, or rewriting an onboarding screen. Conversions improved afterwards, but did the change cause the improvement, or did a campaign, payday, seasonality, or another site change influence the result?
That distinction separates a useful experiment from a hopeful guess. An experiment turns a change into evidence that can support a decision. It gives marketers, product teams, scientists, and policymakers a structured way to learn what caused an outcome, what didn't, and what should happen next.
What an Experiment Really Means Beyond Trying Something
A marketer changes a button from “Learn more” to “Start free”. A week later, conversions look higher. Then the team notices that a new email campaign launched on the same day, and returning visitors may have behaved differently. The result is encouraging, but the cause remains uncertain.
That activity was a trial. It produced an observation, yet observation alone cannot separate the button's effect from campaigns, audience mix, timing, or other changes. An experiment starts before the change goes live, with a comparison designed to show what happened with the change and what would likely have happened without it.
The distinction is easier to see in practice. UK public-sector guidance describes an experiment as testing a new idea, approach, or technology to collect data. For design choices, GOV.UK guidance on A/B testing describes randomised controlled trials. Users are assigned by chance to a control or variant, then their outcomes are compared with the original page.
A casual trial might follow this pattern:
- Change a headline on Monday.
- Check conversions on Friday.
- Attribute any movement to the headline.
A controlled experiment follows a different chain of reasoning:
- Define what you expect the headline to change.
- Keep the original headline as the control.
- Randomly assign comparable visitors to the original or revised version.
- Measure a preselected outcome.
- Decide whether the evidence is strong enough to act.
Randomisation matters when the team needs to make a confident causal claim. Without it, differences between visitors can create a misleading result. If the goal is only to gather an early signal, a trial may be useful. If the team plans to roll out a change because it caused an improvement, a controlled comparison is required.
Practical rule: If you can't explain what the control group experienced, you probably can't claim that your change caused the result.
For CRO teams, an experiment is a decision tool. It gives opinion-led debates a fair comparison, so the next action rests on evidence rather than preference. For broader CRO context, this A/B test definition explains the mechanics in accessible terms.
A reliable workflow begins by defining the idea and its comparison. The team then selects the experiment type, chooses meaningful measures, protects the test from common errors, and interprets the result with suitable caution. Page changes become repeatable learning opportunities, not permanent arguments about taste.
Understanding the Core Idea Behind Every Experiment
Suppose two shop windows stand on the same high street. Window A displays the original arrangement. Window B uses a new display with a clearer product message. If every passer-by sees both windows, you won't know which display influenced a later purchase. Their second impression, the weather, and the order of exposure could all affect the result.
A fairer approach assigns each visitor to one window at random. The two groups should be similar overall because the assignment process doesn't allow the team to choose who sees the original or the new display. The difference in behaviour can then be connected more credibly to the display itself.

Four ingredients create a useful comparison
A hypothesis is an educated prediction about what will change and why. “A shorter checkout form will increase completed purchases because it reduces effort” is stronger than “Let's try a shorter form.” Teams that need help turning observations into testable predictions can use this guide on how to develop a hypothesis.
The control is the original experience. It provides the baseline that answers the counterfactual question, which is, “What would users have done if we hadn't made the change?”
The variant is the new experience. It should contain an intentional change that relates directly to the hypothesis. If the test changes the headline, page layout, trust badges, and form length at the same time, the result may reveal that the variant worked, but not which element created the effect.
Random assignment distributes users between control and variant without deliberate selection. In the GOV.UK model, users are split equally and randomly, so differences in outcomes have a stronger causal interpretation than a simple before-and-after comparison.
An experiment proves causality by holding the relevant conditions steady and comparing randomly assigned groups that differ in the intervention being tested.
Not every attempt to learn qualifies as a randomised experiment. A designer may show a prototype to customers and collect comments. A product manager may release a feature to one team and observe usage. These activities can be valuable, but they don't automatically provide the controlled comparison needed to establish causality.
The answer to “what is experiment?” therefore depends on context. In everyday language, it can mean trying a new approach to gather information. In rigorous CRO, medicine, or policy evaluation, the word usually implies a planned intervention, a comparison group, randomisation where feasible, and a measured outcome. Understanding which meaning applies prevents teams from treating weak evidence as a proven result.
Different Types of Experiments You Will Encounter
Experiment types differ mainly in how they create comparisons. An A/B test changes one experience for a randomly selected group. A split URL test sends comparable traffic to separate page URLs. A multivariate test changes several elements in combinations, which can reveal interactions but demands more traffic and more careful interpretation.
Broader scientific experiments can take place in laboratories, clinics, or controlled environments. Field experiments test an intervention in the setting where behaviour naturally occurs. In product work, qualitative feedback, usability sessions, and observational analysis often sit alongside controlled tests because they help teams identify what to test next, even when they can't prove causality on their own.
Which Experiment Type Fits Your Goal
| Experiment Type | Best For | Key Requirement |
|---|---|---|
| A/B test | Comparing an original page with one principal change | Random assignment and a defined outcome |
| Split URL test | Comparing substantially different page builds or flows | Reliable routing and equivalent audience allocation |
| Multivariate test | Evaluating combinations of several page elements | Enough traffic to separate combinations and interactions |
| Laboratory experiment | Isolating variables under controlled conditions | Consistent procedures and controlled inputs |
| Field experiment | Measuring an intervention in a real setting | A defensible comparison group and protection against outside influences |
For most website CRO programmes, the A/B test is the workhorse. It keeps the question narrow: does version B produce a different outcome from version A for randomly assigned visitors? That structure works well for headlines, calls to action, layouts, forms, and other changes that can be delivered within the same site experience.
A multivariate test can be appropriate when a team has a strong reason to study interactions, such as whether a headline and image work better together than separately. It isn't automatically more valuable. If the audience is limited, a simpler A/B test may produce a clearer decision.
Product teams also need evidence before building a test. Interviews, usability observation, and user feedback for mobile app teams can reveal friction that analytics won't explain. Those methods answer “what are users struggling with?” A controlled experiment answers a different question, “did this specific intervention change the outcome?”
For a deeper look at tests involving several page elements, see what multivariate testing means. The important choice isn't the most complex method. It's the method that matches the decision, audience, risk, and available evidence.
Key Concepts That Make Experiments Trustworthy
A trustworthy experiment turns a proposed change into a decision supported by evidence. The team defines the question, predicts what should happen, assigns people fairly, collects enough observations, and sets the decision rule before reviewing the result. Each layer limits a different source of doubt.
Start with the question and the variables
A hypothesis gives the work direction. It states the proposed change, the expected behaviour, and the reason for expecting it. The independent variable is what the team changes, such as a button label. The dependent variable is what the team measures, such as completed sign-ups.
This distinction keeps the team from treating every tracked event as a success measure. Choose one primary metric that represents the decision, then add secondary metrics to expose side effects. A button change might raise sign-ups while also attracting less qualified enquiries, so both outcomes may matter.
Randomisation and test planning
Randomisation balances factors the team cannot fully observe or control. In website testing, those factors can include device type, returning or new visitors, purchase intent, and source channel. A random split will not produce identical groups in every short period. Across the planned test, it creates a defensible comparison between people who saw each version.
The need for randomisation depends on the decision. Casual feedback may help a team generate ideas, but confident claims about whether one version caused a different outcome require a controlled comparison. Without that structure, a change in traffic quality, seasonality, or audience mix can look like an effect from the intervention.
Sample size and duration should be set before early results influence the plan. Expected change and monthly traffic determine how much exposure a test needs and how long it should run. A small detectable effect needs more evidence than a large difference, while a low-traffic page may require a longer observation period.
Statistical significance as a decision threshold
Many CRO teams use a 95% confidence threshold as a practical standard for judging whether a result is unlikely to reflect random chance. Confidence is not a promise that a variant will perform the same way forever. It describes the strength of the evidence under the chosen statistical model and its assumptions. The explanation of statistical significance in A/B testing describes how a frequentist z-test engine can calculate significance continuously and identify a winner at that threshold.
A platform can handle the arithmetic, but it cannot repair a vague question or a contaminated test.
The final review should answer three questions:
- Did the primary metric move? This shows whether the main objective changed.
- How large and useful was the difference? Statistical evidence does not make an effect commercially worthwhile.
- Did guardrails deteriorate? More sign-ups may come with more support requests, lower satisfaction, or poorer downstream quality.
Trust comes from aligning the hypothesis, design, measurement, and decision rule before the outcome is known.
Real World Examples of Experiments in CRO and Product Testing
A headline experiment can test whether a more specific promise helps visitors understand the offer. A call-to-action experiment can compare “Start free” with “Book a demo” when the audience has different levels of buying intent. A layout experiment might place proof points closer to the form, while a pricing-display experiment tests whether clearer plan distinctions reduce hesitation.
Each test needs an outcome connected to the business. That might be a completed purchase, a qualified enquiry, an activated account, or revenue per visitor. The page element is the intervention, but the business outcome determines whether the change deserves rollout.

GOV.UK illustrates why modest interface changes deserve serious measurement. Its report on A/B testing across two teams recorded 68,000 fewer exits per year and 280,000 more clicks on search results per year. The source explains that changes to common user journeys can create substantial aggregate effects when many people use the service. The GOV.UK account of those measured improvements shows why a small design decision can matter beyond the individual page session.
Measure gains without hiding costs
A primary success metric might be completed checkout. Secondary guardrails could include help-link clicks, form completion, customer satisfaction, refund behaviour, or the quality of leads generated. These measures protect teams from declaring victory because one convenient metric rose while the wider experience worsened.
A product team testing onboarding can measure activation as its primary outcome, then monitor support contacts and early feature abandonment. An ecommerce team testing a promotional message can watch purchases while checking average order value and revenue per visitor. The correct guardrails depend on the user's journey and the risk of the change.
A well-designed experiment also creates an operational record. Teams should capture the hypothesis, audience, allocation, dates, primary metric, guardrails, and final decision. That record turns individual tests into institutional learning instead of isolated dashboard screenshots.
The following video provides a visual introduction to A/B testing and conversion-focused comparisons.
Common Pitfalls That Undermine Your Results
A test can have a polished dashboard and still produce weak evidence. The most damaging errors usually happen in the process, before anyone interprets the final chart.

The errors that create false confidence
Peeking early turns a temporary lead into an apparent win. If the team repeatedly checks the results and stops as soon as the variant looks favourable, it changes the decision process and makes the stated confidence threshold less meaningful.
Using too little evidence creates noisy outcomes. A handful of conversions can move a rate sharply, especially on a low-volume page. Planning the required sample and duration before launch reduces the temptation to interpret random fluctuation as a durable effect.
Changing several important variables at once makes attribution difficult. A redesigned hero section might outperform the original, but the team may not know whether the improvement came from the copy, image, form position, or trust signal.
Ignoring external factors can contaminate a result. Campaigns, product launches, holidays, technical incidents, and unusual traffic sources may affect one period differently from another. Record those conditions and investigate unexpected shifts before making a permanent change.
Stopping too soon can miss delayed behaviour. Some users need time to return, compare options, or complete a longer journey. A short test may overrepresent immediate reactions and underrepresent the outcome that matters commercially.
Correlation creates another trap. If conversions rose after a page change, the timing is compatible with causation, but timing alone doesn't prove it. Random assignment and a control group provide the structure needed to separate the intervention from unrelated movement.
Ethics adds a wider layer of responsibility. Digital-government guidance can frame controlled trials as low-risk ways to try new approaches with limited disruption, but experimentation involving people still raises questions about consent, vulnerability, manipulation, and harm. Those questions are especially visible in the UK, where a January 2025 YouGov poll reported that 70% of Britons supported a law to end animal experiments in medical research by 2035. The poll is reported in the UK government and CMA discussion of experiments, and it demonstrates that the word “experiment” carries ethical weight beyond marketing dashboards.
Before launch, ask whether the change is proportionate, whether users face avoidable harm, and whether the team can explain the intervention. Better evidence doesn't excuse poor treatment of participants.
Putting Experiments to Work for Better Decisions
A practical experiment begins with a decision, not a tool. Write down what you believe, what change will test that belief, which original experience forms the control, and which outcome will decide the result. If the decision affects a high-risk journey or a large audience, define guardrails before launch.
Use randomisation when you need to claim that a change caused a difference. If you're only exploring usability problems, collecting customer language, or diagnosing a broken journey, observation and feedback may be enough at that stage. The method should match the question.
A compact working checklist looks like this:
- State the hypothesis. Use an if-then prediction tied to a reason.
- Choose one primary outcome. Make success clear before exposure begins.
- Keep a valid control. Preserve the original experience for comparison.
- Assign users randomly. Avoid choosing who receives the new treatment.
- Plan sample and duration. Base them on the expected change and available traffic.
- Set the decision threshold. Use the agreed 95% confidence standard where appropriate.
- Review guardrails. Check quality, satisfaction, support demand, and downstream effects.
- Document the decision. Record what the team learned, including an inconclusive result.
Start with one high-impact page element rather than rebuilding the entire customer journey. Lightweight A/B testing software can support variants, traffic allocation, conversion goals, revenue measures, and significance monitoring, provided the implementation doesn't damage loading performance or alter the experience in unintended ways.
The next time a team argues about a headline or button, ask a more useful question than “Which version do we prefer?” Ask whether the proposed change can be tested fairly, what evidence would justify rollout, and what result would make you abandon the idea. That shift turns experimentation into a repeatable way to make better decisions.
Otter A/B lets teams create website split tests, assign traffic to control and variants, and measure goals such as clicks, purchases, revenue, and custom events from a live dashboard. Visit Otter A/B to start a free test without a credit card and replace guesswork with evidence from your next important page change.
Stop guessing
Ready to start testing?
Set up your first A/B test in under five minutes. No credit card required.
- 14-day free trial
- No credit card required
- Cancel anytime