A conversion experiment is worth running when there is a real decision to make, trustworthy measurement, sufficient eligible traffic for a useful effect, and the ability to execute a valid comparison. Testing is not the default response to every page problem. A broken form should be repaired; an unclear buying question may need research; a uncertain commercial tradeoff may justify randomization.

The decision is not simply whether you have “enough traffic.” It is whether you can learn enough, within a stable operating period, to justify the cost and risk of the experiment. Start with the decision the result will change and work backward to the design.

State the decision and competing explanations

Write down what you would do if the result showed a worthwhile benefit, a meaningful harm, or unresolved uncertainty. If the team plans to ship regardless, ask what knowledge the experiment is actually meant to provide. It may still be useful for measuring risk, but that purpose needs to be explicit.

Formulate a mechanism. “Showing the complete delivery cost earlier will help eligible buyers judge the offer before starting checkout” gives you something to investigate. “A different color will increase conversions” is a prediction without much explanation of the buyer’s problem.

Microsoft’s pre-experiment guidance emphasizes a clear hypothesis and appropriate success, guardrail, and data-quality measures. Use that structure to write an experiment brief. Do not treat every intermediate increase as success: more checkout starts can coexist with fewer paid orders.

Decide whether randomization is the right tool

Use this table to choose a next step. The categories are editorial decision aids, not rigid rules.

Situation Better next move Why
Valid enquiries cannot submit Repair and verify There is a confirmed functional defect
Buyers cannot explain the offer Comprehension research You need a plausible alternative before testing impact
Two credible presentations trade clarity against brevity Controlled experiment if feasible The commercial effect is uncertain
Traffic is too sparse to resolve a useful effect Research plus monitored release A weak experiment may consume time without resolving the choice
Tracking disagrees with paid orders Measurement repair A comparison cannot rescue invalid outcomes
A major commercial policy changes simultaneously Redesign or postpone the comparison The treatment and context need to be interpretable

Research and experiments can be sequential. Observe a recurring problem, design a repair, check that buyers understand it, then compare commercial outcomes if uncertainty remains. This is often more informative than randomizing several arbitrary page designs.

Check measurement before power

Confirm that the eligible population is identifiable, assignment persists, exposure is recorded, and the outcome is measured consistently. Follow a visitor across redirects and external payment or booking systems. Determine whether consent choices or blockers create missing data and whether that missingness differs by variant.

Choose an assignment unit compatible with the decision. If the same buyer repeatedly visits the page, session-level assignment can show conflicting versions. If several people decide together, individual assignment may introduce exposure to both experiences. The right design depends on the business and should be reviewed when interference is plausible.

Include a data-quality plan. Microsoft’s sample ratio mismatch guide explains why unexpected allocation imbalance can signal an experiment problem. A large conversion difference is not useful evidence if the eligible groups were created inconsistently.

Choose a worthwhile effect before estimating sample size

Ask what improvement would justify implementation and maintenance. Express it in absolute terms alongside relative terms. Moving from 4% to 4.4% is a 0.4 percentage-point increase and a 10% relative increase. Those descriptions are both correct, but the smaller absolute change determines how many additional conversions occur.

Estimate sample size using the baseline rate, the smallest effect worth resolving, the planned statistical method, power, allocation, and error threshold. Use a calculator or statistical review appropriate to the design. A standard fixed-horizon calculation does not automatically apply to a sequential or clustered experiment.

For a simple binary outcome with equal independent groups, fixed-horizon testing, a two-sided 5% threshold, and roughly 80% power, detecting a 4% to 4.4% difference can require approximately forty thousand observations per group. This is an illustrative approximation, not a universal minimum. A method that accounts for exact proportions, continuity, or design effects may return a different requirement.

Translate the requirement into operating time

Use eligible unique traffic, not total site sessions. Remove people who will not receive the treatment or cannot contribute the chosen outcome. Estimate the allocation to each variant, then account for outcome maturation, weekday patterns, and likely operational changes.

A nominal sample target reached in two days may still fail to represent the normal buying cycle. A target requiring six months may cross campaign changes, seasonality, and offer revisions that make the result hard to use. Plan a stable period and a maximum duration before you start.

Do not solve a slow test by choosing an implausibly large effect merely to make the calculator produce a short duration. That changes the question. If smaller but valuable effects remain unresolved, acknowledge that limit and choose another decision method where appropriate.

Worked example: an experiment that does not fit the calendar

Synthetic example. A page receives 8,000 eligible visitors per month and converts 4% to paid orders. The business would consider a 0.4 percentage-point gain worthwhile. An approximate plan requires about 80,000 total eligible visitors, or around ten months at the current volume.

The team expects a major offer revision in six weeks. This experiment is poorly matched to the decision. Instead, it investigates whether buyers understand delivery costs, tests a prototype for comprehension, and releases a verified clarification with monitoring. It records that the commercial effect was not causally estimated.

If the business later has substantially more stable eligible traffic, it can revisit the experiment. The recommendation is not that small sites should never test; it is that this particular effect and calendar do not support a useful comparison.

Budget the decision, not only the test

Estimate the opportunity cost of waiting as well as the cost of a mistaken release. A reversible clarification with verified accuracy may justify monitored deployment when a long experiment would delay useful information. A change that alters the commitment or could create substantial harm deserves a stronger comparison and rollout plan. Neither choice removes uncertainty. Write down what you will learn, what remains unmeasured, and what event would make you reverse the decision.

Worksheet: experiment readiness

Requirement Your plan
Decision the result changes ______________________________
Buyer problem and proposed mechanism ______________________________
Primary outcome and denominator ______________________________
Guardrails and harm threshold ______________________________
Assignment and exposure unit ______________________________
Baseline and smallest worthwhile effect ______________________________
Method, sample requirement, and assumptions ______________________________
Eligible traffic and stable run window ______________________________
Data-quality checks and stop rules ______________________________
Action if the result remains uncertain ______________________________

Make uncertainty an acceptable result

Commit to an interpretation plan before launch. An inconclusive test is not proof of no effect. It may mean the plausible effects span both useful benefit and unacceptable harm. Report that range and decide what further evidence is worth obtaining. Experimentation earns its cost when it improves a decision, including the decision to avoid an inadequately measured change.

Sources

Send a correction