Combine interviews, session observations, and funnel data by asking each method a different part of the same question. Funnel data describes where measured outcomes differ. Observation shows how people interact with a page or process. Interviews help explain the decision context and the meaning people give to it. Agreement can strengthen a hypothesis, but the methods are not interchangeable.
The aim is a decision with an explicit evidence trail. Avoid a dashboard that says only “users struggle,” or a presentation that turns three quotations into a population statistic. Keep the audience, time period, and limitations beside the interpretation.
Start with one decision question
For example: “Should we explain delivery costs before asking buyers to begin checkout?” This question suggests useful evidence: where exits occur, whether buyers can locate the cost, how they interpret it, and whether the earlier explanation supports the choice.
Define the cohort. Are you investigating first-time mobile visitors from a campaign, returning customers, or every shopper? A useful finding about one group may not apply to another. Align the question with the actual intervention the team could make.
GOV.UK’s research-planning guidance recommends starting with questions and turning assumptions into researchable uncertainties. Use that principle to prevent a favored solution from determining what evidence counts.
Give each method a job
The following table is a working guide, not a ranking of methods. Different questions require different combinations.
| Method | Can help answer | Cannot establish alone |
|---|---|---|
| Funnel data | Where measured progression differs by cohort | Why a person left or whether the page caused it |
| Session observation | What interaction occurred and where a task became difficult | The participant’s unspoken motive or population prevalence |
| Interview | What happened around a recent decision and what mattered | A causal conversion effect or accurate recall of every action |
| Controlled experiment | Whether assignment to a change affects a defined outcome under the design | A complete explanation of the mechanism or universal applicability |
Do not make one method perform the job of another. A cursor moving back and forth does not establish confusion. A participant saying shipping was expensive does not show how frequently cost causes exits. A lower completion rate on mobile does not establish that the layout is the cause.
Check the funnel definitions
Write the start, intermediate events, outcome, unit, and time window. Verify that the events reflect real states. A button event may occur before a request is accepted. A product-view event may be missing on one page version. A cross-domain handoff may make returning purchasers look like new visitors.
Compare meaningful cohorts and show counts. If mobile visitors come mostly from a broad social campaign and desktop visitors come mostly from returning customers, a device comparison also contains an intent difference. That does not make the data useless, but it changes the explanation you can defend.
Identify where additional instrumentation would materially help. A field-error category may be useful; capturing the field’s personal value usually is not. Measure the state necessary to answer the question and review privacy requirements before collecting more detail.
Observe the task without narrating the motive
Choose sessions according to the question rather than selecting only the most dramatic examples. Include successful and unsuccessful journeys, and document the sampling approach. If recordings omit content, mask details, or miss external steps, record those gaps.
Describe events literally: “The visitor opened shipping details, returned to the product page, and ended the recorded session.” Then write the possible interpretations separately. They may have been comparing costs, checking eligibility, or interrupted by something outside the page. The recording does not select one explanation for you.
In moderated work, ask participants to complete a realistic buying task. When they hesitate, invite a neutral explanation rather than suggesting a cause. Record whether the researcher prompted, helped, or changed the task. A successful completion after assistance is different from independent completion.
Use interviews to reconstruct context
Ask about a recent actual decision, including what occurred before and after the site visit. Start with the story before introducing a specific page issue. Learn the purchase deadline, alternatives, budget constraints, and other people involved.
GOV.UK’s interview guidance recommends concrete examples and open questions. Apply that guidance by asking “How did you work out whether delivery would fit your deadline?” instead of “Did our unclear delivery information make you abandon?” The second question embeds both a diagnosis and an outcome.
Treat memory as evidence with limits. People may remember the deciding issue but not the exact sequence of clicks. Their experience can be useful without becoming a precise event log.
Create a finding with contradictions attached
Use a small evidence ledger. State the question, direct observations, interpretation, alternatives, and proposed next action. Link the finding to its supporting records internally, while avoiding unnecessary personal details in reports.
GOV.UK’s analysis guidance provides a useful reference for moving from observations to findings and actions. For this commercial workflow, add a specific column for competing explanations so uncertainty stays attached to the decision.
A finding might be: “First-time buyers need destination-specific delivery information before choosing.” Supporting evidence could be task failures and interview stories. Contradictory evidence might show that repeat buyers already know the delivery policy. That distinction can improve the proposed change rather than weaken the research.
Worked example: evidence that narrows the fix
Synthetic example. Funnel data shows 600 mobile checkout starts and 300 paid orders in a defined window. Recorded sessions show some visitors returning to the product page after the shipping step. Interviews with recent shoppers reveal two different situations: some were checking arrival timing, while others rejected the total cost.
The evidence does not justify saying that half the mobile audience abandoned because delivery was unclear. It suggests separating the timing question from the price question. The team proposes a destination-based delivery estimate and a clearer cost explanation before checkout, while commercial owners review whether the available shipping option fits the audience.
A prototype test checks whether buyers can identify the expected arrival and cost. An experiment, if feasible, can evaluate commercial impact. Each stage answers a narrower question, and the team records which claims remain unproven.
Before the synthesis meeting, ask each observer to bring a direct observation and a separate interpretation. Compare disagreements before merging notes into themes. When several team members read the same incident differently, return to the record or ask a follow-up question. Consensus achieved by deleting contradictions is weaker than a narrower finding that explains where interpretations diverge.
Worksheet: evidence synthesis
| Field | Your notes |
|---|---|
| Decision question | ______________________________ |
| Cohort, period, and outcome | ______________________________ |
| Funnel observation with counts | ______________________________ |
| Direct session observation | ______________________________ |
| Interview context from a real decision | ______________________________ |
| Interpretation shared by the evidence | ______________________________ |
| Contradiction or missing audience | ______________________________ |
| Alternative explanation | ______________________________ |
| Next action and what it will establish | ______________________________ |
Know when to stop collecting
Stop a round when you have enough evidence for the next decision, not when every uncertainty disappears. A reproducible error can justify repair before you estimate its prevalence. A broad commercial claim needs stronger support than one observed incident. Keep unresolved questions that could change the decision, and retire questions that no longer matter. Combining methods should make the explanation more precise, not merely make the research presentation longer.
