Experimentation

Why growth teams should test the whole journey

Many growth moves span offer, page, creative, email, checkout, and follow-up. Testing one touchpoint often misses the real result.

Effective
Last updated
Reading time
7 min

In 2012, a Microsoft engineer tested a small Bing change: combine an ad's second body line into its headline. The test increased annualized U.S. revenue by more than $100 million without hurting monitored user-experience metrics. It became one of Bing's highest-revenue experiment wins.

The same experimentation literature shows the opposite problem. Short-window tests can produce clean local wins that disappear over months. More ads can raise revenue per session and still reduce long-term usage. A page can win the two-week test and hurt the customer journey.

That is the core issue. Most A/B testing tools test a touchpoint: a button, headline, landing page, email, or screen. Growth teams need to know whether a whole journey worked: offer, page, ad, email, retargeting, checkout, subscription, retention, and repeat purchase.

Touchpoint testing is not enough

Touchpoint tests are useful. They answer narrow questions quickly.

But many growth decisions are not narrow. An offer change may require a new landing page, new creative, new email follow-up, new checkout copy, and a new upsell. Testing only one page view does not tell the team whether the full move worked.

The problem gets worse when several tests overlap:

  • A visitor sees variant A on the landing page.
  • They receive an email from another test.
  • They click a retargeting ad with different copy.
  • They enter checkout during a pricing experiment.
  • They return later on another device.

The tool may still report a local lift. But the treatment the customer actually experienced is no longer clean.

Why the unit of testing changed

The old web experiment was often page-shaped. A visitor landed on a page, saw version A or B, converted or did not convert, and the team measured the result.

That world still exists for some product and page tests. It is not the full growth loop anymore.

In e-commerce, a serious growth move usually changes several things at once. A bundle offer needs a landing page, product setup, cart messaging, ad creative, email follow-up, maybe a post-purchase upsell, and a measurement rule that accounts for refunds and repeat purchase. A subscription push may change quiz logic, offer framing, checkout cadence, lifecycle messaging, and cancellation risk. A media-budget test may move spend across audiences while creative and offer fatigue are changing at the same time.

The customer does not experience those as separate tool changes. The customer experiences one journey.

This is why a clean button test can be true and still not answer the business question. A button may increase add-to-cart rate while lowering AOV. A discount may raise conversion while lowering margin. A landing page may win first purchase and lose repeat purchase. A short-window test may look good before refunds arrive. Local lift is useful. It is not the same as journey lift.

What has improved

The experimentation industry did fix one major problem: peeking.

Classical A/B testing assumes the sample size is fixed before the test starts. Operators did not behave that way. They watched the dashboard every morning and stopped when they liked the result. Research from Johari, Pekelis, Walsh, and others showed that this inflates false positives.

Sequential and always-valid methods improved that. Modern platforms can let operators look at results during the test without breaking the math. That is a real advance.

But better statistics on the same wrong unit still leave a gap. If the business question is about the journey, the test has to be designed around the journey.

What teams get wrong in practice

Most bad experiment programs do not fail because nobody understands statistics. They fail because the operating discipline around the test is weak.

The team changes the page after launch. The media buyer shifts budget mid-test. The lifecycle team updates an email that one variant depends on. The audience definition changes. The revenue window ends before refunds or repeat purchases show up. The agency reports blended ROAS while the test was designed around net new customers. The final decision is written in Slack and the next brief never sees it.

None of those mistakes requires bad intent. They happen because the test is managed as a dashboard event instead of as company work.

The discipline that matters is simple:

  • decide the hypothesis before launch
  • decide the primary metric before launch
  • define the audience and assignment before launch
  • record every exposure that matters
  • lock the design after launch
  • keep changes visible
  • tie the result to revenue and margin when the decision is about money
  • write the final decision where the next test can use it

That is not bureaucracy. It is how experimentation avoids becoming opinion with a chart attached.

What journey-level testing requires

A journey-level experiment needs four things.

A durable assignment. The visitor or customer is assigned once, usually at first contact, and that assignment follows them across page, email, ad, checkout, and follow-up.

Exposure logging. Every touchpoint records what the visitor actually saw. Not just "variant A existed," but "this visitor saw this version of this page or email at this time."

Revenue and state outcomes. The result should connect to the business event that matters: order, net revenue, refund, subscription, repeat purchase, LTV, retention, or payback. A click is not enough when the decision is budget allocation.

A locked design. The team should know what was tested, what metric mattered, what window was planned, and what changed after launch. Without a locked design, post-hoc analysis becomes storytelling.

A concrete example

An e-commerce team wants to test whether a bundle offer beats a discount offer.

Variant A is a discount-led journey: discount creative, discount landing page, discount email follow-up, and a cart message focused on savings.

Variant B is a bundle-led journey: bundle creative, bundle landing page, bundle email follow-up, and a cart message focused on value.

The question is not which headline gets more clicks. The question is which journey produces better net revenue, AOV, repeat purchase, refund rate, and payback over the measurement window.

That requires the visitor assignment, every exposure, and the final revenue result to be tied together.

What Lyberty does

Lyberty is built around that loop.

An experiment starts with a brief: hypothesis, audience, primary metric, variants, launch rules, and planned decision window. At launch, the design is locked. The system records exposure events, joins outcomes back to the assigned visitor when possible, and keeps the decision tied to the experiment history.

The point is not to make every test longer. The point is to stop treating a local touchpoint win as proof that the business improved. Some tests should stay small. The expensive growth moves - offers, funnels, creative systems, lifecycle flows, channel shifts - should be measured at the journey level.

What to ask about your testing program

  1. How many wins from the last quarter were measured on revenue or LTV, not just clicks or conversion rate?
  2. Can you see which emails, ads, pages, and checkout states each variant included?
  3. Can you tell whether a visitor stayed in the same variant across the journey?
  4. Can finance trust the revenue result?
  5. Does the next test reuse the learning from the last one?

If the answer is no, the team is running tests, but the learning is not compounding.

Sources

  1. Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd.
  2. Wald, A. (1945). "Sequential Tests of Statistical Hypotheses." Annals of Mathematical Statistics, 16(2), 117-186. https://psycnet.apa.org/record/1945-03252-001
  3. Johari, R., Pekelis, L., and Walsh, D. J. (2015). "Always Valid Inference: Bringing Sequential Analysis to A/B Testing." arXiv:1512.04922. https://arxiv.org/abs/1512.04922
  4. Georgiev, G. Z. (2017). Efficient A/B Testing in Conversion Rate Optimization: The AGILE Statistical Method. https://www.analytics-toolkit.com/pdf/Efficient_AB_Testing_in_Conversion_Rate_Optimization_-_The_AGILE_Statistical_Method_2017.pdf
  5. Kohavi, R., Tang, D., and Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press. https://www.cambridge.org/9781108724265
  6. Kohavi, R., Crook, T., and Longbotham, R. (2009). "Online Experimentation at Microsoft." https://exp-platform.com/Documents/ExP_DMCaseStudies.pdf
  7. Deng, A., Xu, Y., Kohavi, R., and Walker, T. (2013). "Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data." https://www.kdd.org/kdd2016/papers/files/adf0853-dengA.pdf
  8. Vermeer, L., et al. (2017). "Democratizing Online Controlled Experiments at Booking.com." arXiv:1710.08217. https://arxiv.org/abs/1710.08217
  9. Netflix Technology Blog (2019). "Quasi Experimentation at Netflix." https://netflixtechblog.com/quasi-experimentation-at-netflix-566b57d2e362
  10. Convert.com (2024). "A/B Testing & CRO Stats Every Optimizer Should Know." https://www.convert.com/blog/a-b-testing/ab-testing-stats/
  11. Google Analytics Help (January 2023). "Google Optimize sunset." https://support.google.com/analytics/answer/12979939
  12. Datadog press release (May 5, 2025). "Datadog Acquires Eppo." https://www.datadoghq.com/about/latest-news/press-releases/datadog-acquires-eppo/
  13. TechCrunch (January 23, 2025). "Everstone acquires Wingify for $200M." https://techcrunch.com/2025/01/23/everstone-acquires-bootstrapped-indian-startup-wingify-for-200m/