Learn from real experiments

Get help designing a test, browse what other teams tried, and learn how good experiments work, in plain language.

Free Experiment Design

Tell us what you want to learn. We’ll suggest a simple experiment design, free.

Real Experiment Library

Three famous tests that changed how teams think about speed, personalization, and trust. Download the full free library below.

Latency Bing

Does a tiny delay hurt revenue?

What they tried
Slow down search results by just 100 milliseconds for some users.
What happened
Revenue per user fell about 0.6%. At Bing’s scale, that mapped to roughly $100M+ a year.
Takeaway
Speed is a product feature. Small slowdowns can have large business costs.

A classic result from Microsoft Bing (Kohavi et al., 2008).

Personalization Netflix

Should every user see the same thumbnail?

What they tried
Show personalized artwork for each title instead of one static image for everyone.
What happened
Engagement rose about 20-30% on average.
Takeaway
Personalization is not only about what to recommend. How you present it matters too.

From the Netflix Tech Blog (2016). Personalized artwork became part of the core experience.

Trust & quality Airbnb

Do better photos book more stays?

What they tried
Compare listings with professional photos to listings with host-taken photos.
What happened
Professionally photographed listings booked about 2-3× more often.
Takeaway
Sometimes the winning “experiment” is an operations change, not a new algorithm.

Airbnb engineering (2011). Led to a free professional photography program for hosts.

Download everything (free)

Get the full library as a spreadsheet-friendly CSV. Free and open for teaching, research, or product work.

Download CSV

Questions? Email jared@productscience.consulting

Experiment Methodology

Why experiments beat gut feel, what large studies found, and where to go deeper.

Why run an experiment?

An experiment answers a simple question: did this change cause that result? You try one version for some people, another for others, then compare. That works for a website, a classroom policy, or a change to your own habits.

Without random assignment, groups often differ in hidden ways. The people who already signed up, adopted the policy, or started the new habit may have been more motivated to begin with. Those hidden differences are called confounders, and they make it hard to know what actually caused the result.

A famous example: observational studies suggested hormone therapy protected women’s hearts. Later, a large randomized trial found the opposite story once healthier patients were no longer self-selecting into treatment. Randomization is what makes the comparison fair. For health evidence ranked by trustworthiness, see Health Evidence.

Random assignment vs letting people choose

Each shape is a person. Flip between modes to see how unfair groups get when people choose for themselves.

Young
Older
Male
Female
High motivation
Low motivation

What 27,000 headline tests teach us

Upworthy tested nearly every headline they published. Same story, different wording, then they kept the version people clicked most. The public Upworthy Research Archive lets anyone see what actually worked.

27,616
Randomized A/B tests
128,217
Headline & image variations
457M
Impressions served
~1.5%
Average click-through rate
2013-2015
Testing window

Six things the data says

1.7×

The headline decides the outcome

Within a single test, the best headline beat the worst by a median 1.7× in click-through, and by 2× or more in over a third of tests. Same story, same audience; only the words changed.

50/50

Experts can't call the winner

Lining up the archive's measured effects across 50+ language features against what marketing professionals, content writers, and the academic literature predicted, agreement with the real direction was barely better than a coin flip.

#1

The "curiosity gap" is oversold

Curiosity is the single most-recommended tactic in industry guidance, yet headlines coded as curious showed no significant lift, and asking a question in the headline actually lowered clicks.

Anger

Visceral emotion travels

Emotional intensity and negative emotions (anger, fear, even disgust) were among the features that reliably raised click-through. Calm, positive-tone framings tended to underperform.

Concrete

Specific beats clever

Everyday words, real numbers, and named, visible people lifted clicks. Abstract appeals to authority, goals, and comparison framing pushed them down.

1.5%

Small effects, big samples

At a ~1.5% average click rate, telling a real winner from noise takes tens of thousands of impressions per variant, which is exactly why Upworthy tested everything before publishing.

Two real tests from the archive

Aug 2014 · 61,132 impressions · 3 headlines · 2.5× spread
HIV-Positive People Are Living Longer Than Ever. And There's A Big Problem With That.Top performer
0.88%
You Can Live For Years With HIV. Here's The Problem With That.
0.65%
Living With HIV Is Better Than Dying From It. But For How Long?
0.35%
Sep 2014 · 30,712 impressions · 3 headlines · 2.7× spread
Bullies Who Hide Behind The Screen Are Confronted By Kids Who Aren't Afraid To Show Their FaceTop performer
1.00%
Hey, Online Bullies: Take A Look At Some Kids Who Aren't Afraid To Show Their Faces
0.40%
Hey, Online Bullies: These Kids Have A Message For You. And They Aren't Afraid To Show Their Faces
0.36%

Built from the Upworthy Research Archive (Matias, J.N., Munger, K., Le Quere, M.A., & Ebersole, C., 2021), released for research use. The aggregate statistics and the comparison of measured effects to expert, practitioner, and literature predictions are computed from the archive's confirmatory dataset and its companion language analysis; the full archive catalogs 32,487 experiments.

Why big A/B “wins” often shrink

Online, you often see claims like “this button change lifted clicks 50%.” Many of those tests were too small to trust. When researchers re-ran popular changes with millions of users, the big wins usually shrank or disappeared. That pattern is called the winner’s curse.

8
Large-scale replications
4
UI patterns tested
2.4M
Median users per test
2 / 8
Significant in expected direction
KDD ’26
Peer-reviewed summary

Four patterns, high-powered checks

Rounded

Rounded vs. square buttons

An academic study claimed a 55% click-through lift from rounding button corners. High-powered replications found effects roughly two orders of magnitude smaller, often not distinguishable from zero.

Speed

Page performance

Performance still matters, but extreme conversion claims from small case studies look overstated. Even large tests often need surrogate metrics because revenue and purchase conversion stay underpowered.

Coupon

Coupon-code field

Hiding or de-emphasizing the coupon field is a popular “win.” Coop replications at checkout found tiny lifts that were not statistically significant at practical sample sizes.

Sticky

Sticky call-to-action

Sticky CTAs are widely credited with multi-percent conversion gains. Large Coop and Talabat tests showed mixed, mostly small effects, and site-to-site disagreement even for the same treatment.

Lessons for reading experiment libraries

Power

Sample size dominates

Detecting a 2% relative change often needs millions of users. At 20% power, a significant positive estimate can exaggerate the true effect by a factor of ~2.3 on average.

Surrogates

Business metrics are hard

Even at multi-million scale, revenue per user and purchase conversion often stay underpowered. Teams end up using CTR, add-to-cart, or capped counts: useful, but not the decision metric.

Bias

Repositories select winners

Public case banks mix false positives, thresholding bias, and publication bias. Treat large claimed lifts from small tests as hypotheses to re-test, not as settled effects.

Summary of Trustworthy A/B Patterns and the Winner’s Curse: Lessons from Eight Large-Scale Replications (Kohavi, Linowski, Vermeer, Andreev, Dodin, & Furuseth, KDD 2026). Author’s version is shared for personal use via bit.ly/trustworthyABPatternsSummary; the definitive version is the ACM KDD publication. Companion rounded-button analysis: Econ Journal Watch, 2026.

Academic research on experiments

A searchable list of academic papers on A/B testing and online experiments (2000-2025). Use it when you want the deeper research behind the case studies above.

If you use this dataset, please cite:

Return-Aware Platform Experimentation: New Directions for Research, 2025.
Authors: Jacqueline Doremus, Joel Persson, Brian St. Thomas, Carlos A. Flores, Sebastian Ankargren, Mårten Schultzberg, Kyle Kretschman.
Dataset available at: https://github.com/jopersson/online-experimentation-literature-review/

Download full bibliography (CSV)

Loading…

Title Year Venue Field Micro Macro Database

Frequently asked questions

Who is this for?

Anyone planning a test: product teams, teachers, researchers, or curious individuals. Use the library for ideas, then adapt a design to your own question.

Where do these experiments come from?

The library curates experiments from public case studies, talks, and research reports, then standardizes them into a consistent schema with added methodological notes.

Can I use this in teaching or training?

Yes. The commentary is designed to be used as case material for internal training, workshops, or university courses. You can slice the dataset by sector, goal, or method.

Will you add my experiments?

If you run a substantial experiment program and want anonymized entries included, reach out. Contributions can be credited and included in future dataset releases.