Learn from real experiments
Get help designing a test, browse what other teams tried, and learn how good experiments work, in plain language.
Free Experiment Design
Tell us what you want to learn. We’ll suggest a simple experiment design, free.
Real Experiment Library
Three famous tests that changed how teams think about speed, personalization, and trust. Download the full free library below.
Should every user see the same thumbnail?
Do better photos book more stays?
Download everything (free)
Get the full library as a spreadsheet-friendly CSV. Free and open for teaching, research, or product work.
- 69 real experiments with goals, designs, outcomes, and notes.
- Works in Excel, Google Sheets, Python, or R.
- Also available: 18 standout Upworthy headline tests.
Questions? Email jared@productscience.consulting
Experiment Methodology
Why experiments beat gut feel, what large studies found, and where to go deeper.
Why run an experiment?
An experiment answers a simple question: did this change cause that result? You try one version for some people, another for others, then compare. That works for a website, a classroom policy, or a change to your own habits.
Without random assignment, groups often differ in hidden ways. The people who already signed up, adopted the policy, or started the new habit may have been more motivated to begin with. Those hidden differences are called confounders, and they make it hard to know what actually caused the result.
A famous example: observational studies suggested hormone therapy protected women’s hearts. Later, a large randomized trial found the opposite story once healthier patients were no longer self-selecting into treatment. Randomization is what makes the comparison fair. For health evidence ranked by trustworthiness, see Health Evidence.
Random assignment vs letting people choose
Each shape is a person. Flip between modes to see how unfair groups get when people choose for themselves.
What 27,000 headline tests teach us
Upworthy tested nearly every headline they published. Same story, different wording, then they kept the version people clicked most. The public Upworthy Research Archive lets anyone see what actually worked.
Six things the data says
The headline decides the outcome
Within a single test, the best headline beat the worst by a median 1.7× in click-through, and by 2× or more in over a third of tests. Same story, same audience; only the words changed.
Experts can't call the winner
Lining up the archive's measured effects across 50+ language features against what marketing professionals, content writers, and the academic literature predicted, agreement with the real direction was barely better than a coin flip.
The "curiosity gap" is oversold
Curiosity is the single most-recommended tactic in industry guidance, yet headlines coded as curious showed no significant lift, and asking a question in the headline actually lowered clicks.
Visceral emotion travels
Emotional intensity and negative emotions (anger, fear, even disgust) were among the features that reliably raised click-through. Calm, positive-tone framings tended to underperform.
Specific beats clever
Everyday words, real numbers, and named, visible people lifted clicks. Abstract appeals to authority, goals, and comparison framing pushed them down.
Small effects, big samples
At a ~1.5% average click rate, telling a real winner from noise takes tens of thousands of impressions per variant, which is exactly why Upworthy tested everything before publishing.
Two real tests from the archive
Built from the Upworthy Research Archive (Matias, J.N., Munger, K., Le Quere, M.A., & Ebersole, C., 2021), released for research use. The aggregate statistics and the comparison of measured effects to expert, practitioner, and literature predictions are computed from the archive's confirmatory dataset and its companion language analysis; the full archive catalogs 32,487 experiments.
Why big A/B “wins” often shrink
Online, you often see claims like “this button change lifted clicks 50%.” Many of those tests were too small to trust. When researchers re-ran popular changes with millions of users, the big wins usually shrank or disappeared. That pattern is called the winner’s curse.
Four patterns, high-powered checks
Rounded vs. square buttons
An academic study claimed a 55% click-through lift from rounding button corners. High-powered replications found effects roughly two orders of magnitude smaller, often not distinguishable from zero.
Page performance
Performance still matters, but extreme conversion claims from small case studies look overstated. Even large tests often need surrogate metrics because revenue and purchase conversion stay underpowered.
Coupon-code field
Hiding or de-emphasizing the coupon field is a popular “win.” Coop replications at checkout found tiny lifts that were not statistically significant at practical sample sizes.
Sticky call-to-action
Sticky CTAs are widely credited with multi-percent conversion gains. Large Coop and Talabat tests showed mixed, mostly small effects, and site-to-site disagreement even for the same treatment.
Lessons for reading experiment libraries
Sample size dominates
Detecting a 2% relative change often needs millions of users. At 20% power, a significant positive estimate can exaggerate the true effect by a factor of ~2.3 on average.
Business metrics are hard
Even at multi-million scale, revenue per user and purchase conversion often stay underpowered. Teams end up using CTR, add-to-cart, or capped counts: useful, but not the decision metric.
Repositories select winners
Public case banks mix false positives, thresholding bias, and publication bias. Treat large claimed lifts from small tests as hypotheses to re-test, not as settled effects.
Summary of Trustworthy A/B Patterns and the Winner’s Curse: Lessons from Eight Large-Scale Replications (Kohavi, Linowski, Vermeer, Andreev, Dodin, & Furuseth, KDD 2026). Author’s version is shared for personal use via bit.ly/trustworthyABPatternsSummary; the definitive version is the ACM KDD publication. Companion rounded-button analysis: Econ Journal Watch, 2026.
Academic research on experiments
A searchable list of academic papers on A/B testing and online experiments (2000-2025). Use it when you want the deeper research behind the case studies above.
If you use this dataset, please cite:
Return-Aware Platform Experimentation: New Directions for Research, 2025.
Authors: Jacqueline Doremus, Joel Persson, Brian St. Thomas, Carlos A. Flores, Sebastian Ankargren, Mårten Schultzberg, Kyle Kretschman.
Dataset available at: https://github.com/jopersson/online-experimentation-literature-review/
| Title | Year | Venue | Field | Micro | Macro | Database |
|---|
Frequently asked questions
Who is this for?
Anyone planning a test: product teams, teachers, researchers, or curious individuals. Use the library for ideas, then adapt a design to your own question.
Where do these experiments come from?
The library curates experiments from public case studies, talks, and research reports, then standardizes them into a consistent schema with added methodological notes.
Can I use this in teaching or training?
Yes. The commentary is designed to be used as case material for internal training, workshops, or university courses. You can slice the dataset by sector, goal, or method.
Will you add my experiments?
If you run a substantial experiment program and want anonymized entries included, reach out. Contributions can be credited and included in future dataset releases.