Randomized email A/B test — 64,000 customers (Intention-to-Treat)
Intention-to-Treat (ITT): Every one of the 64,000 randomized customers is retained in their assigned arm regardless of open, visit, or purchase status. No conditioning on post-treatment outcomes.
Holm-Bonferroni Correction: Because testing multiple outcomes across multiple subgroups inflates false discovery rates, tests are evaluated in pre-specified families with step-down Holm correction to control Family-Wise Error Rate (α = 0.05).
Confirmed vs. Exploratory: Results marked Confirmed survived Holm multiple comparison adjustment. Results marked Exploratory — did not survive correction did not maintain significance under correction and represent hypotheses, not verified decisions.
The Mens creative produces massive, confirmed visit lifts on Men (+5.8 pp) and dual-category buyers (+6.3 pp; +$1.66 spend lift is exploratory), while performing statistically identically to the Womens creative on Womens-only buyers (observed spend gap is 1.6 cents, p = 0.94). Micro-targeting by purchase history adds routing complexity for an undetectable +$2.86 / 1,000 lift whose 95% bootstrap confidence interval contains zero.
Omnibus Multinomial Logistic LRT: χ²(30) = 27.20, p = 0.613 (Fail to reject balance null).
% of arm visiting site in 2 weeks
% of arm purchasing in 2 weeks
Revenue per customer randomized ($)
Evaluation of treatment interactions across pre-specified customer dimensions (HC3 robust standard errors)
Adjust gross margin and dispatch cost to evaluate break-evens and compare Policy (c) vs. Policy (e).
Why can't we say Womens E-Mail performs better or worse for Womens-only buyers?
Plain-Language Verdict: Because the 1.6¢ gap is far below the $0.37 MDE, this test is underpowered to detect subtle creative nuances on rare spend events. The result is genuinely inconclusive rather than proof of exactly zero true difference.
Crucial caveats and context for interpreting these experimental findings.