Hillstrom E-Mail Campaign A/B Test

Randomized email A/B test — 64,000 customers (Intention-to-Treat)

Methodology & Statistical Guardrails

Intention-to-Treat (ITT): Every one of the 64,000 randomized customers is retained in their assigned arm regardless of open, visit, or purchase status. No conditioning on post-treatment outcomes.

Holm-Bonferroni Correction: Because testing multiple outcomes across multiple subgroups inflates false discovery rates, tests are evaluated in pre-specified families with step-down Holm correction to control Family-Wise Error Rate (α = 0.05).

Confirmed vs. Exploratory: Results marked Confirmed survived Holm multiple comparison adjustment. Results marked Exploratory — did not survive correction did not maintain significance under correction and represent hypotheses, not verified decisions.

Sample Population
64,000
3 Randomized Arms (~21.3k each)
Visit Lift (Mens vs Ctrl) Confirmed
+7.66 pp
+72.1% relative lift (p < 0.0001)
Conv. Lift (Mens vs Ctrl) Confirmed
+0.68 pp
+118.8% relative lift (p < 0.0001)
Spend Lift (Mens vs Ctrl) Confirmed
+$0.77
+117.9% relative lift (p < 0.0001)

Executive Strategy & Recommendation

Send Mens E-Mail to everyone.

The Mens creative produces massive, confirmed visit lifts on Men (+5.8 pp) and dual-category buyers (+6.3 pp; +$1.66 spend lift is exploratory), while performing statistically identically to the Womens creative on Womens-only buyers (observed spend gap is 1.6 cents, p = 0.94). Micro-targeting by purchase history adds routing complexity for an undetectable +$2.86 / 1,000 lift whose 95% bootstrap confidence interval contains zero.

Default Rule: Unconditional Mens Campaign Broadcast
✓

Randomization Balance Succeeded

Omnibus Multinomial Logistic LRT: χ²(30) = 27.20, p = 0.613 (Fail to reject balance null).

Max Absolute SMD
0.0142
Standard Threshold
≤ 0.1000

Visit Rate

% of arm visiting site in 2 weeks

Confirmed
Mens vs Ctrl: +7.66 pp • Womens vs Ctrl: +4.52 pp

Conversion Rate

% of arm purchasing in 2 weeks

Confirmed
Mens vs Ctrl: +0.68 pp • Womens vs Ctrl: +0.31 pp

Mean Spend

Revenue per customer randomized ($)

Confirmed
Mens vs Ctrl: +$0.77 • Womens vs Ctrl: +$0.42
Commercial Impact per 1,000 Emails Dispatched
+77 Visits
Incremental site visitors driven by Mens E-Mail (+45 for Womens E-Mail)
+7 Orders
Incremental purchasing orders generated by Mens E-Mail (+3 for Womens E-Mail)
+$770 Gross Revenue
Incremental top-line customer spend from Mens E-Mail (+$424 for Womens E-Mail)

Treatment Effect Heterogeneity & Subgroups

Evaluation of treatment interactions across pre-specified customer dimensions (HC3 robust standard errors)

Campaign Unit Economics & Policy Decision Simulator

Adjust gross margin and dispatch cost to evaluate break-evens and compare Policy (c) vs. Policy (e).

Illustrative assumptions — not in source data
Gross Margin 40%
10% Default: 40% 100%
Cost per Email Dispatched $0.05
$0.01 Default: $0.05 $0.40
Policy (c): Mens E-Mail to Everyone
$257.93 / 1k
Incremental profit per 1,000 customers
Policy (e): Fixed Womens-only Targeting
$260.79 / 1k
Lift over Mens All: +$2.86 / 1k
Mens Break-Even Cost: $0.308 / email (Conservative 95% low: $0.195)
Womens Break-Even Cost: $0.170 / email (Conservative 95% low: $0.068)
Statistical Reality: The difference between Policy (e) and Policy (c) is +$2.86 per 1,000 customers with a 95% bootstrap confidence interval of [-$71.18, +$76.05]. This difference is not statistically distinguishable from zero at any gross margin ($0 is well within the interval).

Statistical Power & MDE

MDE Sensitivity

Why can't we say Womens E-Mail performs better or worse for Womens-only buyers?

Spend MDE Threshold $0.37 / cust
The minimum spend difference this test had 80% power to detect in Womens-only buyers (74.5% of control spend).
Observed Difference $0.016 / cust
The actual observed sample gap between Mens and Womens email is only 1.6 cents (p = 0.94).

Plain-Language Verdict: Because the 1.6¢ gap is far below the $0.37 MDE, this test is underpowered to detect subtle creative nuances on rare spend events. The result is genuinely inconclusive rather than proof of exactly zero true difference.

Two-sided α = 0.05, 80% Power, N = 19,215

Methodological Limitations

Crucial caveats and context for interpreting these experimental findings.

Always Visible