The 36% number is a gift because it maps directly to variance reduction.
Start by explaining the CUPED variance reduction formula: with a pre-period covariate explaining ρ² of the variance, the variance of the adjusted estimator is reduced by a factor of (1 - ρ²). Then quantify the sample size savings as the inverse of this factor, i.e., you need (1 - ρ²) times the original sample size, which corresponds to a 1/(1 - ρ²) - 1 reduction in required sample size. Finally, discuss practical considerations like covariate choice, correlation estimation, and potential pitfalls.
Pro tip: Mention that the variance reduction factor (1 - ρ²) assumes the covariate is perfectly measured and the relationship is linear; in practice, you might achieve slightly less reduction due to estimation error in the adjustment coefficient. Also, note that CUPED can be combined with other variance reduction techniques like stratification or regression adjustment for even greater gains.
Describe how CUPED uses a pre-period covariate (e.g., pre-experiment metric) to adjust the outcome metric, typically by subtracting a scaled version of the covariate from the outcome. The scaling factor is chosen to minimize variance, often the covariance between covariate and outcome divided by the variance of the covariate.
Show that the variance of the adjusted estimator is reduced by a factor of (1 - ρ²), where ρ is the correlation between the covariate and the outcome. Since ρ² = 0.36, the variance is reduced by 36%, i.e., the new variance is 64% of the original.
Sample size is proportional to variance, so the required sample size with CUPED is 64% of the original. This means a 36% reduction in sample size, or equivalently, you need 1/(1 - ρ²) = 1/0.64 ≈ 1.5625 times the original sample size without CUPED? Wait, careful: If variance is reduced by 36%, then to achieve the same power, you need 64% of the original sample size. So savings = 36%.
Mention that this 36% reduction in sample size can lead to faster experiments, lower costs, or increased power. Also note that the actual savings might be slightly less due to imperfect correlation estimation or if the covariate is not perfectly pre-period.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
I listed the obvious ones: covariate measured pre-treatment, no treatment contamination in the covariate, stable relationship between X and Y across conditions.
Start by explaining the core assumptions of CUPED: that the pre-experiment covariate is unaffected by treatment and is linearly related to the outcome, and that the variance reduction is correctly estimated. Then discuss conditions where these assumptions break down, such as non-linear relationships, covariate imbalance, or when the covariate is affected by treatment. Finally, outline diagnostics like checking covariate balance, residual plots, and sensitivity analyses to validate the assumptions.
Pro tip: Emphasize that CUPED is not a magic bullet—its effectiveness hinges on the quality and relevance of the pre-experiment covariate. Always validate with a placebo test: run CUPED on a pre-experiment period to ensure it doesn't introduce bias.
Clearly list the key assumptions: (1) the covariate is measured before treatment and is unaffected by it, (2) the relationship between covariate and outcome is linear and stable across treatment groups, and (3) the variance reduction factor is correctly estimated from historical data.
Discuss scenarios where CUPED fails: non-linear relationships, covariate affected by treatment (e.g., if treatment influences the pre-period metric), heterogeneous treatment effects on the covariate-outcome relationship, and small sample sizes leading to unstable estimates.
Suggest diagnostics such as: checking covariate balance between groups, plotting residuals vs. covariate, comparing variance reduction across groups, and running placebo tests on pre-experiment data to detect bias.
If assumptions are violated, consider alternatives: using non-linear CUPED (e.g., with splines), stratifying by covariate, or using other variance reduction techniques like stratification or regression adjustment.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.
Said covariate adjustment and they pushed back a bit.
Start by clarifying the experimental context and assumptions, then compare stratified randomization and post-hoc covariate adjustment on bias reduction, variance, and operational feasibility. Conclude with a nuanced recommendation that often favors stratified randomization when key covariates are known pre-treatment, but acknowledges post-hoc adjustment as a fallback for unexpected imbalances.
Pro tip: Emphasize that stratification must happen before randomization and requires pre-registered covariates; post-hoc adjustment can introduce bias if not pre-specified, so always discuss the trade-off between design-time control and analysis-time flexibility.
Ask about the number of covariates, sample size, and whether covariates are known before treatment assignment. This determines the feasibility and necessity of each method.
Explain that stratified randomization balances covariates by design, reducing confounding and increasing precision, while post-hoc adjustment can correct imbalances but may increase variance and risk overfitting.
Discuss implementation complexity: stratification requires pre-registration and can be cumbersome with many covariates, whereas post-hoc adjustment is flexible but must be pre-specified to avoid p-hacking.
State a preference based on typical scenarios (e.g., prefer stratification for few key covariates) and note that post-hoc adjustment is a valid backup when stratification is impractical.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.