I knew the answer involved multiple comparisons but fumbled the specifics under pressure.
Start by acknowledging the multiple comparisons problem and its impact on Type I error inflation. Then discuss correction methods like Bonferroni or FDR, and emphasize the importance of balancing error control with statistical power. Finally, mention practical considerations such as pre-registration and effect size estimation.
Pro tip: At Google, where experimentation is central, interviewers value candidates who can discuss trade-offs between false positives and false negatives in the context of business impact. Highlighting the use of FDR control for exploratory analyses and Bonferroni for confirmatory tests shows practical wisdom.
Explain that running many t-tests inflates the family-wise error rate (FWER), increasing the chance of false positives. Quantify the risk: with 100 tests at α=0.05, expect ~5 false positives by chance.
Discuss options like Bonferroni (controls FWER, conservative) and Benjamini-Hochberg (controls FDR, more power). Justify the choice based on whether the goal is strict error control or discovery.
Note that corrections reduce power, so ensure adequate sample size per test. Mention that with many tests, even small effects may require large samples to detect after correction.
Talk about pre-registering hypotheses to avoid p-hacking, using hierarchical or Bayesian methods as alternatives, and validating findings with holdout data.
Explain how to present corrected results to stakeholders, emphasizing effect sizes and confidence intervals over p-values, and discussing the balance between false positives and false negatives.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.