← Meta Interview Insights

Meta·Data Scientist·Technical Phone Screen·Senior

Senior
Jul 2026

Summary

Meta DS interview with a deep-dive experiment design question on ad insertion in the main feed. Single question but it covered basically every angle of A/B testing you can imagine, from randomization unit to long-term holdback to network spillovers. Dense.

Questions Asked (1)

Q1

You're planning to insert one extra ad every 8 organic posts in the main feed. Walk through the full experiment design: randomization unit and rationale, primary and guardrail metrics, power analysis for a +2% revenue MDE, SRM checks and novelty effect detection, ramp strategy with geographic and age holdouts, a long-term holdback for delayed churn, and how you'd handle network spillovers and bias from dynamic ranking reacting to treatment.

A/B Testing & ExperimentationProduct Analytics & MetricsPricing & Monetization
Author's notes

This one sprawled in a way I wasn't fully ready for.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Structure your answer around the experiment lifecycle: design (randomization, metrics, power), execution (SRM, novelty, ramp), and long-term validation (holdback, spillovers). Emphasize how you'd mitigate biases from dynamic ranking and network effects, showing awareness of Meta's scale and complexity.

Pro tip: Propose using a cluster-randomized design (e.g., by user clusters or geographic regions) to handle network spillovers, and suggest a switchback or holdback to measure long-term effects. Also, mention that you'd pre-register the analysis plan to avoid p-hacking.

1. Randomization Unit and Rationale

Choose the randomization unit (e.g., user, session, or cluster) based on interference risk. For feed ads, user-level randomization is typical, but if network effects are strong, consider cluster randomization (e.g., by social graph clusters) to reduce spillovers.

2. Metrics and Power Analysis

Define primary metric (revenue per user) and guardrail metrics (user engagement, satisfaction, churn). Conduct power analysis for +2% revenue MDE, accounting for variance, traffic, and desired power (80-90%).

3. Validity Checks: SRM, Novelty, and Ramp

Implement SRM checks (chi-squared test) to ensure balanced assignment. Monitor novelty effects via early vs. late period comparisons. Use a ramp strategy with geographic and age holdouts to detect heterogeneous effects and limit risk.

4. Long-Term Holdback and Spillover Mitigation

Maintain a long-term holdback group to measure delayed churn and revenue impact. Address network spillovers by using cluster randomization or measuring spillover via social connections, and adjust for dynamic ranking bias by holding ranking constant or using counterfactual logging.

Key Points to Mention

  • Randomization unit: user-level vs. cluster-level, considering interference and network effects.
  • Primary metric: revenue per user; guardrails: engagement, churn, user satisfaction.
  • Power analysis: sample size calculation for +2% MDE, considering variance and multiple testing corrections.
  • SRM checks: chi-squared test, sequential monitoring; novelty effect: compare early vs. late periods.
  • Ramp strategy: gradual rollout with geographic and age holdouts to detect heterogeneous treatment effects.
  • Long-term holdback: measure delayed churn and LTV; spillover mitigation: cluster randomization, switchback, or social network analysis.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.