This is a five-part question dressed up as one question, which I did not fully appreciate until I was three minutes into talking about Louvain clustering and realized I hadn't even touched metrics yet.
Acknowledge that standard user-level randomization fails due to interference, then propose cluster-based randomization (e.g., by social community or graph partition) to contain spillover. Walk through the design choices, metrics, power analysis, and analysis methods, and discuss mitigation strategies for contamination.
Pro tip: Emphasize that the choice of cluster definition should balance interference reduction with statistical power, and consider using a switchback or ego-cluster design if graph clustering is too complex. Also, pre-register the analysis plan to avoid p-hacking.
Recognize that interactions create spillover, so randomize at a level that contains interference, such as social communities, graph partitions, or ego-networks. Justify the choice based on the social graph structure and expected effect size.
Select primary metrics (e.g., engagement, retention) and guardrail metrics. Compute power using cluster-level variance, accounting for intra-cluster correlation (ICC) and design effect. Determine the number of clusters needed.
Implement cluster randomization, ensuring balanced clusters via stratification or matching. Monitor contamination (e.g., cross-cluster interactions) and consider techniques like graph cuts or temporal separation to minimize it.
Use cluster-robust standard errors, mixed-effects models, or synthetic control to estimate causal effects. If contamination is present, consider instrumental variables or exposure-based analysis.
If contamination exceeds threshold, pause the test, re-randomize, or switch to a switchback design. Alternatively, use a holdout group or adjust analysis to account for spillover.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.