This was one giant prompt broken into five sub-parts and I did not pace myself well.
Start by defining a clear OEC that balances user experience and business goals, then select guardrail metrics to monitor potential harms. Justify the test design based on interference and seasonality, outline a phased ramp with pre-registration and stopping rules, and prepare a quasi-experimental fallback. Finally, debug the mid-test scenario by decomposing the metric movements and checking for interference or novelty effects.
Pro tip: Always pre-register your analysis plan and stopping rules to avoid p-hacking and ensure valid inference. When debugging, segment by user cohorts and time to isolate interference or seasonality effects.
Choose an OEC that captures the recommendation module's impact on long-term user value, and select guardrail metrics to detect negative side effects. Provide formulas for each.
Evaluate user-level RCT, geo-cluster, or switchback based on interference and seasonality. Justify the choice with trade-offs.
Outline a phased rollout with pre-registration of hypotheses, metrics, and stopping rules. Include variance reduction techniques.
If randomization isn't feasible, suggest a quasi-experimental design like synthetic control or difference-in-differences, and discuss assumptions.
Analyze why OEC flatlines while add-to-cart rises and conversion drops. Check for interference, seasonality, or metric definition issues, and propose next steps.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.