I started with click-through rate as the primary metric which felt right, then conversion as secondary.
Start by defining the primary metric as the one that directly measures the carousel's impact on the core product goal, such as user engagement or watch time. Then outline secondary metrics that capture potential trade-offs, guardrails, and long-term effects. Finally, explain a decision framework that balances statistical significance, practical significance, and business impact.
Pro tip: Emphasize that at TikTok, the primary metric should align with the 'For You' experience—like total watch time or session depth—and always include guardrail metrics to catch negative side effects, such as reduced content diversity or user fatigue.
Choose a metric that directly reflects the carousel's intended impact on the core user experience, such as average watch time per user or daily active users.
Select metrics that capture trade-offs (e.g., click-through rate on carousel items) and guardrails (e.g., bounce rate, content diversity, user reports) to ensure no harm.
Ensure proper randomization, sample size, and duration. Analyze both statistical significance (p-value, confidence intervals) and practical significance (effect size, business impact).
Look beyond the experiment window for novelty effects and segment by user cohorts (e.g., new vs. existing users) to see if the carousel benefits some groups more.
Ship if the primary metric improves significantly without guardrail regressions, and if the lift justifies engineering and maintenance costs. Otherwise, iterate or abandon.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.