← TikTok Interview Insights

TikTok·Data Scientist·Technical Phone Screen·Senior

Senior
Jun 2026

Summary

TikTok DS interview with a deep-dive experiment design question around changing the default video length. One question, very open-ended, and they clearly wanted you to get into the weeds on instrumentation and rollout, not just wave at 'run an A/B test'.

Questions Asked (1)

Q1

Product wants to change the default video length from 15 seconds to 60 seconds. Walk through how you'd design this experiment end to end: randomization unit and contamination risks across creators and viewers, primary metrics and guardrails, traffic split and ramp schedule, how you'd handle novelty effects, SUTVA checks, and what a post-experiment rollout looks like.

A/B Testing & ExperimentationProduct Analytics & MetricsSystem Design
Author's notes

This one is genuinely hard because TikTok is a two-sided platform and your randomization choice breaks everything downstream.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clarifying the goal and defining the randomization unit, then systematically address contamination, metrics, ramp schedule, novelty, SUTVA, and rollout. Emphasize the two-sided marketplace dynamics at TikTok and propose a multi-layer experimentation approach with careful guardrails.

Pro tip: Propose a creator-level randomization with viewer-level holdout to isolate supply-side effects, and use a switchback or cluster randomization if contamination is severe. Also, plan for a long-term holdout to measure persistent effects beyond novelty.

1. Clarify Goal and Define Hypotheses

Confirm the product goal (e.g., increase watch time, creator satisfaction) and state testable hypotheses about how longer videos affect key metrics.

2. Choose Randomization Unit and Mitigate Contamination

Decide between creator-level, viewer-level, or cluster randomization. Address contamination via network effects, shared content, and algorithmic spillovers; consider a two-sided experiment design.

3. Select Metrics and Guardrails

Define primary metrics (e.g., total watch time, video completion rate) and guardrails (e.g., creator upload frequency, viewer retention, app performance).

4. Design Ramp Schedule and Handle Novelty/SUTVA

Plan traffic split and gradual ramp-up. Monitor for novelty effects with extended pre/post periods and use SUTVA checks (e.g., compare treatment/control overlap, run A/A tests).

5. Analyze Results and Plan Rollout

After experiment, analyze metrics with appropriate statistical methods, decide on rollout based on trade-offs, and consider a phased launch with continued monitoring.

Key Points to Mention

  • Randomization unit: creator-level to capture supply-side effects, with viewer-level analysis; consider cluster randomization by geography or time zones.
  • Contamination risks: creators in treatment may influence control viewers via shared content; viewers may see both 15s and 60s videos if randomized at viewer level, causing spillover.
  • Primary metrics: total watch time per user, video completion rate, creator uploads; guardrails: viewer retention, app crash rate, content diversity.
  • Traffic split and ramp: start with small % (e.g., 1-5%) and gradually increase; use holdout groups for long-term effects.
  • Novelty effects: run experiment for at least 2-4 weeks, compare early vs. late periods, and use a pre-period to establish baseline.
  • SUTVA checks: ensure no interference between units; run A/A tests, check for spillover via network analysis, and consider switchback designs if needed.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.