← TikTok Interview Insights

TikTok·Data Scientist·Technical Phone Screen·Senior

Senior
May 2026

Summary

TikTok DS interview, one big meaty question about marketplace trust and search ranking. The whole session was basically one extended case study and they went deep on every sub-part, so be ready to stay in one problem for a long time.

Questions Asked (1)

Q1

You want to downrank listings from suspected bad sellers in marketplace search results. Should you launch this? Walk through how you'd design the decision framework and the experiment end-to-end, covering treatment definition, randomization unit, metrics, ramp plan, model uncertainty, and heterogeneity checks.

A/B Testing & ExperimentationProduct Analytics & MetricsTechnical Trade-offs
Author's notes

This was a lot.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by framing the decision as a trade-off between reducing bad seller exposure and potential collateral damage to good sellers and buyer experience. Then walk through a structured experiment design: define treatment precisely, choose randomization unit, select guardrail and success metrics, plan ramp with uncertainty quantification, and pre-specify heterogeneity checks. Emphasize that you would not launch without clear evidence of net positive impact and acceptable risk.

Pro tip: Propose a phased approach: first run a small-scale A/B test with a conservative downranking threshold, then use Bayesian methods to quantify uncertainty and decide whether to expand. This shows you balance speed with rigor and avoid overreacting to noisy early results.

1. Define treatment and randomization unit

Specify exactly how suspected bad sellers are identified (e.g., model score threshold) and what downranking means (e.g., multiply ranking score by 0.5). Choose randomization unit (e.g., seller, listing, or user) based on interference risk and analysis goals.

2. Select metrics and guardrails

Choose primary success metrics (e.g., reduction in bad seller impressions, increase in good seller GMV) and guardrail metrics (e.g., overall search CTR, buyer satisfaction, false positive rate). Ensure metrics are sensitive and aligned with long-term goals.

3. Design experiment and ramp plan

Plan sample size, duration, and ramp stages (e.g., 1% -> 5% -> 20%). Use sequential testing or Bayesian methods to monitor early signals while controlling error rates. Pre-register analysis plan.

4. Quantify model uncertainty and heterogeneity

Account for uncertainty in the bad seller model (e.g., via bootstrap or posterior sampling) and test for heterogeneous treatment effects across seller categories, user segments, and query types. Pre-specify subgroups to avoid p-hacking.

5. Make launch decision and iterate

Evaluate results against pre-defined success criteria, considering both statistical and practical significance. If positive, launch with monitoring; if not, iterate on model or treatment. Document learnings.

Key Points to Mention

  • Randomization unit trade-offs: seller-level avoids contamination but may have low power; user-level enables user experience metrics but risks spillover.
  • Use of guardrail metrics to detect unintended harm to good sellers or overall search quality.
  • Bayesian A/B testing or sequential testing to handle peeking and quantify probability of improvement.
  • Pre-registration of heterogeneity checks to maintain statistical validity.
  • Consideration of false positives/negatives in the bad seller model and their impact on treatment effect.
  • Ramp plan with kill switches and clear rollback criteria.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.