← Upstart Interview Insights

Upstart·Data Scientist·Technical Phone Screen·Senior

Senior
May 2026

Summary

A/B testing deep dive for a Data Scientist role at Upstart. The question was essentially a full experiment design case built around a Pinterest-style video feed change, and it covered a lot of ground in one shot.

Questions Asked (1)

Q1

You're a data scientist at a platform like Pinterest. The team wants to increase the share of video content in the home feed to drive more engagement. Design the A/B test end-to-end, including hypotheses, experiment unit, metrics, segmentation, rollout plan, and stopping rules.

A/B Testing & ExperimentationProduct Analytics & MetricsProduct Sense & Ideation
Author's notes

This question is basically five questions dressed up as one.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by framing the business goal and translating it into a testable hypothesis about video content's impact on engagement. Then walk through the experiment design choices (unit, metrics, segmentation, rollout, stopping rules) while balancing statistical rigor with product constraints. Emphasize how you'd measure both short-term engagement and long-term ecosystem health.

Pro tip: Proactively address potential novelty effects and cannibalization of other content types, and propose guardrail metrics to ensure the change doesn't harm user retention or satisfaction.

1. Define Hypothesis and Success Criteria

State a clear, testable hypothesis (e.g., increasing video share in home feed will increase overall engagement) and define primary and secondary success metrics, including guardrails.

2. Choose Experiment Unit and Randomization

Decide on the randomization unit (e.g., user-level) and ensure it aligns with the metric and avoids interference; discuss why user-level is appropriate for feed changes.

3. Select Metrics and Segmentation

Identify primary metrics (e.g., time spent, saves, shares), secondary metrics (e.g., video views), and guardrails (e.g., hide/report rates). Plan pre-registered segmentation (e.g., new vs. existing users, video affinity).

4. Design Rollout and Stopping Rules

Outline a phased rollout (e.g., 1% -> 5% -> 50%) to catch bugs early, and specify stopping rules: fixed horizon or sequential testing, with early stopping for harm or strong positive effect.

5. Plan Analysis and Decision-Making

Describe how you'll analyze results (e.g., t-test, CUPED variance reduction), check for novelty effects, and make a ship/no-ship decision based on statistical significance and practical significance.

Key Points to Mention

  • Randomization unit: user-level to avoid contamination and ensure consistent experience.
  • Primary metric: overall engagement (e.g., daily active users, time spent) and video-specific metrics (e.g., video views, completion rate).
  • Guardrail metrics: user retention, hide/report rates, and content diversity to prevent negative side effects.
  • Segmentation: analyze by user tenure, video affinity, and device type to understand heterogeneous treatment effects.
  • Stopping rules: use sequential testing or fixed sample size with alpha spending to control false positives; monitor for SRM.
  • Novelty effect: run test for sufficient duration (e.g., 2+ weeks) and analyze trend over time to distinguish novelty from sustained impact.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.