← Yahoo Interview Insights

Yahoo·Data Scientist·Technical Phone Screen·Senior

Senior
May 2026

Summary

Yahoo DS interview with a meaty A/B testing design question around a new recommendation widget. The whole round was basically one big scenario unpacked into several sub-questions, which felt more like a product analytics case than a typical stats quiz.

Questions Asked (1)

Q1

Design an A/B experiment to measure the impact of a new in-app recommendation widget on daily active users and session length. Walk through your choice of primary and guardrail metrics, how you'd size the experiment, and what biases or pitfalls you'd watch out for.

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

This is a lot of question crammed into one prompt and I didn't pace myself well.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clarifying the goal and defining success metrics, then outline the experiment design including randomization, sample size, and duration. Emphasize guardrail metrics and potential biases, and conclude with how you'd analyze and interpret results.

Pro tip: Always tie metrics to business impact and consider novelty effects; run a holdback to measure long-term impact.

1. Define Hypothesis and Metrics

State a clear hypothesis and select primary metric (e.g., DAU) and guardrail metrics (e.g., session length, engagement quality).

2. Design Experiment

Choose randomization unit (user-level), control/treatment groups, and determine sample size using power analysis.

3. Address Pitfalls

Identify potential biases like novelty effect, selection bias, and network effects; plan mitigation strategies.

4. Analyze and Interpret

Use appropriate statistical tests, check for significance, and consider practical significance and business impact.

Key Points to Mention

  • Primary metric: Daily Active Users (DAU); Guardrail metrics: session length, retention, revenue.
  • Sample size calculation: power analysis with expected effect size, significance level, and power.
  • Randomization: user-level randomization to avoid contamination.
  • Novelty effect: run experiment long enough to see steady-state behavior.
  • Network effects: consider if widget affects interactions between users.
  • Statistical tests: t-test or Mann-Whitney U test for non-normal data; consider sequential testing.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.