← Meta Interview Insights

Meta·Data Scientist·Technical Phone Screen·Senior

Senior
Jun 2026

Summary

Meta DS interview focused entirely on the metrics and experimentation side of launching a short-video recommender feed. Four connected sub-questions that built on each other, which I didn't fully anticipate going in.

Questions Asked (4)

Q1

If you had to pick a single metric to evaluate whether Instagram's short-video recommendation system is working, what would it be and why?

Product Analytics & MetricsProduct Sense & Ideation
Author's notes

I went with watch time pretty quickly and then had to defend it against retention.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clarifying the goal of the recommendation system—likely maximizing long-term user engagement and satisfaction—then propose a single North Star metric that balances user value and business value, such as 'time spent per user per day' or 'retention rate'. Justify why this metric is superior to alternatives by linking it to the system's objective and acknowledging potential trade-offs.

Pro tip: Choose a metric that is both sensitive to changes in the recommendation algorithm and aligned with long-term user retention, not just short-term clicks. Mention that you would pair it with guardrail metrics to prevent optimizing for the wrong behavior.

1. Clarify the objective

Confirm that the recommendation system aims to maximize long-term user engagement and satisfaction, not just immediate clicks or views.

2. Propose a single metric

Select a metric like 'daily time spent per user' or '7-day retention rate' that captures sustained engagement and reflects the system's success.

3. Justify the choice

Explain why this metric is better than alternatives (e.g., CTR, likes) by linking it to user value and business goals, and noting its sensitivity to algorithm changes.

4. Acknowledge trade-offs and guardrails

Discuss potential downsides (e.g., optimizing for time spent might reduce content diversity) and suggest guardrail metrics like user satisfaction or report rate.

5. Conclude with measurement plan

Briefly outline how you would measure the metric (e.g., A/B testing) and iterate, showing a data-driven mindset.

Key Points to Mention

  • North Star metric definition and its importance for aligning teams
  • Difference between short-term engagement (clicks) and long-term retention
  • Potential pitfalls of optimizing for a single metric (e.g., clickbait, filter bubbles)
  • Use of guardrail metrics to ensure user well-being and content diversity
  • How the metric ties to business value (e.g., ad revenue, user growth)
  • Sensitivity of the metric to algorithmic changes and ability to measure via experiments

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q2

Sketch the expected distribution of your chosen metric. Label the median, mode, and 95th percentile.

Product Analytics & Metrics
Author's notes

This tripped me up more than it should have.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by selecting a metric relevant to Meta's products, such as daily active users (DAU) or time spent per user, and explicitly state any assumptions about its distribution (e.g., right-skewed). Sketch the distribution on a whiteboard or paper, clearly labeling the median, mode, and 95th percentile, and explain how their relative positions reflect the distribution's shape and what that implies for product decisions.

Pro tip: Always connect the metric's distribution to actionable product insights—for example, a long right tail might indicate a small group of power users driving engagement, which could inform feature development or targeting strategies.

1. Choose a relevant metric

Select a metric that is meaningful for Meta's products, such as daily active users, time spent per user, or number of messages sent. Briefly justify why this metric matters for the business.

2. State distribution assumptions

Describe the expected shape of the distribution (e.g., right-skewed, normal, bimodal) based on domain knowledge or typical user behavior. Mention any factors that could influence the shape.

3. Sketch the distribution

Draw a rough curve on a whiteboard or paper, labeling the x-axis with the metric and the y-axis with frequency or density. Ensure the curve reflects the stated assumptions.

4. Label median, mode, and 95th percentile

Mark the positions of the median, mode, and 95th percentile on the sketch. Explain their relative positions (e.g., mode < median < 95th percentile for right-skewed data) and what they indicate about the data.

5. Interpret implications

Discuss what the distribution and the labeled points imply for product decisions, such as identifying power users, setting performance targets, or designing experiments.

Key Points to Mention

  • The shape of the distribution (e.g., right-skewed, normal, bimodal) and why it is expected for the chosen metric.
  • The relative positions of the median, mode, and 95th percentile and how they relate to skewness.
  • The difference between mean and median in skewed distributions and why median is often more robust.
  • The significance of the 95th percentile for understanding tail behavior and extreme values.
  • How the distribution informs product decisions, such as feature prioritization or user segmentation.
  • Any assumptions made and how they might affect the interpretation.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q3

Suppose your primary metric goes up but a secondary metric drops at the same time. How do you handle that tradeoff?

A/B Testing & ExperimentationProduct Analytics & MetricsCross-functional Alignment
Author's notes

Pretty natural territory for me.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clarifying the metrics and their relationship, then assess the statistical significance and practical impact of both changes. Evaluate the tradeoff using a holistic framework like the OEC or guardrail metrics, and recommend a decision based on the product's strategic priorities and long-term user value.

Pro tip: Show that you understand the difference between a true tradeoff and a temporary dip, and always consider the counterfactual—what would have happened without the change. Demonstrating that you can align stakeholders on the decision criteria is as important as the analysis itself.

1. Clarify metrics and context

Define the primary and secondary metrics, their expected relationship, and the experiment's goal. Confirm if the secondary metric is a guardrail or a driver of long-term value.

2. Validate statistical significance

Check if the changes are statistically significant and not due to noise. Consider confidence intervals, p-values, and practical significance.

3. Assess tradeoff magnitude and direction

Quantify the size of the increase and decrease, and determine if the secondary metric drop is acceptable given the primary gain. Consider segment-level analysis to see if the tradeoff varies across users.

4. Evaluate long-term and strategic impact

Use a holistic metric like the Overall Evaluation Criterion (OEC) or consider long-term proxies. Align with product strategy: does the primary metric align with long-term goals, and is the secondary metric a leading indicator of churn or dissatisfaction?

5. Recommend action and communicate

Decide whether to ship, iterate, or kill the change based on the tradeoff analysis. Communicate the decision and rationale clearly to stakeholders, and propose follow-up experiments if needed.

Key Points to Mention

  • Overall Evaluation Criterion (OEC) and guardrail metrics
  • Statistical significance vs. practical significance
  • Segment-level analysis to uncover heterogeneous treatment effects
  • Long-term impact and leading vs. lagging indicators
  • Stakeholder alignment on decision criteria and risk tolerance
  • Counterfactual analysis and opportunity cost

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q4

Walk me through the full A/B testing process you'd use to launch this recommender feed, from setup to decision.

A/B Testing & ExperimentationProduct Strategy
Author's notes

Covered randomization unit, power calculation, runtime, guardrails, and ship criteria.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Structure your answer around the end-to-end experimentation lifecycle, from defining a clear hypothesis and success metrics to designing the test, analyzing results, and making a ship/no-ship decision. Emphasize statistical rigor, guardrail metrics, and how you'd handle practical challenges like network effects or novelty effects. Tailor your response to Meta's scale and culture by mentioning rapid iteration and cross-functional collaboration.

Pro tip: Show that you think beyond statistical significance by discussing practical significance and the business impact of the decision. Also, mention how you'd handle multiple testing corrections and segment-level analyses to uncover heterogeneous treatment effects.

1. Define Hypothesis and Metrics

Articulate a clear, testable hypothesis about how the new recommender feed will improve user engagement. Define primary success metrics (e.g., CTR, time spent) and guardrail metrics (e.g., user satisfaction, retention) to ensure no harm.

2. Design the Experiment

Choose randomization unit (e.g., user-level), determine sample size and power, set experiment duration, and decide on control/treatment variants. Consider stratification and whether to run a holdback or switchback design if network effects are a concern.

3. Execute and Monitor

Launch the experiment, monitor data quality, and check for sample ratio mismatch (SRM). Track guardrail metrics in real-time to catch any negative impact early and pause if necessary.

4. Analyze Results

Perform statistical tests (e.g., t-test, bootstrap) to measure treatment effects, adjust for multiple comparisons, and conduct subgroup analyses. Assess both statistical and practical significance.

5. Make Decision and Iterate

Based on results, decide to ship, iterate, or abandon. Document learnings, and if shipping, plan for a gradual rollout with continued monitoring. If iterating, refine hypothesis and rerun.

Key Points to Mention

  • Clear hypothesis and success metrics aligned with business goals
  • Randomization unit and sample size calculation
  • Guardrail metrics to detect unintended harm
  • Handling of network effects and interference
  • Statistical methods for analysis (e.g., sequential testing, CUPED)
  • Practical significance and business impact of the decision

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.