← Meta Interview Insights

Meta·Data Scientist·Technical Phone Screen·Senior

Senior
Jun 2026

Summary

This was a deep product analytics case for a Data Scientist role at Meta, structured around designing an end-to-end experiment for a Marketplace feature. Seven sub-questions in one prompt, which felt like a lot to juggle in real time.

Questions Asked (7)

Q1

State your mission for the Verified Seller Badges feature precisely, and write primary and secondary hypotheses that are falsifiable.

Product Sense & IdeationProduct Analytics & Metrics
Author's notes

I fumbled the 'falsifiable' part a bit.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining a precise, measurable mission for the Verified Seller Badges feature that ties directly to Meta's marketplace goals, such as increasing trust and reducing fraud. Then formulate a primary hypothesis about the main effect and a secondary hypothesis about a potential trade-off or boundary condition, ensuring both are falsifiable with clear metrics and thresholds.

Pro tip: Frame your hypotheses in terms of specific metrics (e.g., 'increase verified seller transaction share by X%') and explicitly state what data would refute them, showing you understand the scientific method and business impact.

1. Clarify the feature and business context

Briefly restate what Verified Seller Badges are and how they fit into Meta's marketplace ecosystem, focusing on the problem they solve (e.g., trust, fraud).

2. Define a precise mission statement

Craft a one-sentence mission that is specific, measurable, and aligned with Meta's goals, such as 'Increase trust in Marketplace by incentivizing sellers to verify their identity, leading to a 10% reduction in fraud reports.'

3. Formulate a primary hypothesis

State a testable hypothesis about the main intended effect of the feature, including the metric, direction, and magnitude (e.g., 'Verified badges will increase the share of transactions from verified sellers by 15% within 3 months.').

4. Formulate a secondary hypothesis

Propose a secondary hypothesis that addresses a potential trade-off, unintended consequence, or segment-specific effect (e.g., 'The badges will reduce new seller sign-ups by 5% due to verification friction.').

5. Ensure falsifiability

For each hypothesis, specify the data and thresholds that would disprove it, demonstrating rigor and an experimental mindset.

Key Points to Mention

  • Alignment with Meta's mission and marketplace trust goals
  • Specific metrics: fraud reports, transaction share, seller sign-ups, buyer trust scores
  • Falsifiability: clear thresholds and data sources for refutation
  • Primary vs. secondary hypotheses: main effect vs. trade-off or segment effect
  • Consideration of potential unintended consequences (e.g., seller friction, badge forgery)
  • Use of A/B testing or quasi-experimental methods to validate hypotheses

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q2

Define a North Star Metric and 4 to 6 supporting metrics, including at least two counter or guardrail metrics. Explain why each is diagnostic and how to compute it at user and geo level.

Product Analytics & MetricsA/B Testing & Experimentation
Author's notes

Weekly GMV per active buyer felt like the obvious North Star pick and I went with it.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clarifying the product context (e.g., Facebook Feed) and the North Star Metric (NSM) that captures core user value. Then define 4-6 supporting metrics, including at least two guardrail metrics, and explain how each is diagnostic and computable at user and geo levels. Emphasize the balance between growth and user well-being.

Pro tip: Tie the NSM to a long-term user value metric like 'Meaningful Social Interactions' and show how guardrails prevent optimizing for short-term engagement at the expense of user trust.

1. Clarify Product and Goal

State the product area (e.g., Facebook Feed) and the overarching goal (e.g., maximize long-term user value). This sets the context for metric selection.

2. Define North Star Metric

Choose a single NSM that best captures the core value users get from the product, such as 'Daily Active Users' or 'Meaningful Social Interactions'. Explain why it's diagnostic of long-term success.

3. Select Supporting Metrics

Pick 4-6 metrics that drive the NSM, covering engagement, retention, and monetization. Include at least two guardrail metrics (e.g., user-reported negative experiences, time spent) to monitor unintended consequences.

4. Explain Diagnostic Value and Computation

For each metric, describe why it's diagnostic (e.g., leading indicator of retention) and how to compute it at user level (e.g., average per user) and geo level (e.g., aggregated by region, normalized by population).

5. Summarize and Balance

Conclude by emphasizing the balance between growth and guardrails, and how these metrics together provide a holistic view of product health.

Key Points to Mention

  • North Star Metric should reflect core user value and predict long-term retention.
  • Supporting metrics should be actionable and cover different aspects: acquisition, engagement, retention, monetization.
  • Guardrail metrics (e.g., user reports, time spent) ensure that growth doesn't harm user experience.
  • Diagnostic metrics are leading indicators that can be influenced by product changes.
  • User-level computation: average per user or percentage of users performing an action.
  • Geo-level computation: aggregate by region, normalize by population or DAU, and consider regional differences.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q3

Propose a geo-level clustered A/B test for this feature. Specify the cluster unit, stratification variables, matching strategy, number of clusters per arm, traffic ramp plan, duration, and how you'd handle contamination and spillovers.

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

This was the meatiest part.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining the cluster unit and justifying why geo-level randomization is necessary (e.g., network effects, spillover risks). Then walk through a structured design: stratification and matching to create balanced arms, power analysis to determine cluster count, a phased traffic ramp, and a duration that captures novelty and seasonality. Finally, address contamination and spillovers with mitigation strategies like buffer zones and measurement of interference.

Pro tip: Always tie your design choices back to the specific feature and its potential for interference—interviewers want to see you reason about trade-offs, not just recite textbook steps. Mention that you'd pre-register the analysis plan and run a cluster-level power analysis using historical data to avoid underpowered tests.

1. Define cluster unit and justify geo-level randomization

Choose the appropriate cluster unit (e.g., DMA, city, country) based on the feature's reach and network effects. Explain why individual-level randomization is infeasible due to spillovers or contamination.

2. Stratify and match clusters for balance

Identify key stratification variables (e.g., baseline metric, demographics, region) and use matching techniques (e.g., propensity score matching, paired matching) to create comparable arms. Ensure clusters within pairs are similar on observables.

3. Determine number of clusters and power

Conduct a power analysis at the cluster level, accounting for intra-cluster correlation (ICC) and cluster size variation. Calculate the required number of clusters per arm to detect the minimum detectable effect (MDE) with desired power.

4. Plan traffic ramp and duration

Design a phased rollout (e.g., start with 10% of clusters, then scale) to monitor for early issues. Set duration based on expected effect size, novelty effects, and business cycles (e.g., at least 2-4 weeks).

5. Address contamination and spillovers

Implement buffer zones (exclude clusters near boundaries), use geo-fencing to limit exposure, and measure spillover via control clusters' metrics. Consider techniques like difference-in-differences or synthetic control if interference is detected.

Key Points to Mention

  • Cluster unit selection: DMA, city, or country based on feature scope and network effects.
  • Stratification variables: baseline conversion rate, demographics, region, and pre-experiment metrics.
  • Matching strategy: paired matching or propensity score matching to create balanced arms.
  • Power analysis: account for intra-cluster correlation (ICC) and cluster size variation; use simulation or formulas.
  • Traffic ramp: phased rollout with holdout groups and early stopping rules for harm.
  • Contamination mitigation: buffer zones, geo-fencing, and measuring spillover via control clusters' metrics.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q4

How would you estimate the minimum detectable effect for your North Star Metric, including variance assumptions and design effects from clustering?

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

The design effect piece is what separates a real answer from a textbook one, and I knew it.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining the North Star Metric (NSM) and its unit of analysis, then derive the minimum detectable effect (MDE) using the standard formula that incorporates variance, sample size, power, and significance level. Adjust for clustering by estimating the design effect (DEFF) from intra-cluster correlation (ICC) and average cluster size, and finally validate with historical data or simulations.

Pro tip: Always tie the MDE back to business impact: a statistically detectable effect is meaningless if it's too small to matter. Also, be prepared to discuss how you'd handle multiple testing corrections if the NSM is part of a family of metrics.

1. Define the NSM and unit of analysis

Clarify what the North Star Metric is (e.g., daily active users, revenue per user) and the unit of randomization (e.g., user, session, cluster). This determines the variance structure and clustering adjustments needed.

2. Determine baseline variance and sample size

Estimate the baseline mean and variance of the NSM from historical data or a pilot. Decide on the significance level (α) and power (1-β), and compute the required sample size per variant for a given MDE.

3. Incorporate design effect from clustering

If randomization is at a cluster level (e.g., by user or geography), calculate the design effect (DEFF) using the intra-cluster correlation (ICC) and average cluster size: DEFF = 1 + (m-1)*ICC. Adjust the variance by multiplying by DEFF.

4. Compute MDE with adjusted variance

Use the formula: MDE = (Z_{1-α/2} + Z_{1-β}) * sqrt( (σ^2 * DEFF) * (1/n_t + 1/n_c) ), where n_t and n_c are sample sizes per variant. Alternatively, solve for MDE given fixed sample size.

5. Validate and contextualize

Validate assumptions with historical A/B tests or simulations. Translate the MDE into business terms (e.g., revenue impact) and discuss trade-offs (e.g., longer test duration vs. detecting smaller effects).

Key Points to Mention

  • Definition of MDE and its relationship with sample size, variance, power, and significance level.
  • Variance estimation: use historical data, pilot studies, or conservative assumptions (e.g., maximum variance for proportions).
  • Design effect formula: DEFF = 1 + (m-1)*ICC, where m is average cluster size and ICC is intra-cluster correlation.
  • Clustering sources: user-level randomization, geo-level, or session-level clustering; how to estimate ICC (e.g., ANOVA, mixed models).
  • Trade-offs: increasing sample size or test duration to detect smaller MDEs; balancing statistical power with business relevance.
  • Validation: back-testing with historical experiments, sensitivity analysis on ICC and variance assumptions.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q5

What specific events and attributes would you need in your logs to compute all the metrics and diagnose the mechanism behind any observed lift?

Product Analytics & MetricsData Modeling
Author's notes

Badge impression events, seller profile view events tied to badge visibility, message initiation, and purchase confirmation with session ID linking them all.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clarifying the experiment design and the metric definitions, then work backwards to identify the required log events and attributes. Emphasize that logs must capture both the treatment assignment and the user's post-assignment behavior to compute lift and diagnose its drivers.

Pro tip: Always include a unique experiment ID and a timestamp for every event, and ensure you can join exposure data with outcome data at the user level. This prevents common pitfalls like dilution or misattribution.

1. Clarify the experiment and metrics

Ask clarifying questions about the experiment design, the primary and secondary metrics, and the expected lift mechanism. This ensures you know what to compute and diagnose.

2. Identify necessary events

List the key user actions that define the metrics (e.g., impressions, clicks, conversions) and any intermediate steps that could explain the lift.

3. Define required attributes

For each event, specify attributes like user ID, experiment ID, variant, timestamp, and context (e.g., device, page) that enable segmentation and causal analysis.

4. Ensure data quality and joinability

Verify that events can be linked across sources (e.g., exposure to outcomes) and that data is complete and accurate to avoid biased lift estimates.

5. Plan for diagnostic analysis

Include additional attributes that allow you to test hypotheses about the mechanism, such as user demographics, pre-experiment behavior, and session details.

Key Points to Mention

  • Experiment assignment and exposure events with variant information
  • User-level identifiers and timestamps for all events
  • Metric-specific events (e.g., clicks, purchases) with relevant properties
  • Contextual attributes for segmentation (device, geography, user tenure)
  • Pre-experiment covariates to control for confounding
  • Data pipeline considerations: logging reliability, latency, and join keys

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q6

The test shows a 2% lift on the North Star Metric, a small non-significant drop in sessions per user, and a borderline significant increase in fraud reports. Walk through your launch decision and do a back-of-envelope revenue estimate.

A/B Testing & ExperimentationPricing & MonetizationProduct Strategy
Author's notes

The fraud signal at p=0.06 is the interesting wrinkle here.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clarifying the metric definitions and the experiment's guardrails, then weigh the 2% lift on the North Star against the borderline fraud signal and the non-significant sessions drop. Frame the decision as a risk-adjusted trade-off: quantify the revenue upside with a back-of-envelope estimate, and propose a phased launch or hold based on the fraud risk and statistical uncertainty.

Pro tip: Don't treat the fraud increase as a mere guardrail violation—quantify its expected cost and compare it to the lift's revenue gain; often the fraud cost can erase the apparent win. Also, note that 'borderline significant' means the true effect is uncertain, so a hold-and-monitor or a follow-up experiment is often the mature call.

1. Clarify metrics and experiment validity

Confirm definitions: North Star Metric (e.g., revenue, DAU), sessions per user, and fraud reports. Check sample size, power, and whether the experiment was run correctly (no SRM, etc.).

2. Assess statistical and practical significance

Evaluate the 2% lift's confidence interval and business impact; note the non-significant drop in sessions (likely noise) and the borderline fraud increase (p ~ 0.05) as a potential real risk.

3. Quantify revenue and fraud cost

Back-of-envelope: estimate baseline revenue, apply 2% lift, then subtract expected fraud cost (e.g., fraud rate increase × average fraud loss). Compare net impact.

4. Make a risk-adjusted launch decision

If net positive and fraud risk manageable, launch with monitoring; if fraud cost is high or uncertain, hold and run a follow-up experiment or phased rollout to mitigate risk.

5. Recommend next steps and monitoring

Propose a plan: e.g., launch to a small percentage, set up alerts on fraud metrics, and re-evaluate after a set period. Communicate trade-offs clearly to stakeholders.

Key Points to Mention

  • Statistical significance vs. practical significance: a 2% lift may be meaningful, but the fraud increase is a red flag.
  • Guardrail metrics: fraud reports are a guardrail; a borderline significant increase warrants caution.
  • Back-of-envelope revenue estimate: use baseline revenue, multiply by 2%, and adjust for fraud cost.
  • Risk mitigation: phased rollout, monitoring, or follow-up experiment to resolve uncertainty.
  • Opportunity cost: consider the cost of delaying a potentially winning feature vs. the cost of a bad launch.
  • Stakeholder communication: present a clear recommendation with assumptions and sensitivity analysis.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q7

Name a third-party signal you'd use to improve targeting for this feature, and discuss privacy and compliance considerations plus how you'd validate its incremental value without bias.

Product StrategyA/B Testing & ExperimentationCross-functional Alignment
Author's notes

I went with credit bureau thin-file data as a proxy for seller financial reliability, then immediately second-guessed myself because of FCRA implications in the US.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Choose a concrete third-party signal (e.g., off-platform purchase data from a retail partner) and explain how it fills a targeting gap for the feature. Then systematically address privacy/compliance (consent, data minimization, regulatory frameworks) and outline a bias-aware validation plan (pre-registered holdout, causal inference, fairness audits).

Pro tip: Emphasize that incremental value must be measured against a counterfactual where the signal is absent, and that you'd pre-register the analysis to avoid p-hacking and post-hoc bias.

1. Select and justify the signal

Pick a specific third-party signal (e.g., loyalty card data) and explain why it's relevant to the feature's targeting goal, citing potential lift in key metrics.

2. Address privacy and compliance

Discuss legal bases (consent, legitimate interest), data minimization, anonymization, and adherence to regulations like GDPR/CCPA and Meta's policies.

3. Design bias-aware validation

Propose a randomized controlled trial with a pre-registered analysis plan, ensuring representative sampling and using techniques like propensity score matching to mitigate selection bias.

4. Measure incremental value

Define success metrics (e.g., conversion lift) and compare treatment vs. control groups, using causal inference methods to isolate the signal's effect.

5. Monitor and iterate

Outline ongoing fairness audits, performance monitoring, and a feedback loop to detect and correct any unintended biases or compliance issues.

Key Points to Mention

  • Specific third-party signal (e.g., off-platform purchase data, demographic data from partners)
  • Privacy regulations (GDPR, CCPA) and Meta's data use policies
  • Consent mechanisms and data minimization principles
  • Randomized controlled trials (A/B tests) with pre-registration
  • Bias mitigation techniques (e.g., fairness metrics, representative sampling)
  • Causal inference methods to measure incremental lift

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.