← Pinterest Interview Insights

Pinterest·Data Scientist·Technical Phone Screen·Senior

Senior
Apr 2026

Summary

Pinterest DS interview that was basically a stats exam disguised as a conversation. Heavy on ads measurement theory and then they made you grind through the actual math by hand, which I was not fully prepared for.

Questions Asked (5)

Q1

How would you define Brand Lift Study versus Conversion Lift Study in ads measurement, and what are the main sources of bias and variance in each?

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

I knew the high-level distinction but stumbled when they pushed on specific bias sources.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clearly defining both study types, emphasizing their distinct goals: Brand Lift measures upper-funnel metrics like ad recall and brand awareness, while Conversion Lift measures lower-funnel actions like purchases or sign-ups. Then, systematically discuss the main sources of bias and variance for each, highlighting how they differ due to methodology (survey-based vs. behavioral) and experimental design. Conclude by relating these concepts to Pinterest's context, such as measuring ad effectiveness for visual discovery.

Pro tip: Show depth by mentioning that Brand Lift studies often use randomized control trials with survey-based outcomes, which introduces non-response bias, while Conversion Lift studies rely on user-level randomization and may suffer from dilution bias due to cross-device behavior. Also, note that Pinterest's unique position in visual search can leverage both, but requires careful handling of user privacy and measurement gaps.

1. Define Brand Lift Study

Explain that Brand Lift measures the causal impact of ads on upper-funnel brand metrics such as ad recall, brand awareness, and favorability, typically via surveys with randomized control and treatment groups.

2. Define Conversion Lift Study

Explain that Conversion Lift measures the causal impact of ads on lower-funnel actions like purchases, sign-ups, or app installs, using randomized experiments and behavioral data, often with a ghost ads or intent-to-treat design.

3. Identify Sources of Bias in Brand Lift

Discuss biases such as non-response bias (survey respondents differ from non-respondents), selection bias if randomization fails, and measurement bias due to survey wording or timing.

4. Identify Sources of Bias and Variance in Conversion Lift

Discuss biases like dilution (ads seen by control group via other channels), attribution errors (cross-device, view-through), and variance from low conversion rates, requiring large sample sizes and long test durations.

5. Compare and Contrast

Summarize key differences: Brand Lift is survey-based and prone to self-report biases, while Conversion Lift is behavioral and prone to attribution and dilution biases; both require careful experimental design to minimize variance.

Key Points to Mention

  • Randomized control trials (RCTs) are the gold standard for both, but Brand Lift uses survey-based outcomes while Conversion Lift uses behavioral outcomes.
  • Brand Lift is better for measuring upper-funnel metrics (awareness, recall) and Conversion Lift for lower-funnel (purchases, sign-ups).
  • Non-response bias and survey fatigue are major concerns in Brand Lift studies.
  • Dilution bias and cross-device attribution are major concerns in Conversion Lift studies.
  • Variance in Conversion Lift is often higher due to low base rates, requiring larger sample sizes and longer test windows.
  • Pinterest's visual nature may require innovative measurement approaches, such as combining both study types or using incrementality testing.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q2

When would a Brand Lift Study and a Conversion Lift Study give you conflicting results, and how would you go about reconciling them?

A/B Testing & ExperimentationRoot Cause Analysis
Author's notes

This tripped me up more than it should have.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining what each study measures and why they might diverge, then walk through common scenarios where conflicts arise (e.g., upper-funnel vs. lower-funnel effects, measurement methodologies, attribution windows). Finally, outline a systematic reconciliation process that prioritizes understanding the root cause before deciding which metric to trust or how to combine insights.

Pro tip: Emphasize that conflicting results are often a signal to dig deeper into the data rather than a problem to be resolved by picking a winner. Mention that at Pinterest, where visual discovery often drives both brand and performance outcomes, reconciling these studies requires a nuanced understanding of user behavior across the funnel.

1. Clarify Definitions and Objectives

Explain that Brand Lift measures changes in awareness, ad recall, or favorability, while Conversion Lift measures incremental conversions. Highlight that they serve different purposes and may use different methodologies (e.g., survey-based vs. user-level conversion tracking).

2. Identify Potential Sources of Conflict

List common reasons for divergence: different attribution windows, upper-funnel vs. lower-funnel focus, measurement bias (e.g., survey non-response), or external factors like seasonality. Also consider that brand lift may not immediately translate to conversions.

3. Analyze the Data and Methodology

Dive into the details: check for statistical significance, sample sizes, and potential confounders. Compare the control groups and ensure both studies were properly randomized. Look for segment-level differences that might explain the conflict.

4. Reconcile by Triangulating Insights

If both studies are valid, consider that they capture different aspects of the funnel. Use additional data (e.g., click-through rates, view-through conversions) to build a holistic view. If one study is flawed, address the limitations and possibly rerun with adjustments.

5. Communicate and Recommend Actions

Present a clear narrative to stakeholders: acknowledge the conflict, explain the likely reasons, and recommend how to move forward (e.g., trust the more robust study, run a follow-up, or adjust KPIs). Emphasize the importance of aligning on the primary objective.

Key Points to Mention

  • Difference between brand lift (awareness, favorability) and conversion lift (incremental conversions) metrics.
  • Methodological differences: survey-based brand lift vs. user-level conversion lift, and potential biases.
  • Attribution windows and lag effects: brand impact may take time to convert.
  • Statistical power and significance: ensure both studies are adequately powered.
  • External validity: consider seasonality, competitive activity, or other confounders.
  • Holistic measurement: using both to understand full-funnel impact and avoid siloed decision-making.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q3

Compare Difference-in-Differences and Propensity Score Matching for observational ads data. What are the identification assumptions for each, and what robustness check would you run?

A/B Testing & ExperimentationTechnical Trade-offs
Author's notes

DiD I was comfortable with, parallel trends assumption and all.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by defining each method and its core identification assumptions, then contrast them in the context of observational ads data. Discuss robustness checks that test the validity of these assumptions, and conclude with when to prefer one method over the other.

Pro tip: Emphasize that DiD relies on parallel trends, which is untestable but can be probed with pre-trend tests, while PSM assumes selection on observables, which is also untestable but can be assessed with balance checks. Mention that combining both (e.g., DiD with propensity score weighting) can strengthen causal inference when assumptions are partially met.

1. Define the methods

Briefly explain DiD (compares changes in outcomes over time between treated and control groups) and PSM (matches treated and control units based on propensity scores to create comparable groups).

2. State identification assumptions

For DiD: parallel trends assumption (absent treatment, treated and control groups would have followed parallel outcome trends). For PSM: conditional independence (selection on observables) and common support (overlap in propensity scores).

3. Discuss robustness checks

For DiD: test for pre-treatment trends, placebo tests, sensitivity analysis to violations. For PSM: check covariate balance after matching, sensitivity to unobserved confounders, and common support diagnostics.

4. Compare applicability to ads data

Highlight that DiD is useful when you have pre/post data and a natural control group, while PSM is useful when you have rich covariates but no clear pre-period. Discuss potential biases in ads data (e.g., targeting, seasonality).

5. Conclude with recommendation

Suggest that the choice depends on data structure and assumptions; sometimes combining methods (e.g., DiD with PSM) can be more robust. Mention that in practice, you'd run both and compare results.

Key Points to Mention

  • Parallel trends assumption for DiD and how to test it with pre-period data
  • Conditional independence and common support for PSM
  • Sensitivity analysis for unobserved confounding in both methods
  • Covariate balance checks (e.g., standardized mean differences) after PSM
  • Placebo tests and falsification tests for DiD
  • Combining methods (e.g., propensity score weighting in DiD) to address weaknesses

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q4

In a randomized conversion lift holdout, treatment group had 1200 users with 42 conversions and control had 1180 users with 33 conversions. Compute the point lift, the pooled standard error under the null, and the two-sided z-statistic. Is the result significant at the 5% level?

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

The arithmetic is not hard but doing it live without a calculator while also explaining each step is a different experience.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

First, compute the conversion rates for treatment and control, then calculate the absolute and relative lift. Next, compute the pooled standard error under the null hypothesis of no difference, and use it to calculate the z-statistic. Finally, compare the z-statistic to the critical value for a two-sided test at the 5% significance level (1.96) to determine significance.

Pro tip: Always clarify whether the question asks for absolute or relative lift, and mention that in practice you'd also check for novelty effects and ensure the experiment was run for a sufficient duration to detect meaningful effects.

1. Calculate conversion rates and lift

Compute the conversion rate for treatment (42/1200 = 0.035) and control (33/1180 ≈ 0.02797). Then compute the absolute lift (difference) and relative lift (percentage change).

2. Compute pooled standard error under the null

Under the null hypothesis of no difference, use the pooled conversion rate: (42+33)/(1200+1180) = 75/2380 ≈ 0.03151. Then compute the standard error as sqrt(p_pool*(1-p_pool)*(1/n_t + 1/n_c)).

3. Calculate the z-statistic

Divide the difference in conversion rates by the pooled standard error to get the z-statistic.

4. Determine significance

Compare the absolute value of the z-statistic to the critical value for a two-sided test at α=0.05 (1.96). If |z| > 1.96, the result is statistically significant.

Key Points to Mention

  • Conversion rates: treatment = 3.5%, control ≈ 2.797%
  • Absolute lift = 0.703 percentage points; relative lift ≈ 25.1%
  • Pooled conversion rate ≈ 3.151%
  • Pooled standard error ≈ 0.00716 (or 0.716 percentage points)
  • Z-statistic ≈ 0.982
  • Not significant at 5% level because |z| < 1.96

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.

Q5

You have pre/post Purchase Intent survey data for exposed and control groups with given sample means and variances. Derive the DiD estimator, compute an approximate standard error using the four independent variance terms scaled by their sample sizes, and report the t-statistic. Show your formulas.

A/B Testing & ExperimentationProduct Analytics & Metrics
Author's notes

This was the hardest part of the whole thing.

Create a free account to read the full note

AI HintsAI Generated

Suggested Approach

Start by clearly defining the four sample means and variances, then derive the difference-in-differences (DiD) estimator as the difference of the changes in means between exposed and control groups. Compute the standard error by combining the four independent variance terms scaled by their respective sample sizes, then calculate the t-statistic as the DiD estimate divided by its standard error. Show all formulas step-by-step to demonstrate rigor.

Pro tip: Emphasize that the DiD estimator assumes parallel trends and that the variance calculation treats the four group-time means as independent, which is valid for independent samples. Mention that in practice, you might use regression with interaction terms to obtain the same estimate and robust standard errors.

1. Define notation and assumptions

Label the four groups: exposed pre, exposed post, control pre, control post. Let their sample means be X̄_E,pre, X̄_E,post, X̄_C,pre, X̄_C,post and variances s²_E,pre, s²_E,post, s²_C,pre, s²_C,post with sample sizes n_E,pre, n_E,post, n_C,pre, n_C,post. Assume independent samples and that variances are known or estimated from large samples.

2. Derive the DiD estimator

The DiD estimator is the difference between the change in the exposed group and the change in the control group: Δ = (X̄_E,post - X̄_E,pre) - (X̄_C,post - X̄_C,pre). This can also be written as X̄_E,post - X̄_E,pre - X̄_C,post + X̄_C,pre.

3. Compute the variance of the DiD estimator

Since the four sample means are independent, the variance of Δ is the sum of the variances of each mean, each scaled by its sample size: Var(Δ) = s²_E,post/n_E,post + s²_E,pre/n_E,pre + s²_C,post/n_C,post + s²_C,pre/n_C,pre. The standard error is the square root of this sum.

4. Calculate the t-statistic

The t-statistic is t = Δ / SE(Δ). Under the null hypothesis of no effect, t follows approximately a standard normal distribution for large samples. Report the t-statistic and optionally the p-value.

Key Points to Mention

  • Difference-in-differences estimator formula and its interpretation as the causal effect under parallel trends.
  • Independence assumption for the four sample means and how it simplifies variance calculation.
  • Scaling variances by sample sizes to obtain the standard error of the DiD estimator.
  • The t-statistic formula and its use for hypothesis testing.
  • Potential need for robust standard errors if using regression or if variances are estimated.
  • Assumption of parallel trends and its importance for validity.

AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.