BetterHelp·Data Scientist·Technical Phone Screen
- You need to run around 100 experiments but don't have enough traffic to hit the standard 0.05 significance threshold for each one. How would you adjust your alpha to make sure the features you ship are actually impactful?
- After seeing early results from those experiments, how would you bring in an adaptive method like a multi-armed bandit to update your decision thresholds as you go?
- For a new product feature, what primary metric would you choose to measure success, why does it matter, and how would you prevent that metric from being gamed or crowding out other important signals?
“This is where I stumbled a bit.”