Start by tying metrics to the experiment's hypothesis and business goal, then categorize them into success, guardrail, and diagnostic metrics. Explain that good metrics are actionable, sensitive to change, and aligned with long-term objectives, while bad metrics are vanity, ambiguous, or easily gamed.
Pro tip: Emphasize the importance of guardrail metrics to catch unintended consequences, and mention that at Amazon, metrics should be customer-centric and tied to input measures that teams can directly influence.
Clearly articulate what you're testing and the expected outcome. This ensures metrics are directly tied to the experiment's purpose.
Select one key metric that best represents the desired outcome and is sensitive enough to detect a meaningful change.
Identify metrics that should not degrade, such as latency, error rates, or customer satisfaction, to monitor unintended side effects.
Pick additional metrics to help explain why the primary metric moved, such as click-through rates or conversion funnels.
Assess each metric against criteria like actionability, sensitivity, and alignment with long-term goals to separate good from bad.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.