Xometry·Machine Learning Engineer·Onsite - System Design / Architecture
Jun 2026
Xometry ML Engineer interview, system design round focused on productionizing a model under strict latency constraints. Pretty thorough question that covers a lot of ground at once, from profiling to hardware to post-launch monitoring.
- You're handed a trained model and need to deploy it as a real-time inference service with a 200ms latency budget per prediction. Walk through how you'd clarify requirements, measure baseline performance, optimize the model and serving stack, pick hardware, validate accuracy vs. latency tradeoffs, and set up post-launch monitoring.
“This one is a full gauntlet disguised as a single question.”