Experian·Data Scientist·Technical Phone Screen
Apr 2026
Interviewed for a Data Scientist role at Experian, and the technical round leaned heavily into big-data infrastructure, Spark internals, and cloud pipeline stuff. More engineering-flavored than I expected for a DS title, but not impossible.
- What are the differences between Spark RDDs, DataFrames, and Spark SQL, and what are the advantages of each?
- What advantages does Spark have over traditional MapReduce?
- How does lazy evaluation work in Spark, and how does it help with execution efficiency?
- Walk me through how you would submit and monitor Spark jobs on AWS EMR or a similar managed cluster service.
“I started with RDDs being the low-level building block, then moved to DataFrames having schema awareness and the Catalyst optimizer kicking in.” The rest of the author's notes on Data Scientist interview at Experian, Technical Phone Screen round, covers how they worked through the question, what the panel pushed back on, and what they would do differently.
View Post