I could rattle off linearity and independence fine, then kind of fumbled through homoscedasticity.
Start by listing the core assumptions of linear regression (linearity, independence, homoscedasticity, normality of errors, no multicollinearity, and no endogeneity). Then, for each assumption, explain the specific consequences of violation on coefficient estimates, standard errors, and predictions, and briefly mention detection and remedies. Finally, tie it back to practical implications for model reliability and decision-making.
Pro tip: Emphasize that some violations (e.g., non-normality) matter mainly for inference in small samples, while others (e.g., endogeneity) bias coefficients even in large samples—showing you understand the nuance between statistical and practical significance.
Clearly state the key assumptions: linearity, independence of errors, homoscedasticity, normality of errors, no multicollinearity, and no endogeneity (or exogeneity of regressors).
For each assumption, describe the impact: e.g., non-linearity leads to biased predictions; heteroscedasticity or autocorrelation inflates standard errors; multicollinearity inflates coefficient variance; endogeneity biases coefficients.
Mention common diagnostic tools (residual plots, VIF, Durbin-Watson, Breusch-Pagan) and potential fixes (transformations, robust standard errors, regularization, instrumental variables).
Connect the violations to real-world consequences: unreliable feature importance, poor generalization, incorrect confidence intervals, and misguided business decisions.
AI-generated suggestions, not part of the candidate's original notes. May be inaccurate — verify before relying on them.