Training a model is only half the job. You also need to know whether it works, and "it got 95% accuracy" can be misleading: a model can look great on the data it was trained on and fail on anything new, or score well overall while missing the cases that matter most. Model evaluation is the set of tools for measuring performance honestly.
The lessons cover cross validation, which tests a model on data it hasn't seen; the confusion matrix and the metrics built from it (sensitivity and specificity); ROC curves and AUC, for comparing classifiers across thresholds; and the bias-variance tradeoff, which explains why models underfit or overfit.
Watch the lessons in order. If a video runs long, feel free to treat it as a reference you dip back into later rather than something to finish in one sitting.
StatQuest on how to decide which model suits your data: split it into training and testing blocks, rotate them, and compare methods on data they haven't seen.
Video 6minHow to summarize a classifier's results in a table of true and false positives and negatives, and use it to compare different models.
Video 7minTwo metrics built from the confusion matrix: how well a model finds the true positives, and how well it avoids false alarms.
Video 12minHow ROC curves show a classifier's performance across every threshold, and how AUC boils that curve down to a single number for comparing models.
Video 16minWhy some models are too simple and others too flexible, and how the bias-variance tradeoff explains underfitting and overfitting.
Video 6min