Real datasets are messy: missing values, inconsistent types, outliers, and columns on wildly different scales. Before any model can learn something useful, that data has to be loaded, inspected, cleaned, and reshaped into a form the algorithm can actually work with. This module covers the practical toolkit for that stage — NumPy, Pandas, and scikit-learn — rather than the theory behind any one algorithm.
The lessons start with the single most common preprocessing task — handling missing values and fixing datatypes in Pandas — then move to scikit-learn's preprocessing tools for scaling, encoding, and pipelines. The last video is an optional deep dive: a full, project-based Pandas course if you want more practice cleaning a genuinely messy, real-world dataset.
Watch the lessons in order. If a video runs long, feel free to treat it as a reference you dip back into later rather than something to finish in one sitting.
Corey Schafer's focused Pandas lesson on the single most common preprocessing task: spotting missing values, choosing between dropping and filling them, and casting columns to the right datatype.
Video 31minVincent D. Warmerdam walks through scikit-learn's preprocessing tools — scaling, encoding, and pipelines — showing how preprocessing plugs directly into a model.
Video 2h 9minSantiago Basulto's project-based Pandas course, with a hands-on data cleaning project on the Google Playstore dataset — spotting and fixing null values, duplicates, and outliers in a genuinely messy dataset.
Video 5h