Introduction to Statistical-Learning Methods for Data Classification
To Know
About this Class
This introductory lecture presents data classification as a statistical-learning topic at the intersection of statistics, computer science, and engineering. It emphasizes the predictive-performance culture of classification while introducing core concepts, such as predictor variables (or features) and output variables (or labels), binary and multiclass classification, class-probability prediction, model fitting, and model validation. The lecture will cover practical performance assessment using accuracy, sensitivity, specificity, and related notions. It will also briefly introduce some frequently used classifier families, such as the K-nearest-neighbors (KNN) classifier and generalized linear models (including logistic regression). Attendees should have a beginner level of statistical knowledge, intermediate is preferred.