Learning Outcomes
Upon successful completion of the course, students will:
1. understand the notions of prediction and statistical errors.
2. be able to study parametric and non-parametric models.
3. be able to compare and contrast techniques of supervised and unsupervised learning.
4. be able to measure the accuracy of a model.
5. be able to identify strong and weak aspects of various statistical learning methods.
6. understand methods for data analysis.
Course Content (Syllabus)
This course concerns statistical inference from data. We focus on basic principles of supervised and unsupervised learning, as well as in the implementation and in applications of the models in real-world datasets. We also focus on the evaluation of the results obtained from the analysis of data. We will cover topics such as: the notion of distance in statistics. Classification methods, clustering, and dimensionality reduction. Resampling methods, cross-validation, bootstrap, support vector machines, model selection methods.
Keywords
classification, clustering, supervised learning, unsupervised learning
Additional bibliography for study
T. Hastie, R. Tibshirani, J. Friedman, "Elements of Statistical Learning: Data mining, Inference and Prediction", 2nd edition, Springer (2009).
G. James, D. Witten, T. Hastie, R. Tibshirani, "An introduction to statistical learning: with applications in R", Springer texts in Statistics (2017).
N. Cesa-Bianchi, G. Lugosi, "Prediction, learning, and games", Cambridge university press (2006).
Kevin Patrick Murphy, "Probabilistic Machine Learning", MIT Press, 2022.
C.M. Bishop, "Pattern Recognition and Machine Learning", Springer 2006.
Shai Shalev-Shwartz and Shai Ben-David, "Understanding Machine Learning: From Theory to Algorithms", Cambridge University Press. 2014