Welcome to the upgraded MacSphere! We're putting the finishing touches on it; if you notice anything amiss, email macsphere@mcmaster.ca

An Efficient Implementation of a Robust Clustering Algorithm

dc.contributor.advisorMcNicholas, Paul D.
dc.contributor.authorBlostein, Martin
dc.contributor.departmentMathematics and Statisticsen_US
dc.date.accessioned2016-10-05T18:37:47Z
dc.date.available2016-10-05T18:37:47Z
dc.date.issued2016
dc.description.abstractClustering and classification are fundamental problems in statistical and machine learning, with a broad range of applications. A common approach is the Gaussian mixture model, which assumes that each cluster or class arises from a distinct Gaussian distribution. This thesis studies a robust, high-dimensional extension of the Gaussian mixture model that automatically detects outliers and noise, and a computationally efficient implementation thereof. The contaminated Gaussian distribution is a robust elliptic distribution that allows for automatic detection of ``bad points'', and is used to make robust the usual factor analysis model. In turn, the mixtures of contaminated Gaussian factor analyzers (MCGFA) algorithm allows high-dimesional, robust clustering, classification and detection of bad points. A family of MCGFA models is created through the introduction of different constraints on the covariance structure. A new, efficient implementation of the algorithm is presented, along with an account of its development. The fast implementation permits thorough testing of the MCGFA algorithm, and its performance is compared to two natural competitors: parsimonious Gaussian mixture models (PGMM) and mixtures of modified t factor analyzers (MMtFA). The algorithms are tested systematically on simulated and real data.en_US
dc.description.degreeMaster of Science (MSc)en_US
dc.description.degreetypeThesisen_US
dc.identifier.urihttp://hdl.handle.net/11375/20598
dc.language.isoenen_US
dc.subjectclusteringen_US
dc.subjectclassificationen_US
dc.subjectstatistical learningen_US
dc.subjectmachine learningen_US
dc.subjectrobusten_US
dc.subjectcomputational statisticsen_US
dc.subjectmixture modelsen_US
dc.titleAn Efficient Implementation of a Robust Clustering Algorithmen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
blostein_martin_201609_msc.pdf
Size:
183.32 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.78 KB
Format:
Item-specific license agreed upon to submission
Description: