gmcluster: Gaussian mixture clustering#
gmcluster estimates the parameters of a Gaussian mixture model from sample data, and picks the number of clusters for you.
Key features:
Estimates the number of clusters automatically by minimum description length (MDL), so you do not have to guess K.
One-line
fiton your data, thenclassify,posterior,log_likelihood, orsample.Full or diagonal cluster covariances.
Optional coordinate whitening to better condition the problem.
Models overlapping clusters with an EM Gaussian mixture, not hard k-means.
Pure NumPy, with no compiled dependencies.
A modern Python rewrite of Bouman’s classic Cluster C program.
MDL selects the number of clusters that best fits the data.
An EM Gaussian mixture separates clusters even when they overlap.
No compiled dependencies; installs and runs anywhere NumPy does.
What gmcluster does.
Install from source and run a demo.
The GaussianMixture class and its methods.
Worked examples on synthetic data.