gmcluster: Gaussian mixture clustering

gmcluster: Gaussian mixture clustering#

gmcluster estimates the parameters of a Gaussian mixture model from sample data, and picks the number of clusters for you.

gmcluster overview: sample data on the left, fitted Gaussian mixture on the right

Key features:

  • Estimates the number of clusters automatically by minimum description length (MDL), so you do not have to guess K.

  • One-line fit on your data, then classify, posterior, log_likelihood, or sample.

  • Full or diagonal cluster covariances.

  • Optional coordinate whitening to better condition the problem.

  • Models overlapping clusters with an EM Gaussian mixture, not hard k-means.

  • Pure NumPy, with no compiled dependencies.

  • A modern Python rewrite of Bouman’s classic Cluster C program.

Choose K automatically

MDL selects the number of clusters that best fits the data.

Overlapping clusters

An EM Gaussian mixture separates clusters even when they overlap.

Pure NumPy

No compiled dependencies; installs and runs anywhere NumPy does.

Overview

What gmcluster does.

Overview
Installation

Install from source and run a demo.

Installation
API

The GaussianMixture class and its methods.

API Documentation
Demos

Worked examples on synthetic data.

Demo Details