Clear, practical technology insights BSOD Code Lookup · Windows Error Code Lookup · Wi-Fi Troubleshooting · PC Troubleshooting Checklist

Data Clustering in Machine Learning: Methods and Examples

Understand what clusters are, how unsupervised algorithms such as k-means, hierarchical clustering, and DBSCAN group data, and how correlation differs from clustering.

Table of Contents

Clustering is an unsupervised machine-learning technique that groups data points according to a chosen measure of similarity. The aim is to place points that resemble one another in the same cluster while separating points that differ.

A cluster is not an objective fact hidden in every dataset. The result depends on the features, scaling, distance measure, algorithm, and parameters. A useful analysis therefore explains how the groups were produced and whether they remain meaningful outside a chart.

 

What a data cluster looks like

When data has two measurable features, a scatter plot may reveal dense groups visually. The example below appears to contain three groups.

Scatter plot showing three visible data clusters

Real datasets often have more than two dimensions, overlapping groups, outliers, or irregular shapes. Visualization is a valuable first check, but an algorithm is needed to apply a consistent grouping rule.

Clustering is unsupervised learning

In supervised learning, training examples include known target labels. In clustering, the algorithm receives features without a correct cluster label for each point. It must discover a structure according to its mathematical objective.

Common uses include customer segmentation, grouping similar documents, detecting unusual observations, organizing images, and exploring measurements before building another model. A cluster should not automatically be interpreted as a demographic, diagnosis, or causal category.

Prepare the data before clustering

  1. Select relevant features. Irrelevant variables can dominate the grouping.
  2. Handle missing values. Choose an appropriate method rather than silently treating missing data as zero.
  3. Encode categories carefully. Numeric codes such as 1, 2, and 3 may create a false ordering.
  4. Scale numeric features. A variable measured in thousands can overwhelm one measured between 0 and 1.
  5. Choose a similarity measure. Euclidean distance is common, but it is not appropriate for every type of data.
  6. Inspect outliers. They may be errors, important rare cases, or a separate structure.

Major clustering approaches

Partitioning methods

Partitioning algorithms divide observations into a specified number of groups. k-means is the best-known example: it assigns points to the nearest centroid and updates the centroids until the assignments stabilize.

K-means is fast and easy to interpret, but it works best for roughly compact, similarly sized clusters. You must choose k, and the result can be sensitive to scaling, outliers, and initialization.

Hierarchical methods

Hierarchical clustering creates a tree of nested groups called a dendrogram. Agglomerative methods begin with each observation separate and repeatedly merge the closest groups; divisive methods begin with one group and split it.

This approach lets you inspect several possible cluster counts without rerunning a single fixed k. Its behavior depends on the distance metric and linkage rule. BIRCH and CURE are examples designed for particular large or irregular datasets.

Density-based methods

DBSCAN and OPTICS identify dense regions separated by sparser areas. They can find irregularly shaped clusters and mark isolated points as noise. Unlike k-means, DBSCAN does not require a cluster count in advance, but it does require density parameters and can struggle when cluster densities vary greatly.

 

Grid-based methods

Grid-based methods divide the feature space into cells and analyze populated regions rather than comparing every pair of observations. This can improve efficiency for large spatial datasets. STING and CLIQUE are established examples, with CLIQUE also addressing subspaces in high-dimensional data.

How to choose a clustering method

SituationMethod to considerMain caution
Compact groups and a known cluster countk-meansSensitive to scale and outliers
Need a nested view of possible groupingsHierarchical clusteringLinkage choice changes the tree
Irregular shapes and meaningful noise pointsDBSCAN or OPTICSDensity parameters require care
Large spatial or high-dimensional searchGrid/subspace methodsGrid resolution affects the result

Evaluate whether the clusters are useful

There is rarely one perfect score. Combine several checks:

  • Internal separation: are points closer to their own cluster than to other clusters?
  • Stability: do similar groups appear after resampling or small parameter changes?
  • Interpretability: can a domain expert explain what distinguishes the groups?
  • External usefulness: do the groups improve a real decision or downstream task?
  • Fairness: could the grouping encode sensitive attributes or produce harmful treatment?

A visually neat cluster can still be meaningless, and a useful structure may not appear clearly in a two-dimensional projection.

Correlation is not clustering

The Pearson correlation coefficient, usually written as r, measures the direction and strength of a linear relationship between two numeric variables. Its value ranges from -1 to +1. It does not assign observations to groups.

Value of rTypical interpretationLinear direction
-1.00Perfect linear relationshipNegative
-0.70Strong linear relationshipNegative
-0.50Moderate linear relationshipNegative
-0.30Weak linear relationshipNegative
0No linear relationshipNone
+0.30Weak linear relationshipPositive
+0.50Moderate linear relationshipPositive
+0.70Strong linear relationshipPositive
+1.00Perfect linear relationshipPositive

These labels are only rough conventions; what counts as “strong” depends on the field, measurement reliability, and purpose.

A perfect positive linear relationship:

Scatter plot with a perfect positive correlation

A perfect negative linear relationship:

Scatter plot with a perfect negative correlation

 

A moderately strong positive linear relationship:

Scatter plot with a positive correlation around 0.61

No apparent linear relationship:

Scatter plot with little or no linear correlation

Correlation does not imply causation, and r can be near zero even when two variables have a strong nonlinear relationship. Always inspect the scatter plot and look for outliers.

Key distinction

Use clustering when you want to group observations. Use correlation when you want to summarize a linear relationship between two variables. Both can support exploratory analysis, but they answer different questions and should not be treated as interchangeable techniques.

Discussion

Reader Comments 0

Sign in with email or Google to join the discussion.