Subscribe Bookmark RSS Feed

What distance is saved when I click "save clusters" in the K-means clustering report

likunz0

Community Trekker

Joined:

Jan 15, 2016

I am wondering what distance is saved for each row when I click "save clusters" in the Kmeans clustering report. I used the K-means method to participate my data table of 10000 rows into 50 clusters. When I clicked "save clusters", I saved two columns. One is the cluster column, which indicate which cluster the row is assigned to; the other one is called "Distance". I am wondering what distance is the "Distance". I found that the distance between each row and the cluster center  is much smaller than the "Distance".

2 REPLIES
txnelson

Super User

Joined:

Jun 22, 2012

This is taken from the Multivariate Methods book available in JMP under Help==>Books==>Multivariate Methods11050_Documentation on Distance.jpg

Jim
likunz0

Community Trekker

Joined:

Jan 15, 2016

Thanks, Jim. It looks like that the distance is calculated as the Euclidean length between two vectors. Then which two vectors are used to calculate the "Distance"? Is it the distance between each row and the center of the cluster of that row? Or is it the distance between each row and the mean of all rows? However, rrom my calculation, the "Distance" is larger than both cases.