Feature Scaling Impact on K-Means
Xem dạng PDF
Gửi bài giải
Điểm:
100,00
Giới hạn thời gian:
2.0s
Giới hạn bộ nhớ:
256M
Tác giả:
Dạng bài
Ngôn ngữ cho phép
Python
Problem Statement
Demonstrate how feature scaling affects K-Means clustering quality by computing the Within-Cluster Sum of Squares (WCSS) under three scaling strategies:
- No scaling — use raw data.
- Min-Max scaling — scale each feature to [0, 1]:
x' = (x - min) / (max - min). - Standard scaling — zero mean, unit variance:
x' = (x - mean) / std(use population std,ddof=0).
For each strategy, run K-Means with the same random seed for 10 iterations and compute the final WCSS.
K-Means details: random initialization (pick k unique points), run for n_iters=10 iterations, assign to nearest centroid by squared Euclidean distance.
Function signature:
def scaling_wcss(X: list, k: int, seed: int) -> tuple:
# returns (wcss_no_scaling, wcss_minmax, wcss_standard)
Input Format
Line 1: n k — number of points and number of clusters
Lines 2..n+1: space-separated floats — feature values (2D points)
Output Format
Three lines: WCSS for no scaling, min-max scaling, and standard scaling (10 significant figures each)
Example
Input:
6 2
0.0 0.0
1.0 0.0
0.0 1.0
100.0 100.0
100.0 101.0
101.0 100.0
Output:
2.666666667
0.0002614122798
0.00106657186
Notes
- When all values in a feature dimension are identical, set the range/std to 1 to avoid division by zero.
- Scaling typically improves clustering when features have very different magnitudes.
- Use
seed=0as passed to the function (the harness always calls withseed=0).
Bình luận