Feature Scaling Impact on K-Means

Xem dạng PDF

Gửi bài giải

Điểm: 100,00
Giới hạn thời gian: 2.0s
Giới hạn bộ nhớ: 256M

Tác giả:
Dạng bài
Ngôn ngữ cho phép
Python

Problem Statement

Demonstrate how feature scaling affects K-Means clustering quality by computing the Within-Cluster Sum of Squares (WCSS) under three scaling strategies:

  1. No scaling — use raw data.
  2. Min-Max scaling — scale each feature to [0, 1]: x' = (x - min) / (max - min).
  3. Standard scaling — zero mean, unit variance: x' = (x - mean) / std (use population std, ddof=0).

For each strategy, run K-Means with the same random seed for 10 iterations and compute the final WCSS.

K-Means details: random initialization (pick k unique points), run for n_iters=10 iterations, assign to nearest centroid by squared Euclidean distance.

Function signature:

def scaling_wcss(X: list, k: int, seed: int) -> tuple:
    # returns (wcss_no_scaling, wcss_minmax, wcss_standard)

Input Format

Line 1: n k — number of points and number of clusters
Lines 2..n+1: space-separated floats — feature values (2D points)

Output Format

Three lines: WCSS for no scaling, min-max scaling, and standard scaling (10 significant figures each)

Example

Input:

6 2
0.0 0.0
1.0 0.0
0.0 1.0
100.0 100.0
100.0 101.0
101.0 100.0

Output:

2.666666667
0.0002614122798
0.00106657186

Notes

  • When all values in a feature dimension are identical, set the range/std to 1 to avoid division by zero.
  • Scaling typically improves clustering when features have very different magnitudes.
  • Use seed=0 as passed to the function (the harness always calls with seed=0).

Bình luận

Hãy đọc nội quy trước khi bình luận.


Không có bình luận tại thời điểm này.