Silhouette Score: Per-Point Derivation (1D)

Xem dạng PDF

Gửi bài giải

Điểm: 100,00
Giới hạn thời gian: 2.0s
Giới hạn bộ nhớ: 256M

Tác giả:
Dạng bài
Ngôn ngữ cho phép
Python

Problem Statement

Compute the per-point silhouette scores and their mean for a 1D clustering.

For each point x_i in cluster c_i:

  • a(i) = mean absolute distance to all other points in cluster c_i (excluding i itself)
  • b(i) = minimum over all other clusters c ≠ c_i of the mean absolute distance to all points in c
  • If i is the only point in its cluster: s(i) = 0
  • Otherwise: s(i) = (b(i) - a(i)) / max(a(i), b(i))

Function signature:

def silhouette_detailed(X: np.ndarray, labels: np.ndarray) -> tuple:
    # returns (per_point_scores, mean_score)

Input Format

Line 1: n k — number of points and number of clusters
Lines 2..n+1: value label — one point and its cluster label per line

Output Format

Line 1: n per-point silhouette scores (6 decimal places)
Line 2: mean silhouette score (6 decimal places)

Example

Input:

6 2
1.0 0
1.5 0
2.0 0
5.0 1
5.5 1
6.0 1

Output:

0.833333 0.875000 0.785714 0.785714 0.875000 0.833333
0.831349

Derivation

For point x=1.0 (cluster 0, members {1.0, 1.5, 2.0}):

  • a(0) = mean(|1.0-1.5|, |1.0-2.0|) = mean(0.5, 1.0) = 0.75
  • b(0) = mean(|1.0-5.0|, |1.0-5.5|, |1.0-6.0|) = mean(4, 4.5, 5) = 4.5
  • s(0) = (4.5 - 0.75) / max(0.75, 4.5) = 3.75 / 4.5 ≈ 0.8333

The silhouette score ranges from -1 to +1. High scores (near 1) indicate the point is well-matched to its cluster and poorly matched to neighboring clusters.

Notes

  • Use index-based exclusion of self when computing a(i) (not value-based, in case of duplicate values).
  • Absolute distance |x_i - x_j| is used here (1D data).
  • The mean score summarizes overall clustering quality.

Bình luận

Hãy đọc nội quy trước khi bình luận.


Không có bình luận tại thời điểm này.