Silhouette Score: Per-Point Derivation (1D)
Xem dạng PDF
Gửi bài giải
Điểm:
100,00
Giới hạn thời gian:
2.0s
Giới hạn bộ nhớ:
256M
Tác giả:
Dạng bài
Ngôn ngữ cho phép
Python
Problem Statement
Compute the per-point silhouette scores and their mean for a 1D clustering.
For each point x_i in cluster c_i:
a(i)= mean absolute distance to all other points in clusterc_i(excludingiitself)b(i)= minimum over all other clustersc ≠ c_iof the mean absolute distance to all points inc- If
iis the only point in its cluster:s(i) = 0 - Otherwise:
s(i) = (b(i) - a(i)) / max(a(i), b(i))
Function signature:
def silhouette_detailed(X: np.ndarray, labels: np.ndarray) -> tuple:
# returns (per_point_scores, mean_score)
Input Format
Line 1: n k — number of points and number of clusters
Lines 2..n+1: value label — one point and its cluster label per line
Output Format
Line 1: n per-point silhouette scores (6 decimal places)
Line 2: mean silhouette score (6 decimal places)
Example
Input:
6 2
1.0 0
1.5 0
2.0 0
5.0 1
5.5 1
6.0 1
Output:
0.833333 0.875000 0.785714 0.785714 0.875000 0.833333
0.831349
Derivation
For point x=1.0 (cluster 0, members {1.0, 1.5, 2.0}):
a(0)= mean(|1.0-1.5|, |1.0-2.0|) = mean(0.5, 1.0) = 0.75b(0)= mean(|1.0-5.0|, |1.0-5.5|, |1.0-6.0|) = mean(4, 4.5, 5) = 4.5s(0)= (4.5 - 0.75) / max(0.75, 4.5) = 3.75 / 4.5 ≈ 0.8333
The silhouette score ranges from -1 to +1. High scores (near 1) indicate the point is well-matched to its cluster and poorly matched to neighboring clusters.
Notes
- Use index-based exclusion of self when computing
a(i)(not value-based, in case of duplicate values). - Absolute distance
|x_i - x_j|is used here (1D data). - The mean score summarizes overall clustering quality.
Bình luận