Adjusted Rand Index
Xem dạng PDF
Gửi bài giải
Điểm:
100,00
Giới hạn thời gian:
2.0s
Giới hạn bộ nhớ:
256M
Tác giả:
Dạng bài
Ngôn ngữ cho phép
Python
Problem Statement
Compute the Adjusted Rand Index (ARI) between two cluster label assignments for n data points.
The ARI measures clustering similarity adjusted for chance agreement. ARI = 1 means perfect agreement; ARI = 0 means no better than random.
Algorithm:
- Build a contingency table
CwhereC[i,j]= number of points in true clusteriand predicted clusterj. - Compute
sum_comb = Σ C(n_ij, 2)over all cells. - Compute
sum_row = Σ C(a_i, 2)over row sums,sum_col = Σ C(b_j, 2)over column sums. expected = sum_row * sum_col / C(n, 2)ARI = (sum_comb - expected) / (0.5*(sum_row + sum_col) - expected)
If the denominator is 0: return 1.0 if sum_comb == expected, else 0.0.
Here C(x, 2) = x*(x-1)//2.
Function signature:
def adjusted_rand_index(true: list, pred: list) -> float:
Input Format
Line 1: n k — number of points and number of clusters
Line 2: n space-separated integers — true cluster labels
Line 3: n space-separated integers — predicted cluster labels
Output Format
One line: the ARI value (10 significant figures)
Example
Input:
6 2
0 0 0 1 1 1
0 0 0 1 1 1
Output:
1
Notes
- Label values do not need to be contiguous; use
numpy.uniqueto find them. - Perfect agreement yields ARI = 1.0 regardless of label naming.
Bình luận