Adjusted Rand Index

Xem dạng PDF

Gửi bài giải

Điểm: 100,00
Giới hạn thời gian: 2.0s
Giới hạn bộ nhớ: 256M

Tác giả:
Dạng bài
Ngôn ngữ cho phép
Python

Problem Statement

Compute the Adjusted Rand Index (ARI) between two cluster label assignments for n data points.

The ARI measures clustering similarity adjusted for chance agreement. ARI = 1 means perfect agreement; ARI = 0 means no better than random.

Algorithm:

  1. Build a contingency table C where C[i,j] = number of points in true cluster i and predicted cluster j.
  2. Compute sum_comb = Σ C(n_ij, 2) over all cells.
  3. Compute sum_row = Σ C(a_i, 2) over row sums, sum_col = Σ C(b_j, 2) over column sums.
  4. expected = sum_row * sum_col / C(n, 2)
  5. ARI = (sum_comb - expected) / (0.5*(sum_row + sum_col) - expected)

If the denominator is 0: return 1.0 if sum_comb == expected, else 0.0.

Here C(x, 2) = x*(x-1)//2.

Function signature:

def adjusted_rand_index(true: list, pred: list) -> float:

Input Format

Line 1: n k — number of points and number of clusters
Line 2: n space-separated integers — true cluster labels
Line 3: n space-separated integers — predicted cluster labels

Output Format

One line: the ARI value (10 significant figures)

Example

Input:

6 2
0 0 0 1 1 1
0 0 0 1 1 1

Output:

1

Notes

  • Label values do not need to be contiguous; use numpy.unique to find them.
  • Perfect agreement yields ARI = 1.0 regardless of label naming.

Bình luận

Hãy đọc nội quy trước khi bình luận.


Không có bình luận tại thời điểm này.