Target Encoding Leakage Quantification

Xem dạng PDF

Gửi bài giải

Điểm: 100,00
Giới hạn thời gian: 2.0s
Giới hạn bộ nhớ: 256M

Tác giả:
Dạng bài
Ngôn ngữ cho phép
Python

Task

When a category has only one training sample, naive mean target encoding leaks the label directly: the encoding equals the label itself. Leave-one-out (LOO) encoding avoids this by falling back to the global mean.

Given a dataset description, compute for each singleton category (exactly 1 sample):

  • naive: naive encoding = the sample's label (0 or 1)
  • loo: LOO encoding = global mean of all training labels
  • gap: |naive - loo| (the leakage amount)

Also output the global mean.

Input

  • Line 1: integer n — total training samples
  • Line 2: float pos_rate — fraction of positive labels (global mean)
  • Line 3: integer k — number of singleton categories to evaluate
  • Lines 4 to 3+k: cat_label where label is 0 or 1 (the single sample for that category)

Output

  • Line 1: global mean (10 sig figs)
  • Then for each singleton category in input order, print: naive loo gap

Print all floats with 10 significant figures ({:.10g}).

Example

Input

100
0.3
3
A 1
B 0
C 1

Output

0.3
1 0.3 0.7
0 0.3 0.3
1 0.3 0.7

Notes

  • Track M: pure Python, no NumPy required.
  • For a singleton category (nc = 1), naive encoding = sc / n_c = label (0 or 1).
  • LOO for n_c = 1 falls back to global mean (no in-group samples remain after leaving one out).
  • The leakage gap = |naive - globalmean| = |label - globalmean|.

Scaffolding

def leakage_quant(n: int, pos_rate: float, singletons: list[tuple[str, int]]) -> tuple[float, list[tuple[float, float, float]]]:
    """
    Returns (global_mean, [(naive, loo, gap), ...]).
    """
    pass

Bình luận

Hãy đọc nội quy trước khi bình luận.


Không có bình luận tại thời điểm này.