Target Encoding Leakage Quantification
Xem dạng PDF
Gửi bài giải
Điểm:
100,00
Giới hạn thời gian:
2.0s
Giới hạn bộ nhớ:
256M
Tác giả:
Dạng bài
Ngôn ngữ cho phép
Python
Task
When a category has only one training sample, naive mean target encoding leaks the label directly: the encoding equals the label itself. Leave-one-out (LOO) encoding avoids this by falling back to the global mean.
Given a dataset description, compute for each singleton category (exactly 1 sample):
naive: naive encoding = the sample's label (0 or 1)loo: LOO encoding = global mean of all training labelsgap: |naive - loo| (the leakage amount)
Also output the global mean.
Input
- Line 1: integer
n— total training samples - Line 2: float
pos_rate— fraction of positive labels (global mean) - Line 3: integer
k— number of singleton categories to evaluate - Lines 4 to 3+k:
cat_labelwhere label is 0 or 1 (the single sample for that category)
Output
- Line 1: global mean (10 sig figs)
- Then for each singleton category in input order, print:
naive loo gap
Print all floats with 10 significant figures ({:.10g}).
Example
Input
100
0.3
3
A 1
B 0
C 1
Output
0.3
1 0.3 0.7
0 0.3 0.3
1 0.3 0.7
Notes
- Track M: pure Python, no NumPy required.
- For a singleton category (nc = 1), naive encoding = sc / n_c = label (0 or 1).
- LOO for n_c = 1 falls back to global mean (no in-group samples remain after leaving one out).
- The leakage gap = |naive - globalmean| = |label - globalmean|.
Scaffolding
def leakage_quant(n: int, pos_rate: float, singletons: list[tuple[str, int]]) -> tuple[float, list[tuple[float, float, float]]]:
"""
Returns (global_mean, [(naive, loo, gap), ...]).
"""
pass
Bình luận