Leaky vs Correct Scaling Pipeline
Xem dạng PDF
Gửi bài giải
Điểm:
100,00
Giới hạn thời gian:
2.0s
Giới hạn bộ nhớ:
256M
Tác giả:
Dạng bài
Ngôn ngữ cho phép
Python
Task
A common data leakage mistake is fitting a scaler on the combined train+test data before splitting. This leaks test-set statistics into the training process.
Given train and test feature arrays, compute:
std_correct: standard deviation of train features when scaler is fit only on train (ddof=0)std_leaky: standard deviation of the combined train+test features (ddof=0) — what a leaky pipeline would useleaks: "yes" if stdleaky != stdcorrect (i.e., test data changes the scale), else "no"
Input
- Line 1: integer
n_train - Line 2:
n_trainspace-separated floats (training feature values) - Line 3: integer
n_test - Line 4:
n_testspace-separated floats (test feature values)
Output
Three lines:
std_correct(10 sig figs)std_leaky(10 sig figs)leaks— "yes" or "no"
Example
Input
4
1.0 2.0 3.0 4.0
2
10.0 20.0
Output
1.118033989
6.624868971
yes
Notes
- Track A: pure Python / NumPy only — no scipy, no sklearn.
- Use population standard deviation (ddof=0) throughout.
- Two std values are equal when their absolute difference is less than 1e-9.
Scaffolding
def pipeline_check(train: list[float], test: list[float]) -> tuple[float, float, str]:
"""
Returns (std_correct, std_leaky, leaks) where:
std_correct — std of train only
std_leaky — std of train+test combined
leaks — "yes" if they differ, else "no"
"""
pass
Bình luận