Missing Value Imputation Strategy
Xem dạng PDF
Gửi bài giải
Điểm:
100,00
Giới hạn thời gian:
2.0s
Giới hạn bộ nhớ:
256M
Tác giả:
Dạng bài
Ngôn ngữ cho phép
Python
Task
Given a feature column with missing values (marked as NA), automatically select an imputation strategy based on the distribution's skewness, then impute.
Strategy:
- Compute the sample skewness of non-missing values
- If |skewness| < 1.0: impute with mean (distribution is approximately symmetric)
- If |skewness| ≥ 1.0: impute with median (distribution is skewed)
Sample skewness formula: skew = [n / ((n-1)(n-2))] × [∑(xᵢ - x̄)³ / s³]
where n = number of non-missing values, x̄ = mean, s = sample standard deviation (ddof=1).
Edge cases:
- If n < 3 or s < 1e-15, treat skewness as 0.0 (use mean strategy).
- For median with even n: average the two middle values.
Input
- Line 1: integer
n— total number of values (including NAs) - Lines 2 to n+1: either a float or the string
NA
Output
- Line 1:
strategy— eithermeanormedian - Line 2: skewness of non-missing values (formatted with
.10g) - Lines 3 to n+2: imputed column values in order (original non-missing values unchanged; NAs replaced), each formatted with
.10g
Example
Input
5
1.0
NA
3.0
2.0
NA
Output
mean
0
1
2
3
2
2
(Non-missing: [1, 3, 2]; mean = 2.0; skew ≈ 0 → mean strategy; NAs replaced with 2.0)
Notes
- Track A: pure Python, stdlib only — no numpy, no scipy, no pandas.
- Output floats using Python's
format(v, '.10g')— this strips trailing zeros.
Scaffolding
Submit a Python file defining:
def impute(values: list) -> tuple[str, float, list[float]]:
"""
Parameters
----------
values : list of float | None
Column values; None represents a missing entry (NA).
Returns
-------
strategy : str
'mean' or 'median'
skewness : float
Sample skewness of non-missing values (0.0 if n < 3 or std ≈ 0).
imputed_values : list of float
Full column with NAs replaced by the chosen statistic.
"""
...
Bình luận