GroupBy Log Aggregation
Xem dạng PDF
Gửi bài giải
Điểm:
100,00
Giới hạn thời gian:
2.0s
Giới hạn bộ nhớ:
256M
Tác giả:
Dạng bài
Ngôn ngữ cho phép
Python
Task
Given a list of HDFS log events as (BlockId, EventTemplate) pairs, compute per-block statistics using a groupby aggregation:
count: total number of events for that blockunique: number of distinct event templates seen for that blocktop_event: the most frequent event template (ties broken alphabetically)
Input
- Line 1: integer
n— number of log entries - Lines 2 to n+1:
block_id event_template(space-separated; blockid is an integer, eventtemplate is a single word with no spaces)
Output
For each block in ascending order of block_id, print one line:
block_id count unique top_event
Example
Input
7
1 E1
1 E2
1 E1
2 E3
2 E3
2 E1
3 E5
Output
1 3 2 E1
2 3 2 E3
3 1 1 E5
Notes
- Track A: pure Python / NumPy only — no pandas, no scipy, no sklearn.
Scaffolding
Submit a Python file defining:
def groupby_agg(events: list[tuple[int, str]]) -> list[tuple[int, int, int, str]]:
...
Receives list of (blockid, eventtemplate) tuples. Returns list of (blockid, count, unique, topevent) tuples sorted by block_id ascending.
Bình luận