Python Coding Practice for Data Engineers
Every solution is executed against its tests by scripts/build.py. Run them yourself in the browser with the web platform (Pyodide), or locally: copy the starter, implement, paste the tests.
Read first: The data engineering coding round · Algorithm patterns · Classic DSA problems by pattern: Algorithms track
Easy
| # | Problem | Topics |
|---|---|---|
| 01 | Top-K Most Frequent Search Terms | hash-map, heap, counting |
| 02 | Flatten Nested JSON Records | recursion, json, schema |
| 03 | Deduplicate Records Keeping the Latest Version | hash-map, deduplication, cdc |
| 04 | Parse Web Server Logs into Hourly Status Counts | parsing, regex, aggregation |
| 05 | Merge Two Sorted Event Streams Lazily | generators, two-pointers, streaming |
| 06 | Implement GROUP BY with Multiple Aggregates | hash-map, aggregation, one-pass |
| 07 | Validate Records Against a Schema | validation, data-quality, schema |
| 22 | Batch a Stream into Fixed-Size Chunks | generators, itertools, streaming |
Medium
Hard
| # | Problem | Topics |
|---|---|---|
| 24 | Uniform Sample from a Stream of Unknown Length | sampling, streaming, probability |
| 25 | Consistent Hashing Ring for Sharding | hashing, distributed-systems, bisect |
| 26 | Bloom Filter for Streaming Deduplication | probabilistic, hashing, deduplication |
| 27 | Sort Data Larger Than Memory (External Merge Sort) | external-sort, heap, generators |
| 32 | Usage Credit Ledger With Expiring Grants | heap, ledger, design |