Skip to content
Reliable Data Engineering
Lesson
Open in the interactive app

The Data Engineering Coding Round

DE coding rounds are rarely about dynamic programming puzzles. They test whether you can process records correctly and efficiently: parse, group, dedupe, merge, window, schedule, and stream data with clean, testable Python.


1. What to expect

FormatTypical promptWhat they grade
LeetCode-style (45 min)Merge intervals, top-K, LRU cache, sliding windowCorrectness, complexity, edge cases
Data manipulation (45–60 min)“Given these log lines / JSON records, compute…”Parsing robustness, grouping logic, clean code
Mini-pipeline / take-home”Build an ingestion step that dedupes, validates and writes Parquet”Structure, tests, idempotency, error handling
PySpark/pandas live”Compute 7-day rolling revenue per user in PySpark”API fluency, understanding of shuffles

Difficulty calibration: mostly LeetCode easy/medium with a data flavour. Hard algorithmic DP is rare; streaming, hashing and heap problems are common.


2. The Python toolkit to know cold

from collections import Counter, defaultdict, deque, OrderedDict, namedtuple
from itertools import groupby, islice, chain, accumulate, pairwise   # pairwise: 3.10+
from heapq import heappush, heappop, heapify, nlargest, nsmallest, merge
from bisect import bisect_left, bisect_right, insort
from functools import lru_cache, reduce, wraps
from dataclasses import dataclass, field
from datetime import datetime, timedelta, timezone
import json, csv, re
NeedToolComplexity
Count thingsCounter(iterable), .most_common(k)O(n), O(n log k)
Group into listsdefaultdict(list)O(n)
Sliding window / queuedeque(maxlen=k), popleft()O(1) per op
Top-K / K-way merge / schedulingheapq (min-heap; negate for max)O(log n) per op
Sorted lookups / as-of searchbisectO(log n)
Ordered recency (LRU)OrderedDict.move_to_end, popitem(last=False)O(1)
Lazy streamsgenerators (yield), itertoolsO(1) memory
Group consecutive equal keysitertools.groupby (input must be sorted by key)O(n)

Gotchas that cost offers


3. The eight recurring patterns (with DE framing)

#PatternDE framingPractice
1Hash map counting / groupingAggregate logs, dedupe records, join two datasets in memory01, 03, 06
2Sorting + sweep / merge intervalsMerge sessions, subscription coverage, peak concurrency08, 09
3Sliding windowRate limiting, moving averages, sessionization11, 12, 16
4HeapTop-K, K-way merge of sorted files, running median, scheduling10, 19, 27
5Binary search on sorted dataAs-of joins, point-in-time lookups14
6Graphs / topological sortDAG scheduling, dependency resolution, lineage15
7Streaming / generatorsProcess files bigger than memory, batching, retries05, 22, 23
8Probabilistic / systems structuresSampling, dedup at scale, sharding24, 25, 26

4. How to answer (the 5-step loop)

  1. Clarify the data: size (fits in memory?), sortedness, duplicates, nulls/malformed rows, time zones, ties.
  2. State the approach and complexity before coding: “Sort by start then sweep: O(n log n) time, O(n) space.”
  3. Write clean code: small functions, type hints, descriptive names, no premature cleverness.
  4. Test out loud: happy path, empty input, single element, ties, boundary values, malformed input.
  5. Discuss scale: “If the file is 500 GB, I’d stream it with a generator / do an external sort / push it to Spark with groupBy + window.”

What “production-quality” looks like in an interview

from dataclasses import dataclass
from datetime import datetime
from typing import Iterable, Iterator

@dataclass(frozen=True)
class Event:
    user_id: str
    ts: datetime
    kind: str

def parse_events(lines: Iterable[str]) -> Iterator[Event]:
    """Parse 'user_id,iso_ts,kind' lines; skip malformed lines but count them."""
    for line_no, line in enumerate(lines, 1):
        try:
            user_id, ts, kind = line.rstrip("\n").split(",")
            yield Event(user_id, datetime.fromisoformat(ts), kind)
        except ValueError:
            # In production: log + send to a dead-letter file with line_no
            continue

Typed records, a generator (constant memory), explicit handling of bad input. That’s the signal they want.


5. pandas and PySpark equivalents (often asked as follow-ups)

Taskpure PythonpandasPySpark
Group & sumdefaultdict(int)df.groupby("k")["v"].sum()df.groupBy("k").agg(F.sum("v"))
Dedup keep latestdict keyed by id with max tsdf.sort_values("ts").drop_duplicates("id", keep="last")row_number() over window, filter 1
Rolling meandeques.rolling(7).mean()F.avg().over(w.rowsBetween(-6, 0))
As-of joinbisectpd.merge_asofwindow + filter, or range join
Explode list columnloopdf.explode("items")F.explode("items")
Top-K per groupheap per groupdf.groupby("g").head(k) after sortrow_number() ≤ k

Scale talking point: pure Python dicts are fine up to a few GB of keys; beyond that, the same logic becomes a groupBy (shuffle by key) in Spark. Explaining that mapping is a strong senior signal.