Free · no account needed
Data Engineering Interview Prep
A study platform for senior data engineering interviews: system design with real diagrams, SQL and Python that run and auto-check in your browser, a step-by-step Python debugger, algorithms by pattern, Spark and Databricks deep dives, and long-form model answers to the questions that decide senior loops.
- lessons
- 43+
- runnable problems
- 122+
- design case studies
- 44+
- interview questions
- 341+
Your progress (solved problems, flashcard schedule, notes) is saved only in your own browser and never sent to a server. Clearing your browser data resets it; the app's Progress page lets you export a backup.
1. Learn the concepts
System Design
Interview framework, estimation, building blocks, reliability, hot keys and design patterns.
7 lessons
Data Architecture
Lakehouse, streaming, CDC, table formats, orchestration, quality, governance, AI, mesh and the big picture.
11 lessons
SQL
Window functions, advanced query patterns, performance and dialects.
3 lessons
Python
The coding round, 16 algorithm patterns, and production Python: decorators, OOP, generators, testing.
3 lessons
Data Modeling
Grain, dimensional modeling, SCD types, Data Vault and modern modeling.
5 lessons
Spark
Internals, memory and OOM debugging, joins and broadcast, caching, explain plans, shuffle, spill and skew, tuning.
9 lessons
Cloud Platforms
Databricks in depth: compute and cost, Delta internals, Unity Catalog, pipelines, jobs, DevOps and security.
5 lessons
2. Practise
System Design
Full designs with diagrams, trade-offs, failure modes, follow-ups and rubrics.
23 problems
SQL
Problems verified in CI and auto-checked in your browser.
41 problems
Python
Data engineering coding problems with tests, a step debugger and per-test feedback.
34 problems
Algorithms
Classic DSA problems grouped by pattern: sliding window, two pointers, heaps, monotonic stacks, DP and more.
47 problems
Data Modeling
Case studies with ER diagrams and rubrics.
10 problems
Spark and Databricks
Query plans, executor sizing, skew, UDFs, partitioning, DPP and an end-to-end Databricks design.
11 problems
Popular: Design a Real-Time Clickstream Analytics Platform · Design a Usage Metering and Billing Pipeline (Cloud/SaaS) · Hot Keys: The Complete Guide Across Kafka, Spark, Flink, Databases and Caches · Shuffle, Spill and Salting: Complete Internals · Minimum Window Containing All Required Characters
3. Drill interview questions
Data Engineering Fundamentals Q&A
Core data engineering interview questions: concepts, SQL and databases, big data, warehousing, cloud, Python and modeling basics.
SQL Interview Questions
Conceptual SQL questions asked in data engineering loops: joins, NULLs, window functions, performance and correctness traps.
Spark and Databricks Interview Questions
Grilling questions on Spark internals, performance, Delta Lake, Databricks, Unity Catalog and Structured Streaming.
Data Modeling Interview Questions
Dimensional modeling, grain, facts, dimensions, SCDs, Data Vault, OBT and modern modeling trade-offs.
Architecture and Streaming Interview Questions
Lakehouse, medallion, table formats, batch vs streaming, Kafka, exactly-once, watermarks, CDC and orchestration questions.
System Design Rapid-Fire Questions
Short trade-off questions interviewers ask during and after a data system design round.
AI Data Engineering Interview Questions
RAG pipelines, embeddings, vector search, evaluation, feature stores, LLM observability and agentic systems from a data engineering perspective.
Python for Data Engineering: Interview Questions
Python language and ecosystem questions for data engineers: generators, memory, concurrency, typing, testing, pandas and PySpark.
Orchestration, Data Quality and Governance Questions
Airflow/Dagster/dbt, data quality, contracts, observability, lineage, access control, PII and GDPR questions.
Senior Deep Dive: Pipeline Reliability and Correctness
In-depth model answers on building reliable pipelines: idempotency, exactly-once, late-arriving data, backfills, backpressure, consistency, deduplication, schema evolution, retries and DLQs, ordering, replay, deletes and reconciliation.
Senior Deep Dive: Operating Data Pipelines at Scale
In-depth model answers on observability and SLOs, data quality strategy, incident response, testing, CI/CD, scaling 10×, cost, multi-tenancy, dependencies, freshness trade-offs, time zones, batch/stream consistency, migrations and build-vs-buy.
Senior Deep Dive: Advanced Data Pipeline Architecture
Staff-level architecture questions with model answers: CDC end to end, stream enrichment and joins, serving-layer choices, semantic layers, reverse ETL, multi-region DR, online/offline feature consistency, event-driven patterns, PII architecture, catalogs and table-format interoperability.
Databricks Interview Questions
The Databricks questions interviewers ask most, with crisp model answers: architecture, compute and cost, Delta Lake, Unity Catalog, Auto Loader, Declarative Pipelines, Jobs, streaming, performance, DevOps and security.