Skip to content
Reliable Data Engineering
Overview

Cloud Platforms: Learning Path

Many senior data engineering roles name a platform in the job description, and the interview probes it in depth: not just “have you used it?” but how it works, what it costs and how you’d operate it safely. This track covers platform-specific knowledge. It starts with Databricks; AWS, Azure and GCP data services will follow.

Databricks

#ModuleYou will be able to
1Platform architecture & computeExplain control vs compute plane, pick the right compute (all-purpose, jobs, serverless, SQL warehouses), reason about Photon and DBU cost
2Delta Lake internalsExplain the transaction log, optimistic concurrency and conflicts, MERGE internals, VACUUM and time travel, liquid clustering, deletion vectors, CDF and clones
3Unity Catalog & governanceDesign catalogs and permissions, implement row filters, masks and ABAC, use lineage and system tables, share data with Delta Sharing and federation
4Ingestion, pipelines, jobs & streamingChoose Auto Loader vs COPY INTO, build Declarative Pipelines with expectations and AUTO CDC, orchestrate with Lakeflow Jobs, run Structured Streaming well
5DevOps, security, serving & AIShip with Asset Bundles and CI/CD, secure the deployment, serve BI with SQL warehouses, support MLflow, Model Serving and Vector Search

Then: Databricks interview questions · End-to-end Databricks design scenario · Related: Spark internals & tuning

Databricks evolves quickly and renames products often (Delta Live Tables → Lakeflow Declarative Pipelines, Workflows → Lakeflow Jobs). Interviewers accept either name; knowing both shows you’re current.