Skip to content
Reliable Data Engineering
Lesson
Open in the interactive app

Databricks 1: Platform Architecture and Compute

Databricks interviews start with “how does the platform work?” and quickly move to “which compute would you use for this workload, and what will it cost?”. This module gives you the architecture vocabulary and the decision rules.


1. The architecture: control plane and compute plane

flowchart TB
    subgraph ACC["Databricks account (one per organisation)"]
        IAM[Identity: users, groups, service principals<br/>SCIM from Entra ID / Okta]
        UC[Unity Catalog metastore<br/>one per region]
        WS1[Workspace: prod]
        WS2[Workspace: dev]
    end
    subgraph CP["Control plane (Databricks-managed)"]
        WEB[Web app, notebooks, jobs scheduler,<br/>cluster manager, REST APIs, query history]
    end
    subgraph CMP["Compute plane"]
        CLASSIC["Classic compute<br/>VMs in YOUR cloud account / VNet"]
        SLS["Serverless compute<br/>in Databricks' account, isolated per workspace"]
    end
    STOR[("Cloud object storage in your account<br/>S3 / ADLS / GCS: Delta tables, volumes")]
    WS1 --> CP
    CP --> CLASSIC & SLS
    CLASSIC --> STOR
    SLS --> STOR
    UC -. governs .-> STOR

Why it matters in interviews: security questions (“does data leave our account?”), networking (private link, no public IPs), and the classic vs serverless trade-off all follow from this split.


2. Compute types and when to use each

ComputeWhat it isUse it forAvoid for
All-purpose (interactive) clusterLong-running cluster shared by notebooksDevelopment, exploration, collaborationProduction jobs (more expensive DBU rate, shared state, “it worked on my cluster”)
Jobs compute (job cluster)Created for a job run, terminated afterScheduled production pipelines on classic computeInteractive work
Serverless compute for notebooks/jobsManaged, instant-start computeMost jobs and notebooks when the workload fits serverless limits; spiky or short workloadsWorkloads needing custom VM images, specific instance types, GPUs, or unsupported libraries/configs
Serverless Lakeflow Declarative PipelinesManaged pipeline compute with enhanced autoscalingDeclarative ETL (streaming tables, materialised views)Custom low-level Spark tuning
SQL warehouse (serverless / pro / classic)SQL-optimised endpoint with Photon, caching, concurrency scalingBI dashboards, SQL analytics, dbt runsPython/Scala ETL
Instance poolsPre-warmed idle VMs for classic clustersFaster classic cluster start, reduced cloud API throttlingServerless (not needed)
GPU / ML runtime clustersClusters with ML libraries and GPUsDeep learning, fine-tuning, batch inferenceGeneral ETL

Access modes (Unity Catalog era):


3. Databricks Runtime and Photon


4. Autoscaling, spot and cluster policies


5. How cost works: DBUs

Cost levers interviewers expect: jobs instead of all-purpose compute, serverless for spiky/short workloads (no idle), auto-termination, right-sizing (memory vs compute-optimised), spot for batch, Photon where it pays back, liquid clustering and predictive optimisation to reduce scanned data, SQL warehouse auto-stop and sizing, and cluster policies to prevent oversized clusters.


6. Choosing compute: worked decisions

A nightly dbt project with 300 SQL models, run by a scheduler. Which compute?

A serverless SQL warehouse (or pro SQL warehouse) sized for the workload: Photon, result and disk caching, and intelligent workload management for concurrent models; it auto-stops after the run. Trigger it from a Lakeflow Job (dbt task) or your orchestrator.

A PySpark pipeline using a proprietary JVM library and a custom init script, every 2 hours.

Classic jobs compute with a cluster policy, because custom init scripts and specific libraries/instance types may not be supported on serverless. Use an LTS runtime, autoscaling, spot workers with fallback, and an instance pool if start time matters.

Ad-hoc data science exploration by 10 people.

A shared all-purpose cluster in standard access mode (Unity Catalog governed) with auto-termination and a policy limiting size, or serverless notebooks if the libraries fit. Use dedicated-mode or ML runtime clusters only for users who need GPUs or ML libraries.

A streaming pipeline that must run 24/7 with low latency.

A continuous Lakeflow Declarative Pipeline or a Structured Streaming job on jobs compute with on-demand workers (no spot for stateful streaming), RocksDB state store, autoscaling sized for peak, and alerts on processing lag. Serverless pipelines are an option if supported features fit.


Interview questions

Explain the Databricks control plane and compute plane. Where does my data live?

The control plane (Databricks-managed) runs the web app, notebooks, job scheduler, cluster manager and APIs. The compute plane processes data: classic clusters run in your cloud account/network; serverless compute runs in Databricks-managed infrastructure with workspace isolation. Your table data lives in your cloud object storage (directly or as Unity Catalog managed storage), governed by Unity Catalog; the control plane holds metadata such as notebook source and job configs.

All-purpose vs job clusters vs serverless: how do you choose?

All-purpose for interactive development (higher DBU rate, shared); job clusters for scheduled production on classic compute (cheaper DBU rate, isolated per run, reproducible configuration); serverless when you want instant start and no infrastructure management, and the workload fits serverless constraints (supported languages, libraries and configs). SQL warehouses for SQL/BI and dbt.

What is Photon and when doesn't it help?

Photon is Databricks’ vectorised C++ execution engine for Spark SQL and DataFrame operations. It accelerates scans, joins, aggregations, writes and MERGE on columnar batches. It doesn’t help Python UDFs, RDD code or workloads dominated by I/O wait or external calls, and its higher DBU rate must be justified by the speed-up.

How would you control and reduce Databricks costs across 20 teams?

Visibility first: mandatory tags via cluster policies, billing system tables, dashboards and budgets per team. Then guardrails: policies (instance types, max workers, auto-termination, spot), production on jobs compute or serverless, SQL warehouse auto-stop and sizing. Then efficiency: Photon where it pays back, liquid clustering and predictive optimisation, incremental processing instead of full refreshes, and deleting unused tables and jobs. Review top spenders monthly with owners.

What do cluster access modes mean for Unity Catalog?

Standard (shared) access mode supports multiple users with isolation and enforces Unity Catalog fine-grained permissions, row filters and column masks; dedicated (single user) mode assigns compute to one user or group and supports workloads needing full machine access (ML runtimes, RDD APIs); no-isolation mode is legacy and can’t access Unity Catalog data securely. Production pipelines typically run as a service principal on jobs compute or serverless.