Orchestration, Data Quality and Governance Questions
Operational excellence questions. Senior candidates are expected to own reliability, quality and compliance, not just pipelines.
Tags: [core] = expected at every level · [senior] = expected at senior/staff level.
Orchestration and dbt
[core] What does an orchestrator do and not do?
Schedules and orders tasks, handles retries, dependencies, parameters, alerting and run history. It shouldn’t do heavy data processing itself; it triggers Spark/dbt/warehouse jobs.
[senior] Airflow vs Dagster vs Databricks Workflows?
Airflow: task-centric, huge ecosystem, good for heterogeneous orchestration. Dagster: asset-centric with lineage, partitions and testing built in. Workflows: managed and lakehouse-native with zero infrastructure. Choose by ecosystem breadth vs asset awareness vs platform integration.
[senior] dbt incremental strategies?
append (inserts only), merge (upsert on unique_key), delete+insert, insert_overwrite (replace partitions). Choose by update patterns and engine. Add a lookback window for late data, and use full refresh as a safety valve.
[senior] What is slim CI in dbt?
Build and test only modified models and their downstream dependants (state:modified+) in a CI schema, deferring unchanged upstream references to production artifacts. Fast, cheap PR validation.
Data quality
[core] What data quality dimensions do you monitor?
Completeness, uniqueness, validity, consistency, accuracy and timeliness/freshness, plus volume and schema changes. Pipeline health: lag, duration, failures.
[senior] Where should quality checks live and what happens on failure?
Shift left: contracts in producer CI, schema validation at ingestion, row-level expectations in silver (quarantine), dataset-level business rules and reconciliation in gold (block publish via WAP). Severity decides warn vs quarantine vs fail.
[senior] What is a data contract?
An agreement between a producer and its consumers covering schema, semantics, SLAs, quality rules, ownership and evolution policy, enforced automatically (schema registry, CI tests, ingestion validation).
[senior] How do you avoid alert fatigue in data observability?
Tier tables by criticality, seasonal baselines instead of static thresholds, consecutive-failure rules, dedupe incidents by root table using lineage, route to owners, feedback loops to tune sensitivity, and track alert precision.
Governance and privacy
[core] RBAC vs ABAC?
RBAC grants permissions to roles/groups on objects. ABAC evaluates policies on attributes/tags (data classification, user region), so new tagged tables are protected automatically. Scales better for fine-grained, many-domain governance.
[senior] How do row filters and column masks work in a lakehouse catalog?
Functions attached to tables evaluated at query time with the caller’s identity: masks transform column values (e.g. hide email unless in pii_readers); row filters restrict visible rows (e.g. by region membership). Applied uniformly across engines that go through the catalog.
[senior] How do you implement GDPR deletion in Delta tables?
Find subject data via catalog tags/lineage; batched DELETE/anonymise per table; physically remove with deletion-vector purge/OPTIMIZE and VACUUM within the legal window (time travel retention below the SLA); crypto-shredding for immutable stores; prevent resurrection on backfills/late data; audit evidence.
[senior] Pseudonymisation vs anonymisation?
Pseudonymised data can be re-identified with additional information (keys, mapping tables): still personal data under GDPR. Anonymised data cannot reasonably be re-identified (aggregation, k-anonymity) and falls outside GDPR scope.
[senior] What is data lineage used for?
Impact analysis before changes, root-cause analysis in incidents, compliance (where PII flows), trust (where numbers come from), and cost cleanup (unused tables). Column-level lineage is most valuable for PII and metric tracing.