Data Architecture: Learning Path
| # | Module | Key ideas |
|---|---|---|
| 1 | Lakehouse and medallion | Warehouse vs lake vs lakehouse, layer contracts, anti-patterns |
| 2 | Batch vs streaming, Lambda vs Kappa | Choosing latency tiers, the lakehouse hybrid |
| 3 | Streaming deep dive | Kafka, windows, watermarks, state, joins, exactly-once |
| 4 | Ingestion and CDC | Watermark extraction, Debezium, outbox, AutoLoader, APIs |
| 5 | Storage and table formats | Parquet, Delta vs Iceberg vs Hudi, small files, clustering |
| 6 | Orchestration and backfills | DAG design, Airflow vs Dagster, dbt in prod, safe backfills |
| 7 | Data quality and observability | Contracts, WAP, anomaly detection, data SLOs |
| 8 | Governance, security, privacy | RBAC/ABAC, masks, PII, GDPR deletion |
| 9 | AI data architecture | RAG pipelines, vector stores, evals, feature stores, agents |
| 10 | Data mesh and platform | Domains, data products, self-serve platform, FinOps |
| 11 | The big picture | Explain how lakehouse, medallion, mesh, contracts, quality, observability, catalog and the semantic layer fit into one platform, and walk through it in 2 minutes |
Hero diagrams live in assets/diagrams.