Data Platform Engineer Path¶
From containers and cloud storage to streaming, orchestration, lakehouse tables and the operations that keep them running.
Who it is for: you build and run the systems that move, store and process data, and you are accountable for their reliability and cost.
Labs in this path: 10, 04, 05, 03, 09 and the Lab 06 capstone, about 9 to 12 hours in total. Labs 04, 05 and 10 need Docker. The labs need no cloud account.
Stage 1 — Working foundations¶
Goal: be fluent in the tools everything else runs on.
- Read: Linux and Bash, Git, Docker, Cloud Storage and Terraform.
- Also: DE Concepts for the vocabulary, if any term is new.
Checkpoint: write a Compose file for a service with a health check and a volume, and explain why the health check matters for the services that depend on it.
Stage 2 — Getting data in¶
Goal: move data from operational systems into the platform without losing or duplicating it.
- Read: Data Ingestion and CDC, then Apache Kafka.
- Do: Lab 04, Kafka Streaming (90–120 minutes): dead-letter topics, deduplication, event-time windows and watermarks.
- Do: Lab 10, CDC with Debezium (90–120 minutes): Postgres to Kafka, applying changes idempotently through duplicates, reordering and restarts.
Checkpoint: explain what a replication slot is, what happens to the source database when a connector stays down for a day, and how you would monitor it.
Stage 3 — Orchestration¶
Goal: schedule, retry and backfill pipelines safely.
- Read: Apache Airflow, then Dagster or Prefect for a second model.
- Do: Lab 05, Airflow Orchestration (90–120 minutes): backfills, an idempotent load, a quality gate, pools and asset scheduling.
Checkpoint: a task failed halfway through a load. Explain what you clear, what you rerun, and why a rerun cannot duplicate data.
Stage 4 — Processing and lakehouse tables¶
Goal: process data at scale and store it in tables with transactions.
- Read: PySpark, Delta Lake and Apache Iceberg. Trino covers interactive SQL over the same tables.
- Do: Lab 03, Spark Lakehouse (90–120 minutes), then Lab 09, Iceberg Lakehouse (90–120 minutes). They build the same pipeline in two table formats, so you can compare them.
- Run it on shared infrastructure: Kubernetes for Data Workloads.
Checkpoint: a MERGE of 0.1% of a table's rows rewrote the whole table. Explain why, and name two ways to reduce the cost.
Stage 5 — Reliability and operations¶
Goal: prove changes are safe, see failures early, and respond well.
- Read: Testing and CI/CD, Pipeline Observability, DataOps, Data Security and Privacy and Cost Optimization.
Checkpoint: write a runbook for one alert (a pipeline that is late), including who is paged, what they check first, and when they escalate.
Stage 6 — Architecture¶
Goal: turn requirements into a design, and defend the trade-offs.
- Read: System Design, then Choosing a Stack.
- Practise: the system design case studies.
- Do: Lab 06, Capstone: Dagster pipeline (90–120 minutes) to combine ingestion, modelling, quality gates and orchestration.
You are done when you can take a set of requirements (freshness, volume, budget, team size), sketch a platform for them, and say which failure you expect first and how you would notice it.
Going further¶
- Add streaming analytics with Apache Flink or Streaming SQL, and serve results from a real-time analytics database.
- Read Apache Beam and Dataflow for the unified batch and streaming model.
- Add the data-facing skills with the analytics engineer path.