Skip to content

Data Platform Engineer Path

From containers and cloud storage to streaming, orchestration, lakehouse tables and the operations that keep them running.

Who it is for: you build and run the systems that move, store and process data, and you are accountable for their reliability and cost.

Labs in this path: 10, 04, 05, 03, 09 and the Lab 06 capstone, about 9 to 12 hours in total. Labs 04, 05 and 10 need Docker. The labs need no cloud account.

Stage 1 — Working foundations

Goal: be fluent in the tools everything else runs on.

Checkpoint: write a Compose file for a service with a health check and a volume, and explain why the health check matters for the services that depend on it.

Stage 2 — Getting data in

Goal: move data from operational systems into the platform without losing or duplicating it.

Checkpoint: explain what a replication slot is, what happens to the source database when a connector stays down for a day, and how you would monitor it.

Stage 3 — Orchestration

Goal: schedule, retry and backfill pipelines safely.

Checkpoint: a task failed halfway through a load. Explain what you clear, what you rerun, and why a rerun cannot duplicate data.

Stage 4 — Processing and lakehouse tables

Goal: process data at scale and store it in tables with transactions.

Checkpoint: a MERGE of 0.1% of a table's rows rewrote the whole table. Explain why, and name two ways to reduce the cost.

Stage 5 — Reliability and operations

Goal: prove changes are safe, see failures early, and respond well.

Checkpoint: write a runbook for one alert (a pipeline that is late), including who is paged, what they check first, and when they escalate.

Stage 6 — Architecture

Goal: turn requirements into a design, and defend the trade-offs.

You are done when you can take a set of requirements (freshness, volume, budget, team size), sketch a platform for them, and say which failure you expect first and how you would notice it.

Going further