Roadmap
What is planned for the handbook, in what order, and how to influence it.
Overview
The handbook already covers the core of a modern data platform: foundations, storage, processing, orchestration, streaming, quality, infrastructure, architecture and AI engineering, with seven hands-on labs. The roadmap focuses on three things: filling the remaining coverage gaps, keeping existing guides accurate, and making it easy for others to contribute.
This page states intent, not commitments. Items move as priorities change, and any item marked help wanted is open for contribution.
flowchart LR
N["Now<br/>New labs, lab CI<br/>and review of existing guides"] --> A["Next<br/>Site features"]
A --> B["Then<br/>Launch"]
B --> V["1.0<br/>Complete coverage,<br/>all labs in CI,<br/>all guides reviewed"]
Now
Item
Status
Review dates and lab-tested versions on every guide
Done (how it works )
Contributor path: first-contribution guide, code owners, changelog, citation file
Done
Coverage round A: five new guides (see below)
Done
Coverage round B: four new guides (see below)
Done
Coverage round C: four new guides (see below)
Done
Review each existing guide against current vendor documentation, so the review dates reflect real checks
Done
Coverage
New guides follow the same template as the existing ones: Basic to Advanced, pitfalls, cheat sheet, interview questions and a Mermaid diagram.
Round A (published)
Guide
Scope
Testing and CI/CD for Data Pipelines
Unit and property tests, idempotency and backfills, CI design, data diff, write-audit-publish, promotion
Trino and Query Federation
Distributed SQL over lakes and databases, connectors, pushdown, Iceberg, fault-tolerant execution
Real-Time Analytics Databases
ClickHouse, Apache Druid and Apache Pinot: when to use each, ingestion and modelling
Kubernetes for Data Workloads
Jobs, resources, node pools and spot capacity, Spark, Airflow and Flink on Kubernetes
Azure and Microsoft Fabric
OneLake, capacity, lakehouse and warehouse, shortcuts and mirroring, Event Hubs, CI/CD
Round B (published)
Guide
Scope
Apache Beam and Dataflow
The unified batch and streaming model, windows and triggers, testing, runners, Dataflow
Streaming SQL
Materialize and RisingWave: incremental views over streams
BI Tools
Apache Superset and Metabase: modelling, performance, row-level security, embedding, operations
NoSQL and Operational Stores
DynamoDB, MongoDB, Valkey and Redis, and Cassandra: modelling, CDC and exports, serving data back
Round C (published)
Guide
Scope
DataOps
Severity levels, on-call, runbooks, incident response, blameless postmortems, operating metrics
MCP and Text-to-SQL
A tested read-only SQL server for assistants, SQL validation, evaluation of generated SQL
Data Catalogs in Practice
DataHub and OpenMetadata: ingestion, metadata model, catalog as code
Choosing a Stack
Requirements first, reference architectures, stack review checks, decision records
Labs
Item
Status
Run Labs 04 and 05 (Kafka, Airflow) in CI, not only validate their Compose files
Done
A dev container so every lab starts in one click
Done
Data quality lab with Great Expectations or Soda
Done (Lab 08 )
Apache Iceberg lab
Done (Lab 09 )
Change data capture lab with Debezium
Done (Lab 10 )
Site
Item
Status
PDF export: one PDF per guide and one for the whole handbook
Done (how it works )
EPUB export
Not planned: the guides are code and diagram heavy, which suits fixed-layout PDF better. Open a topic request if you need it.
Print-friendly guides and cheat sheets
Done
Role-based learning paths (analytics engineer, platform engineer, AI engineer)
Done (paths )
How to influence the roadmap
Ask for a topic with the topic request form. Add a 👍 to an existing request to show demand; requests with more support move up.
Pick up an item. Filter the issues by help wanted or good first issue , and comment to claim one so work is not duplicated.
Discuss direction in Discussions .
Completed work is recorded in the changelog .