Skip to content

Roadmap

What is planned for the handbook, in what order, and how to influence it.

Overview

The handbook already covers the core of a modern data platform: foundations, storage, processing, orchestration, streaming, quality, infrastructure, architecture and AI engineering, with seven hands-on labs. The roadmap focuses on three things: filling the remaining coverage gaps, keeping existing guides accurate, and making it easy for others to contribute.

This page states intent, not commitments. Items move as priorities change, and any item marked help wanted is open for contribution.

flowchart LR
    N["Now<br/>New labs, lab CI<br/>and review of existing guides"] --> A["Next<br/>Site features"]
    A --> B["Then<br/>Launch"]
    B --> V["1.0<br/>Complete coverage,<br/>all labs in CI,<br/>all guides reviewed"]

Now

Item Status
Review dates and lab-tested versions on every guide Done (how it works)
Contributor path: first-contribution guide, code owners, changelog, citation file Done
Coverage round A: five new guides (see below) Done
Coverage round B: four new guides (see below) Done
Coverage round C: four new guides (see below) Done
Review each existing guide against current vendor documentation, so the review dates reflect real checks Done

Coverage

New guides follow the same template as the existing ones: Basic to Advanced, pitfalls, cheat sheet, interview questions and a Mermaid diagram.

Round A (published)

Guide Scope
Testing and CI/CD for Data Pipelines Unit and property tests, idempotency and backfills, CI design, data diff, write-audit-publish, promotion
Trino and Query Federation Distributed SQL over lakes and databases, connectors, pushdown, Iceberg, fault-tolerant execution
Real-Time Analytics Databases ClickHouse, Apache Druid and Apache Pinot: when to use each, ingestion and modelling
Kubernetes for Data Workloads Jobs, resources, node pools and spot capacity, Spark, Airflow and Flink on Kubernetes
Azure and Microsoft Fabric OneLake, capacity, lakehouse and warehouse, shortcuts and mirroring, Event Hubs, CI/CD

Round B (published)

Guide Scope
Apache Beam and Dataflow The unified batch and streaming model, windows and triggers, testing, runners, Dataflow
Streaming SQL Materialize and RisingWave: incremental views over streams
BI Tools Apache Superset and Metabase: modelling, performance, row-level security, embedding, operations
NoSQL and Operational Stores DynamoDB, MongoDB, Valkey and Redis, and Cassandra: modelling, CDC and exports, serving data back

Round C (published)

Guide Scope
DataOps Severity levels, on-call, runbooks, incident response, blameless postmortems, operating metrics
MCP and Text-to-SQL A tested read-only SQL server for assistants, SQL validation, evaluation of generated SQL
Data Catalogs in Practice DataHub and OpenMetadata: ingestion, metadata model, catalog as code
Choosing a Stack Requirements first, reference architectures, stack review checks, decision records

Labs

Item Status
Run Labs 04 and 05 (Kafka, Airflow) in CI, not only validate their Compose files Done
A dev container so every lab starts in one click Done
Data quality lab with Great Expectations or Soda Done (Lab 08)
Apache Iceberg lab Done (Lab 09)
Change data capture lab with Debezium Done (Lab 10)

Site

Item Status
PDF export: one PDF per guide and one for the whole handbook Done (how it works)
EPUB export Not planned: the guides are code and diagram heavy, which suits fixed-layout PDF better. Open a topic request if you need it.
Print-friendly guides and cheat sheets Done
Role-based learning paths (analytics engineer, platform engineer, AI engineer) Done (paths)

How to influence the roadmap

  • Ask for a topic with the topic request form. Add a 👍 to an existing request to show demand; requests with more support move up.
  • Pick up an item. Filter the issues by help wanted or good first issue, and comment to claim one so work is not duplicated.
  • Discuss direction in Discussions.

Completed work is recorded in the changelog.