About this handbook¶
Notes on how the handbook is written, kept current, and how to use it.
Who writes it¶
I'm Sarang Ambekar, a data engineer. This handbook is the reference I wanted while learning and working: one place that connects the concepts, the tools and the production patterns, from a first query to a running pipeline, including the AI and LLM engineering that now sits next to it.
- GitHub: sarangambekar1997
- Issues and corrections: open an issue
How the guides work¶
Every guide follows the same shape, so you can find things without relearning the layout:
| Part | What it gives you |
|---|---|
| Overview | The problem, the solution, and where it sits in a data platform |
| Basic → Intermediate → Advanced | Working code, in order of difficulty |
| Common Pitfalls | What goes wrong in production, and the fix |
| Cheat Sheet | Commands and syntax to copy |
| Interview Questions | Short answers with the trade-off |
| Further Reading | Vendor documentation |
The labs and the projects turn the guides into practice.
Keeping it current¶
Tools change faster than books. Three things keep this one honest:
- Every code block is parsed in CI, and the labs run end to end, so examples don't silently rot.
- Model IDs and prices are not hardcoded. The AI guides link to the vendor pages and load prices from config. A scheduled check flags any retired model ID.
- Every guide shows when it was last reviewed against the vendor documentation, and a guide that is more than six months overdue says so on the page. Links are checked weekly.
How it is maintained explains what each label means and how you can help. If you find something out of date, please tell me.
The roadmap lists what is planned next and where help is wanted.
License¶
Content and code are released under the MIT License. See Contributing to help.