Official code companion for the eBook PySpark: 1,000 Examples: A Practical Reference for Data Engineers - available on Amazon Kindle.
This repository contains the reference environment and dataset generator for the book. Combined with the Docker image, it is everything you need to practice the 1,000 examples locally - matching the exact environment they were verified against.
The book itself (narrative, explanations, worked solutions, common mistakes, pattern insights) is not included here. Get it on Amazon Kindle.
Prerequisite: Docker Desktop installed.
Mac / Linux:
mkdir practice
docker run --rm -p 8888:8888 -p 4040:4040 \
-v $(pwd)/practice:/workspace/practice \
bilearner/pyspark1000-practice:1.0Windows PowerShell:
mkdir practice -Force
docker run --rm -p 8888:8888 -p 4040:4040 `
-v ${PWD}/practice:/workspace/practice `
bilearner/pyspark1000-practice:1.0Open http://localhost:8888 - JupyterLab opens with 22 chapter folders ready for practice.
That is the entire setup. See docs/SETUP.md for the detailed walkthrough.
- Open the book on Kindle, pick an example (say Chapter 10, Example 388)
- In JupyterLab, open
chapter-10-aggregations/scratch.ipynb - Run the Chapter setup cell at the top (once per session)
- Type the example's Solution code into a new cell
- Shift+Enter - compare your output to the book's Output section
- Read the book's Explanation, Common Mistake, Recommendation, and Pattern Insight
Why type the code yourself? Reading code is not the same as writing it. Typing F.col("x"), catching your own typos, exploring autocomplete - that is where the learning happens.
| Folder | What's inside |
|---|---|
docs/SETUP.md |
Detailed step-by-step setup guide |
docs/chapters/ |
Index of every chapter and example in the book |
docs/function-index.md |
Alphabetical index of PySpark functions covered |
docs/errata.md |
Corrections found after publication (report yours via Issues) |
conf/ |
Spark configuration (log4j2, spark-defaults) |
scripts/ |
Utilities: environment verifier, chapter-index generator |
datasets/ |
Dataset generator (generate.py) - produces the data every example references |
Dockerfile, Dockerfile.allinone, docker-compose.yml, requirements.txt |
The reference environment |
build-and-push.ps1 |
Script that builds the practice image and pushes it to Docker Hub |
Every example in the book was verified against this exact stack:
| Component | Version |
|---|---|
| Python | 3.11 |
| PySpark | 3.5.3 |
| Delta Lake | 3.2.0 |
| Apache Iceberg | 1.5.2 |
| pandas | 2.2.3 |
| pyarrow | 17.0.0 |
| Java | OpenJDK 17 |
The all-in-one Docker image at bilearner/pyspark1000-practice:1.0 contains this exact stack plus all the datasets.
The book is organised in 8 parts covering 22 chapters and 1,000 examples:
The practice image bakes in all datasets the examples reference - 7 core CSVs, 15 raw messy files for Chapter 4, and Parquet/ORC/Avro/Delta variants for Chapters 5 and 21. Everything is seeded from a fixed random value, so every reader works with byte-identical data.
See datasets/README.md for the full list, schemas, and the generator script if you want to regenerate at a different scale.
PySpark: 1,000 Examples is a practical reference for data engineers who work with Apache Spark daily. Each example is short, focused, and independently runnable. Rather than long walkthroughs, the book packs each example into a consistent seven-part structure:
- Problem Statement - what you are solving
- Solution - the code
- Output - verified output from the reference environment
- Explanation - why this works
- Common Mistake - the trap most people fall into
- Recommendation - best practice
- Pattern Insight - how this generalizes
Get the book: Amazon Kindle
Found a typo? Something doesn't work in the code? Have a question?
- Report an errata - corrections we add to
docs/errata.md - Ask a question - for anything community-relevant
- Suggest a topic - for a future edition
Please check docs/errata.md before reporting to avoid duplicates.
- Code in
scripts/,conf/,datasets/generate.py, and the Dockerfiles: MIT License - do whatever you want, commercial use OK. - The book itself (narrative, explanations, worked solutions, structure, cover, images): Copyright (c) 2026 Bi Learner. All rights reserved. Available on Amazon Kindle.
- Chapter indexes in
docs/chapters/: CC BY 4.0 - please attribute if you republish.
If this book saved you time, the best thing you can do is:
- Leave an honest review on Amazon - reviews genuinely help other engineers find the book
- Star this repository
- Share with your team or on LinkedIn/X
Made with care by Bi Learner. Reach out on YouTube for tutorials and updates.