Skip to content

About

Official companion code for the eBook "PySpark: 1,000 Examples — A Practical Reference for Data Engineers" on Amazon Kindle. 22 chapter notebooks, 1,000 runnable examples, reproducible Docker environment.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

PySpark: 1,000 Examples - Companion Code

Official code companion for the eBook PySpark: 1,000 Examples: A Practical Reference for Data Engineers - available on Amazon Kindle.

This repository contains the reference environment and dataset generator for the book. Combined with the Docker image, it is everything you need to practice the 1,000 examples locally - matching the exact environment they were verified against.

The book itself (narrative, explanations, worked solutions, common mistakes, pattern insights) is not included here. Get it on Amazon Kindle.


Quick start - 2 minutes to running

Prerequisite: Docker Desktop installed.

Mac / Linux:

mkdir practice
docker run --rm -p 8888:8888 -p 4040:4040 \
  -v $(pwd)/practice:/workspace/practice \
  bilearner/pyspark1000-practice:1.0

Windows PowerShell:

mkdir practice -Force
docker run --rm -p 8888:8888 -p 4040:4040 `
  -v ${PWD}/practice:/workspace/practice `
  bilearner/pyspark1000-practice:1.0

Open http://localhost:8888 - JupyterLab opens with 22 chapter folders ready for practice.

That is the entire setup. See docs/SETUP.md for the detailed walkthrough.


How to use

  1. Open the book on Kindle, pick an example (say Chapter 10, Example 388)
  2. In JupyterLab, open chapter-10-aggregations/scratch.ipynb
  3. Run the Chapter setup cell at the top (once per session)
  4. Type the example's Solution code into a new cell
  5. Shift+Enter - compare your output to the book's Output section
  6. Read the book's Explanation, Common Mistake, Recommendation, and Pattern Insight

Why type the code yourself? Reading code is not the same as writing it. Typing F.col("x"), catching your own typos, exploring autocomplete - that is where the learning happens.


What's in this repository

Folder What's inside
docs/SETUP.md Detailed step-by-step setup guide
docs/chapters/ Index of every chapter and example in the book
docs/function-index.md Alphabetical index of PySpark functions covered
docs/errata.md Corrections found after publication (report yours via Issues)
conf/ Spark configuration (log4j2, spark-defaults)
scripts/ Utilities: environment verifier, chapter-index generator
datasets/ Dataset generator (generate.py) - produces the data every example references
Dockerfile, Dockerfile.allinone, docker-compose.yml, requirements.txt The reference environment
build-and-push.ps1 Script that builds the practice image and pushes it to Docker Hub

Reference environment

Every example in the book was verified against this exact stack:

Component Version
Python 3.11
PySpark 3.5.3
Delta Lake 3.2.0
Apache Iceberg 1.5.2
pandas 2.2.3
pyarrow 17.0.0
Java OpenJDK 17

The all-in-one Docker image at bilearner/pyspark1000-practice:1.0 contains this exact stack plus all the datasets.


Chapter index

The book is organised in 8 parts covering 22 chapters and 1,000 examples:

Part I - Getting Started

Part II - DataFrame Fundamentals

Part III - Shaping Data

Part IV - Combining Data

Part V - Rich Structures

Part VI - Handling Variety

Part VII - Advanced Techniques

Part VIII - Production Patterns


Datasets

The practice image bakes in all datasets the examples reference - 7 core CSVs, 15 raw messy files for Chapter 4, and Parquet/ORC/Avro/Delta variants for Chapters 5 and 21. Everything is seeded from a fixed random value, so every reader works with byte-identical data.

See datasets/README.md for the full list, schemas, and the generator script if you want to regenerate at a different scale.


About the book

PySpark: 1,000 Examples is a practical reference for data engineers who work with Apache Spark daily. Each example is short, focused, and independently runnable. Rather than long walkthroughs, the book packs each example into a consistent seven-part structure:

  • Problem Statement - what you are solving
  • Solution - the code
  • Output - verified output from the reference environment
  • Explanation - why this works
  • Common Mistake - the trap most people fall into
  • Recommendation - best practice
  • Pattern Insight - how this generalizes

Get the book: Amazon Kindle


Reporting errata or asking questions

Found a typo? Something doesn't work in the code? Have a question?

Please check docs/errata.md before reporting to avoid duplicates.


License

  • Code in scripts/, conf/, datasets/generate.py, and the Dockerfiles: MIT License - do whatever you want, commercial use OK.
  • The book itself (narrative, explanations, worked solutions, structure, cover, images): Copyright (c) 2026 Bi Learner. All rights reserved. Available on Amazon Kindle.
  • Chapter indexes in docs/chapters/: CC BY 4.0 - please attribute if you republish.

Support the work

If this book saved you time, the best thing you can do is:

  1. Leave an honest review on Amazon - reviews genuinely help other engineers find the book
  2. Star this repository
  3. Share with your team or on LinkedIn/X

Made with care by Bi Learner. Reach out on YouTube for tutorials and updates.

About

Official companion code for the eBook "PySpark: 1,000 Examples — A Practical Reference for Data Engineers" on Amazon Kindle. 22 chapter notebooks, 1,000 runnable examples, reproducible Docker environment.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages