100 million sales go into 4 partitions with one lazy plan, and nothing runs until the action. Then store 3 wins with $479.6 million.
python3 --version.pip install "pyspark>=3.5"brew install openjdk@21git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "pyspark>=3.5"
brew install openjdk@21 # macOS sudo apt install openjdk-21-jdk # Debian / Ubuntu java -version
cd data-engineering/13-pyspark python3 src/sales.py
from pyspark.sql import SparkSession
spark = SparkSession.builder.getOrCreate()
df = (spark.range(100_000_000, numPartitions=4)
.selectExpr("id % 50 AS store",
"abs(hash(id)) % 500 AS amt"))
print("partitions:", df.rdd.getNumPartitions())
top = (df.where("amt >= 100").groupBy("store")
.sum("amt").toDF("store", "rev"))
print(top.orderBy("rev", ascending=False).first())Too much data for one machine? Spark splits it. Each core takes a slice, all at once. A hundred million sales, in four partitions.
Each sale gets a store and an amount. Filter, group by store, sum. Nothing runs yet. Spark only builds a plan.
The work starts at the action: first(). Each partition sums its own stores. Then the shuffle: a store's totals meet in one place. Run it: four partitions, and store 3 wins with 479.
6 million. Split. Plan. Act once.
That's how Spark scales.
Read the lesson on GitHub →