Streaming caught the fraud after 2 s and $1,442; batch after 11 h 52 m and $4,202.
python3 --version.git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
cd data-engineering/26-batch-vs-streaming python3 src/stream.py
from swipes import day, NIGHTLY, LAG
from swipes import suspicious, report
batch, stream = {}, {}
for e in day: # streaming: swipe by swipe, live
if suspicious(stream, e):
report("stream", e, e.ts + LAG)
for e in day: # batch: the whole day at 02:00
if suspicious(batch, e):
report("batch", e, NIGHTLY)Some data can't wait for the nightly job. One fraud rule: three card swipes in 60 seconds, from two countries. Streaming, simulated here, checks every swipe the moment it lands. Batch runs the exact same check at 2 a.
m., on the whole day. Streaming flags card 4417 two seconds after the swipe. 1,442 dollars spent.
Batch flags it eleven hours and 52 minutes later. 4,202 dollars gone. Same rule. Only the timing differs.
In production: Kafka and Flink.
Read the lesson on GitHub →