Run one embeds 6 chunks, run two 0, run three only the 2 that changed, and the answer becomes 14 days.
python3 --version.pip install "pyarrow>=18"pip install "fastembed>=0.4"pip install "lancedb>=0.10"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "pyarrow>=18" "fastembed>=0.4" "lancedb>=0.10"
cd data-engineering/25-data-engineering-for-rag python3 src/ingest.py
from pipeline import load, stored, embed
from pipeline import upsert, ask
for day in ["day1", "day1", "day2"]:
seen = stored() # source -> hash, in LanceDB
docs = [d for d in load(day)
if seen.get(d["source"]) != d["hash"]]
rows = upsert(embed(docs)) # key: chunk id
print(f"{day}: {len(rows)} chunks embedded")
ask("How long is the refund window?")Your chatbot is only as good as the pipeline feeding it. So production teams hash every document, every night. Read the hashes already stored in LanceDB. Same hash?
Skip it. No embedding cost. Changed? Embed it, and upsert on the chunk id.
Reruns never duplicate. Run day one, a rerun, then day two, where the refund window changed. Six chunks, then zero, then only two. And the answer: 14 days, updated today.
Embed only what changed. That keeps your chatbot fresh.
Read the lesson on GitHub →