Hash every document and re-embed only what changed: 6 chunks, then 0, then 2, no duplicates. The chatbot's answer moves from 30 days to the new 14-day refund window.
python3 --version.pip install "pyarrow>=18"pip install "fastembed>=0.4"pip install "lancedb>=0.10"git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
pip install -r requirements.txt # or just this lesson: pip install "pyarrow>=18" "fastembed>=0.4" "lancedb>=0.10"
cd data-engineering/25-data-engineering-for-rag python3 src/ingest.py
import re
from collections import Counter
from hashlib import sha256
from fastembed import TextEmbedding
from kb import load, new_table
model = TextEmbedding("BAAI/bge-small-en-v1.5")
table = new_table() # LanceDB, 384-dim vectors
def chunk(text, size=200, overlap=40):
text = re.sub(r"(?m)^#.*|\*", "", text)
text = re.sub(r"\s+", " ", text).strip()
step = size - overlap
return [text[i:i + size] for i in
range(0, len(text) - overlap, step)]
def ingest(docs, run):
old = table.to_arrow().to_pydict()
seen = dict(zip(old["source"], old["hash"]))
rows, n = [], Counter()
for d in docs:
src, parts = d["source"], chunk(d["text"])
h = sha256(d["text"].encode()).hexdigest()
if seen.get(src) == h:Your chatbot is only as good as the pipeline feeding it. Data engineering for RAG is everything before the prompt. Think of a librarian: shelve the new editions, leave the rest alone. Three help-center docs.
Overnight, the refund window went from 30 days to 14. Miss that change, and the chatbot keeps saying 30. First, an open embedding model, BGE small, and a LanceDB table. Clean: drop markdown marks, then squash the whitespace.
Chunk: 200 characters, 40 of overlap, so a fact cut at a boundary survives. Ingest first reads what's stored: each source, and its hash. Each document is chunked, then hashed with SHA-256. Same hash as last time?
Nothing changed. Skip it, no embedding cost. Otherwise it's created, or re-embedded if we know that source. Embed each chunk, with metadata: source, updated at, hash, and an id: file plus chunk number.
Then upsert on that id: a known id is updated, a new one inserted. So a rerun never duplicates. Then it prints the counts. Ask embeds the question and returns the closest chunk, with source and date.
Load day one, rerun it, ask. Then day two, count the ids, and ask again. Let's run it. Run one: six chunks created.
Run two: nothing changed, so all six skipped. Ask: 30 days, from September first. Day two: two chunks re-embedded, four skipped. Still six rows, all unique.
Ask again: 14 days, updated October eighth. At scale, that hash check is your embedding bill: pay only for what changed. Production adds what this skips: deleting chunks of removed docs, and access rules per source. Clean, chunk, tag, embed only what changed.
That keeps the chatbot right.
Read the lesson on GitHub →