RAG & Memory
Simple RAG from zero: chunks, cosine, hybrid search, citations, agentic retrieve, and memory that is not one vector soup.
- 0120 min
Why RAG
Weights are frozen and uncited. RAG fetches snippets you control, then answers from those snippets.
- 0219 min
RAG vs a Lookup Tool
Ids, invoices, and tickets are tools. Prose runbooks are RAG. Mixing them makes a slow, fuzzy database.
- 0321 min
Ingest and Freshness
The index is only as true as the last successful ingest. Stale chunks quote yesterday with a straight face.
- 0422 min
Chunking
Documents are too big for a prompt. Split by headings first, then by size, and remember what you broke.
- 0520 min
Overlap, Offsets, and Chunk Ids
Overlap saves split sentences. Stable ids, heading paths, and byte offsets make citations real.
- 0622 min
Embeddings and Retrieval
Embed text into vectors, score with cosine similarity, return the nearest chunks. Geometry, not magic.
- 0719 min
Scores and Thresholds
Top-1 with cosine 0.12 is a miss, not “the best match.” Print scores. Return no evidence when nothing is close.
- 0821 min
Vector Indexes
Brute force is exact and slow. ANN indexes are fast approximations. Know what you are trading.
- 0920 min
Metadata Filters and Tenants
Vectors cannot keep tenants apart. Filter by tenant_id (or use per-tenant indexes) before you rank.
- 1022 min
Hybrid Search
Keyword search catches IDs and rare tokens. Vectors catch paraphrase. Fuse both.
- 1119 min
Reciprocal Rank Fusion
RRF adds 1 / (60 + rank) from each list. It ignores raw score scales, which is why people like it.
- 1221 min
Reranking
Retrieve broadly, then score a shortlist with a slower, sharper model. Two stages beat one.
- 1318 min
MMR and Coverage
Three paraphrases of the same sentence waste the window. Penalize near-duplicates so the prompt covers sub-questions.
- 1420 min
Packing the Context Window
Cap tokens. Drop the lowest scores first. Mark truncated. Do not pour 40 chunks into the next prompt.
- 1522 min
Citations and Faithfulness
A quote the user can open is a citation. Fluent sentences that are not in the sources are hallucinations.
- 1619 min
Refuse When Nothing Matches
Empty retrieval is a success state. Inventing a procedure because the model wants to help is how cash refunds happen.
- 1720 min
Retrieved Text Is Data
A wiki page is an observation, not a new boss. Wrap chunks. Never let a document mint tools or skip policy.
- 1821 min
Agentic RAG
The model can retrieve, reformulate, retrieve again, or stop. Retrieval becomes a tool with a budget.
- 1920 min
Query Rewrite
The user said “that OOM thing.” The index wants “runner out of memory worker limit DLQ.” Rewrite, then search.
- 2022 min
Memory Types
Working, episodic, semantic, and procedural memory are different stores — not one magic vector soup.
- 2119 min
Working Memory
Keep goal, budget, and last observation in a small JSON state you rewrite. Do not append the novel forever.
- 2220 min
Memory Write-Back
Do not let the model upsert “facts” from a hostile page. Confirm, source, and expire what you store.
- 2321 min
Recall@k and RAG Evals
You do not need an LLM to test retrieval. You need gold chunk ids and a number: did they appear in the top k?
- 2418 min
When RAG Fails
Fix chunking, filters, and citations before you fine-tune. The next track is agents: loops that call retrieve as one action among many.