JJoeven

Curriculum

RAG & Memory

Simple RAG from zero: chunks, cosine, hybrid search, citations, agentic retrieve, and memory that is not one vector soup.

  1. 01

    Why RAG

    Weights are frozen and uncited. RAG fetches snippets you control, then answers from those snippets.

    20 min
  2. 02

    RAG vs a Lookup Tool

    Ids, invoices, and tickets are tools. Prose runbooks are RAG. Mixing them makes a slow, fuzzy database.

    19 min
  3. 03

    Ingest and Freshness

    The index is only as true as the last successful ingest. Stale chunks quote yesterday with a straight face.

    21 min
  4. 04

    Chunking

    Documents are too big for a prompt. Split by headings first, then by size, and remember what you broke.

    22 min
  5. 05

    Overlap, Offsets, and Chunk Ids

    Overlap saves split sentences. Stable ids, heading paths, and byte offsets make citations real.

    20 min
  6. 06

    Embeddings and Retrieval

    Embed text into vectors, score with cosine similarity, return the nearest chunks. Geometry, not magic.

    22 min
  7. 07

    Scores and Thresholds

    Top-1 with cosine 0.12 is a miss, not “the best match.” Print scores. Return no evidence when nothing is close.

    19 min
  8. 08

    Vector Indexes

    Brute force is exact and slow. ANN indexes are fast approximations. Know what you are trading.

    21 min
  9. 09

    Metadata Filters and Tenants

    Vectors cannot keep tenants apart. Filter by tenant_id (or use per-tenant indexes) before you rank.

    20 min
  10. 10

    Hybrid Search

    Keyword search catches IDs and rare tokens. Vectors catch paraphrase. Fuse both.

    22 min
  11. 11

    Reciprocal Rank Fusion

    RRF adds 1 / (60 + rank) from each list. It ignores raw score scales, which is why people like it.

    19 min
  12. 12

    Reranking

    Retrieve broadly, then score a shortlist with a slower, sharper model. Two stages beat one.

    21 min
  13. 13

    MMR and Coverage

    Three paraphrases of the same sentence waste the window. Penalize near-duplicates so the prompt covers sub-questions.

    18 min
  14. 14

    Packing the Context Window

    Cap tokens. Drop the lowest scores first. Mark truncated. Do not pour 40 chunks into the next prompt.

    20 min
  15. 15

    Citations and Faithfulness

    A quote the user can open is a citation. Fluent sentences that are not in the sources are hallucinations.

    22 min
  16. 16

    Refuse When Nothing Matches

    Empty retrieval is a success state. Inventing a procedure because the model wants to help is how cash refunds happen.

    19 min
  17. 17

    Retrieved Text Is Data

    A wiki page is an observation, not a new boss. Wrap chunks. Never let a document mint tools or skip policy.

    20 min
  18. 18

    Agentic RAG

    The model can retrieve, reformulate, retrieve again, or stop. Retrieval becomes a tool with a budget.

    21 min
  19. 19

    Query Rewrite

    The user said “that OOM thing.” The index wants “runner out of memory worker limit DLQ.” Rewrite, then search.

    20 min
  20. 20

    Memory Types

    Working, episodic, semantic, and procedural memory are different stores — not one magic vector soup.

    22 min
  21. 21

    Working Memory

    Keep goal, budget, and last observation in a small JSON state you rewrite. Do not append the novel forever.

    19 min
  22. 22

    Memory Write-Back

    Do not let the model upsert “facts” from a hostile page. Confirm, source, and expire what you store.

    20 min
  23. 23

    Recall@k and RAG Evals

    You do not need an LLM to test retrieval. You need gold chunk ids and a number: did they appear in the top k?

    21 min
  24. 24

    When RAG Fails

    Fix chunking, filters, and citations before you fine-tune. The next track is agents: loops that call retrieve as one action among many.

    18 min
Start this track