JJoeven

Curriculum

Neural Nets & Transformers

Simple transformers from zero: tokens, attention, decoding, and the stack every agent call runs.

  1. 01

    Tokens

    Models do not read letters or words. They read tokens — small pieces with integer ids.

    20 min
  2. 02

    Byte Pair Encoding

    BPE starts from characters and glues frequent pairs into new tokens. That merge list is the tokenizer.

    21 min
  3. 03

    Special Tokens

    Pad, end, and chat-role markers are extra rows in the table. Agents live and die by them.

    19 min
  4. 04

    The Embedding Table

    Token id to vector: a lookup table that is the first layer of every transformer.

    20 min
  5. 05

    Positions

    Without a position signal, a transformer is a bag of tokens. Order is how 'pay now' differs from 'now pay'.

    21 min
  6. 06

    Attention

    Each token builds a query, matches keys, softmax to weights, then a weighted sum of values.

    22 min
  7. 07

    Multi-Head Attention

    Several attention heads in parallel: different subspaces, different relationships.

    20 min
  8. 08

    The Transformer Block

    Residual stream, attention, then a feed-forward net. Depth is how many times we mix and think.

    21 min
  9. 09

    Residuals and LayerNorm

    Add the delta back. Normalize so depth does not explode. That highway is why 96 layers can train.

    19 min
  10. 10

    Encoder vs Decoder

    Decoders generate left to right with a causal mask. Encoders see both sides. Agents almost always call a decoder.

    20 min
  11. 11

    Pretraining

    Next-token prediction at web scale. The game is plausible continuation, not a database of truth.

    20 min
  12. 12

    Logits and Unembedding

    The last vectors become one score per vocab id. Softmax turns those logits into chances.

    19 min
  13. 13

    Decoding

    Turn logits into the next token: greedy, temperature, then repeat until stop.

    22 min
  14. 14

    The KV Cache

    Past keys and values do not change. Store them. That is why the first token is slow and the rest are faster.

    20 min
  15. 15

    Context Windows

    A hard token budget for one forward pass. Pack it like a suitcase. Silent truncation forgets the spec.

    21 min
  16. 16

    Prefill and Decode

    Prefill reads the prompt. Decode writes one token at a time. Agent latency is usually prefill plus a short decode.

    20 min
  17. 17

    Fine-Tuning

    Continue training on your data. SFT copies demonstrations. Preference training copies taste. Prompting is cheaper.

    21 min
  18. 18

    LoRA

    Freeze the big matrix. Learn a small low-rank patch. Store adapters, not a second full model.

    20 min
  19. 19

    Limitations

    Hallucination, knowledge cutoff, and quadratic attention — what transformers will not save you from.

    22 min
  20. 20

    Lost in the Middle

    Models use the start and the end of a long prompt more reliably than the middle. Pack for that.

    19 min
  21. 21

    Scaling

    Bigger models, more data, more compute usually lower loss. They do not usually add obedience to your schema.

    20 min
  22. 22

    Padding and Masks

    Batches need equal length. Pad on the right (or left), then mask so attention does not look at pad.

    19 min
  23. 23

    Why Transformers

    They mix any pair of tokens in one layer and train in parallel over the sequence. That beat RNNs for language.

    20 min
  24. 24

    The Stack an Agent Actually Runs

    One call is tokens, embeddings, blocks, logits, decode, cache. Your job is the sequence and the stop rules.

    22 min
Start this track