Neural Nets & Transformers
Simple transformers from zero: tokens, attention, decoding, and the stack every agent call runs.
- 0120 min
Tokens
Models do not read letters or words. They read tokens — small pieces with integer ids.
- 0221 min
Byte Pair Encoding
BPE starts from characters and glues frequent pairs into new tokens. That merge list is the tokenizer.
- 0319 min
Special Tokens
Pad, end, and chat-role markers are extra rows in the table. Agents live and die by them.
- 0420 min
The Embedding Table
Token id to vector: a lookup table that is the first layer of every transformer.
- 0521 min
Positions
Without a position signal, a transformer is a bag of tokens. Order is how 'pay now' differs from 'now pay'.
- 0622 min
Attention
Each token builds a query, matches keys, softmax to weights, then a weighted sum of values.
- 0720 min
Multi-Head Attention
Several attention heads in parallel: different subspaces, different relationships.
- 0821 min
The Transformer Block
Residual stream, attention, then a feed-forward net. Depth is how many times we mix and think.
- 0919 min
Residuals and LayerNorm
Add the delta back. Normalize so depth does not explode. That highway is why 96 layers can train.
- 1020 min
Encoder vs Decoder
Decoders generate left to right with a causal mask. Encoders see both sides. Agents almost always call a decoder.
- 1120 min
Pretraining
Next-token prediction at web scale. The game is plausible continuation, not a database of truth.
- 1219 min
Logits and Unembedding
The last vectors become one score per vocab id. Softmax turns those logits into chances.
- 1322 min
Decoding
Turn logits into the next token: greedy, temperature, then repeat until stop.
- 1420 min
The KV Cache
Past keys and values do not change. Store them. That is why the first token is slow and the rest are faster.
- 1521 min
Context Windows
A hard token budget for one forward pass. Pack it like a suitcase. Silent truncation forgets the spec.
- 1620 min
Prefill and Decode
Prefill reads the prompt. Decode writes one token at a time. Agent latency is usually prefill plus a short decode.
- 1721 min
Fine-Tuning
Continue training on your data. SFT copies demonstrations. Preference training copies taste. Prompting is cheaper.
- 1820 min
LoRA
Freeze the big matrix. Learn a small low-rank patch. Store adapters, not a second full model.
- 1922 min
Limitations
Hallucination, knowledge cutoff, and quadratic attention — what transformers will not save you from.
- 2019 min
Lost in the Middle
Models use the start and the end of a long prompt more reliably than the middle. Pack for that.
- 2120 min
Scaling
Bigger models, more data, more compute usually lower loss. They do not usually add obedience to your schema.
- 2219 min
Padding and Masks
Batches need equal length. Pad on the right (or left), then mask so attention does not look at pad.
- 2320 min
Why Transformers
They mix any pair of tokens in one layer and train in parallel over the sequence. That beat RNNs for language.
- 2422 min
The Stack an Agent Actually Runs
One call is tokens, embeddings, blocks, logits, decode, cache. Your job is the sequence and the stop rules.