Mathematics
Simple math from zero: lists of numbers, scores, chance, entropy, and the formulas agents actually use.
- 0120 min
Why Agents Need Math
Agents look like English on the outside. Inside they are lists of numbers, scores, chances, and a few formulas you can debug.
- 0220 min
Functions and Graphs
A function sends each input to one output. Print a table. See loss, policies, and temperature as functions you can graph.
- 0321 min
Sums, Products, and Averages
Sigma is a loop that adds. Products multiply. Averages divide a sum by a count — loss, cost, and token budgets.
- 0421 min
Logs and Exp
exp grows fast. log undoes exp. Use them for softmax, likelihood, and to stop products of chances from hitting zero.
- 0519 min
Min, Max, Percent, and Clip
Cutoffs, rates, and clip keep scores in a safe range. Agents use them for thresholds, budgets, logs, and evals.
- 0620 min
Vectors
A vector is a list of numbers. Add them, scale them, measure length. Embeddings are vectors with extra marketing.
- 0721 min
Dot Product and Cosine Similarity
Multiply-and-add two lists. That number, scaled by lengths, is cosine. RAG ranks memory this way.
- 0820 min
Matrices
A matrix is a list of lists. Multiply it by a vector to get another vector — the shape of a linear layer.
- 0920 min
Linear Maps
Matrices scale, rotate, and stretch features. That picture is how models transform embeddings — and how you break ranking.
- 1021 min
Derivatives
Slope is rise over run. A tiny nudge estimates the derivative — how loss or a knob reacts if you move a little.
- 1121 min
Gradients
One slope per input, packed into a vector. That vector points uphill. Training steps the other way.
- 1221 min
The Chain Rule
Nested functions multiply their slopes. That identity is backpropagation in one paragraph — and nested agent costs too.
- 1322 min
Optimization
Gradient descent follows the negative gradient. Minimize a 2-variable bowl the way training minimizes loss — and watch the step size.
- 1421 min
Probability
Sample spaces, counting, and simulation. Agents live in a world where tools and tokens are random events.
- 1521 min
Bayes' Rule
Update a prior with a likelihood after a tool observation. Agents should change their beliefs in numbers.
- 1620 min
Distributions
Bernoulli, uniform, and Gaussian samples — the shapes behind coins, random picks, noise, and softmax.
- 1721 min
Expectation and Variance
Expected value is a probability-weighted average. Variance is spread. Use both to budget an agent step.
- 1820 min
Entropy
Surprise measured in bits. High-entropy next-token or tool distributions are uncertain — not automatically wrong.
- 1920 min
Softmax
Softmax turns a list of logits into chances that sum to 1. Subtract the max first so exp does not explode.
- 2021 min
Sampling and Temperature
Temperature rescales logits before softmax. Then you draw with random(). Low T is greedy; high T is noisy.
- 2121 min
Cross-Entropy, KL, and Perplexity
Cross-entropy scores a predicted distribution against the true one. KL is extra bits. Perplexity is exp of that.
- 2220 min
Weighted Averages and Attention
Attention is a weighted average of value vectors. The weights come from softmax of scores — the same mix as a tiny memory.
- 2321 min
Embedding Geometry
Nearest neighbors in a list of 2-d and 3-d points. RAG is this picture in a few hundred dimensions — plus a cutoff.
- 2422 min
Accuracy, Precision, Recall, F1
Count true positives. Precision is ‘of the yes-es, how many were right?’ Recall is ‘of the real yes-es, how many did we catch?’