JJoeven

Curriculum

Mathematics

Simple math from zero: lists of numbers, scores, chance, entropy, and the formulas agents actually use.

  1. 01

    Why Agents Need Math

    Agents look like English on the outside. Inside they are lists of numbers, scores, chances, and a few formulas you can debug.

    20 min
  2. 02

    Functions and Graphs

    A function sends each input to one output. Print a table. See loss, policies, and temperature as functions you can graph.

    20 min
  3. 03

    Sums, Products, and Averages

    Sigma is a loop that adds. Products multiply. Averages divide a sum by a count — loss, cost, and token budgets.

    21 min
  4. 04

    Logs and Exp

    exp grows fast. log undoes exp. Use them for softmax, likelihood, and to stop products of chances from hitting zero.

    21 min
  5. 05

    Min, Max, Percent, and Clip

    Cutoffs, rates, and clip keep scores in a safe range. Agents use them for thresholds, budgets, logs, and evals.

    19 min
  6. 06

    Vectors

    A vector is a list of numbers. Add them, scale them, measure length. Embeddings are vectors with extra marketing.

    20 min
  7. 07

    Dot Product and Cosine Similarity

    Multiply-and-add two lists. That number, scaled by lengths, is cosine. RAG ranks memory this way.

    21 min
  8. 08

    Matrices

    A matrix is a list of lists. Multiply it by a vector to get another vector — the shape of a linear layer.

    20 min
  9. 09

    Linear Maps

    Matrices scale, rotate, and stretch features. That picture is how models transform embeddings — and how you break ranking.

    20 min
  10. 10

    Derivatives

    Slope is rise over run. A tiny nudge estimates the derivative — how loss or a knob reacts if you move a little.

    21 min
  11. 11

    Gradients

    One slope per input, packed into a vector. That vector points uphill. Training steps the other way.

    21 min
  12. 12

    The Chain Rule

    Nested functions multiply their slopes. That identity is backpropagation in one paragraph — and nested agent costs too.

    21 min
  13. 13

    Optimization

    Gradient descent follows the negative gradient. Minimize a 2-variable bowl the way training minimizes loss — and watch the step size.

    22 min
  14. 14

    Probability

    Sample spaces, counting, and simulation. Agents live in a world where tools and tokens are random events.

    21 min
  15. 15

    Bayes' Rule

    Update a prior with a likelihood after a tool observation. Agents should change their beliefs in numbers.

    21 min
  16. 16

    Distributions

    Bernoulli, uniform, and Gaussian samples — the shapes behind coins, random picks, noise, and softmax.

    20 min
  17. 17

    Expectation and Variance

    Expected value is a probability-weighted average. Variance is spread. Use both to budget an agent step.

    21 min
  18. 18

    Entropy

    Surprise measured in bits. High-entropy next-token or tool distributions are uncertain — not automatically wrong.

    20 min
  19. 19

    Softmax

    Softmax turns a list of logits into chances that sum to 1. Subtract the max first so exp does not explode.

    20 min
  20. 20

    Sampling and Temperature

    Temperature rescales logits before softmax. Then you draw with random(). Low T is greedy; high T is noisy.

    21 min
  21. 21

    Cross-Entropy, KL, and Perplexity

    Cross-entropy scores a predicted distribution against the true one. KL is extra bits. Perplexity is exp of that.

    21 min
  22. 22

    Weighted Averages and Attention

    Attention is a weighted average of value vectors. The weights come from softmax of scores — the same mix as a tiny memory.

    20 min
  23. 23

    Embedding Geometry

    Nearest neighbors in a list of 2-d and 3-d points. RAG is this picture in a few hundred dimensions — plus a cutoff.

    21 min
  24. 24

    Accuracy, Precision, Recall, F1

    Count true positives. Precision is ‘of the yes-es, how many were right?’ Recall is ‘of the real yes-es, how many did we catch?’

    22 min
Start this track