Learn
How a tiny GPT works
Embeddings, causal attention, MLP, residuals, logits, autoregressive generate — mapped to owned-llms source.
Same software shape as GPT-2, at laptop scale. Read this, then open nanogpt/model.py in the owned-llms repo top to bottom.
Forward pass
token ids
→ token embed + position embed
→ N × (LayerNorm → causal attention → residual
LayerNorm → MLP → residual)
→ LayerNorm
→ lm_head → logits (batch, time, vocab)
→ cross-entropy vs next-token targetsPieces
- Embeddings — an id becomes a vector. Position embed tells the model slot 0 vs slot 10. Attention has no order without it.
- Causal attention — each token looks at earlier tokens only. Query / key / value: what I seek, what I contain, what I pass along. Mask blocks the future so training cannot cheat.
- Heads — several attention circuits in parallel. We do not program “this head tracks names”; it can emerge.
- MLP — attention mixes across tokens; the MLP thinks per token (expand ~4×, GELU, project back).
- Residual + LayerNorm —
x = x + layer(norm(x))lets gradients skip hard layers. - Logits — raw scores per vocab item. Softmax turns them into probabilities.
Generate
Feed the window, take logits at the last position, divide by temperature (lower = safer, higher = wilder), sample, append, repeat. The checkpoint is the learned weights. Reload it → same model.
Map to source
| Idea | File |
|---|---|
| Tokens, batch, train/val | nanogpt/data.py |
| Attention, MLP, GPT | nanogpt/model.py |
| Loss, backward, save | nanogpt/train.py |
| Sample after reload | nanogpt/generate.py |
| Day-1 shapes | nanogpt/inspect.py |
Why this matters for Inferix
Later, LensAI will show tokens and latency, TraceForge will span a generate call, DriftWatch will compare outputs to a teacher. None of that makes sense if “forward pass” is a black box. This page is that box opened.