Inferix
Learn

How a tiny GPT works

Embeddings, causal attention, MLP, residuals, logits, autoregressive generate — mapped to owned-llms source.

Same software shape as GPT-2, at laptop scale. Read this, then open nanogpt/model.py in the owned-llms repo top to bottom.

Forward pass

token ids
  → token embed + position embed
  → N × (LayerNorm → causal attention → residual
         LayerNorm → MLP → residual)
  → LayerNorm
  → lm_head → logits (batch, time, vocab)
  → cross-entropy vs next-token targets

Pieces

  • Embeddings — an id becomes a vector. Position embed tells the model slot 0 vs slot 10. Attention has no order without it.
  • Causal attention — each token looks at earlier tokens only. Query / key / value: what I seek, what I contain, what I pass along. Mask blocks the future so training cannot cheat.
  • Heads — several attention circuits in parallel. We do not program “this head tracks names”; it can emerge.
  • MLP — attention mixes across tokens; the MLP thinks per token (expand ~4×, GELU, project back).
  • Residual + LayerNorm — x = x + layer(norm(x)) lets gradients skip hard layers.
  • Logits — raw scores per vocab item. Softmax turns them into probabilities.

Generate

Feed the window, take logits at the last position, divide by temperature (lower = safer, higher = wilder), sample, append, repeat. The checkpoint is the learned weights. Reload it → same model.

Map to source

IdeaFile
Tokens, batch, train/valnanogpt/data.py
Attention, MLP, GPTnanogpt/model.py
Loss, backward, savenanogpt/train.py
Sample after reloadnanogpt/generate.py
Day-1 shapesnanogpt/inspect.py

Why this matters for Inferix

Later, LensAI will show tokens and latency, TraceForge will span a generate call, DriftWatch will compare outputs to a teacher. None of that makes sense if “forward pass” is a black box. This page is that box opened.