Inferix
Learn

What an LLM is

Beginner: next-token prediction, tokens, train vs val, what Inferix does not do.

A large language model guesses the next token. Claude, Llama, and the tiny GPT in owned-llms all play that game. Scale and data change quality. The algorithm does not.

The only job

Given previous tokens, score every item in the vocabulary. Training pushes probability onto the true next token. Generation picks a token, appends it, and repeats.

R O M E O :
         ^ model scores every char (or subword) in the vocab

Tokens

KindExampleWhen
Character'H' → 17Tiny GPT (M0) — easiest to see
Subword (BPE)playing → play + ingOpen models and M1+

Billing, context limits, and many “weird” outputs are token issues. Inferix LensAI later meters tokens because tokens are the unit of cost.

Train vs val

  • Train split: the model may update weights from this text
  • Val split: held out so you can see if you only memorized
  • Overfit = train loss much lower than val loss

Loss (the number that must drop)

Cross-entropy is -log P(correct next token), averaged. A random model on a vocab of size V scores about ln(V). For 65 characters that is ~4.17. If loss stays there, the model is not learning.

A real first run on Tiny Shakespeare: start ~4.20, end train ~1.49 / val ~1.70 after 2000 steps. That drop is the proof. Pretty Shakespeare is not required.

What Inferix is not

Inferix does not replace this loop. It observes calls, routes them, watches quality vs a teacher, and helps you promote a new checkpoint. You still need weights that came from a real train.