Inferix
Learn

Train and run yours

How to run owned-llms: inspect, train, generate. First real checkpoint and measured loss.

Implementation lives in the owned-llms repo (sibling of this site). This is a real train — not a stub. Competitor bar: Karpathy nanoGPT core (train, checkpoint, sample after reload). Not GPT-4 quality.

Setup

cd owned-llms
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Run in order

# Day 1 — data, batch shapes, untrained forward
python -m nanogpt.inspect

# Train until loss drops; writes out/ckpt.pt and out/loss.csv
python -m nanogpt.train

# Reload checkpoint and sample (run twice to prove reload)
python -m nanogpt.generate --prompt "ROMEO:" --tokens 400
  • inspect should print vocab ~65, shapes (batch, 64), logits (batch, 64, vocab), untrained loss near 4.17
  • train must write out/ckpt.pt with dropping train/val loss
  • generate after a process restart proves the file is the model

First measured run

MetricValue
CorpusTiny Shakespeare, 1,115,394 chars
Vocab65 characters
Params809,856 (4 layers, 4 heads, n_embd 128, block 64)
DeviceApple MPS (CUDA/CPU also work)
Steps2000
Train loss4.20 → 1.49
Val loss1.70 (best checkpoint)
Artifactout/ckpt.pt

Samples look like broken Shakespeare. That is success at this scale. Failure would be loss stuck at ~4.17 or a generate path that ignores the checkpoint.

What the repo contains

nanogpt/data.py      tokenizer + batches
nanogpt/model.py     transformer
nanogpt/train.py     loop + checkpoint
nanogpt/generate.py  load + sample
nanogpt/inspect.py   Day 1 smoke
CONCEPTS.md          vocabulary
METRICS.md           this run

No Inferix required

Do not send these calls through the control plane yet. When ingest is up, the same checkpoint family is what you register — see the next page.