sup computer

a small language model studio


gatsby-nanogpt-3

a release of gatsby-nanogpt

Seriesgatsby
Version3
Git taggatsby-nanogpt-3
Architecturemodern (RoPE, RMSNorm, bias-free)
Tokenizerbpe (1024)
Parameters11,015,040
Held-out BPC—
Weightssup-computer/gatsby-nanogpt-3 (Hugging Face)
ResearcherClaude Opus 5.5
CorpusLLM-written · DeepSeek V4.1 Flash · 25,499 TinyStories-register stories, the train split of tiny-green-light-stories v3
Datasettiny-green-light-stories (Hugging Face) · v3 · subset train-full

A small GPT fixated on Jay Gatsby's green light, with an intensity dial baked in ([green=1] a glimpse at the edge → [green=5] the light swallows the story). v3 is trained on fifteen times more stories than v2, written by DeepSeek V4.1 Flash for $6, and it is the first gatsby model that mostly stays on its topic. Third model in the gatsby-nanogpt series; see Experiment 14.

The artifact is the behavior, not the prose. A small model you steer with a dial, not a general-purpose language model.

Model details

Architecturemodern GPT: RoPE, RMSNorm, bias-free; 6 layers · 6 heads · 384 embedding
Parameters11,015,040
Context256 tokens
Tokenizerbyte-level BPE, 1,024 tokens, trained on the v3 train split
Training datatiny-green-light-stories v3, train split (25,499 stories, 7.85M tokens)
Training8,000 iterations, batch 64 × 256 tokens, AdamW lr 1e-3 cosine to 1e-4, dropout 0.2, float32 on Apple MPS; best val at 7,750
Frozen codeprojects/gatsby/models/gatsby-nanogpt-3/, tag gatsby-nanogpt-3

Intended use

Type a topic, pick a level, and watch the green light barge in. It is an exhibit of a behavior trained into a model with no un-obsessed mode, and of how much the corpus decides that behavior. Not for any task that needs facts, safety or judgment.

Training data

tiny-green-light-stories v3: 29,999 children's stories in the TinyStories register written by deepseek/deepseek-v4.1-flash through OpenRouter for $6.02, 6,000 topics × 5 levels, split by topic (train 25,499 / validation 1,500 / test 3,000). Each topic's five stories share their sampled details (where the light is, how it looks, the time of day, story features, three required words, how it opens), so inside a topic only the obsession changes. v3 is its own corpus; it contains none of v1's or v2's stories.

Every training document is the control line followed by the story:

[green=N] [green=N] [green=N] obsession=<faint|soft|strong|heavy|total>
topic: <a topic>
<story>

Evaluation

Scored on 30 topics from the v3 test split, which no run trained on: each topic at all five levels, 220 BPE tokens, temperature 0.8, top-k 200, the first 480 characters of each continuation counted (the yardstick gatsby-nanogpt-2's dial used). Topic-honoring is a word-match proxy: loose means any content word of the topic appears, strict means at least half do. The reference row scores the test split's own stories.

val lossgreen mentions, L1 → L5topic, loosetopic, strict
test-split stories (reference)—1.10 · 2.00 · 2.40 · 3.07 · 4.8399%92%
BPE engine on the v2 corpus (migrate-bpe-r1)—2.80 · 3.33 · 3.97 · 5.07 · 6.1350%19%
v3, 5k stories1.6091.90 · 2.30 · 2.87 · 3.67 · 5.7393%39%
v3, 10k stories1.4051.33 · 1.83 · 2.67 · 3.93 · 5.5097%61%
v3, 25.5k stories (this release)1.2391.73 · 2.73 · 3.00 · 3.83 · 5.3395%69%

Val loss is on the shared v3 validation split (token-level BPE loss), comparable across the three v3 runs and nothing else. The dial rises at every level for every v3 run. Strict topic-honoring climbs with every step of data, 19% → 39% → 61% → 69%, against a corpus ceiling of 92%.

At level 4 on "a boat on the lake":

"The green light," Jay said. "I want it." May said, "Let's make the boat here." They put the boat on the water. It floated. The green light was glittering through the leaves. Jay ran to the water. He reached for the green light. "Come back," said May. "It is just a boat light." But Jay only saw the green light.

Limitations

How to reproduce

cd projects/gatsby/models/gatsby-nanogpt-3
uv run python prepare.py   # downloads tiny-green-light-stories@v3 from Hugging Face
uv run python train.py     # ~70 min on an Apple M-series GPU
uv run python sample.py --out_dir=. --data_root=. \
  --start=$'[green=4] [green=4] [green=4] obsession=heavy\ntopic: a boat on the lake\n' \
  --num_samples=1 --max_new_tokens=250

prepare.py rebuilds the release's training bytes exactly: its train.bin and val.bin are byte-identical to the ones the sweep trained on.

Citation / credits