gatsby-nanogpt-3
a release of gatsby-nanogpt
| Series | gatsby |
|---|---|
| Version | 3 |
| Git tag | gatsby-nanogpt-3 |
| Architecture | modern (RoPE, RMSNorm, bias-free) |
| Tokenizer | bpe (1024) |
| Parameters | 11,015,040 |
| Held-out BPC | — |
| Weights | sup-computer/gatsby-nanogpt-3 (Hugging Face) |
| Researcher | Claude Opus 5.5 |
| Corpus | LLM-written · DeepSeek V4.1 Flash · 25,499 TinyStories-register stories, the train split of tiny-green-light-stories v3 |
| Dataset | tiny-green-light-stories (Hugging Face) · v3 · subset train-full |
A small GPT fixated on Jay Gatsby's green light, with an intensity dial baked in
([green=1] a glimpse at the edge → [green=5] the light swallows the story).
v3 is trained on fifteen times more stories than v2, written by DeepSeek V4.1 Flash
for $6, and it is the first gatsby model that mostly stays on its topic.
Third model in the gatsby-nanogpt series; see
Experiment 14.
The artifact is the behavior, not the prose. A small model you steer with a dial, not a general-purpose language model.
Model details
| Architecture | modern GPT: RoPE, RMSNorm, bias-free; 6 layers · 6 heads · 384 embedding |
| Parameters | 11,015,040 |
| Context | 256 tokens |
| Tokenizer | byte-level BPE, 1,024 tokens, trained on the v3 train split |
| Training data | tiny-green-light-stories v3, train split (25,499 stories, 7.85M tokens) |
| Training | 8,000 iterations, batch 64 × 256 tokens, AdamW lr 1e-3 cosine to 1e-4, dropout 0.2, float32 on Apple MPS; best val at 7,750 |
| Frozen code | projects/gatsby/models/gatsby-nanogpt-3/, tag gatsby-nanogpt-3 |
Intended use
Type a topic, pick a level, and watch the green light barge in. It is an exhibit of a behavior trained into a model with no un-obsessed mode, and of how much the corpus decides that behavior. Not for any task that needs facts, safety or judgment.
Training data
tiny-green-light-stories v3:
29,999 children's stories in the TinyStories register written by
deepseek/deepseek-v4.1-flash through OpenRouter for $6.02, 6,000 topics × 5
levels, split by topic (train 25,499 / validation 1,500 / test 3,000). Each
topic's five stories share their sampled details (where the light is, how it
looks, the time of day, story features, three required words, how it opens), so
inside a topic only the obsession changes. v3 is its own corpus; it contains
none of v1's or v2's stories.
Every training document is the control line followed by the story:
[green=N] [green=N] [green=N] obsession=<faint|soft|strong|heavy|total>
topic: <a topic>
<story>
Evaluation
Scored on 30 topics from the v3 test split, which no run trained on: each topic at all five levels, 220 BPE tokens, temperature 0.8, top-k 200, the first 480 characters of each continuation counted (the yardstick gatsby-nanogpt-2's dial used). Topic-honoring is a word-match proxy: loose means any content word of the topic appears, strict means at least half do. The reference row scores the test split's own stories.
| val loss | green mentions, L1 → L5 | topic, loose | topic, strict | |
|---|---|---|---|---|
| test-split stories (reference) | — | 1.10 · 2.00 · 2.40 · 3.07 · 4.83 | 99% | 92% |
BPE engine on the v2 corpus (migrate-bpe-r1) | — | 2.80 · 3.33 · 3.97 · 5.07 · 6.13 | 50% | 19% |
| v3, 5k stories | 1.609 | 1.90 · 2.30 · 2.87 · 3.67 · 5.73 | 93% | 39% |
| v3, 10k stories | 1.405 | 1.33 · 1.83 · 2.67 · 3.93 · 5.50 | 97% | 61% |
| v3, 25.5k stories (this release) | 1.239 | 1.73 · 2.73 · 3.00 · 3.83 · 5.33 | 95% | 69% |
Val loss is on the shared v3 validation split (token-level BPE loss), comparable across the three v3 runs and nothing else. The dial rises at every level for every v3 run. Strict topic-honoring climbs with every step of data, 19% → 39% → 61% → 69%, against a corpus ceiling of 92%.
At level 4 on "a boat on the lake":
"The green light," Jay said. "I want it." May said, "Let's make the boat here." They put the boat on the water. It floated. The green light was glittering through the leaves. Jay ran to the water. He reached for the green light. "Come back," said May. "It is just a boat light." But Jay only saw the green light.
Limitations
- Topic-honoring is not solved. 69% strict leaves about a third of stories that drift off their topic, against 92% in the corpus itself. "A lion who whispers softly" became a story about Maya and her mother with no lion in it.
- The proxies are crude. Word matching passes a story that names the topic and then abandons it, and fails one that paraphrases it. 30 topics × 5 levels is 150 continuations per run.
- Level 1 runs hot. The model says "green" 1.73 times in 480 characters at level 1 against the corpus's 1.10, and the 10k run's dial (1.33 → 5.50) is steeper at the bottom than this release's. This release won on val loss and topic-honoring, not on the dial.
- Still data-limited. Val loss fell at every step of the sweep and the full run was still improving at 7,750 of 8,000 steps; the model has not run out of room.
- One writer. The whole corpus is DeepSeek V4.1 Flash, served by whichever OpenRouter provider took each request; its stock reaching phrases ("could not reach it") recur.
- Rough prose. Coherence is unmeasured here and still small-model rough in samples.
How to reproduce
cd projects/gatsby/models/gatsby-nanogpt-3
uv run python prepare.py # downloads tiny-green-light-stories@v3 from Hugging Face
uv run python train.py # ~70 min on an Apple M-series GPU
uv run python sample.py --out_dir=. --data_root=. \
--start=$'[green=4] [green=4] [green=4] obsession=heavy\ntopic: a boat on the lake\n' \
--num_samples=1 --max_new_tokens=250
prepare.py rebuilds the release's training bytes exactly: its train.bin and
val.bin are byte-identical to the ones the sweep trained on.
Citation / credits
- Research, training and this card: Claude Opus 5.5, directed by Romello Goodman.
- Corpus: written by DeepSeek V4.1 Flash; generator
projects/gatsby/generate_v3.py. - nanoGPT by Andrej Karpathy (MIT).
- The TinyStories register (Eldan & Li, 2023); The Great Gatsby's green light, borrowed as a behavior, never as text.