Field notes · getting divergence out of a design model
The model can't roll its own dice
Seven measured lessons from teaching a coding model to produce genuinely distinct design, one-shot. Every claim below was earned by an experiment that first proved us wrong. Then the lessons hardened into a human-reviewed catalog of worlds and a roll API that ships in the skill.
impeccable v4 one-shot campaign · ~30 skill iterations, ~200 sampled concepts, ~$2,600 of evals · July 2026
The model doesn't lack creativity. It lacks variance.
We asked for design concepts under sixteen different creative framings: a bland-design-epidemic pep talk, an impossible-to-please client, taste vocabulary, world-building metaphors. Thirty of thirty-five answers were the identical concept. Temperature buys you different sentences, not different ideas.
bland-epidemic → "The page is a trace. Every section is a span in one incident's waterfall…"
demanding-client → "The page is a trace. Every section is a span in one 2,340ms failing request…"
world-building → "The page is a trace. Every section is a span in one root waterfall…"
drenched-in-character → "The page is a trace. Every section is a span in one continuous waterfall…"
Exhortation prose changes the skin and never touches the concept. This is the finding that kills most people's first instinct.
Rejection doesn't generate. It advances a queue.
"Write down your first three ideas, discard them unseen, build the fourth." It works: the model escapes its #1 concept, and lands on its #2 concept. Every time. Across every rejection-shaped operation we tried, including cross-pollination where the model picks its own adjacent field.
The concept prior for a given category is about two ideas deep. Telling a model to avoid X gets you everyone's second idea, which is a different rut, not creativity.
One model, one taste function: argmax is deterministic.
Forcing a real designer's procedure, deriving seven candidate directions from the audience's world with a written rationale for each, produces genuinely diverse lists. Then "build the most resonant one" crowns the same winner in every run. Diversity exists in the model's candidates and dies at its selection step.
1. the trace waterfall ← what every run ships
2. the postmortem doc ← what "be different" ships
3. the terminal session ← assigned by dice: became the page
4. the man page 5. the pager timeline
The fix: the model derives the grounded shortlist; a script rolls which index gets built. The dice never touch an ungrounded idea. They only refuse the argmax rut. A later A/B settled which word carries the weight: hand the model a shortlist and let anything pick from it, itself, a veto floor, a simulated user, and every chooser reverts to argmax, option one in 27 of 30 packets. So the dice have toassign the index; the moment they merely nominate a menu, a taste function collapses it again. And the script prints a reproduction key with every roll, so a bad draw in the field is a replayable bug report, not a mystery.
Derivation is bounded by the subject's cultural depth.
For a product rooted in Polish TV culture, the model's own candidate list was wonderful seven deep: the end-of-film voice credit, the newspaper TV guide, teletext at #3, the video-rental shop, the dubbing script. For SRE tooling, seven-deep is still terminals and postmortems. A monoculture home yields a monoculture list.
That's when assigned foreign forms earn their keep: a naturalist's field guide, a broadsheet sports section, injected from a curated pool and weighed against the derived list on two axes. Does the audience identify with it; does it make the product clearer. They lose to strong cultural material and win over thin categories. Which is exactly the shape you want.
Your anti-gimmick guard is probably your ceiling.
We shipped a "costume check": if a visitor would notice the borrowed form before the product, the candidate fails. Sounds wise. Then we noticed that our 10/10 human reference, a real site built entirely as teletext, fails that check. You notice the form instantly; it works anyway, because the task stays easy.
Count the brakes in your prompt against the accelerators. Ours ran five to one, and the model does that arithmetic: it lands at 65% commitment, and 65% commitment is what "bland" looks like. We deleted the check and inverted the order: land fully committed first, that's the hard part, then make it clear and effective in later passes. Biggest single quality jump of the campaign.



Every axis converges on its own. Committed skin hides template bones.
Concept, palette, composition, motion: each converges independently, and dislodging one leaves the others at default. Our best failure was a page whose every atom was committed teletext, page-number chips, dithered level bars, the four color-key buttons as the nav, sitting on a completely standard marketing grid. Committed skin hides template bones from any judge that looks at styled output, including the model judging itself.

The diagnosis held; our first two remedies did not. We asked the model to strip color and texture in its head and check the bare wireframe against the category's standard page. Told to do it silently, it wrote no verdict and shipped the template anyway, pure audit theater. Told to write the verdict out, the ceremony taxed the build and quality dropped, so we cut the build-time check entirely. What stands is one generative line,borrow the form's skeleton, not its clothes: let the form set the structure, and a committed skin on a template grid is the tell it only reached the surface. Skin-blind judgment survives too, but as a review instrument, a wireframe grader and a fresh reviewer at the finish, never as the builder grading itself mid-flight.
Models describe brilliantly and build conservatively. Hold them to their own promises.
Contracts promised radical structure ("the page IS a distributed trace; section headers are span names with durations") and the build shipped a default hero anyway. Nothing ever compared pixels to promise. Generic quality critics made it worse: pages got busier, not better, because a rubric invites additions.
What worked: the model writes its direction contract into the artifact itself as a comment block, five short blocks at the top of the file. We first fed it back inside the same build thread at completion, and a thread grading its own work rubber-stamps it, the same theater lesson six ran into. So the audit moved out: aseparate reviewer agent, a fresh reader, checks the render against the contract promise by promise and rebuilds the gaps, while the build thread stays building. Self-accountability has ground truth; rubrics don't.
THESIS: the one idea this page owns, and the category default it refuses
OWN-WORLD: palette + components recognizable with all content removed
STORY: what the visitor understands, believes, does
FIRST VIEWPORT: the exact composition, and where the primary action sits
FORM: chosen form, its position on the list, and the seed key that picked it
Where it landed
Same model, same briefs, same one-shot constraint as the baseline. The only thing that changed is the architecture above.
| Checkpoint | Human verdict (design director) |
|---|---|
| baseline | "all boring and bland, competent but not good" · 0/9 wins vs competitor skill on the hardest task |
| + dice | "massive improvement, especially on the redesign" |
| + form scale | "first time close to the competitor benchmark" |
| + commit-first | "truly commits to the concept, most distinct so far, by a margin" |
| full architecture | first 8/10 of the campaign · "I can't believe we got something so unique and good" · category beat the competitor skill AND the bare model |


Cross-model, the pairwise judge that had scored every previous iteration in the 30–40% range against the strongest competitor skill scored the final architecture at 100% of decisive pairs on the deep-validated tasks, and positive across four fresh categories it had never been tuned on.
Part two · what the dice became
From seven lessons to a catalog you can roll
The campaign proved the architecture. Production needed the dice to have faces worth landing on. So the curated pool from lesson four grew into a full catalog: 375 authored worlds and counting, every one through human review, wired to a public roll API. 185 are approved; 83 carry the three-star flagship rating that doubles their odds in the draw. Here is what survived contact with a reviewer.
A world is a graphic system, not a mood.
Every catalog entry passes the born-designed test: its source is a produced 2D or display artifact with an existing graphic system, never a material, a mood, or a place. Era-and-school specific, so "1950s Blue Note session sleeve", never "record covers". Each entry carries five system rules (palette and material, type and composition, topology, controls and state, responsive motion) plus one vivid spark scene, and declares its strength: a world (it brings the visual identity), a composition (it brings the page's structure), or dual. Compositions are a separate deck, one per surface mode: persuade, operate, read, experience.
The sharpest predictor of a flagship rating turned out to be a verb: the winning systems argue or perform. The Du Bois data portraits argue; a Massin stage page performs its dialogue; a variety-show telop field performs the speech it captions. Systems that merely exist, pure notation, minimal seriality, tasteful restraint, die in review no matter how legitimate their lineage. Names from the approved pool, the way the API exposes them: Fillmore Handbill. Wax Print Market. Teletext Service, the world lesson five's human reference pointed at all along. Du Bois Data Portraits. eBoy Pixorama.
Nothing ships without a human verdict and a render.
Authoring is agent work; approval is not. New entries land as pending and pass a render gate before review: a specimen board first, then a desktop hero generated with the board attached as binding reference, so both images stay one system. The reviewer sees the world rendered, then decides: approve or reject, star ratings on approvals, a note on anything instructive.
360 entries have been reviewed by hand so far. 188 approved, 83 of those rated three-star flagship, 172 rejected outright. The dominant rejection reads "doesn't translate to interface": material worlds without an intrinsic 2D pattern system never recover, and re-authoring them fails again. Rework hit about a third, and only ever for taste fixes, never translation fixes. A later round added a rejection dimension no style guide saw coming: an aesthetic current AI output has colonized is dead on arrival regardless of its history. Frutiger Aero has a real 2004 lineage; it died in review because this month's glassmorphism slop wears its face.
Proven seams saturate fast.
The obvious growth strategy, mine deeper where the flagships came from, failed on contact. A depth round of twelve second-tier artifacts from already-mined veins scored 3/12 with zero flagships; the reviewer's recurring note was "too similar to others we already have". A different artifact with the same green-phosphor system is a duplicate, whatever its name.
Breadth wins. New rounds stay small, target unmined territory at its canonical peak, and expect a 25–40% hit rate now that the initial harvest is done. Exhausted families are retired outright and never receive new entries. The subtler trap took longer to name: asking an authoring agent to justify each candidate by its nearest existing neighbor selects for redundancy, because only near-duplicates have a neighbor to cite. Dedup now happens at shelf level. A candidate names the shelf it would sit on; if that shelf already holds an approved entry, it needs a higher peak or it is out. The same gate produces a strange, useful side effect: with the familiar middle pruned away, review outcomes go bimodal. One round came back four approvals from seven entries, every single one flagship.
The roll is weighted, chained, and reproducible.
A roll deals six challengers, two from each translation tier (graphic systems, instrument languages, atmosphere worlds), plus three compositions matched to the surface's mode. Ratings feed the draw: a three-star world carries double odds, a one-star world sits out entirely unless its tier has nothing else. Re-rolls chain: every world already dealt under a key is excluded from the next hand, so "deal again" explores instead of reshuffling.
And lesson three's reproduction key survived productionization: the same key against the same pool revision reproduces a roll bit for bit, in the skill and in the API. Each dealt world now travels with its rendered cards, a design-system board and a desktop hero, framed to the model as a quality bar rather than a mockup: this is the finish level the build must reach, in its own direction. In a controlled build run, the two pages whose models saw the cards came back the strongest of the campaign; the one page whose hand happened to carry no cards came back the flattest. One accidental control is not proof. It repeated.
The dice are now a public API.
GET /api/roll deals from the approved pool. The catalog never ships to clients in full; a response exposes exactly the entries it deals, nothing more. The request itself is the impression record, and an anonymous choice ping notes which dealt world a project committed to. No identity, no project data, and senders honor DO_NOT_TRACK before calling. That choice stream is the next review signal: worlds that keep winning commits argue for their rating.
You can watch the dice work: the homepage deals a live hand from the flagship tier, and Neo Mirai is the world that took a commit all the way to a shipped build.
The model always had the taste. Our job turned out to be rolling dice it can't roll itself, deleting our own brakes, and holding it to its own promises.
Method notes: all concept-level findings were measured with contract probes (sampling design intent as ~200-word contracts, pennies per variant) before any build was paid for; human verdicts are per-sample ratings by the project's design director; the pairwise judge is screenshot-based with both-orderings position-bias cancellation. Lesson three's assignment A/B ran both selection mechanisms on matched dice keys with identical guidance in both arms. Catalog figures reflect the July 2026 reviews. Produced with impeccable v4.