Baseline
One frozen prompt per modality, rerun on every frontier release. Not a leaderboard. A control test, where the prompt is the instrument and never changes.
Why this exists
Text models are measured to death. The generative media models that actually sit inside a production pipeline are measured by vibes, by a vendor reel, and by whichever cherry-picked output went viral that week.
So this runs the other way around. Four prompts, one per modality, each stacking that modality’s known failure modes into a single generation. They were written once, frozen in version control, and are never edited. Every model gets the same text, first generation only, no rerolls.
The value is not any single round. It is the time axis. Two years of the same prompt through the same rubric shows you what actually improved and what a release note only claimed.
The four control prompts
The watchmaker's bench
Text rendering, optics, hands, relational composition.
4 runs · 4 modelsView →
The pour behind the pass
Object permanence, fluid physics, camera control, identity hold.
5 runs · 5 modelsView →
The shipment call
Prosodic turns, non-verbals, code-switching, silence.
Awaiting first roundView →
The ice cream truck
Organic meets hard surface, occluded sides, thin parts, texture fidelity, real scale.
5 runs · 4 modelsView →
Rubric · six axes, 1–5
- Prompt adherence
- Did it render what was asked, including the exclusions?
- 3 = Most stated elements present. One or two clauses ignored or approximated.
- Physical plausibility
- Does light, fluid, anatomy and material behave the way the world does?
- 3 = Reads correct at a glance. Breaks down on inspection of one system.
- Aesthetic control
- Does the direction survive, or does house style overwrite it?
- 3 = Attractive but generic. The model's default look overrides the direction.
- Artifact rate
- Frequency and severity of visible failure in the output.
- 3 = One clear artifact a viewer would notice, but a retouch could fix.
- Cost + latency
- What it costs in money and time to reach one usable output.
- 3 = One usable output within three attempts at typical market rate.
- Production readiness
- Could this enter a real pipeline, and at what cleanup cost?
- 3 = Usable after targeted manual fixing. Under an hour of work.
Settings policy
- Fixed seed where the tool exposes one. Runs without a seed are marked as not reproducible.
- Default settings otherwise. No custom LoRAs, no upscaling, no post-processing.
- First generation only. No cherry-picking from a batch, no rerolls for a better take.
- Same operator, same sitting, for every model in a release round.
- Cost and latency measured on the published consumer tier, not on enterprise pricing.