#006 · Index · BaselineRev. R.07 · APR 2026

Baseline

One frozen prompt per modality, rerun on every frontier release. Not a leaderboard. A control test, where the prompt is the instrument and never changes.

Why this exists

Text models are measured to death. The generative media models that actually sit inside a production pipeline are measured by vibes, by a vendor reel, and by whichever cherry-picked output went viral that week.

So this runs the other way around. Four prompts, one per modality, each stacking that modality’s known failure modes into a single generation. They were written once, frozen in version control, and are never edited. Every model gets the same text, first generation only, no rerolls.

The value is not any single round. It is the time axis. Two years of the same prompt through the same rubric shows you what actually improved and what a release note only claimed.

The four control prompts

Rubric · six axes, 1–5

Prompt adherence
Did it render what was asked, including the exclusions?
3 = Most stated elements present. One or two clauses ignored or approximated.
Physical plausibility
Does light, fluid, anatomy and material behave the way the world does?
3 = Reads correct at a glance. Breaks down on inspection of one system.
Aesthetic control
Does the direction survive, or does house style overwrite it?
3 = Attractive but generic. The model's default look overrides the direction.
Artifact rate
Frequency and severity of visible failure in the output.
3 = One clear artifact a viewer would notice, but a retouch could fix.
Cost + latency
What it costs in money and time to reach one usable output.
3 = One usable output within three attempts at typical market rate.
Production readiness
Could this enter a real pipeline, and at what cleanup cost?
3 = Usable after targeted manual fixing. Under an hour of work.

Settings policy

  • Fixed seed where the tool exposes one. Runs without a seed are marked as not reproducible.
  • Default settings otherwise. No custom LoRAs, no upscaling, no post-processing.
  • First generation only. No cherry-picking from a batch, no rerolls for a better take.
  • Same operator, same sitting, for every model in a release round.
  • Cost and latency measured on the published consumer tier, not on enterprise pricing.