Methodology
Same prompt. Same rules. Public receipts.
Every benchmark version preserves its prompt, tool policy, execution route, run date, model release date, model metadata, artifacts, and notes. Manual/native website runs are labeled separately from API runs.
Active Prompts
| Version | Tool policy | Prompt |
|---|---|---|
| svg-v1-no-web | no web | Create a standalone SVG portrait of Gary Busey's face. Output only valid SVG markup. Do not wrap the SVG in Markdown fences. Do not use external images, links, scripts, CSS imports, or remote assets. Make the portrait recognizable as Gary Busey using vector shapes only. Include face shape, hair, eyes, eyebrows, nose, mouth, teeth, and expressive features. Use a 1024 by 1024 viewBox. Use detailed SVG-native vector techniques: layered paths, gradients, masks, clipping paths, shadows, highlights, blur filters, opacity, and fine strokes. The portrait should be as recognizable and detailed as possible. |
| image-v1 | manual unknown | A recognizable editorial portrait of actor Gary Busey, centered face, expressive eyes, broad smile, detailed hair, neutral studio background, high detail, square composition. Do not include text, logos, watermarks, or extra people. |
Your Eyes First, Scores Second
The point of BuseyBench is visual comparison over time: how model releases improve, degrade, or simply get stranger under the same prompt. The AI judge scores described below exist to make that progress sortable and chartable — they are a structured version of the same subjective call your own eyes make, not a claim of objective truth. When a score and your eyes disagree, trust your eyes; the gallery is always the ground truth.
What To Compare
| Model release date | Newest and oldest model sorting uses the model release date, not the date the prompt was run. |
|---|---|
| Prompt version | Runs are only apples-to-apples when they use the same prompt version and tool policy. |
| Execution route | OpenRouter, direct API, native website, and manual uploads are labeled because those surfaces can behave differently. |
| Artifact source | SVG runs expose stored source so the output can be inspected beyond the rendered preview. |
Safety Handling
Model SVG output is treated as untrusted code. The app extracts SVG markup, rejects scripts, event handlers, remote assets, and embedded objects, then stores the sanitized artifact for review before publication.
AI Judge Scoring
Every eligible published output is scored by a panel of vision models from 3 different labs (gpt-5.2, gemini-3.1-pro-preview, claude-sonnet-4.6) — an ensemble, so no single model gets to flatter its own family. Each judge sees the rendered image (SVGs are rasterized at 1024px) and an anchored rubric, and scores four dimensions from 0-10: prompt adherence, face coherence, Busey likeness, and aesthetics. Each judge scores 3× at temperature 0 and we take its median; judges are then averaged. Failed calls, actual sample counts, judge-to-judge agreement, and score dispersion are retained. Partial panels are provisional and automatically remain in the scoring queue.
The overall score is a fixed weighted composite computed by the site, not by the judges: likeness 50%, face coherence 20%, aesthetics 15%, prompt adherence 15%. Each model contributes one canonical published output. Once its full judge panel completes, that output's weighted absolute score is the model's official score; later duplicate generations cannot replace it with a luckier result.
Head-to-head judging uses 3 judges. Every judge sees both A/B and B/A; orientation disagreement becomes a draw, then the panel majority is recorded. Any two models whose current raw absolute scores differ by no more than 0.25 points are eligible for a matchup. A neutral Bradley–Terry rating turns that evidence into a bounded ranking adjustment: at most ±0.125 points, ramping with coverage until it reaches full weight after eight games. The official ranking score is the raw absolute score plus that adjustment. The displayed raw score never changes; W/L/D, game count, rating, and uncertainty remain public.
Every result is attached to an immutable evaluation configuration hash covering the exact prompt text, output requirements, tool policy, renderer, rubric (busey-v2), weights, judge models, sample counts, and calibration version. Scoring inputs are isolated from ranking policy so a ranking-policy change does not invalidate an already judged image. Incompatible scoring configurations are never mixed.
Calibration and limitations
AI judges can share training-data biases, prefer polished realism, and disagree about celebrity likeness. Before a judge-panel change is promoted, a fixed owner-reviewed calibration set should be rescored and checked for rank reversals, systematic provider favoritism, orientation bias, and agreement drift. The configured calibration version is part of the evaluation hash so those checks are reproducible.
Ratings are comparative evidence, not scientific truth. Release dates may be incomplete, provider routing can change without notice, and API outputs can differ from consumer products. BuseyBench exposes route, prompt, sample count, variability, artifacts, and methodology so readers can judge those caveats for themselves.