Skip to main content
BuseyBenchcreative model benchmarkSVG models tested
SVG testImage testLeaderboardTimelineMethodology
Browse
SVG testImage testLeaderboardTimelineMethodology
BuseyBenchSame prompt. Same rules. Maximum facial voltage.
Score trendsTimelineMethodologyAboutPrivacyGovernanceRSSAdmin
306 SVG models tested
BuseyBench SVG TestImage Generator Portrait Test

Leaderboard

Judged. Ranked. Busey'd.

One output per model, independent judge scores, and close-score pairwise tie-breaks. The best Gary Busey portraits rise to the top.

Explore the dataDrawing ability over time→
306 models ranked7103 official head-to-head verdictsPairwise tie-break active

Official standings

The rankings

301–306 of 306 results
#3011500 rating ±94
Qwen3 VL 30B A3B Instruct output

Qwen3 VL 30B A3B Instruct

Qwen / Oct 6, 2025 / openrouter

Busey likeness0.0
Face coherence0.4
Aesthetics1.4
Prompt adherence0.5

1 judged output · judge spread σ 0.04 · W 0 / L 0 / D 14

0.4overall
#3021539 rating ±94
GPT-4o output

GPT-4o

OpenAI / May 13, 2024 / openrouter

Busey likeness0.0
Face coherence0.2
Aesthetics1.2
Prompt adherence0.4

1 judged output · judge spread σ 0.14 · W 2 / L 0 / D 12

0.3overall
#3031500 rating ±94
Granite 4.0 Micro output

Granite 4.0 Micro

IBM / Oct 2, 2025 / openrouter

Busey likeness0.0
Face coherence0.4
Aesthetics1.3
Prompt adherence0.5

1 judged output · judge spread σ 0.05 · W 0 / L 0 / D 14

0.3overall
#3041500 rating ±94
MythoMax 13B output

MythoMax 13B

Gryphe / Jul 2, 2023 / openrouter

Busey likeness0.0
Face coherence0.3
Aesthetics0.9
Prompt adherence0.4

1 judged output · judge spread σ 0.07 · W 0 / L 0 / D 14

0.2overall
#3051409 rating ±101
Llama 3.1 8B Instruct output

Llama 3.1 8B Instruct

Meta / Jul 23, 2024 / openrouter

Busey likeness0.0
Face coherence0.0
Aesthetics0.6
Prompt adherence0.5

1 judged output · judge spread σ 0.05 · W 0 / L 4 / D 8

0.2overall
#3061498 rating ±101
Aion-RP 1.0 (8B) output

Aion-RP 1.0 (8B)

AionLabs / Feb 4, 2025 / openrouter

Busey likeness0.0
Face coherence0.0
Aesthetics0.9
Prompt adherence0.1

1 judged output · judge spread σ 0.06 · W 0 / L 0 / D 12

0.1overall
← PreviousPage 13 of 13Next →
Ranking statusDisplayed score + pairwise tie-break

Overall = likeness 50% + face 20% + aesthetics 15% + prompt 15%. Each model contributes one canonical judged output. The visible overall score, rounded to one decimal, always sets the score band. Pairwise results only order models tied within that displayed band, so a lower displayed score can never leapfrog a higher one.

Read the full methodology →View score trends →
1canonical output judged per model
Closest-scorepairwise matchmaking strategy
busey-v2current scoring rubric