30 models on OpenUI, plus a 6-model comparison across 3 generative UI formats

OpenUIBenchmarks

Generative UI Benchmark

Model comparison

Provider-coloured dots with one Pareto frontier. Key references are labeled; hover any point for its name and values.

Compares OpenUI structural validity and cost across 30 models using provider-coloured dots and one Pareto frontier.Compares validity and cost across OpenUI, A2UI, and json-render for six models.
Models20 / 30
20 of 30 models selected
20 selected
xAI1
OpenAI3
Anthropic4
Google6
Moonshot1
Meta1
Alibaba4
Zhipu1
Thinking Machines2
DeepSeek2
Microsoft1
Mistral1
IBM1
Liquid1
InclusionAI1
70%80%90%100%$0.10$0.01$0.001$0.0001Structural validity vs cost, by modelStructural validitybetter ↗Cost per task · USDGrok 4.6GPT-5.6 SolGemini 3.7 FlashClaude Opus 4.8Claude Sonnet 5GPT-5.6 TerraMuse Spark 1.2Kimi K3Claude Opus 5Claude Sonnet 4.6Qwen3.8 2.4TGLM-5.3Inkling SmallDeepSeek V4 FlashDeepSeek V4 ProGPT-5.6 LunaQwen3.8 27BGemini 3.6 FlashGemini 3.5 Flash-LiteInkling
Pareto frontierxAIOpenAIAnthropicGoogleMoonshotMetaAlibabaZhipuThinking MachinesDeepSeek
View chart data

Model comparison data

OpenUI model-board structural validity and measured cost for all models
ModelProviderFamilyValidCost per taskPricingFrontier (all 30)Shown in chart
Grok 4.6xAI99.5%$0.0185List priceYesYes
GPT-5.6 SolOpenAIGPT-5.699.5%$0.0476List priceNoYes
Claude Opus 4.8AnthropicClaude98.9%$0.0493List priceNoYes
Gemini 3.7 FlashGoogleGemini Flash98.9%$0.0100List priceYesYes
Claude Sonnet 5AnthropicClaude98.4%$0.0202List priceNoYes
GPT-5.6 TerraOpenAIGPT-5.698.4%$0.0241List priceNoYes
Kimi K3Moonshot96.2%$0.0333List priceNoYes
Claude Opus 5AnthropicClaude96.2%$0.0774List priceNoYes
Muse Spark 1.2Meta96.2%$0.0157List priceNoYes
Claude Sonnet 4.6AnthropicClaude92.9%$0.0450List priceNoYes
Qwen3.8 2.4TAlibabaQwen3.891.8%$0.0170List priceNoYes
GLM-5.3Zhipu90.8%$0.0120List priceNoYes
Inkling SmallThinking MachinesInkling88.6%$0.0033List priceYesYes
DeepSeek V4 FlashDeepSeekDeepSeek V485.8%$0.0004List priceYesYes
DeepSeek V4 ProDeepSeekDeepSeek V484.2%$0.0089List priceNoYes
GPT-5.6 LunaOpenAIGPT-5.683.7%$0.0022List priceNoYes
Qwen3.8 27BAlibabaQwen3.878.8%$0.0054List priceNoYes
Gemini 3.6 FlashGoogleGemini Flash78.8%$0.0091List priceNoYes
Gemini 3.5 Flash-LiteGoogleGemini Flash78.3%$0.0041List priceNoYes
InklingThinking MachinesInkling73.4%$0.0080List priceNoYes
Qwen3.6 27BAlibabaQwen3.668.5%$0.0063List priceNoNo
Qwen3.6 35B-A3BAlibabaQwen3.661.4%$0.0017List priceNoNo
Gemma 4 31BGoogleGemma 446.7%$0.0009List priceNoNo
Phi-4Microsoft44.0%$0.0004List priceNoNo
Gemma 4 26B-A4BGoogleGemma 429.9%$0.0007List priceNoNo
Ministral 8BMistral27.2%$0.0009List priceNoNo
Granite 4.1 8BIBM14.7%$0.0002List priceYesNo
LFM 2.5 2.6BLiquid3.3%$0 / freeFreeYesNo
DiffusionGemma 26B-A4BGoogle13.0%Not comparableSelf-hostedNoNo
Ling 3.0 TinyInclusionAI9.8%Not comparableSelf-hostedNoNo

All 30models remain in this table and the downloads, including models hidden by the chart’s default filter. Frontier membership here is computed over the full board, so it does not shift with the chart’s selection; the drawn line is the frontier of the models currently shown. Self-hosted cost is unknown, not zero. Filter state is preserved in this page’s URL. Focused benchmark: language and model results. Full dataset: JSON, CSV, or agent Markdown.

Format comparison data

Structural validity and cost per 46-screen pass for each model and format
ModelFormatValidCost per pass
GPT-5.6 SolOpenUI99.5%$3.08
Claude Opus 4.8OpenUI98.9%$2.27
Kimi K3OpenUI96.2%$1.53
Qwen3.8 2.4TOpenUI91.8%$0.81
Muse Spark 1.2OpenUI96.2%$0.71
Gemini 3.7 FlashOpenUI98.9%$0.42
GPT-5.6 SolA2UI96.2%$6.84
Claude Opus 4.8A2UI99.5%$5.85
Kimi K3A2UI95.7%$3.16
Qwen3.8 2.4TA2UI91.3%$1.90
Muse Spark 1.2A2UI96.7%$1.38
Gemini 3.7 FlashA2UI94.0%$0.95
GPT-5.6 Soljson-render82.6%$6.30
Claude Opus 4.8json-render87.5%$4.88
Kimi K3json-render71.2%$2.91
Qwen3.8 2.4Tjson-render81.0%$1.46
Muse Spark 1.2json-render81.5%$1.34
Gemini 3.7 Flashjson-render92.9%$0.85

Focused benchmark: framework comparison. Download this comparison as JSON, CSV, or agent Markdown.

Headline results

Blank screens vs. renders

Did the generation actually render?

1,104 runs per format across 6 models.
Counted as each SDK’s own renderer produced them.

OpenUI had 1 blank screen in 1,104 runs, compared with 37 for A2UI and 4 for json-render.

Render rate

Runs that rendered against runs that came back blank, out of 184 per model.

OpenUI99.9%
0184
0184
0184
0184
1183
0184
A2UI96.6%
5179
0184
5179
6178
16168
5179
json-render99.6%
0184
0184
0184
0184
4180
0184
RenderedCame back blank
GPT-5.6 SolClaude Opus 4.8Kimi K3Gemini 3.7 FlashQwen3.8 2.4TMuse Spark 1.2
View render data
Rendered and blank runs for every model and generative UI format
ModelFormatRenderedBlankRender rate
GPT-5.6 SolOpenUI1840100.0%
Claude Opus 4.8OpenUI1840100.0%
Kimi K3OpenUI1840100.0%
Gemini 3.7 FlashOpenUI1840100.0%
Qwen3.8 2.4TOpenUI183199.5%
Muse Spark 1.2OpenUI1840100.0%
GPT-5.6 SolA2UI179597.3%
Claude Opus 4.8A2UI1840100.0%
Kimi K3A2UI179597.3%
Gemini 3.7 FlashA2UI178696.7%
Qwen3.8 2.4TA2UI1681691.3%
Muse Spark 1.2A2UI179597.3%
GPT-5.6 Soljson-render1840100.0%
Claude Opus 4.8json-render1840100.0%
Kimi K3json-render1840100.0%
Gemini 3.7 Flashjson-render1840100.0%
Qwen3.8 2.4Tjson-render180497.8%
Muse Spark 1.2json-render1840100.0%

Structural validity vs. render success

Was the output valid, and did anything render?

Valid: every part it refers to exists, and every setting is a real one. Render success: something appeared at all. 1,104 runs per format.

OpenUI leads on both; A2UI and json-render each trade one off against the other.

Structural validity and render success

Share of 1,104 runs per format; longer is better. Scales start at 75% and 95%, not zero. Arrows compare with OpenUI.

OpenUI96.9%99.9%
A2UI95.6%1.496.6%3.3
json-render82.8%14.199.6%0.3
View validity and render data
Structural validity and render success by format
FormatValid runsPartial runsBlank runsValidityRender success
OpenUI107033196.9%99.9%
A2UI1055123795.6%96.6%
json-render914186482.8%99.6%

What counts as valid. Parses, has a root, every reference resolves, nothing orphaned or invented, no missing or out-of-range props, not truncated, and at least as many components as the brief has requirements. That last one is a count floor, not a check that each requirement was addressed: this measures structure, not coverage.

Structural validity by model

How often the component graph holds together

Passes only if nothing is left dangling, nothing is invented, and every setting is valid. 184 runs per model, per format.

OpenUI leads overall: 96.9% vs 95.6% for A2UI. The lead changes by model, with two ties.

Structural validity by model

Runs whose component graph parses, resolves and validates, judged by each format's own SDK.

Structural validity percentage for each model and generative UI format
ModelOpenUIA2UIjson-render
SolOpenAI99.596.282.6
Claude Opus 4.8Anthropic98.999.587.5
Kimi K3Moonshot96.295.771.2
Gemini 3.7 FlashGoogle98.994.092.9
Qwen3.8 2.4TAlibaba91.891.381.0
Muse Spark 1.2Meta96.296.781.5
Average96.995.682.8

Token consumption and cost

Tokens and dollars for the same screens

46 screens,
priced at provider list prices.

OpenUI uses fewer tokens and costs 1.82.6× less across every priced model.

Token consumption

System prompt + output for one screen

OpenUI5,0311,362
A2UI12,6102.5×2,8232.1×
json-render7,6511.5×3,2582.4×
View token data
System prompt and mean output tokens per screen by format
FormatSystem promptMean outputCombined
OpenUI5,0311,3626,393
A2UI12,6102,82315,433
json-render7,6513,25810,909

Cost of one benchmark pass

The same 46 screens at list prices

Cost in USD for one 46-screen benchmark pass by model and generative UI format
ModelOpenUIA2UIjson-render
Gemini 3.7 FlashGoogle$0.42$0.952.3×$0.852.0×
Muse Spark 1.2Meta$0.71$1.381.9×$1.341.9×
Qwen3.8 2.4TAlibaba$0.81$1.902.3×$1.461.8×
Kimi K3Moonshot$1.53$3.162.1×$2.911.9×
Claude Opus 4.8Anthropic$2.27$5.852.6×$4.882.1×
SolOpenAI$3.08$6.842.2×$6.302.0×

Streaming speed

How long a screen takes to render

Mean output per screen,
decoded at 50 tokens per second.

A screen streams in about half the time, because there is about half as much to write.

Time to stream a screen

Mean output per screen, decoded at 50 tokens per second.

OpenUI~27s
A2UI2.1×~56s
json-render2.4×~65s
View streaming data
Estimated streaming time from mean output tokens at 50 tokens per second
FormatMean output tokensDecode rateEstimated seconds
OpenUI1,36250 tok/s27.2s
A2UI2,82350 tok/s56.5s
json-render3,25850 tok/s65.2s

Structural validity by screen complexity

Validity as requirements increase

46 briefs across 5 complexity bands,
about 9 per band.

As screens get harder, OpenUI and A2UI stay above 90%; json-render falls to 72%.

Structural validity by screen complexity

The more a brief asks for, the less every format delivers. OpenUI and A2UI track each other closely at every size; json-render falls fastest and furthest.

70809010091%93%72%2–34–67–911–1316–18requirements per screen
OpenUI91%A2UI93%json-render72%
View complexity data
Structural validity by brief requirement count and generative UI format
Requirements per screenRunsOpenUIA2UIjson-render
2–3240100.0%99.6%89.6%
4–624099.6%98.8%81.7%
7–924098.8%94.2%85.4%
11–1319291.1%91.7%68.2%
16–1819290.6%93.2%71.9%

Production repair

What gets fixed before users see it

Based on real production data:
OpenUI Cloud traffic, not benchmark runs.

Only 0.9% of generations reach a user broken. Of the ones that fail validation, 88% are repaired via incremental editing.

Repair funnel

What happens to a broken generation before it can reach a user.

100%7%0.9%
100%All generationsNo model callParser fixes syntax issues first
7%Fail validationNo model callStructural issues like dangling refs and bad enums
0.9%Reach a user brokenOne model call88% are repaired via incremental editing
View repair data
Production repair funnel as a share of all streaming generations
StageShare of generationsWhat happensModel call
All generations100%Parser fixes syntax issues firstNo model call
Fail validation7%Structural issues like dangling refs and bad enumsNo model call
Reach a user broken0.9%88% are repaired via incremental editingOne model call

Production traffic, not benchmark runs.

Improve your Generative UI reliability with OpenUI Cloud.