Internal evidenceNot a leaderboard

UI-Resolve benchmark

Evidence before rank.

These are reproducible internal checkpoints, not a global model, skill, or harness ranking. We publish the denominator, the loss, and the uncertainty beside every improvement.

Current publication state

Internal
Preview
Verified
Evidence date
2026-08-10
Latest exact block
51 cells

Latest checkpoint · internal

High is the measured default. It is not a winner declaration.

Codex configuration-routed Luna/Terra/Sol completed three fixed tasks once at every supported effort. The result sets a bounded OmD routing default while preserving explicit supported choices.

Task design
3 tasks × 1 trial
Attribution
Configuration only
Terminal + valid
51/51

Exact block completed

Objective resolved
38/51

Critical product gates

Objective + proof
34/51

Both gates passed

Observed tokens
36.89M

50/51 cells reported usage

Semantic effort order

Same task block, different effort

Rows stay low → ultra. They are not sorted into a leaderboard.

low

9 fixed cells

Objective
7/9
+ proof
6/9

3.98M

tokens · 9/9 observed

medium

9 fixed cells

Objective
6/9
+ proof
6/9

4.19M

tokens · 9/9 observed

highOmD default

9 fixed cells

Objective
8/9
+ proof
8/9

5.51M

tokens · 9/9 observed

xhigh

9 fixed cells

Objective
7/9
+ proof
6/9

6.38M

tokens · 9/9 observed

max

9 fixed cells · explicit only

Objective
7/9
+ proof
5/9

9.85M

tokens · 9/9 observed

ultra

6 fixed cells · explicit only

Objective
3/6
+ proof
3/6

6.98M

tokens · 5/6 observed

Separate questions

Three tracks. No blended winner.

01Not published

Model Track

What does the model deliver without a third-party UI skill?

02Not published

Skill Lift

How much does a portable skill change a fixed model?

03Evidence below

Harness Track

Does the complete workflow resolve the task within a fixed budget?

Harness Track · internal

The bounded harness resolved more runs. The interval still includes zero.

Claude Opus 4.8xhigh18/18 complete

UI-Resolved@1

All critical product gates must pass.

Portable workflow5/9
Bounded repair harness8/9
+33.33pp

Observed resolved-rate lift

4/4/1

Paired win / tie / loss

-22.22100pp

95% interval; includes zero

Median wall time

7.8mvs 9.1m

Median uncached tokens

132.9kvs 116.5k

The loss is part of the result

operations-t3-harness · 77/85

The harness lost one pair on design.primary_action and design.card_radius. Task, state, responsive, accessibility, and evidence gates still passed.

1 candidate loss

Recovery record

We did not turn a failed run green.

The original accessibility failure remains 79/85. We changed the static contract, tested a focused control, then required a fresh model run to prove the recovery.

Fresh recovery · 1 candidate cell
85/85
Axe serious / critical
0

single candidate recovery cell · no general locale-lift claim · no model, skill, efficiency, or frontier claim.

  1. 1.9.34

    A real failure stayed visible

    The command region scrolled horizontally but could not receive keyboard focus. Axe and traversal checks failed on constrained viewports.

    79/85
  2. 1.9.35

    The rule changed, not the old score

    A provider-free calibration required useful scroll regions to contain a reachable control or become explicit focus targets.

    79 → 85 control
  3. 1.9.36

    A fresh run proved the recovery

    A new exact-model run passed clipboard, locale state, responsive geometry, keyboard traversal, and axe checks without a replacement verifier.

    85/85

What 85 means

Polish cannot average away a broken journey.

UI-Resolved requires every critical product gate. A higher visual score cannot compensate for a failed action, hidden overflow, inaccessible control, or invented fact.

01

Task contract

Required content and actions exist and work.

02

State journey

Click, keyboard, and localized status transitions agree.

03

Responsive

Desktop, 390px, 320px, and 200% geometry stay usable.

04

Accessibility

Keyboard focus remains visible and serious or critical axe findings stay at zero.

05

Design grounding

The surface follows the frozen product and design contract.

06

Evidence & Unknown

Protected facts remain intact and unsupported claims stay absent.

Publication ladder

Internal is a stage, not a softer word for verified.

Internal

Runner and task calibration. May change; no public rank.

Current

Preview

Reproducible packages and examples, but limited tasks, trials, or review.

Next

Verified

24+ tasks × 10 runs, 10+ practitioners, and independent audit.

Gate

After the evidence

Inspect first. Then continue in your own project.

This result remains Internal. These actions let you inspect the workflow or start from a reference; they do not prove installation, resolved UI, or reuse.

Install the local workflow

Copy the exact command, then run it in your project terminal.

$ npx oh-my-design-cli@latest

See a complete run

Follow the six-step docs runbook before trying your own brief.

Open live demo

Choose a reference

Open Builder, select a brand, and inspect its verified preview.

Open Builder