Design System benchmark
Each metric is the Benchmark result across all capabilities, followed by the percent change from Control in parentheses. Models use equal-weight ranks across shared checks with complete, error-free values and a scoring direction. Implementation-agent usage breaks ties. Checks average per-check, per-trial pass percentages or measurement means; units and directions stay separate. Skips and errors are shown separately. Latest results: .
| Model | Effort | Checks | Output tokens | Premium requests | AI credits | Session time | API time |
|---|---|---|---|---|---|---|---|
| claude-opus-5.5 | medium | 75.0% (+114.3%) | 0 (0%) | 75 (0%) | 167.283 (+63.9%) | 5m 28.7s (+59.1%) | 3m 51.0s (+69.6%) |
| gpt-6-astra | medium | 71.1% (+101.1%) | 0 (0%) | 5 (0%) | 641.303 (+111.5%) | 11m 30.5s (+112.0%) | 9m 29.2s (+114.2%) |
| claude-sonnet-5 | medium | 72.5% (+53.8%) | 0 (0%) | 5 (0%) | 190.338 (-1.7%) | 7m 59.7s (-21.1%) | 6m 11.4s (-6.7%) |
| gpt-6-sol | medium | 50.0% (+52.2%) | 0 (0%) | 5 (0%) | 127.545 (+46.9%) | 6m 8.4s (+23.0%) | 4m 53.7s (+22.9%) |
| gpt-6-luna | medium | 32.9% (-7.1%) | 0 (0%) | 5 (0%) | 6.031 (-15.5%) | 5m 47.7s (-17.6%) | 4m 48.3s (-12.6%) |
Trends
Strong lines show Benchmark results and muted lines show Control over time. Check charts average each check's per-trial values; skipped outcomes and errors are excluded.
Premium requests
AI credits
Session time
API time
004-agent-setup-nextjs / node-tests
003-agent-uses-form-from-primer / node-tests
005-agent-enables-theme-switching / node-tests
002-agent-uses-octicon-from-primer / node-tests
001-agent-uses-button-from-primer / node-tests
Change from Control
Benchmark results are shown first, followed by the percent change from Control in parentheses.
| Model | |||||
|---|---|---|---|---|---|
| claude-opus-5 (medium) | 0 (0%) | 0 (0%) | 0 (0%) | N/A | N/A |
| claude-opus-5.5 (medium) | N/A | N/A | N/A | 0 (0%) | 0 (0%) |
| claude-sonnet-5 (medium) | 0 (0%) | 0 (0%) | 0 (0%) | 0 (0%) | 0 (0%) |
| gpt-5.6-sol (medium) | 0 (0%) | 0 (0%) | 0 (0%) | N/A | N/A |
| gpt-5.6-terra (medium) | 0 (0%) | 0 (0%) | 0 (0%) | N/A | N/A |
| gpt-6-astra (medium) | N/A | N/A | N/A | 0 (0%) | 0 (0%) |
| gpt-6-luna (medium) | N/A | N/A | N/A | 0 (0%) | 0 (0%) |
| gpt-6-sol (medium) | N/A | N/A | N/A | 0 (0%) | 0 (0%) |
View raw trend data
| Date | Model | Output tokens | Premium requests | AI credits | Session time | API time | 004-agent-setup-nextjs / node-tests | 003-agent-uses-form-from-primer / node-tests | 005-agent-enables-theme-switching / node-tests | 002-agent-uses-octicon-from-primer / node-tests | 001-agent-uses-button-from-primer / node-tests |
|---|---|---|---|---|---|---|---|---|---|---|---|
| claude-opus-5 (medium) | 0 (0%) | 75 (0%) | N/A (N/A) | 13m 51.5s (+84.1%) | 9m 28.1s (+75%) | 75.0% (0%) | 100.0% (N/A) | 77.8% (N/A) | 100.0% (0%) | 100.0% (N/A) | |
| claude-sonnet-5 (medium) | 0 (0%) | 5 (0%) | N/A (N/A) | 5m 50.2s (-41.9%) | 3m 57.0s (-36.6%) | 75.0% (-14.3%) | 85.7% (0%) | 0.0% (0%) | 100.0% (0%) | 100.0% (N/A) | |
| gpt-5.6-sol (medium) | 0 (0%) | 5 (0%) | N/A (N/A) | 8m 9.7s (+63.9%) | 6m 15.5s (+71.4%) | 62.5% (+150%) | 100.0% (+600%) | 5.6% (N/A) | 100.0% (0%) | 0.0% (0%) | |
| gpt-5.6-terra (medium) | 0 (0%) | 5 (0%) | N/A (N/A) | 4m 8.4s (+39.2%) | 2m 32.7s (+37.3%) | 62.5% (+66.7%) | 100.0% (+600%) | 0.0% (0%) | 100.0% (0%) | 0.0% (0%) | |
| claude-opus-5 (medium) | 0 (0%) | 75 (0%) | N/A (N/A) | 13m 22.6s (+61%) | 9m 45.1s (+50.9%) | 75.0% (+50%) | 100.0% (N/A) | 88.9% (N/A) | 100.0% (0%) | 100.0% (+33.3%) | |
| claude-sonnet-5 (medium) | 0 (0%) | 5 (0%) | N/A (N/A) | 7m 33.3s (-4.21%) | 5m 44.2s (-3.78%) | 75.0% (+20%) | 100.0% (+75%) | 50.0% (N/A) | 100.0% (0%) | 100.0% (N/A) | |
| gpt-5.6-sol (medium) | 0 (0%) | 5 (0%) | N/A (N/A) | 6m 15.1s (-74.9%) | 4m 49.9s (-79.7%) | 50.0% (-33.3%) | 100.0% (+600%) | 0.0% (0%) | 100.0% (0%) | 0.0% (0%) | |
| gpt-5.6-terra (medium) | 0 (0%) | 5 (0%) | N/A (N/A) | 5m 35.4s (+105%) | 3m 51.4s (+114%) | 62.5% (+150%) | 100.0% (+600%) | 0.0% (0%) | 100.0% (0%) | 100.0% (N/A) | |
| claude-opus-5 (medium) | 0 (0%) | 75 (0%) | 524.524 (+52.2%) | 11m 47.4s (+17.6%) | 8m 13.6s (+28.1%) | 75.0% (0%) | 100.0% (+16.7%) | 0.0% (0%) | 100.0% (0%) | 100.0% (N/A) | |
| claude-sonnet-5 (medium) | 0 (0%) | 5 (0%) | 256.779 (+34.6%) | 10m 25.2s (-9.08%) | 7m 0.6s (+8.25%) | 75.0% (0%) | 100.0% (N/A) | 0.0% (0%) | 100.0% (0%) | 100.0% (N/A) | |
| gpt-5.6-sol (medium) | 0 (0%) | 5 (0%) | 305.16 (+74.9%) | 7m 49.7s (+72.8%) | 6m 28.8s (+86.2%) | 50.0% (-20%) | 85.7% (+500%) | 5.6% (N/A) | 100.0% (0%) | 0.0% (0%) | |
| gpt-5.6-terra (medium) | 0 (0%) | 5 (0%) | 119.457 (+71.8%) | 5m 53.1s (+54.7%) | 4m 23.5s (+56.6%) | 62.5% (+150%) | 100.0% (+600%) | 0.0% (0%) | 100.0% (0%) | 0.0% (0%) | |
| claude-opus-5.5 (medium) | 0 (0%) | 5 (0%) | 188.336 (+50.8%) | 6m 38.8s (+37.5%) | 4m 42.4s (+49.7%) | 75.0% (0%) | 0.0% (-100%) | 83.3% (N/A) | 100.0% (0%) | 100.0% (N/A) | |
| claude-sonnet-5 (medium) | 0 (0%) | 5 (0%) | 230.011 (-8.74%) | 10m 53.3s (-30.4%) | 7m 51.9s (-27.2%) | 75.0% (-14.3%) | 100.0% (0%) | 0.0% (0%) | 100.0% (0%) | 100.0% (N/A) | |
| gpt-6-astra (medium) | 0 (0%) | 5 (0%) | 659.181 (+62.3%) | 14m 41.0s (-29.1%) | 12m 10.7s (+30.6%) | 50.0% (-20%) | 100.0% (+600%) | 88.9% (N/A) | 100.0% (0%) | 100.0% (N/A) | |
| gpt-6-luna (medium) | 0 (0%) | 5 (0%) | 5.717 (+22.5%) | 8m 10.3s (+7.55%) | 6m 37.1s (+15.2%) | 50.0% (-42.9%) | 14.3% (0%) | 0.0% (0%) | 100.0% (0%) | 0.0% (0%) | |
| gpt-6-sol (medium) | 0 (0%) | 5 (0%) | 143.37 (+57.4%) | 10m 41.5s (+42.3%) | 8m 29.1s (+42.6%) | 50.0% (0%) | 100.0% (+600%) | 0.0% (0%) | 100.0% (0%) | 100.0% (N/A) | |
| claude-opus-5.5 (medium) | 0 (0%) | 75 (0%) | 167.283 (+63.9%) | 5m 28.7s (+59.1%) | 3m 51.0s (+69.6%) | 75.0% (0%) | 100.0% (N/A) | 0.0% (0%) | 100.0% (0%) | 100.0% (N/A) | |
| claude-sonnet-5 (medium) | 0 (0%) | 5 (0%) | 190.338 (-1.7%) | 7m 59.7s (-21.1%) | 6m 11.4s (-6.66%) | 62.5% (+25%) | 100.0% (+16.7%) | 0.0% (0%) | 100.0% (0%) | 100.0% (N/A) | |
| gpt-6-astra (medium) | 0 (0%) | 5 (0%) | 641.303 (+112%) | 11m 30.5s (+112%) | 9m 29.2s (+114%) | 50.0% (-20%) | 100.0% (+600%) | 5.6% (N/A) | 100.0% (0%) | 100.0% (N/A) | |
| gpt-6-luna (medium) | 0 (0%) | 5 (0%) | 6.031 (-15.5%) | 5m 47.7s (-17.6%) | 4m 48.3s (-12.6%) | 50.0% (-20%) | 14.3% (0%) | 0.0% (0%) | 100.0% (0%) | 0.0% (0%) | |
| gpt-6-sol (medium) | 0 (0%) | 5 (0%) | 127.545 (+46.9%) | 6m 8.4s (+23%) | 4m 53.7s (+22.9%) | 50.0% (0%) | 100.0% (+600%) | 0.0% (0%) | 100.0% (0%) | 0.0% (0%) |
Experiments
View all experimentsChecks show per-trial pass percentages or measurement means, averaged across trials. Skips, errors, and missing values are reported separately. Resource usage is the average per trial. Treatments are grouped by model and reasoning effort; compare scenario and trial counts before comparing performance.
MCP with server instructions
Comparing `@primer/mcp` with and without server instructions to determine if server instructions improve the model's ability to follow instructions and complete tasks effectively. The experiment will involve two groups: one using `@primer/mcp` with server instructions and another using `@primer/mcp` without server instructions.
No results have been recorded for this experiment yet.
MCP
Compare MCP versus local instructions performance for Primer usage.
No results have been recorded for this experiment yet.
noop
A fast experiment for testing agent-eval
No results have been recorded for this experiment yet.