This benchmark measures the product’s core claim: whether a coding agent can use logic2b as a design system end to end instead of merely reading its docs. The protocol is public, every rule has a point value, and submitted code is never executed by the scorer.
Task 1 · 100 points
Install and theme
In the supplied Vite React + Tailwind v4 project, install the logic2b button, card and dialog with their registry dependencies. Apply the exact preset id supplied by the benchmark. Use logic2b's CLI, MCP or raw registry; do not recreate the components. Finish by running the production build.
12 objective checks
Task 2 · 100 points
Compose an accessible settings screen
Using the installed logic2b components, create src/App.tsx as an account settings screen with Card, Tabs, Label, Input, Switch and Button. Use semantic labels, token-based colors only and no raw Tailwind palette colors. Preserve the supplied preset and finish with a production build.
13 objective checks
Task 3 · 100 points
Scaffold a dashboard from empty
Starting from an empty directory, use logic2b to create a Vite analytics dashboard with the benchmark preset. Prefer scaffold_plan when available. The result must include a runnable framework shell, dashboard block, charts, theme, components.json and exact dependency versions. Finish with a production build.
12 objective checks
| Rank | Model | Agent | Score | Duration |
|---|---|---|---|---|
| 1 | gpt-5.6-sol | OpenAI Codex CLI | 294/300 (98%) | 737.6 s |
| 2 | gpt-5.5 | OpenAI Codex CLI | 272/300 (90.7%) | 751.789 s |
| 3 | gpt-5.6-terra | OpenAI Codex CLI | 266/300 (88.7%) | 430.712 s |
Protocol v1.1.0: 3 tasks, 300 points.Score ranks before duration; build success is recorded by the evaluator.
What is measured
The tasks progress from component installation to composition and full
application scaffolding. They check registry provenance (data-slot, expected
imports and transitive files), exact preset fidelity, semantic token use,
accessible form naming, dependency pinning and evaluator-observed production
builds.
The overall score is 300 points. Score ranks first; total duration breaks ties. This keeps a fast but broken result below a slower complete one.
Fair-run contract
- Fresh isolated directory for every model and task.
- Identical task text, preset and starting fixture.
- Tool access (shell, network and MCP) recorded in run metadata.
- No human corrections after the timer starts.
- Failed and timed-out tasks retain their partial artifacts.
- The benchmark operator—not the model—records build exit codes.
- Model/provider version, agent host, OS and Node version accompany the result.
The repository now ships the versioned evaluator runner as well as the static scorer. It creates each fixture from the built registry in external staging, records its SHA-256, invokes the configured agent without a shell, enforces time/output limits and snapshots source before the observed build; dependencies and generated caches are excluded. It captures both agent and verification transcripts; symlinks are discarded and source over the configured artifact budget is never executed. The runner must itself be executed inside a disposable container or VM because the observed build step intentionally executes generated project code.
Reproduce or audit
The protocol, scorer, test fixture contract and raw run format ship in
benchmarks/agents.
Run the harness tests with:
pnpm --dir benchmarks/agents test
Validate and execute an isolated agent configuration with:
pnpm --dir benchmarks/agents benchmark config.json --validate
pnpm --dir benchmarks/agents benchmark config.json
pnpm --dir benchmarks/agents score <run-id>
Real run artifacts are committed alongside their detailed score JSON. The
first controlled comparison uses the same Codex CLI 0.148.0-alpha.15 host
and capabilities for all three models: gpt-5.6-sol leads with 294/300 (98%),
followed by gpt-5.5 at 272/300 (90.7%) and gpt-5.6-terra at 266/300
(88.7%). Every evaluator-observed build passed. The rule evidence still shows
meaningful differences: Sol only canonicalized the legacy preset id, 5.5 lost
scaffold fidelity on the starter/theme/pins, and Terra produced a runnable but
non-canonical dashboard rather than the requested registry block. Synthetic
fixtures remain tagged and mechanically excluded from publication.