Skip to content
Logic2BUI

Search docs

Search components and documentation

New
Menu

Agent Benchmarks

A public, reproducible benchmark for agents installing, theming and composing logic2b ui.

This benchmark measures the product’s core claim: whether a coding agent can use logic2b as a design system end to end instead of merely reading its docs. The protocol is public, every rule has a point value, and submitted code is never executed by the scorer.

Task 1 · 100 points

Install and theme

In the supplied Vite React + Tailwind v4 project, install the logic2b button, card and dialog with their registry dependencies. Apply the exact preset id supplied by the benchmark. Use logic2b's CLI, MCP or raw registry; do not recreate the components. Finish by running the production build.

12 objective checks

Task 2 · 100 points

Compose an accessible settings screen

Using the installed logic2b components, create src/App.tsx as an account settings screen with Card, Tabs, Label, Input, Switch and Button. Use semantic labels, token-based colors only and no raw Tailwind palette colors. Preserve the supplied preset and finish with a production build.

13 objective checks

Task 3 · 100 points

Scaffold a dashboard from empty

Starting from an empty directory, use logic2b to create a Vite analytics dashboard with the benchmark preset. Prefer scaffold_plan when available. The result must include a runnable framework shell, dashboard block, charts, theme, components.json and exact dependency versions. Finish with a production build.

12 objective checks

RankModelAgentScoreDuration
1gpt-5.6-solOpenAI Codex CLI294/300 (98%)737.6 s
2gpt-5.5OpenAI Codex CLI272/300 (90.7%)751.789 s
3gpt-5.6-terraOpenAI Codex CLI266/300 (88.7%)430.712 s

Protocol v1.1.0: 3 tasks, 300 points.Score ranks before duration; build success is recorded by the evaluator.

What is measured

The tasks progress from component installation to composition and full application scaffolding. They check registry provenance (data-slot, expected imports and transitive files), exact preset fidelity, semantic token use, accessible form naming, dependency pinning and evaluator-observed production builds.

The overall score is 300 points. Score ranks first; total duration breaks ties. This keeps a fast but broken result below a slower complete one.

Fair-run contract

  • Fresh isolated directory for every model and task.
  • Identical task text, preset and starting fixture.
  • Tool access (shell, network and MCP) recorded in run metadata.
  • No human corrections after the timer starts.
  • Failed and timed-out tasks retain their partial artifacts.
  • The benchmark operator—not the model—records build exit codes.
  • Model/provider version, agent host, OS and Node version accompany the result.

The repository now ships the versioned evaluator runner as well as the static scorer. It creates each fixture from the built registry in external staging, records its SHA-256, invokes the configured agent without a shell, enforces time/output limits and snapshots source before the observed build; dependencies and generated caches are excluded. It captures both agent and verification transcripts; symlinks are discarded and source over the configured artifact budget is never executed. The runner must itself be executed inside a disposable container or VM because the observed build step intentionally executes generated project code.

Reproduce or audit

The protocol, scorer, test fixture contract and raw run format ship in benchmarks/agents. Run the harness tests with:

pnpm --dir benchmarks/agents test

Validate and execute an isolated agent configuration with:

pnpm --dir benchmarks/agents benchmark config.json --validate
pnpm --dir benchmarks/agents benchmark config.json
pnpm --dir benchmarks/agents score <run-id>

Real run artifacts are committed alongside their detailed score JSON. The first controlled comparison uses the same Codex CLI 0.148.0-alpha.15 host and capabilities for all three models: gpt-5.6-sol leads with 294/300 (98%), followed by gpt-5.5 at 272/300 (90.7%) and gpt-5.6-terra at 266/300 (88.7%). Every evaluator-observed build passed. The rule evidence still shows meaningful differences: Sol only canonicalized the legacy preset id, 5.5 lost scaffold fidelity on the starter/theme/pins, and Terra produced a runnable but non-canonical dashboard rather than the requested registry block. Synthetic fixtures remain tagged and mechanically excluded from publication.