
# How Beamline UI is tested

What we have checked, with dates and scope. The runs below are separate and some cover earlier versions; they are not
one all-green run of the current release. What has not been checked yet is on [known limitations](https://beamline.io/docs/limitations).

| check | what it covers | latest result |
|---|---|---|
| Contrast | text, edges and focus rings, 16 pairs, in all 9 looks, dark and light | all pairs pass in all 18 look and theme pairs, 10 October 2026 |
| Accessibility | Lighthouse (lab) on beamline.io: the home page, the library, a component page, about | accessibility 100 on the home page, 93 on the library, 100 on a component page, 96 on about; 10 October 2026 |
| Speed and motion | every example of every part, frame by frame at 120 Hz | 581 scenarios with recorded limits, 10 October 2026 |
| Next.js | a fresh `create-next-app` (App Router) installs parts, builds, starts | 13 of 13 checks, Next.js 16.4, 10 October 2026 |
| Vite | the release installed into a Vite app with plain npm, every screen driven in a browser | 49 of 49 checks, dark, light and phone, 0.2.0-rc.31, 10 October 2026 |
| Coding agents | Claude and Codex build apps from a written brief | 29 recorded trials, 6 to 8 October 2026 (not a pass rate) |
| Browsers | Chromium for the checks above; Firefox 155 and WebKit 26.6 on Linux | 220 gallery routes at desktop and phone width in each engine, 10 October 2026; limits below |

## Contrast in every look

A script opens the library in Chromium and measures 16 colour pairs in each of the 9 looks, in dark and in light: body
and secondary text on the page and on cards, text on the accent and status colours, and edges and focus rings. Text
must reach 4.5:1 and edges and focus rings 3:1, the WCAG 2.2 AA levels. A brand colour you set keeps its text at 4.5:1
or more in that token calculation ([Customise](https://beamline.io/docs/customize)); rendered controls still need checking in context.

These are dated, partial automated checks, not a claim of WCAG conformance for the whole library; a wider automated
accessibility audit was still running on 10 October. Focused reruns confirmed named repairs: disclosure names, toast
announcements, combobox and popover names, and focus return from phone dialogs.

## Speed and motion, frame by frame

The bench renders every example of every component and screen, plus interactions and the site's own pages, in a
browser that draws every frame itself at 120 Hz on a virtual clock. Frame counts, layout work, style work and bytes are
recorded for each one, so a change that slows a part down, alters how it looks or how it moves shows up as a number
before it ships. Parts move only when their state changes, so a page left open does almost no work. What the first
full run found is in [what measuring every example taught us](https://beamline.io/guides/react-performance-lessons).

## Installs into real apps

A fresh `create-next-app` (App Router, TypeScript, Tailwind) installs 15 parts, builds with every page prerendered, and
starts with no hydration warning or console error. The release itself is copied to an empty folder, installed into a
Vite app with plain npm, type-checked, built, and every screen of the reference app is driven in a browser in dark,
light and at phone width. Beamline is also a shadcn registry ([Install](https://beamline.io/docs/install)).

## Firefox and WebKit: 10 October 2026

Playwright 1.63.0 on Linux, Firefox 155.0 and WebKit 26.6, starting in dark at 1512×982 desktop and 390×844 with touch.
There were 220 unique gallery routes (189 entries, 15 categories, 14 Foundations pages and 2 indexes) and 162 unique
interaction scenarios, each exercised at both widths. Some workflows deliberately resize.

| engine | desktop routes | phone routes | desktop interactions | phone interactions |
|---|---|---|---|---|
| Firefox 155.0 | 220/220 | 220/220 | 162/162 | 161/162 |
| WebKit 26.6 | 220/220 | 220/220 | 162/162 | 159/162 |

This combines complete sweeps on the rc.32 base with targeted replays after three fixes: missing speech-recognition
fallback, phone accent retention and a combobox closing error. It is not one clean run of a later integrated release,
and entries added afterwards are outside its scope. Remaining phone failures were Foundations overflow in both
engines, a WebKit dialog opening about 5 px from its intended origin, and a 90 ms Stepper caption animation under
reduced motion. Glass dialog/background readability also remains a visual finding.

The same day's public-site checks rendered all 40 non-library pages at both widths on rc.32. Desktop passed 40/40 in
each engine; phone passed 33/40 in Firefox and 32/40 in WebKit, with horizontal overflow on the remaining pages. These
are historical results, not a claim that later repairs have been retested in that matrix.

Actual Safari on macOS/iOS and real iPhones were not tested. Native speech recognition was absent in both Linux
engines: the checks covered manual-text and missing-feature fallbacks, demo voice/no-device states and local audio
playback, not microphone-present transcription, paid providers, signed-in account operations or purchases.

## Coding agents building from a brief

Each trial gives an agent a real app and a written brief, with Beamline connected the way a buyer connects it. Build
and audit results say the code is sound; they do not prove every screen works.

- **Codex**, 7 October: turned a plain Next.js invoicing app into a dashboard, an invoice list, an invoice editor and
  settings in 12 minutes; build, lint, its own tests and the audit passed. A second brief, "dark, Prism, violet, rounder,
  calmer", reached every screen by changing one root setting.
- **Claude Sonnet 5.5 and Haiku 5.5**, 8 October: the same invoicing brief; both built all four screens with the build
  and the audit passing. One of Haiku's screens failed when opened; the part it misused now names the fix.
- **Claude Opus 5.5**, 6 October: a fitness intake form from 37 different parts, with no native controls or hand-set
  colours.

Every finding went back into the guide, the parts or the audit. These are single, dated runs, not a success rate or a
model ranking; a larger evaluation is under way and its numbers are not published yet.
