Skip to content

test: contract tests and a benchmark harness (no src changes) - #68

Merged
chiefcll merged 17 commits into
mainfrom
main-bench-tests
Oct 4, 2026
Merged

chiefcll merged 17 commits into
mainfrom
main-bench-tests

Conversation

@chiefcll

@chiefcll chiefcll commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Why

This is the part of the 1.7 work that does not depend on 1.7: the contract
tests and the CPU benchmark harness written for it, adapted to run on main
(solid 1.6.4 on @solidtv/renderer 1.9). Landing them on main first shrinks
the 1.7 diff to the behaviour and performance changes themselves, and gives
1.7 a baseline: every test here passes (or is skipped with a note) on 1.6.4,
so whatever 1.7 changes shows up as a test it has to update or unskip, and
the benchmark measures 1.7 against the 1.6.4 release on the same scenarios.

What

  • Contract tests (tests/contract-*.test.tsx, tests/publicApi.test.ts),
    through the public surface only: key handling and focus, states and styles,
    Row/Column/Grid/VirtualRow/VirtualGrid/Lazy navigation and scrolling, flex
    on both engines (flex.ts and flexLayout.ts), text props on the DOM
    renderer, node behaviour, the runtime export names of every entry point and
    the package.json exports map. +432 tests.
  • Type contract (tests/contract-types.tsx): the public types apps augment
    with declare module '@solidtv/solid', and every type-only export name.
    pnpm tsc now runs tsc && tsc -p tests/tsconfig.contract.json, so the
    build, CI and the release (prepack → build) type-check it.
  • Known 1.6.4 bugs and gaps are skipped with a BUG: or CONTRACT GAP
    note
    and the correct behaviour (23 skipped, 1 todo): fixing one means
    unskipping its test.
  • Flex with text on the real WebGL renderer (tests/webgl,
    pnpm test:webgl): headless Chromium, renderer 1.9, an SDF font; final
    positions of text in flex containers, where SDF and DOM measurement differ.
  • Test hygiene: spatialNavigation disposes its render and useHold.spec
    mocks a fresh module, so files pass in any order in one worker; the jsdom
    and browser configs exclude tests/webgl and bench, and the browser run
    leaves out the contract files that pin the jsdom DOM renderer.
  • npm run bench (bench/, methodology in docs/perf/README.md):
    builds a bench app for arm A (the 1.6.4 release on npm renderer 1.9.3)
    and arm C (this working tree on the installed renderer), and measures
    key-press, state-change, node-creation and text-in-flex scenarios in
    headless Chromium under 6x CPU throttling: time per op (handler, microtask
    tail, frame), bytes allocated per op by owner, CPU profile self and
    inclusive time, and counts (flex passes, renderer writes, text layouts,
    loaded events).
  • devDependencies: playwright pinned to 1.56.1 (the version
    @vitest/browser-playwright already used), terser, and
    @lightningtv/vite-hex-transform (it brings a second vite, 5.x, into the
    lockfile). Nothing new ships to npm: no runtime dependency, no src/
    change.

How to run

npx vitest run                                   # jsdom suite incl. the contract tests
VITE_USE_NEW_FLEX=true npx vitest run tests/contract-flex.test.tsx   # the JSX part on flexLayout.ts
pnpm tsc                                         # also type-checks tests/contract-types.tsx
npx playwright install chromium                  # once, for the next two
pnpm test:webgl                                  # WebGL flex/text tests
npm run bench -- --scenarios smoke --runs 1 --quick   # harness check
npm run bench                                    # full series: arms A,C, 3 runs, all modes

CI runs build, test and lint only: test:webgl and bench are manual.

Notes

  • Arm A uses npm renderer 1.9.3; arm C uses whatever is installed, and the
    lockfile resolves the ^1.9.0 peer to 1.9.0. The two differ in small
    ways (loop pause/resume, the WebGL power preference, an enableClear
    branch). The summary lists each arm's Solid and renderer, and the A/C ratio
    is a noise floor only when they match.
  • docs/perf/README.md is not linked from the docs sidebar.
  • The Roboto (Apache-2.0) and Lato (OFL-1.1) MSDF atlases are copies of
    solid-demo-app's; each font folder has a README with the license and source.

Size and what to read

76 files, about +34,000 lines, but most of it is data: three MSDF font atlas
JSONs (+17,552 lines; with their PNG pages about 570 KB of fonts), 12 small
placeholder images for the bench (48 KB), and the lockfile (+633). The code to review is the test files
(about +10,700 lines, mostly tables of expected values), bench/ (about
+4,600 lines: run.mjs, harness/*.mjs, vite.config.ts,
prepare-arms.mjs, the scenarios) and docs/perf/README.md. Suggested
order: docs/perf/README.md, then one contract file (contract-keys or
contract-focus), then bench/run.mjs and bench/harness/probe.mjs.

Not included

No src/ change and no bug fix: the BUG: tests stay skipped. No renderer
2.0 code (the bench's renderer 2.x count hooks are kept, labelled, and never
install on 1.x). No benchmark results: summary.md is meant to be committed
with a measured series; raw JSON and profiles are gitignored.

Gates: npx vitest run 654 passed / 23 skipped / 1 todo (also in 3 shuffled
file orders and with VITE_USE_NEW_FLEX=true), pnpm tsc clean, pnpm lint
0 errors, pnpm test:webgl 14/14 on both flex engines.

🤖 Generated with Claude Code

chiefcll and others added 17 commits October 4, 2026 12:06
The suite runs without isolation. spatialNavigation left its focus
manager listening on document, and useHold.spec's mock missed a
useHold module another file had loaded. Exclude bench/ (the benchmark
arms carry their own copies of the tests) and the browser attachments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…erer

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pnpm tsc now also type-checks tests/contract-types.tsx, which augments
each public type with declare module '@solidtv/solid'.

On main the type-only export lists are the ones of renderer 1.9: the
renderer's names re-exported through src/core/index.ts are its 99, and
the Shader*Props names and WebGlShader are Solid's own (src/core/shaders.ts).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pnpm test:webgl runs tests/webgl in headless Chromium against
@solidtv/renderer (renderer 1.9 on main: WebGL engine, SDF text) with
an SDF font, and pins the final positions of text in flex containers:
DOM and SDF measurement differ. The numbers were read on the renderer v2
lockstep and hold unchanged on renderer 1.9, on both flex engines.

On renderer 1.9, settle() waits on Stage.hasSceneUpdates(), and the
late-font test starts the font load before the render: 1.9 creates a
text only for a family that is loaded or loading.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The harness (next commits) builds with terser and the hex-colour
transform, as solid-demo-app does, and drives Chromium through
Playwright, pinned to the version @vitest/browser-playwright already
uses. Its arms (bench/.arms), builds (bench/dist) and raw results stay
out of git, prettier and eslint; its Node scripts get Node and browser
globals (they hold functions that run in the page).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 1.7 branch's harness, on main's two arms: A is the pinned 1.6.4
release (71c170f) on npm @solidtv/renderer 1.9.3, C is the working
tree on its installed renderer (rendererMajor read from that renderer's
package.json; src/arm-v1.ts is the only bootstrap). Each arm is built
with Solid from source, chunked by owner (framework, reactivity,
renderer, user), and run in a fresh headless Chromium with CDP CPU
throttling, in time, alloc, profile and count modes; --save-profiles
and harness/inclusive.mjs give inclusive time per function, --flex
old|new picks the flex engine, BENCH_TERSER_COMPRESS tries terser
options. The summary reports A/C ratios.

The count hooks for renderer 1.x now also find the shader props:
renderer 1.x defines them as non-enumerable accessors, which the
props swap missed, so shader writes read 0 and a shader updated after
the swap threw (page-mount's count runs failed). The renderer 2.x hook
paths in harness/probe.mjs stay, labelled v2; they never install on 1.x.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
nav: a Column of Rows with every scroll mode, a $focus Thumbnail row,
VirtualRow and VirtualGrid windows, a NavDrawer states toggle,
onFocusChanged driving text colours, Poster and whole-page mount and
swap; each probes its state, so the summary flags a run or an arm that
ends somewhere else. text: text in flex containers (a Row of text tiles
remounted with the same or new strings, a details panel whose text
changes, a VirtualRow of text tiles); the runner adds the text metrics
for these. All use only the public API both arms share.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
How npm run bench builds and runs the two arms, what one operation is,
what each mode measures and how (time, alloc, profile, count and its
hooks on renderer 1.x), why times are means under 6x throttling, and
where the noise comes from.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… out of the browser run

npm run test:browser picked up tests/webgl (run by vitest.webgl.config.ts)
and the copies of the tests in bench/.arms. Six contract files pin the DOM
renderer that the jsdom run sets up (SOLIDTV_DOM_RENDERING) and fail when
Solid renders through WebGL, and publicApi needs node:fs: those are left
to the jsdom run. The other contract files pass in the browser.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The contract tests came from the 1.7 branch, and their comments named
its plan ("Phase 1", "arm B", "the brief", "renderer v2"). Reworded in
place, same line counts: they pin solid 1.6.4 on renderer 1.9. The
late-font WebGL test no longer claims the texts wait for the font: the
files are already fetched, so the font can arrive before the first
frame; only the layout after it is asserted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…main

The A/C note no longer assumes the two renderers match: the summary
header now lists each arm's Solid revision and renderer version (arm C
records the installed version, not its local path). The renderer 2.x
count hooks in probe.mjs say they are inactive on 1.x, and the scenario
headers no longer cite the 1.7 brief.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The bench's Roboto (Apache-2.0) and the WebGL tests' Lato (OFL-1.1)
MSDF atlases are copies of solid-demo-app's; each folder now says what
they are, under which license, and where the fonts come from.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With isolate: false a later file in the same worker loaded useHold
against the mock, so its holds never suppressed the key
(keySuppression's webOS Back hold failed in CI's file order).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@chiefcll
chiefcll merged commit 42dd504 into main Oct 4, 2026
1 check passed
chiefcll added a commit that referenced this pull request Oct 6, 2026
Main brought the contract tests and benchmark harness from 1.7, adapted to
1.6.4 on renderer 1.9 (#68), the per-side boundsMargin tuple on the DOM
renderer (#67), restoreFocus skipping destroyed elements (#69), the
renderer 1.10.1 requirement (#70) and version 1.6.5. Resolution, keeping
1.7's behaviour and the smallest diff against main:

- Tests: merged against the Phase 1 commits main started from (a12c778),
  so only main's rewording applied. 1.7's assertions and expected values
  win. Main's headers are kept where true on 1.7; "1.6.4 (renderer 1.9)"
  becomes "1.7 (renderer 2.0)". contract-types.tsx keeps 1.7's renderer 2
  type names (main added no Solid-side name that 1.7 lacks);
  tests/webgl/setup.ts and the late-font WebGL test stay 1.7's (renderer 2
  lets a text wait for its font). tests/focusStack.test.tsx from main
  passes on 1.7.
- bench/: main's files, plus arm B (f1c8ab0 on renderer faf4b9f), arm C on
  the linked ../renderer-v2-solid (main's installed-renderer lookup, plus
  1.7's rebuild of its dist and its revision in the record), skipC for the
  demo scripts, the A/B and B/C ratio table, default --arms A,B,C.
  Main's probe fix (renderer 1.x shader props are non-enumerable accessors)
  and the summary's per-arm Solid/renderer lines are kept; the renderer 2
  hooks install on arms B and C. bench/src/arm-v2.ts, bench/demo/* and
  bench/micro/* stay.
- docs/perf/README.md: main's text plus arm B, the per-major count hooks,
  the 120 Hz note and a section for bench/demo and bench/micro.
- package.json: version 1.6.5; @solidtv/renderer stays 1.7's (peer
  ^2.0.0-alpha.0, devDependency link:../renderer-v2-solid). pnpm install
  regenerated the lockfile (unchanged from 1.7's).
- .gitignore, eslint.config.js: union (main's recursive results ignore,
  1.7's .superpowers ignores).
- src/: #67 and #69's createFocusStack change auto-merged. #69's DOMNode
  `destroyed` field and markDestroyed() are dropped: 1.7's DOMNode already
  has a `destroyed` getter (this node or an ancestor destroyed), which
  restoreFocus reads through ElementNode.destroyed. Renderer 2.0 still
  takes one boundsMargin number (the tuple's widest edge, with a warning),
  so the tuple works on the DOM renderer only.
- MIGRATION-1.7.md: the boundsMargin tuple is a break inherited from
  renderer 2.0; the DOM renderer's `destroyed` was undefined before 1.6.5.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@chiefcll
chiefcll deleted the main-bench-tests branch October 6, 2026 14:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant