A test is a ci.splash beside what it tests; the script decides every input and the model only ever judges one picture against one acceptance text. mod.ci: launch (hidden, --remote, user_seq preserved), key, type_text, click, get, snap, wait_log (a * is a gap inside one line), no_errors, grab, quit; step, sleep, check, run; cargo, check_targets (the cargo makepad check matrix, check only for platforms we are not on, a test fails if the two tables drift), test, build, machine (another box over the makepad tunnel), exclusive; judge, accept, ask. The watcher polls git ls-remote once a minute for work and any extra branches, syncs a checkout the CI owns, runs the root script first and alone, then the rest up to a parallel limit behind one shared model judge. The window is a wall of squares, one per script: green passed, orange warnings, red failures, with a detail panel for the selected one.
Scripts: the root ci.splash (workspace check with core warnings denied, the tests), apps/wm (desktop up, switch to macOS by Cmd+Space / type / Return, launch the terminal and the browser, each waited for by the WM's own first-frame line), and one per main app in the default shape. Proven here: apps/wm/ci.splash green in 280 s, fifteen target checks and seven vision verdicts.
Models come from Hugging Face through the hub: registry entries qwen3.5-4b-vision and qwen3.5-9b-vision with exact revisions, sizes and digests, and hub-install, a command line over LocalModels::start_install. vlm-probe reads PNG.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>