Commit graph

9 commits

Author SHA1 Message Date
Admin
76627b6f19 ci: a driven app takes every key once, and the storage module builds for the browser -- the mini's wm script went red one run in three at the browser launch, and the shell menu's new log lines said why: the filter read "bbrowser" and the menu had opened twice. The remote bridge applied every wait=1 input first and, when the frame after it could not be sealed (the window was busy presenting the warm browsers), answered "requested input frame could not be submitted; retry"; the CI driver took that at its word and sent the input again, so a busy app took the Cmd+Space and the first letter twice. The bridge now keeps the waiters of an applied input and asks for the frame on the next beat, as it already did for a drawable still being acquired; the driver never asks an input route twice, whatever the answer, and still retries a grab it could not place. The workspace row was orange and apps/scope red on wasm32 since 8d7246231: the public volume_available_bytes had been put between the not(wasm32) guard and the native module it guarded, so the module compiled in the browser with nothing using it (17 warnings) and the function it exported was missing there; the guard is back on the module, and the browser has a volume_available_bytes that says the free space is not known 2026-09-22 19:04:04 +02:00
Admin
611a4e2035 ci: the tests are the platform's own, in release, in two minutes -- the user's law: "we mostly only care about the platform tests", "a max of about 2 minutes of tests", "test everything the CI does in release builds; debug builds are uselessly slow for this". ci.test now builds every test binary of its selection ONCE in release and runs them several at a time (half the cores, at least two), each one timed and capped, so a hung test costs its own cap and never the hour, and the step's detail names the slowest binaries with their counts; the doc tests follow through cargo for the packages that have a library. A selection is dirs (the packages whose manifests live under those directories, by cargo metadata), packages, package or workspace: true; a run over budget_secs (120 by default) turns the block orange and says so. The root script runs platform, draw, widgets and tools/ci by default: 49 binaries, 2535 tests, 22 s on a laptop. Every other crate's tests are the deep run, opt-in with deep_tests = true in ci.toml or ci --deep (ci.deep in scripts). The wall's cards lose a line: the time sits beside the count, 1 passed · 12s
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-22 09:50:38 +02:00
Admin
c20cf99c89 ci: a judgement is never lost to the way it was phrased, and every "retry" of the bridge is retried -- the small vision model, asked to describe the desk's terminal, locked onto one phrase and repeated it to the token cap, so its answer had no verdict line and the window manager tile went red over a terminal that was fine; the decoder now stops a looping answer (the tail of the text is one short pattern over and over), and when the free answer has no verdict line the conversation is continued with "VERDICT:" already written and the model finishes that line, so the judgement is asked for outright instead of being lost. The harness retries every answer of the bridge that ends in "; retry" (a grab it could not arm as well as an input frame it could not place), where it matched one wording only and failed the browser step on the other. The workspace test suite gets three hours instead of one: it runs beside the app scripts now, and a hung test was worth a whole hour before
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-22 08:46:45 +02:00
Admin
863ca812a3 ci: the workspace script gives the wall back after warming -- after a push that touches a core crate the workspace script spent ten minutes and more alone (fifteen target checks, a release build of every app, the workspace check and the whole test suite) while every other tile sat grey, which read as a hang. It still goes first and alone for the part only it may do alone, warming the cache every other script reads, and then calls ci.shared(): its permit turns from exclusive into an ordinary one and the waiting scripts start beside its checks and tests. The runner sends it through the same pool as the rest and hands out nothing else until it holds the wall, instead of running it to the end before the pool even started. It also builds the pty helper while it is alone, so the terminal, director and window manager scripts get a cache answer instead of a cargo run that would queue behind the test build
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-22 00:07:23 +02:00
Admin
c091430d32 ci: a quiet header, a build that visibly moves, and scripts that leave the machine alone -- the window's top is one band: the branch and its tip, the test progress large (17 / 29) with failed and warning counts only when there are any, the run's elapsed time, Run now and Stop; the line under it is the run progress bar and the tiles start right below it, the rows of counts, "now:" and "Watching" are gone and next poll and the model's state moved to the one footer line; a running step carries what its command is doing ("compiled makepad_draw · 212 crates", told at most twice a second from cargo's artifact messages) and the running tile shows it, so a ten-minute warm build no longer looks stuck; director embeds a terminal, so its script builds the pty helper as the terminal's does; ci.launch takes app_env, and Files runs on its synthetic tree (MAKEPAD_FILES_DEMO) instead of walking the real home, which made macOS ask the person at the CI box for access to Downloads on behalf of "release"; and the workspace tests run with --no-fail-fast, so a run names every failing test target instead of the first
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 23:41:02 +02:00
Admin
7ffcdab2ca ci: a stop or a restart is not a test result -- restarting the CI app in the middle of a run painted the whole wall red and then left it red: every script still going failed with "stopped by user", a script that had not launched yet got nil from ci.launch and raised method-not-found errors that counted as real failures, and the interrupted tip was recorded as tested, so the new process had nothing to run. Now a failure recorded while the run is already stopped is a skipped step, never a failed one; ci.launch hands a stopped script the inert app; a run the process shutdown interrupts leaves the last finished run as the record and its tip untested, so the next start runs it again; and a run the user stops keeps what finished, leaves the rest untested and says "stopped before it finished" instead of passing
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 23:15:29 +02:00
Admin
5165eaf349 ci: the terminal gets its pty helper, director is a desktop tool, and a refused input is asked again -- on macOS the terminal starts its shell through makepad-screen, which lives next to it and which an app build alone does not produce, so the terminal and window manager scripts build it (the CI box had a terminal that could not spawn a shell and a desk whose terminal window stayed blank); director drives cargo builds and child processes and most of it is compiled out on web and mobile, so its script asks for the desktop targets only; and when an app answers that it could not place an input on a frame and says retry, the harness asks again, up to five times, instead of failing the step
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 22:44:50 +02:00
Admin
0458f0c2b7 ci: a run is one target dir and one cargo batch per target, and the wall draws -- the root script warms a run-level cache with ONE cargo check per target covering every app package (the matrix's BuildTy per row, diagnostics attributed by package_id, --keep-going then per-package fallback when one package fails) and one release build of all bins, so the 24 app scripts answer check_targets and build from the cache; the tile grid registers its draw type on DrawQuad and merges its shader into the widget, which is what makes the green / orange / red / grey tiles appear at all; the window takes the left half of the main display (system_profiler, overridable with window = [x, y, w, h]); only apps/<name>/Cargo.toml is an app, so a private app's nested deps are its libraries; scope is no longer skipped by default; a failed launch hands the script an inert app instead of nil and a failure a step already carries is said once; the log keeps the first 20 diagnostics per cargo invocation and hub-install counts a present file once
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 21:52:43 +02:00
Admin
229cdd4d2e ci: tools/ci, a Splash runtime that watches a branch and runs every ci.splash it finds -- each script checks its app on every other platform, builds and runs it here, drives it over --remote and has a small local vision model judge the grabs
A test is a ci.splash beside what it tests; the script decides every input and the model only ever judges one picture against one acceptance text. mod.ci: launch (hidden, --remote, user_seq preserved), key, type_text, click, get, snap, wait_log (a * is a gap inside one line), no_errors, grab, quit; step, sleep, check, run; cargo, check_targets (the cargo makepad check matrix, check only for platforms we are not on, a test fails if the two tables drift), test, build, machine (another box over the makepad tunnel), exclusive; judge, accept, ask. The watcher polls git ls-remote once a minute for work and any extra branches, syncs a checkout the CI owns, runs the root script first and alone, then the rest up to a parallel limit behind one shared model judge. The window is a wall of squares, one per script: green passed, orange warnings, red failures, with a detail panel for the selected one.

Scripts: the root ci.splash (workspace check with core warnings denied, the tests), apps/wm (desktop up, switch to macOS by Cmd+Space / type / Return, launch the terminal and the browser, each waited for by the WM's own first-frame line), and one per main app in the default shape. Proven here: apps/wm/ci.splash green in 280 s, fifteen target checks and seven vision verdicts.

Models come from Hugging Face through the hub: registry entries qwen3.5-4b-vision and qwen3.5-9b-vision with exact revisions, sizes and digests, and hub-install, a command line over LocalModels::start_install. vlm-probe reads PNG.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 21:23:11 +02:00