Whole-tree sync: cad-core/cad-ui split sources, nigig-build construction_frame migration, pdf port progress, mpesa/pay/uikit/doc updates, workspace members/profiles/lock, CI workflows and reviews. See individual file history for details.
20 KiB
CAD benchmark baseline
Reproduce with either:
# needs the Makepad desktop stack (wayland/X11/GL/alsa/polkit)
cargo test --locked --release -p nigig-build --lib -- \
--ignored --nocapture --test-threads=1 profile_benchmarks
# host-only: isolated toolchain, no desktop, self-cleaning
CAD_BENCH=1 ./tools/test-cad-coverage.sh
The second form exists because these benchmarks were unrunnable for ten
review tranches — REVIEWS/PAY_CAD_IMPLEMENTATION_STATUS.md recorded the
blocker as "no Cargo toolchain is available in this execution
environment". profile_benchmarks.rs turns out to import nothing from
Makepad but the math types, so it builds in the same host-only harness
the coverage run uses.
Release mode matters. Debug builds are 3–5× slower here and the ratios between arms shift, so a debug number is not comparable to anything below.
These are measurements, not assertions — the benchmarks are #[ignore]-d
and deliberately contain no timing assert!s (a wall-clock assertion in
the normal test run flakes under CI load). Compare against this file by
hand when changing the hot paths.
The frame-submission rows below were recorded on the Phase 5 tree,
x86-64, --release. The wall-clock rows above are retained machine
baselines and are not comparable across hosts; the submission rows are
structural counts and are intentionally stable.
| Benchmark | Result |
|---|---|
bench_scene_cache_hit_vs_rebuild |
cold 25.9 µs / warm 53 ns (100 parts) — 488× |
bench_scene_cache_scaling |
10/50/100/500 parts: cold scales linearly, warm stays ~150 ns |
bench_mesh_cache_hit_cost |
11.1 µs/frame, 110 ns per part (100 parts, all hits) |
bench_size_parametric_vs_mesh_derived |
parametric 2 ns vs mesh-derived 886 ns — 443× |
bench_pick_broadphase_mesh_bounds_vs_size |
63.3 µs vs 131.0 µs per frame — 2.1× saved |
bench_pick_broadphase_world_aabb_recompute_vs_cache |
230 µs → 39 µs per pick at 500 parts — 5.9×. Transforming 8 corners per part per pick, versus a PlacedHash lookup. Hover picking runs on mouse-move, so this is per-frame. Cached in SceneCache::world_aabb_for (Phase 3.2). |
bench_parts_script_regeneration_per_drag_frame |
42 µs (50 parts) / 170 µs (200) / 425 µs (500) per drag frame, before the editor's own set_text. script_dirty was set on every motion event; it is now set once on commit (Phase 3.6). |
bench_param_hash_cost_per_frame |
Box 50 ns/node (25 µs/frame at 500) — not worth a cache. Extruded 64-gon 987 ns/node; that is the vertex data, and bulk-hashing the slice measured no faster. Phase 3.3 left undone deliberately. |
bench_parallel_threshold_warm_vs_cold_cache |
200 nodes: cold 8.80 ms seq / 6.21 ms par (1.42×), warm 5.81 / 5.73 (1.01×). The 32-node proxy is imprecise but harmless; Phase 3.8 rewrite not worth it. |
bench_glb_export_with_and_without_cache |
cold 1.42 ms / warm 1.35 ms (50 parts) |
bench_parallel_vs_sequential_export |
sequential 4.83 ms / parallel 3.58 ms (100 parts) — 1.3× |
bench_command_execute_overhead |
0.8 µs per execute, 0.5 µs per undo (1000 commands over a 1000-node scene). Was 254 µs / 246 µs: CadCommandCtx::new eagerly built a scene snapshot no production command reads. Now built on first scene() call — roughly 300×. |
bench_gpu_upload_mesh_source |
boxes: rebuild 0.27 µs/part vs warm cache 0.178 µs — 2×; extruded 24-gons: 1.11 µs vs 0.516 µs — 2×. Cold cache 0.79 µs/part is slower than rebuilding (hash + insert) |
bench_delete_invalidation_clear_vs_evict |
deleting 1 of 100 parts: clear+rebuild 52.9 µs vs evict+reuse 14.5 µs — 4× |
bench_viewport_snapshot_sync_per_frame |
deep copies 21.2 µs/frame + forced rebuilds 77.2 µs/frame = 98.5 µs/frame; generation-checked 0.051 µs — 1930× (100 parts) |
Frame submission budget
bench_frame_submission_budget, added by Phase 0 of
REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md and extended through Phase 5.
Every other benchmark in this file measures CPU work behind the
renderer; none measures a live GPU frame, so the submission report stays
as an explicit structural count rather than pretending to be frame time.
Counts, not timings — counts do not flake under CI load. 1920×1080, parts on a lattice across a 200 m site (deliberately larger than the viewport at either zoom: a scene that fits on screen is the case where culling can do nothing, and reporting only that would flatter it).
2D, after Phase 1 culling, Phase 4 stroke batching, and Phase 5
sub-pixel LOD. "Before" is the Phase 0 baseline: every part in the
document submitted, one stroke() per part and one per grid line.
| Zoom | Parts | Visible | Full outlines | LOD points | Part calls | Tessellations | Before |
|---|---|---|---|---|---|---|---|
| 5 m (170 grid lines) | 100 | 1 | 1 | 0 | 1 stroke | 3 (2 grid + 1 parts) | 270 |
| 500 | 4 | 4 | 0 | 2 strokes | 4 | 670 | |
| 2,000 | 12 | 12 | 0 | 2 strokes | 4 | 2,170 | |
| 200 m (270 grid lines) | 100 | 100 | 1 | 99 | 1 stroke + 1 fill | 4 | 370 |
| 500 | 500 | 1 | 499 | 1 stroke + 1 fill | 4 | 770 | |
| 2,000 | 2,000 | 1 | 1,999 | 1 stroke + 1 fill | 4 | 2,270 | |
| 500 m (170 grid lines) | 100 | 100 | 1 | 99 | 1 stroke + 1 fill | 4 | 270 |
| 500 | 500 | 1 | 499 | 1 stroke + 1 fill | 4 | 670 | |
| 2,000 | 2,000 | 1 | 1,999 | 1 stroke + 1 fill | 4 | 2,170 |
The benchmark models one material plus one selected part. Selection is intentionally kept at full detail, so the first visible part remains an outline while the ordinary parts cross the LOD threshold. At 5 m, a 1 m part is comfortably resolvable and every survivor stays full detail. At 200 m, the same part is only about 2.7 logical pixels across; at 500 m it is about 1.1 pixels, so the ordinary survivors become 2×2 logical-pixel filled markers.
Phase 5 does not claim another draw-call win. The 2D scene remains one
DrawVector::end() submission, and the one material still makes one
outline stroke plus one marker fill. The win is geometry volume: a filled
marker is bounded and does not carry a full rectangle/circle outline for
a part whose shape cannot be resolved. FrameBudget reports full
outlines and LOD markers separately so this change cannot be mistaken for
another tessellation reduction.
3D, frustum culling (Phase 1), instanced batching (Phase 3), and bounding-box LOD (Phase 5), 60° fov, camera orbiting the site centre. "Shapes" is how many distinct full-detail solids the visible parts use — 6 for a catalogue-built model, all distinct for the worst case. LOD boxes are all instances of one shared unit-cube proxy batch:
| Camera | Parts | Visible | Full meshes | LOD boxes | Draw calls, 6 shapes | Draw calls, all distinct | Before |
|---|---|---|---|---|---|---|---|
| 20 m | 100 | 39 | 39 | 0 | 6 | 39 | 100 |
| 500 | 198 | 198 | 0 | 6 | 198 | 500 | |
| 2,000 | 733 | 733 | 0 | 6 | 733 | 2,000 | |
| 80 m | 100 | 76 | 76 | 0 | 6 | 76 | 100 |
| 500 | 386 | 386 | 0 | 6 | 386 | 500 | |
| 2,000 | 1,459 | 1,459 | 0 | 6 | 1,459 | 2,000 | |
| 400 m | 100 | 100 | 20 | 80 | 7 | 21 | 100 |
| 500 | 500 | 63 | 437 | 7 | 64 | 500 | |
| 2,000 | 2,000 | 380 | 1,620 | 7 | 381 | 2,000 |
The bottom row is the one to look at. At 400 m the whole site is on screen, so culling removes nothing. LOD replaces 1,620 unresolvable meshes with instances of one proxy, and the six full-shape groups plus that proxy draw in 7 calls instead of 2,000. The "all distinct" column is the same scene with every full-detail part a different size; 380 remain full and the 1,620 tiny parts still share the proxy, so the count is 381. The selected/hovered exemption is not present in this synthetic 3-D benchmark; the policy tests pin that decorated parts remain full detail.
Where culling does work it is decisive on its own: drafting at a 5 m zoom over a 200 m site, the 2,000-part scene queues 12 outlines instead of 2,000. Where it cannot help — everything on screen — batching still holds the call count at 4.
Two different costs, deliberately not summed:
- Tessellations are CPU calls into
tessellate_path_stroke. The whole 2D scene is one draw call fromDrawVector::end()no matter how many parts, so this — not draw calls — was the number that scaled there. Since Phase 4 it scales with colour groups instead. - Draw calls are GPU submissions. 3D used to be one per uploaded, visible part; since Phase 3 full-detail parts use one call per distinct shape, with the parts riding inside as instance rows. Since Phase 5, all tiny parts use one additional shared proxy batch.
The grid line count is unchanged by culling, as it should be: it was
already virtualised, and stays at 170–270 lines across the measured zoom
range because the step adapts. Phase 4 changed what those lines cost —
two stroke() calls rather than one each — and Phase 5 changes only the
representation of tiny parts. The tests culling_shows_up_as_fewer_part_tessellations,
the_grid_costs_two_strokes_at_any_zoom, and the LOD policy tests pin
those separate costs.
Phase 2 (sharing geometry between parts of the same shape, keyed by a
new ShapeHash — not ParamHash, which includes the node id) and
Phase 3 (instancing) brought the draw-call column down in the zoomed-out
case where culling cannot; Phase 4 did the same for the 2-D tessellation
column. Phase 5 addresses the remaining geometry volume in both paths:
the 2,000-part rows at 200 m and 500 m keep one selected outline but
replace 1,999 unresolvable outlines with bounded point markers, while the
400 m 3-D row replaces 1,620 unresolvable meshes with one shared proxy
batch. A real GPU/window visual check is still required before treating
the structural counts as frame-time results.
GPU geometry buffers, keyed by shape
bench_geometry_buffers_shared_by_shape, added by Phase 2. Counts, not
timings. The number it replaces was never measured because it was never
in doubt: one buffer per part, however many of those parts were the
same column.
| Model | Parts | Buffers | Ratio |
|---|---|---|---|
200 identical walls (build_house_scene(200)) |
200 | 1 | 200× |
| 420-part repetitive model (columns, walls in three lengths, two opening types) | 420 | 6 | 70× |
| 420 parts, every one a different size | 420 | 420 | 1× |
Read the last row too. Sharing is a property of the model, not of the code: a scene where every part differs gets nothing from this phase. That is the honest ceiling, and it is why the draw-call column is untouched here — Phase 3 (instancing) is what collapses draw calls, and it needs the shared buffers this phase produces.
The middle row is the one to plan against. Real architectural models are catalogue-built: a few column types, a few wall lengths, a door and a window repeated everywhere. 70× fewer GPU buffers, and 70× fewer mesh uploads on the first frame after a load.
Note: two benchmarks had drifted off the caching path
If you ran these between the Phase 5.1 generation-tracked store and the
commit that added this note, bench_scene_cache_hit_vs_rebuild reported
1.2× against the 488× in the table above, and
bench_scene_cache_scaling reported 1× at every part count. That
looked exactly like a catastrophic cache regression. It was not.
Both benchmarks called SceneCache::scene(&[CadNode]), which is
documented as always rebuilding — it takes a bare slice, so it has no
generation to compare and cannot cache. The caching entry point the
editor uses is SceneCache::scene_for(&PartsStore). When the store
landed, these two benchmarks were not moved with it, so their "warm"
sample was a second full rebuild and the printed speedup was allocator
noise.
Repointed at scene_for, they reproduce the table: cold 24.6 µs / warm
48 ns, 512×, and the warm read stays flat at ~55 ns from 10 parts
to 500. bench_scene_cache_hit_vs_rebuild now asserts Arc::ptr_eq on
the two samples, so it fails loudly rather than quietly measuring two
rebuilds if it is ever pointed at a non-caching path again.
SceneCache::scene() is not the bug and was left alone: its docstring
says what it does, and nine tests legitimately use it for exactly the
case it describes.
Note on bench_gpu_upload_mesh_source
Phase 5.3 routed the GPU upload path through MeshCache instead of
calling build_solid() directly. The measured speed win is 2× on a warm
cache and negative on a cold one — a first-ever build now pays a
ParamHash and a map insert on top of the meshing.
That trade is worth taking, but not primarily for speed:
ensure_part_geometryonly builds for parts with no uploaded buffer, so the cold case is a scene load or an undo/redo, not a frame in a loop.- The warm case is every part that was already exported, hovered or previously uploaded — which, during editing, is most of them.
- The real gain is that there is now one node-to-triangles path. Two
independent
matcharms over the same enum were free to drift, and nothing would have caught it: the preview would simply have disagreed with the export.
Do not "optimise" this back by special-casing cheap primitives. The 0.5 µs is not worth reintroducing the second pipeline.
Note on bench_viewport_snapshot_sync_per_frame
This benchmark measures the primitives the sync path is built from --
Vec<CadNode> clones and SceneCache::scene() rebuilds -- not the widget
method itself, which needs a live Cx and three real viewports. Its
printed numbers therefore do not move when the sync path is fixed;
they quantify what one frame of syncing costs and what the same frame
costs when a generation check is allowed to short-circuit it.
Read it as the size of the prize, not as a regression gate. The Phase 5.2
change is verified by behaviour tests
(replace_all_bumps_the_generation_so_destinations_can_compare,
mesh_cache_self_invalidates_on_parameter_and_transform_edits,
geometry_is_retained_for_surviving_ids_only), each negative-tested.
Reading these
The scene cache is doing its job. 488× on a warm hit, and the scaling run shows the warm path is O(1) while the cold path is linear. No work needed here.
MeshCache hits are not free. 186 ns per part is ParamHash::from_node
being recomputed on every lookup — it hashes the solid's every field, the
transform and the material, even when the entry is present and valid. At
100 parts that is ~19 µs per hover frame. Not the dominant cost today, but
it is the next thing to fix if picking needs to get faster: memoise the
hash on CadNode and invalidate it on mutation.
CadNode::size() is a trap. 2 ns for a parametric solid, 886 ns for a
CSG or extruded one, because those have no closed-form extent and the
function meshes the solid to derive a bounding box (added in Phase 1.7 to
fix unpickable parts). It reads like a cheap field access at the call site.
Anything calling it in a loop should take bounds from a cached mesh
instead — which is exactly what pick_part now does.
Export is already fast enough. 1.4 ms for 50 parts, and the mesh cache barely moves it because JSON/binary serialisation dominates, not meshing. The parallel path wins only 1.3× at 100 parts. Neither is worth tuning before the UI hot paths.
What Phase 3 changed
| Change | Effect |
|---|---|
pick_part takes AABB bounds from the mesh it already holds instead of calling size() |
2.1× less broad-phase work per hover frame; also more correct — see below |
Hover picking throttled by HOVER_PICK_MIN_MOVE_PX (3 px) |
Sub-pixel jitter no longer triggers a full ray cast against every triangle |
Removed cx.redraw_all() from the part-drag MouseMove handler |
No full-widget-tree relayout per motion event during a drag |
The AABB change also fixed a latent correctness bug. size() returns a
symmetric extent about the origin, but an extruded polygon grows along
+Y from its base plane — so the old broad-phase box was in the wrong
place and could reject a ray that actually hits. pick_bounds_tests:: mesh_bounds_are_asymmetric_where_size_is_not pins this.
Not done, and why
Memoising ParamHash on CadNode (~186 ns/part/frame). It needs a
cache field on the node plus invalidation on every mutation path, which is
the same ownership problem Phase 5.1 solves properly with a generation
counter. Doing it now would add a second thing to keep in sync by hand —
exactly the pattern this codebase already suffers from.
A BVH or spatial index for picking. Current cost is linear in part count with a cheap early-out. At the 10–100 parts these models actually contain, a BVH would cost more to build than it saves. Revisit if scenes grow an order of magnitude.
The other 81 cx.redraw_all() calls. Most are on discrete user
actions (keypress, button, tool change) where a full redraw is once-per-
gesture and harmless. Only the per-motion ones were worth removing, and
orbit_3d_by_screen_delta / pan_3d were already correct.
bench_command_execute_overhead — resolved
Recorded here because the number moved 300× and the reason is worth keeping.
CadCommandCtx::new eagerly built a scene snapshot:
let built_scene = scene_cache.scene_for(parts); // O(nodes)
CommandContext::scene() needs one, but no production command reads
it — MoveNode, ResizeNode, RotateNode, YawNode, ModifyNode,
CreateNode and DeleteNode all address a node by id. Every command
mutation bumps the PartsStore generation, so the next context
construction was a guaranteed cache miss that cloned every node and
re-registered every material. move_selected issues one command per
selected part per frame, so a drag was O(commands × nodes).
The snapshot is now built on the first scene() call and memoised, and
dropped by every mutating method (paired with the existing
mark_dirty() calls, so a new mutator cannot silently miss it).
The eager build was also a latent correctness bug: it captured the
scene before the command ran, so a command that mutated and then read
scene() saw its own edit missing. Pinned by
a_lazily_built_scene_reflects_edits_made_earlier_in_the_same_context.
Two measurement notes:
- The old 0.4 µs figure predated
PartsStore; the generation counter turned a previously accidental cache hit into a guaranteed miss. The regression was real, not a measurement artefact. - The benchmark itself called
ctx.scene()each iteration to read the start position, which no production path does —move_selectedandfinish_part_dragboth readp.pos()off the parts list. Fixing the lazy build alone moved undo 228 → 0.5 µs but left execute at 231 µs, because the benchmark was timing its own scene read. It now mirrors the production callers.
Engine baselines (doc + spreadsheet)
Added in PLAN_PERF_OPTIMIZATION.md Phase 0. Release-only, #[ignore]-d
benches that print, never assert wall-clock numbers (CI-load flakes).
Host: x86-64, macOS, --release, target-cpu not set. Reproduce:
cargo test --release -p spreadsheet-engine --lib -- \
--ignored --nocapture --test-threads=1 bench_
cargo test --release -p doc-engine --lib -- \
--ignored --nocapture --test-threads=1 bench_
| Benchmark | Baseline (this host) | What it measures |
|---|---|---|
bench_grid_set_get_10k |
1.65 µs per get_raw+display | HashMap-backed grid read hot loop |
bench_formula_chain_recalc_1k |
22.6 ms per full recalc | 1000-cell dependency-chain rebuild+eval |
bench_column_sum_10k |
91.8 ms per recalc | SUM(A1:A10000) — aggregate hot loop |
bench_workbook_roundtrip_10k |
97.7 ms per roundtrip | serialize+deserialize 10k cells |
bench_crdt_materialize_2k |
29.9 ms per materialize | whole-op-log → projection, 2000 chars |
bench_crdt_save_wire_2k |
9.9 ms per to_json+from_json | full-doc save wire, 2000 chars |
bench_insert_char_1k |
3.65 ms per keystroke | insert atom + re-materialize per key |
The doc rows confirm the Phase 6 targets are real: typing is ~3.65 ms/key (materialize dominates), and each save rewrites the whole op log.