nigig-org/BENCH_BASELINE.md
andodeki ac8f8aa002
Some checks failed
repo hygiene / hygiene (push) Has been cancelled
PDF engine / engine (push) Has been cancelled
PDF engine / makepad-integration (push) Has been cancelled
PDF engine / fuzz (push) Has been cancelled
nigig-build (CAD) / supply-chain (push) Has been cancelled
nigig-build (CAD) / cad-module (push) Has been cancelled
nigig-build (CAD) / full-crate-check (push) Has been cancelled
nigig-build (CAD) / cad-engine-coverage (push) Has been cancelled
nigig-build (CAD) / doc-workspace-coverage (push) Has been cancelled
nigig-build (CAD) / cad-widget-coverage (push) Has been cancelled
traffic / gates (push) Has been cancelled
traffic / nigig-traffic (push) Has been cancelled
traffic / supply-chain (push) Has been cancelled
doc-engine / engine (push) Has been cancelled
doc-engine / coverage (push) Has been cancelled
doc-engine / consumer (push) Has been cancelled
email / gates (push) Has been cancelled
email / email-domain (push) Has been cancelled
email / nigig-email (push) Has been cancelled
email / supply-chain (push) Has been cancelled
nigig-map / test (push) Has been cancelled
sms / gates (push) Has been cancelled
sms / robius-sms (push) Has been cancelled
sms / android (push) Has been cancelled
sms / nigig-sms (push) Has been cancelled
sms / supply-chain (push) Has been cancelled
spreadsheet / engine-coverage (push) Has been cancelled
spreadsheet / ui-controller-coverage (push) Has been cancelled
p2p-intel / engine (push) Has been cancelled
p2p-intel / notifications (push) Has been cancelled
p2p-intel / coverage (push) Has been cancelled
p2p-intel / makepad-app (push) Has been cancelled
p2p-intel / exchange-tab (push) Has been cancelled
Payment domain, storage, platform and UI / isolated-payment-tests (push) Has been cancelled
Payment domain, storage, platform and UI / payment-ui-tests (push) Has been cancelled
chore: sync full working tree to gitdab
Whole-tree sync: cad-core/cad-ui split sources, nigig-build
construction_frame migration, pdf port progress, mpesa/pay/uikit/doc
updates, workspace members/profiles/lock, CI workflows and reviews.
See individual file history for details.
2026-09-12 07:15:24 +03:00

20 KiB
Raw Permalink Blame History

CAD benchmark baseline

Reproduce with either:

# needs the Makepad desktop stack (wayland/X11/GL/alsa/polkit)
cargo test --locked --release -p nigig-build --lib -- \
  --ignored --nocapture --test-threads=1 profile_benchmarks

# host-only: isolated toolchain, no desktop, self-cleaning
CAD_BENCH=1 ./tools/test-cad-coverage.sh

The second form exists because these benchmarks were unrunnable for ten review tranches — REVIEWS/PAY_CAD_IMPLEMENTATION_STATUS.md recorded the blocker as "no Cargo toolchain is available in this execution environment". profile_benchmarks.rs turns out to import nothing from Makepad but the math types, so it builds in the same host-only harness the coverage run uses.

Release mode matters. Debug builds are 3–5× slower here and the ratios between arms shift, so a debug number is not comparable to anything below.

These are measurements, not assertions — the benchmarks are #[ignore]-d and deliberately contain no timing assert!s (a wall-clock assertion in the normal test run flakes under CI load). Compare against this file by hand when changing the hot paths.

The frame-submission rows below were recorded on the Phase 5 tree, x86-64, --release. The wall-clock rows above are retained machine baselines and are not comparable across hosts; the submission rows are structural counts and are intentionally stable.

Benchmark Result
bench_scene_cache_hit_vs_rebuild cold 25.9 µs / warm 53 ns (100 parts) — 488×
bench_scene_cache_scaling 10/50/100/500 parts: cold scales linearly, warm stays ~150 ns
bench_mesh_cache_hit_cost 11.1 µs/frame, 110 ns per part (100 parts, all hits)
bench_size_parametric_vs_mesh_derived parametric 2 ns vs mesh-derived 886 ns — 443×
bench_pick_broadphase_mesh_bounds_vs_size 63.3 µs vs 131.0 µs per frame — 2.1× saved
bench_pick_broadphase_world_aabb_recompute_vs_cache 230 µs → 39 µs per pick at 500 parts — 5.9×. Transforming 8 corners per part per pick, versus a PlacedHash lookup. Hover picking runs on mouse-move, so this is per-frame. Cached in SceneCache::world_aabb_for (Phase 3.2).
bench_parts_script_regeneration_per_drag_frame 42 µs (50 parts) / 170 µs (200) / 425 µs (500) per drag frame, before the editor's own set_text. script_dirty was set on every motion event; it is now set once on commit (Phase 3.6).
bench_param_hash_cost_per_frame Box 50 ns/node (25 µs/frame at 500) — not worth a cache. Extruded 64-gon 987 ns/node; that is the vertex data, and bulk-hashing the slice measured no faster. Phase 3.3 left undone deliberately.
bench_parallel_threshold_warm_vs_cold_cache 200 nodes: cold 8.80 ms seq / 6.21 ms par (1.42×), warm 5.81 / 5.73 (1.01×). The 32-node proxy is imprecise but harmless; Phase 3.8 rewrite not worth it.
bench_glb_export_with_and_without_cache cold 1.42 ms / warm 1.35 ms (50 parts)
bench_parallel_vs_sequential_export sequential 4.83 ms / parallel 3.58 ms (100 parts) — 1.3×
bench_command_execute_overhead 0.8 µs per execute, 0.5 µs per undo (1000 commands over a 1000-node scene). Was 254 µs / 246 µs: CadCommandCtx::new eagerly built a scene snapshot no production command reads. Now built on first scene() call — roughly 300×.
bench_gpu_upload_mesh_source boxes: rebuild 0.27 µs/part vs warm cache 0.178 µs — 2×; extruded 24-gons: 1.11 µs vs 0.516 µs — 2×. Cold cache 0.79 µs/part is slower than rebuilding (hash + insert)
bench_delete_invalidation_clear_vs_evict deleting 1 of 100 parts: clear+rebuild 52.9 µs vs evict+reuse 14.5 µs — 4×
bench_viewport_snapshot_sync_per_frame deep copies 21.2 µs/frame + forced rebuilds 77.2 µs/frame = 98.5 µs/frame; generation-checked 0.051 µs — 1930× (100 parts)

Frame submission budget

bench_frame_submission_budget, added by Phase 0 of REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md and extended through Phase 5. Every other benchmark in this file measures CPU work behind the renderer; none measures a live GPU frame, so the submission report stays as an explicit structural count rather than pretending to be frame time.

Counts, not timings — counts do not flake under CI load. 1920×1080, parts on a lattice across a 200 m site (deliberately larger than the viewport at either zoom: a scene that fits on screen is the case where culling can do nothing, and reporting only that would flatter it).

2D, after Phase 1 culling, Phase 4 stroke batching, and Phase 5 sub-pixel LOD. "Before" is the Phase 0 baseline: every part in the document submitted, one stroke() per part and one per grid line.

Zoom Parts Visible Full outlines LOD points Part calls Tessellations Before
5 m (170 grid lines) 100 1 1 0 1 stroke 3 (2 grid + 1 parts) 270
500 4 4 0 2 strokes 4 670
2,000 12 12 0 2 strokes 4 2,170
200 m (270 grid lines) 100 100 1 99 1 stroke + 1 fill 4 370
500 500 1 499 1 stroke + 1 fill 4 770
2,000 2,000 1 1,999 1 stroke + 1 fill 4 2,270
500 m (170 grid lines) 100 100 1 99 1 stroke + 1 fill 4 270
500 500 1 499 1 stroke + 1 fill 4 670
2,000 2,000 1 1,999 1 stroke + 1 fill 4 2,170

The benchmark models one material plus one selected part. Selection is intentionally kept at full detail, so the first visible part remains an outline while the ordinary parts cross the LOD threshold. At 5 m, a 1 m part is comfortably resolvable and every survivor stays full detail. At 200 m, the same part is only about 2.7 logical pixels across; at 500 m it is about 1.1 pixels, so the ordinary survivors become 2×2 logical-pixel filled markers.

Phase 5 does not claim another draw-call win. The 2D scene remains one DrawVector::end() submission, and the one material still makes one outline stroke plus one marker fill. The win is geometry volume: a filled marker is bounded and does not carry a full rectangle/circle outline for a part whose shape cannot be resolved. FrameBudget reports full outlines and LOD markers separately so this change cannot be mistaken for another tessellation reduction.

3D, frustum culling (Phase 1), instanced batching (Phase 3), and bounding-box LOD (Phase 5), 60° fov, camera orbiting the site centre. "Shapes" is how many distinct full-detail solids the visible parts use — 6 for a catalogue-built model, all distinct for the worst case. LOD boxes are all instances of one shared unit-cube proxy batch:

Camera Parts Visible Full meshes LOD boxes Draw calls, 6 shapes Draw calls, all distinct Before
20 m 100 39 39 0 6 39 100
500 198 198 0 6 198 500
2,000 733 733 0 6 733 2,000
80 m 100 76 76 0 6 76 100
500 386 386 0 6 386 500
2,000 1,459 1,459 0 6 1,459 2,000
400 m 100 100 20 80 7 21 100
500 500 63 437 7 64 500
2,000 2,000 380 1,620 7 381 2,000

The bottom row is the one to look at. At 400 m the whole site is on screen, so culling removes nothing. LOD replaces 1,620 unresolvable meshes with instances of one proxy, and the six full-shape groups plus that proxy draw in 7 calls instead of 2,000. The "all distinct" column is the same scene with every full-detail part a different size; 380 remain full and the 1,620 tiny parts still share the proxy, so the count is 381. The selected/hovered exemption is not present in this synthetic 3-D benchmark; the policy tests pin that decorated parts remain full detail.

Where culling does work it is decisive on its own: drafting at a 5 m zoom over a 200 m site, the 2,000-part scene queues 12 outlines instead of 2,000. Where it cannot help — everything on screen — batching still holds the call count at 4.

Two different costs, deliberately not summed:

  • Tessellations are CPU calls into tessellate_path_stroke. The whole 2D scene is one draw call from DrawVector::end() no matter how many parts, so this — not draw calls — was the number that scaled there. Since Phase 4 it scales with colour groups instead.
  • Draw calls are GPU submissions. 3D used to be one per uploaded, visible part; since Phase 3 full-detail parts use one call per distinct shape, with the parts riding inside as instance rows. Since Phase 5, all tiny parts use one additional shared proxy batch.

The grid line count is unchanged by culling, as it should be: it was already virtualised, and stays at 170–270 lines across the measured zoom range because the step adapts. Phase 4 changed what those lines cost — two stroke() calls rather than one each — and Phase 5 changes only the representation of tiny parts. The tests culling_shows_up_as_fewer_part_tessellations, the_grid_costs_two_strokes_at_any_zoom, and the LOD policy tests pin those separate costs.

Phase 2 (sharing geometry between parts of the same shape, keyed by a new ShapeHash — not ParamHash, which includes the node id) and Phase 3 (instancing) brought the draw-call column down in the zoomed-out case where culling cannot; Phase 4 did the same for the 2-D tessellation column. Phase 5 addresses the remaining geometry volume in both paths: the 2,000-part rows at 200 m and 500 m keep one selected outline but replace 1,999 unresolvable outlines with bounded point markers, while the 400 m 3-D row replaces 1,620 unresolvable meshes with one shared proxy batch. A real GPU/window visual check is still required before treating the structural counts as frame-time results.

GPU geometry buffers, keyed by shape

bench_geometry_buffers_shared_by_shape, added by Phase 2. Counts, not timings. The number it replaces was never measured because it was never in doubt: one buffer per part, however many of those parts were the same column.

Model Parts Buffers Ratio
200 identical walls (build_house_scene(200)) 200 1 200×
420-part repetitive model (columns, walls in three lengths, two opening types) 420 6 70×
420 parts, every one a different size 420 420 1×

Read the last row too. Sharing is a property of the model, not of the code: a scene where every part differs gets nothing from this phase. That is the honest ceiling, and it is why the draw-call column is untouched here — Phase 3 (instancing) is what collapses draw calls, and it needs the shared buffers this phase produces.

The middle row is the one to plan against. Real architectural models are catalogue-built: a few column types, a few wall lengths, a door and a window repeated everywhere. 70× fewer GPU buffers, and 70× fewer mesh uploads on the first frame after a load.

Note: two benchmarks had drifted off the caching path

If you ran these between the Phase 5.1 generation-tracked store and the commit that added this note, bench_scene_cache_hit_vs_rebuild reported 1.2× against the 488× in the table above, and bench_scene_cache_scaling reported 1× at every part count. That looked exactly like a catastrophic cache regression. It was not.

Both benchmarks called SceneCache::scene(&[CadNode]), which is documented as always rebuilding — it takes a bare slice, so it has no generation to compare and cannot cache. The caching entry point the editor uses is SceneCache::scene_for(&PartsStore). When the store landed, these two benchmarks were not moved with it, so their "warm" sample was a second full rebuild and the printed speedup was allocator noise.

Repointed at scene_for, they reproduce the table: cold 24.6 µs / warm 48 ns, 512×, and the warm read stays flat at ~55 ns from 10 parts to 500. bench_scene_cache_hit_vs_rebuild now asserts Arc::ptr_eq on the two samples, so it fails loudly rather than quietly measuring two rebuilds if it is ever pointed at a non-caching path again.

SceneCache::scene() is not the bug and was left alone: its docstring says what it does, and nine tests legitimately use it for exactly the case it describes.

Note on bench_gpu_upload_mesh_source

Phase 5.3 routed the GPU upload path through MeshCache instead of calling build_solid() directly. The measured speed win is 2× on a warm cache and negative on a cold one — a first-ever build now pays a ParamHash and a map insert on top of the meshing.

That trade is worth taking, but not primarily for speed:

  • ensure_part_geometry only builds for parts with no uploaded buffer, so the cold case is a scene load or an undo/redo, not a frame in a loop.
  • The warm case is every part that was already exported, hovered or previously uploaded — which, during editing, is most of them.
  • The real gain is that there is now one node-to-triangles path. Two independent match arms over the same enum were free to drift, and nothing would have caught it: the preview would simply have disagreed with the export.

Do not "optimise" this back by special-casing cheap primitives. The 0.5 µs is not worth reintroducing the second pipeline.

Note on bench_viewport_snapshot_sync_per_frame

This benchmark measures the primitives the sync path is built from -- Vec<CadNode> clones and SceneCache::scene() rebuilds -- not the widget method itself, which needs a live Cx and three real viewports. Its printed numbers therefore do not move when the sync path is fixed; they quantify what one frame of syncing costs and what the same frame costs when a generation check is allowed to short-circuit it.

Read it as the size of the prize, not as a regression gate. The Phase 5.2 change is verified by behaviour tests (replace_all_bumps_the_generation_so_destinations_can_compare, mesh_cache_self_invalidates_on_parameter_and_transform_edits, geometry_is_retained_for_surviving_ids_only), each negative-tested.


Reading these

The scene cache is doing its job. 488× on a warm hit, and the scaling run shows the warm path is O(1) while the cold path is linear. No work needed here.

MeshCache hits are not free. 186 ns per part is ParamHash::from_node being recomputed on every lookup — it hashes the solid's every field, the transform and the material, even when the entry is present and valid. At 100 parts that is ~19 µs per hover frame. Not the dominant cost today, but it is the next thing to fix if picking needs to get faster: memoise the hash on CadNode and invalidate it on mutation.

CadNode::size() is a trap. 2 ns for a parametric solid, 886 ns for a CSG or extruded one, because those have no closed-form extent and the function meshes the solid to derive a bounding box (added in Phase 1.7 to fix unpickable parts). It reads like a cheap field access at the call site. Anything calling it in a loop should take bounds from a cached mesh instead — which is exactly what pick_part now does.

Export is already fast enough. 1.4 ms for 50 parts, and the mesh cache barely moves it because JSON/binary serialisation dominates, not meshing. The parallel path wins only 1.3× at 100 parts. Neither is worth tuning before the UI hot paths.


What Phase 3 changed

Change Effect
pick_part takes AABB bounds from the mesh it already holds instead of calling size() 2.1× less broad-phase work per hover frame; also more correct — see below
Hover picking throttled by HOVER_PICK_MIN_MOVE_PX (3 px) Sub-pixel jitter no longer triggers a full ray cast against every triangle
Removed cx.redraw_all() from the part-drag MouseMove handler No full-widget-tree relayout per motion event during a drag

The AABB change also fixed a latent correctness bug. size() returns a symmetric extent about the origin, but an extruded polygon grows along +Y from its base plane — so the old broad-phase box was in the wrong place and could reject a ray that actually hits. pick_bounds_tests:: mesh_bounds_are_asymmetric_where_size_is_not pins this.


Not done, and why

Memoising ParamHash on CadNode (~186 ns/part/frame). It needs a cache field on the node plus invalidation on every mutation path, which is the same ownership problem Phase 5.1 solves properly with a generation counter. Doing it now would add a second thing to keep in sync by hand — exactly the pattern this codebase already suffers from.

A BVH or spatial index for picking. Current cost is linear in part count with a cheap early-out. At the 10–100 parts these models actually contain, a BVH would cost more to build than it saves. Revisit if scenes grow an order of magnitude.

The other 81 cx.redraw_all() calls. Most are on discrete user actions (keypress, button, tool change) where a full redraw is once-per- gesture and harmless. Only the per-motion ones were worth removing, and orbit_3d_by_screen_delta / pan_3d were already correct.

bench_command_execute_overhead — resolved

Recorded here because the number moved 300× and the reason is worth keeping.

CadCommandCtx::new eagerly built a scene snapshot:

let built_scene = scene_cache.scene_for(parts);   // O(nodes)

CommandContext::scene() needs one, but no production command reads it — MoveNode, ResizeNode, RotateNode, YawNode, ModifyNode, CreateNode and DeleteNode all address a node by id. Every command mutation bumps the PartsStore generation, so the next context construction was a guaranteed cache miss that cloned every node and re-registered every material. move_selected issues one command per selected part per frame, so a drag was O(commands × nodes).

The snapshot is now built on the first scene() call and memoised, and dropped by every mutating method (paired with the existing mark_dirty() calls, so a new mutator cannot silently miss it).

The eager build was also a latent correctness bug: it captured the scene before the command ran, so a command that mutated and then read scene() saw its own edit missing. Pinned by a_lazily_built_scene_reflects_edits_made_earlier_in_the_same_context.

Two measurement notes:

  • The old 0.4 µs figure predated PartsStore; the generation counter turned a previously accidental cache hit into a guaranteed miss. The regression was real, not a measurement artefact.
  • The benchmark itself called ctx.scene() each iteration to read the start position, which no production path does — move_selected and finish_part_drag both read p.pos() off the parts list. Fixing the lazy build alone moved undo 228 → 0.5 µs but left execute at 231 µs, because the benchmark was timing its own scene read. It now mirrors the production callers.

Engine baselines (doc + spreadsheet)

Added in PLAN_PERF_OPTIMIZATION.md Phase 0. Release-only, #[ignore]-d benches that print, never assert wall-clock numbers (CI-load flakes). Host: x86-64, macOS, --release, target-cpu not set. Reproduce:

cargo test --release -p spreadsheet-engine --lib -- \
  --ignored --nocapture --test-threads=1 bench_
cargo test --release -p doc-engine --lib -- \
  --ignored --nocapture --test-threads=1 bench_
Benchmark Baseline (this host) What it measures
bench_grid_set_get_10k 1.65 µs per get_raw+display HashMap-backed grid read hot loop
bench_formula_chain_recalc_1k 22.6 ms per full recalc 1000-cell dependency-chain rebuild+eval
bench_column_sum_10k 91.8 ms per recalc SUM(A1:A10000) — aggregate hot loop
bench_workbook_roundtrip_10k 97.7 ms per roundtrip serialize+deserialize 10k cells
bench_crdt_materialize_2k 29.9 ms per materialize whole-op-log → projection, 2000 chars
bench_crdt_save_wire_2k 9.9 ms per to_json+from_json full-doc save wire, 2000 chars
bench_insert_char_1k 3.65 ms per keystroke insert atom + re-materialize per key

The doc rows confirm the Phase 6 targets are real: typing is ~3.65 ms/key (materialize dominates), and each save rewrites the whole op log.