# CAD benchmark baseline Reproduce with either: ```bash # needs the Makepad desktop stack (wayland/X11/GL/alsa/polkit) cargo test --locked --release -p nigig-build --lib -- \ --ignored --nocapture --test-threads=1 profile_benchmarks # host-only: isolated toolchain, no desktop, self-cleaning CAD_BENCH=1 ./tools/test-cad-coverage.sh ``` The second form exists because these benchmarks were unrunnable for ten review tranches — `REVIEWS/PAY_CAD_IMPLEMENTATION_STATUS.md` recorded the blocker as "no Cargo toolchain is available in this execution environment". `profile_benchmarks.rs` turns out to import nothing from Makepad but the math types, so it builds in the same host-only harness the coverage run uses. **Release mode matters.** Debug builds are 3–5× slower here and the ratios between arms shift, so a debug number is not comparable to anything below. These are measurements, not assertions — the benchmarks are `#[ignore]`-d and deliberately contain no timing `assert!`s (a wall-clock assertion in the normal test run flakes under CI load). Compare against this file by hand when changing the hot paths. The frame-submission rows below were recorded on the Phase 5 tree, x86-64, `--release`. The wall-clock rows above are retained machine baselines and are not comparable across hosts; the submission rows are structural counts and are intentionally stable. | Benchmark | Result | |---|---| | `bench_scene_cache_hit_vs_rebuild` | cold 25.9 µs / warm 53 ns (100 parts) — **488×** | | `bench_scene_cache_scaling` | 10/50/100/500 parts: cold scales linearly, warm stays ~150 ns | | `bench_mesh_cache_hit_cost` | 11.1 µs/frame, **110 ns per part** (100 parts, all hits) | | `bench_size_parametric_vs_mesh_derived` | parametric 2 ns vs mesh-derived 886 ns — **443×** | | `bench_pick_broadphase_mesh_bounds_vs_size` | 63.3 µs vs 131.0 µs per frame — **2.1× saved** | | `bench_pick_broadphase_world_aabb_recompute_vs_cache` | **230 µs → 39 µs per pick at 500 parts — 5.9×.** Transforming 8 corners per part per pick, versus a `PlacedHash` lookup. Hover picking runs on mouse-move, so this is per-frame. Cached in `SceneCache::world_aabb_for` (Phase 3.2). | | `bench_parts_script_regeneration_per_drag_frame` | **42 µs (50 parts) / 170 µs (200) / 425 µs (500)** per drag frame, before the editor's own `set_text`. `script_dirty` was set on every motion event; it is now set once on commit (Phase 3.6). | | `bench_param_hash_cost_per_frame` | Box **50 ns/node** (25 µs/frame at 500) — not worth a cache. Extruded 64-gon **987 ns/node**; that is the vertex data, and bulk-hashing the slice measured no faster. Phase 3.3 left undone deliberately. | | `bench_parallel_threshold_warm_vs_cold_cache` | 200 nodes: cold **8.80 ms seq / 6.21 ms par (1.42×)**, warm **5.81 / 5.73 (1.01×)**. The 32-node proxy is imprecise but harmless; Phase 3.8 rewrite not worth it. | | `bench_glb_export_with_and_without_cache` | cold 1.42 ms / warm 1.35 ms (50 parts) | | `bench_parallel_vs_sequential_export` | sequential 4.83 ms / parallel 3.58 ms (100 parts) — 1.3× | | `bench_command_execute_overhead` | **0.8 µs per execute, 0.5 µs per undo** (1000 commands over a 1000-node scene). Was 254 µs / 246 µs: `CadCommandCtx::new` eagerly built a scene snapshot no production command reads. Now built on first `scene()` call — roughly **300×**. | | `bench_gpu_upload_mesh_source` | boxes: rebuild 0.27 µs/part vs warm cache 0.178 µs — **2×**; extruded 24-gons: 1.11 µs vs 0.516 µs — **2×**. Cold cache 0.79 µs/part is **slower** than rebuilding (hash + insert) | | `bench_delete_invalidation_clear_vs_evict` | deleting 1 of 100 parts: clear+rebuild 52.9 µs vs evict+reuse 14.5 µs — **4×** | | `bench_viewport_snapshot_sync_per_frame` | deep copies 21.2 µs/frame + forced rebuilds 77.2 µs/frame = **98.5 µs/frame**; generation-checked 0.051 µs — **1930×** (100 parts) | ### Frame submission budget `bench_frame_submission_budget`, added by Phase 0 of `REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md` and extended through Phase 5. Every other benchmark in this file measures CPU work *behind* the renderer; none measures a live GPU frame, so the submission report stays as an explicit structural count rather than pretending to be frame time. Counts, not timings — counts do not flake under CI load. 1920×1080, parts on a lattice across a 200 m site (deliberately larger than the viewport at either zoom: a scene that fits on screen is the case where culling can do nothing, and reporting only that would flatter it). **2D, after Phase 1 culling, Phase 4 stroke batching, and Phase 5 sub-pixel LOD.** "Before" is the Phase 0 baseline: every part in the document submitted, one `stroke()` per part and one per grid line. | Zoom | Parts | Visible | Full outlines | LOD points | Part calls | Tessellations | Before | |---|---|---:|---:|---:|---:|---:|---:| | 5 m (170 grid lines) | 100 | 1 | 1 | 0 | 1 stroke | **3** (2 grid + 1 parts) | 270 | | | 500 | 4 | 4 | 0 | 2 strokes | **4** | 670 | | | 2,000 | 12 | 12 | 0 | 2 strokes | **4** | 2,170 | | 200 m (270 grid lines) | 100 | 100 | 1 | 99 | 1 stroke + 1 fill | **4** | 370 | | | 500 | 500 | 1 | 499 | 1 stroke + 1 fill | **4** | 770 | | | 2,000 | 2,000 | 1 | 1,999 | 1 stroke + 1 fill | **4** | 2,270 | | 500 m (170 grid lines) | 100 | 100 | 1 | 99 | 1 stroke + 1 fill | **4** | 270 | | | 500 | 500 | 1 | 499 | 1 stroke + 1 fill | **4** | 670 | | | 2,000 | 2,000 | 1 | 1,999 | 1 stroke + 1 fill | **4** | 2,170 | The benchmark models one material plus one selected part. Selection is intentionally kept at full detail, so the first visible part remains an outline while the ordinary parts cross the LOD threshold. At 5 m, a 1 m part is comfortably resolvable and every survivor stays full detail. At 200 m, the same part is only about 2.7 logical pixels across; at 500 m it is about 1.1 pixels, so the ordinary survivors become 2×2 logical-pixel filled markers. Phase 5 does **not** claim another draw-call win. The 2D scene remains one `DrawVector::end()` submission, and the one material still makes one outline stroke plus one marker fill. The win is geometry volume: a filled marker is bounded and does not carry a full rectangle/circle outline for a part whose shape cannot be resolved. `FrameBudget` reports full outlines and LOD markers separately so this change cannot be mistaken for another tessellation reduction. **3D, frustum culling (Phase 1), instanced batching (Phase 3), and bounding-box LOD (Phase 5)**, 60° fov, camera orbiting the site centre. "Shapes" is how many distinct full-detail solids the visible parts use — 6 for a catalogue-built model, all distinct for the worst case. LOD boxes are all instances of one shared unit-cube proxy batch: | Camera | Parts | Visible | Full meshes | LOD boxes | Draw calls, 6 shapes | Draw calls, all distinct | Before | |---|---:|---:|---:|---:|---:|---:|---:| | 20 m | 100 | 39 | 39 | 0 | **6** | 39 | 100 | | | 500 | 198 | 198 | 0 | **6** | 198 | 500 | | | 2,000 | 733 | 733 | 0 | **6** | 733 | 2,000 | | 80 m | 100 | 76 | 76 | 0 | **6** | 76 | 100 | | | 500 | 386 | 386 | 0 | **6** | 386 | 500 | | | 2,000 | 1,459 | 1,459 | 0 | **6** | 1,459 | 2,000 | | 400 m | 100 | 100 | 20 | 80 | **7** | 21 | 100 | | | 500 | 500 | 63 | 437 | **7** | 64 | 500 | | | 2,000 | 2,000 | 380 | 1,620 | **7** | 381 | 2,000 | The bottom row is the one to look at. At 400 m the whole site is on screen, so culling removes nothing. LOD replaces 1,620 unresolvable meshes with instances of one proxy, and the six full-shape groups plus that proxy draw in **7 calls instead of 2,000**. The "all distinct" column is the same scene with every full-detail part a different size; 380 remain full and the 1,620 tiny parts still share the proxy, so the count is 381. The selected/hovered exemption is not present in this synthetic 3-D benchmark; the policy tests pin that decorated parts remain full detail. Where culling does work it is decisive on its own: drafting at a 5 m zoom over a 200 m site, the 2,000-part scene queues 12 outlines instead of 2,000. Where it cannot help — everything on screen — batching still holds the call count at 4. Two different costs, deliberately not summed: - **Tessellations** are CPU calls into `tessellate_path_stroke`. The whole 2D scene is **one** draw call from `DrawVector::end()` no matter how many parts, so this — not draw calls — was the number that scaled there. Since Phase 4 it scales with colour groups instead. - **Draw calls** are GPU submissions. 3D used to be one per uploaded, visible part; since Phase 3 full-detail parts use one call per distinct shape, with the parts riding inside as instance rows. Since Phase 5, all tiny parts use one additional shared proxy batch. The grid line count is unchanged by culling, as it should be: it was already virtualised, and stays at 170–270 lines across the measured zoom range because the step adapts. Phase 4 changed what those lines cost — two `stroke()` calls rather than one each — and Phase 5 changes only the representation of tiny parts. The tests `culling_shows_up_as_fewer_part_tessellations`, `the_grid_costs_two_strokes_at_any_zoom`, and the LOD policy tests pin those separate costs. Phase 2 (sharing geometry between parts of the same shape, keyed by a new `ShapeHash` — **not** `ParamHash`, which includes the node id) and Phase 3 (instancing) brought the draw-call column down in the zoomed-out case where culling cannot; Phase 4 did the same for the 2-D tessellation column. Phase 5 addresses the remaining geometry volume in both paths: the 2,000-part rows at 200 m and 500 m keep one selected outline but replace 1,999 unresolvable outlines with bounded point markers, while the 400 m 3-D row replaces 1,620 unresolvable meshes with one shared proxy batch. A real GPU/window visual check is still required before treating the structural counts as frame-time results. ### GPU geometry buffers, keyed by shape `bench_geometry_buffers_shared_by_shape`, added by Phase 2. Counts, not timings. The number it replaces was never measured because it was never in doubt: **one buffer per part**, however many of those parts were the same column. | Model | Parts | Buffers | Ratio | |---|---|---|---| | 200 identical walls (`build_house_scene(200)`) | 200 | **1** | 200× | | 420-part repetitive model (columns, walls in three lengths, two opening types) | 420 | **6** | 70× | | 420 parts, every one a different size | 420 | **420** | 1× | Read the last row too. Sharing is a property of the model, not of the code: a scene where every part differs gets nothing from this phase. That is the honest ceiling, and it is why the draw-call column is untouched here — Phase 3 (instancing) is what collapses draw calls, and it needs the shared buffers this phase produces. The middle row is the one to plan against. Real architectural models are catalogue-built: a few column types, a few wall lengths, a door and a window repeated everywhere. 70× fewer GPU buffers, and 70× fewer mesh uploads on the first frame after a load. ### Note: two benchmarks had drifted off the caching path If you ran these between the Phase 5.1 generation-tracked store and the commit that added this note, `bench_scene_cache_hit_vs_rebuild` reported **1.2×** against the **488×** in the table above, and `bench_scene_cache_scaling` reported **1× at every part count**. That looked exactly like a catastrophic cache regression. It was not. Both benchmarks called `SceneCache::scene(&[CadNode])`, which is documented as *always rebuilding* — it takes a bare slice, so it has no generation to compare and cannot cache. The caching entry point the editor uses is `SceneCache::scene_for(&PartsStore)`. When the store landed, these two benchmarks were not moved with it, so their "warm" sample was a second full rebuild and the printed speedup was allocator noise. Repointed at `scene_for`, they reproduce the table: cold 24.6 µs / warm **48 ns**, **512×**, and the warm read stays flat at ~55 ns from 10 parts to 500. `bench_scene_cache_hit_vs_rebuild` now asserts `Arc::ptr_eq` on the two samples, so it fails loudly rather than quietly measuring two rebuilds if it is ever pointed at a non-caching path again. `SceneCache::scene()` is not the bug and was left alone: its docstring says what it does, and nine tests legitimately use it for exactly the case it describes. ### Note on `bench_gpu_upload_mesh_source` Phase 5.3 routed the GPU upload path through `MeshCache` instead of calling `build_solid()` directly. The measured speed win is **2× on a warm cache and negative on a cold one** — a first-ever build now pays a `ParamHash` and a map insert on top of the meshing. That trade is worth taking, but not primarily for speed: - `ensure_part_geometry` only builds for parts with no uploaded buffer, so the cold case is a scene load or an undo/redo, not a frame in a loop. - The warm case is every part that was already exported, hovered or previously uploaded — which, during editing, is most of them. - The real gain is that there is now **one** node-to-triangles path. Two independent `match` arms over the same enum were free to drift, and nothing would have caught it: the preview would simply have disagreed with the export. Do not "optimise" this back by special-casing cheap primitives. The 0.5 µs is not worth reintroducing the second pipeline. ### Note on `bench_viewport_snapshot_sync_per_frame` This benchmark measures the *primitives* the sync path is built from -- `Vec` clones and `SceneCache::scene()` rebuilds -- not the widget method itself, which needs a live `Cx` and three real viewports. Its printed numbers therefore do **not** move when the sync path is fixed; they quantify what one frame of syncing costs and what the same frame costs when a generation check is allowed to short-circuit it. Read it as the size of the prize, not as a regression gate. The Phase 5.2 change is verified by behaviour tests (`replace_all_bumps_the_generation_so_destinations_can_compare`, `mesh_cache_self_invalidates_on_parameter_and_transform_edits`, `geometry_is_retained_for_surviving_ids_only`), each negative-tested. --- ## Reading these **The scene cache is doing its job.** 488× on a warm hit, and the scaling run shows the warm path is O(1) while the cold path is linear. No work needed here. **`MeshCache` hits are not free.** 186 ns per part is `ParamHash::from_node` being recomputed on every lookup — it hashes the solid's every field, the transform and the material, even when the entry is present and valid. At 100 parts that is ~19 µs per hover frame. Not the dominant cost today, but it is the next thing to fix if picking needs to get faster: memoise the hash on `CadNode` and invalidate it on mutation. **`CadNode::size()` is a trap.** 2 ns for a parametric solid, 886 ns for a CSG or extruded one, because those have no closed-form extent and the function meshes the solid to derive a bounding box (added in Phase 1.7 to fix unpickable parts). It reads like a cheap field access at the call site. Anything calling it in a loop should take bounds from a cached mesh instead — which is exactly what `pick_part` now does. **Export is already fast enough.** 1.4 ms for 50 parts, and the mesh cache barely moves it because JSON/binary serialisation dominates, not meshing. The parallel path wins only 1.3× at 100 parts. Neither is worth tuning before the UI hot paths. --- ## What Phase 3 changed | Change | Effect | |---|---| | `pick_part` takes AABB bounds from the mesh it already holds instead of calling `size()` | **2.1×** less broad-phase work per hover frame; also *more correct* — see below | | Hover picking throttled by `HOVER_PICK_MIN_MOVE_PX` (3 px) | Sub-pixel jitter no longer triggers a full ray cast against every triangle | | Removed `cx.redraw_all()` from the part-drag `MouseMove` handler | No full-widget-tree relayout per motion event during a drag | The AABB change also fixed a latent correctness bug. `size()` returns a *symmetric* extent about the origin, but an extruded polygon grows along +Y from its base plane — so the old broad-phase box was in the wrong place and could reject a ray that actually hits. `pick_bounds_tests:: mesh_bounds_are_asymmetric_where_size_is_not` pins this. --- ## Not done, and why **Memoising `ParamHash` on `CadNode`** (~186 ns/part/frame). It needs a cache field on the node plus invalidation on every mutation path, which is the same ownership problem Phase 5.1 solves properly with a generation counter. Doing it now would add a second thing to keep in sync by hand — exactly the pattern this codebase already suffers from. **A BVH or spatial index for picking.** Current cost is linear in part count with a cheap early-out. At the 10–100 parts these models actually contain, a BVH would cost more to build than it saves. Revisit if scenes grow an order of magnitude. **The other 81 `cx.redraw_all()` calls.** Most are on discrete user actions (keypress, button, tool change) where a full redraw is once-per- gesture and harmless. Only the per-motion ones were worth removing, and `orbit_3d_by_screen_delta` / `pan_3d` were already correct. ## `bench_command_execute_overhead` — resolved Recorded here because the number moved 300× and the reason is worth keeping. `CadCommandCtx::new` eagerly built a scene snapshot: ```rust let built_scene = scene_cache.scene_for(parts); // O(nodes) ``` `CommandContext::scene()` needs one, but **no production command reads it** — `MoveNode`, `ResizeNode`, `RotateNode`, `YawNode`, `ModifyNode`, `CreateNode` and `DeleteNode` all address a node by id. Every command mutation bumps the `PartsStore` generation, so the next context construction was a guaranteed cache miss that cloned every node and re-registered every material. `move_selected` issues one command per selected part per frame, so a drag was O(commands × nodes). The snapshot is now built on the first `scene()` call and memoised, and dropped by every mutating method (paired with the existing `mark_dirty()` calls, so a new mutator cannot silently miss it). The eager build was also a latent **correctness** bug: it captured the scene *before* the command ran, so a command that mutated and then read `scene()` saw its own edit missing. Pinned by `a_lazily_built_scene_reflects_edits_made_earlier_in_the_same_context`. Two measurement notes: - The old 0.4 µs figure predated `PartsStore`; the generation counter turned a previously accidental cache hit into a guaranteed miss. The regression was real, not a measurement artefact. - The benchmark itself called `ctx.scene()` each iteration to read the start position, which no production path does — `move_selected` and `finish_part_drag` both read `p.pos()` off the parts list. Fixing the lazy build alone moved undo 228 → 0.5 µs but left execute at 231 µs, because the benchmark was timing its own scene read. It now mirrors the production callers. --- ## Engine baselines (doc + spreadsheet) Added in PLAN_PERF_OPTIMIZATION.md Phase 0. Release-only, `#[ignore]`-d benches that **print, never assert** wall-clock numbers (CI-load flakes). Host: x86-64, macOS, `--release`, `target-cpu` not set. Reproduce: ```bash cargo test --release -p spreadsheet-engine --lib -- \ --ignored --nocapture --test-threads=1 bench_ cargo test --release -p doc-engine --lib -- \ --ignored --nocapture --test-threads=1 bench_ ``` | Benchmark | Baseline (this host) | What it measures | |---|---|---| | `bench_grid_set_get_10k` | 1.65 µs per get_raw+display | HashMap-backed grid read hot loop | | `bench_formula_chain_recalc_1k` | 22.6 ms per full recalc | 1000-cell dependency-chain rebuild+eval | | `bench_column_sum_10k` | 91.8 ms per recalc | SUM(A1:A10000) — aggregate hot loop | | `bench_workbook_roundtrip_10k` | 97.7 ms per roundtrip | serialize+deserialize 10k cells | | `bench_crdt_materialize_2k` | **29.9 ms per materialize** | whole-op-log → projection, 2000 chars | | `bench_crdt_save_wire_2k` | **9.9 ms per to_json+from_json** | full-doc save wire, 2000 chars | | `bench_insert_char_1k` | **3.65 ms per keystroke** | insert atom + re-materialize per key | The doc rows confirm the Phase 6 targets are real: typing is ~3.65 ms/key (materialize dominates), and each save rewrites the whole op log.