Some checks failed
repo hygiene / hygiene (push) Has been cancelled
PDF engine / engine (push) Has been cancelled
PDF engine / makepad-integration (push) Has been cancelled
PDF engine / fuzz (push) Has been cancelled
nigig-build (CAD) / supply-chain (push) Has been cancelled
nigig-build (CAD) / cad-module (push) Has been cancelled
nigig-build (CAD) / full-crate-check (push) Has been cancelled
nigig-build (CAD) / cad-engine-coverage (push) Has been cancelled
nigig-build (CAD) / doc-workspace-coverage (push) Has been cancelled
nigig-build (CAD) / cad-widget-coverage (push) Has been cancelled
traffic / gates (push) Has been cancelled
traffic / nigig-traffic (push) Has been cancelled
traffic / supply-chain (push) Has been cancelled
doc-engine / engine (push) Has been cancelled
doc-engine / coverage (push) Has been cancelled
doc-engine / consumer (push) Has been cancelled
email / gates (push) Has been cancelled
email / email-domain (push) Has been cancelled
email / nigig-email (push) Has been cancelled
email / supply-chain (push) Has been cancelled
nigig-map / test (push) Has been cancelled
sms / gates (push) Has been cancelled
sms / robius-sms (push) Has been cancelled
sms / android (push) Has been cancelled
sms / nigig-sms (push) Has been cancelled
sms / supply-chain (push) Has been cancelled
spreadsheet / engine-coverage (push) Has been cancelled
spreadsheet / ui-controller-coverage (push) Has been cancelled
p2p-intel / engine (push) Has been cancelled
p2p-intel / notifications (push) Has been cancelled
p2p-intel / coverage (push) Has been cancelled
p2p-intel / makepad-app (push) Has been cancelled
p2p-intel / exchange-tab (push) Has been cancelled
Payment domain, storage, platform and UI / isolated-payment-tests (push) Has been cancelled
Payment domain, storage, platform and UI / payment-ui-tests (push) Has been cancelled
Whole-tree sync: cad-core/cad-ui split sources, nigig-build construction_frame migration, pdf port progress, mpesa/pay/uikit/doc updates, workspace members/profiles/lock, CI workflows and reviews. See individual file history for details.
375 lines
20 KiB
Markdown
375 lines
20 KiB
Markdown
# CAD benchmark baseline
|
||
|
||
Reproduce with either:
|
||
|
||
```bash
|
||
# needs the Makepad desktop stack (wayland/X11/GL/alsa/polkit)
|
||
cargo test --locked --release -p nigig-build --lib -- \
|
||
--ignored --nocapture --test-threads=1 profile_benchmarks
|
||
|
||
# host-only: isolated toolchain, no desktop, self-cleaning
|
||
CAD_BENCH=1 ./tools/test-cad-coverage.sh
|
||
```
|
||
|
||
The second form exists because these benchmarks were unrunnable for ten
|
||
review tranches — `REVIEWS/PAY_CAD_IMPLEMENTATION_STATUS.md` recorded the
|
||
blocker as "no Cargo toolchain is available in this execution
|
||
environment". `profile_benchmarks.rs` turns out to import nothing from
|
||
Makepad but the math types, so it builds in the same host-only harness
|
||
the coverage run uses.
|
||
|
||
**Release mode matters.** Debug builds are 3–5× slower here and the ratios
|
||
between arms shift, so a debug number is not comparable to anything below.
|
||
|
||
These are measurements, not assertions — the benchmarks are `#[ignore]`-d
|
||
and deliberately contain no timing `assert!`s (a wall-clock assertion in
|
||
the normal test run flakes under CI load). Compare against this file by
|
||
hand when changing the hot paths.
|
||
|
||
The frame-submission rows below were recorded on the Phase 5 tree,
|
||
x86-64, `--release`. The wall-clock rows above are retained machine
|
||
baselines and are not comparable across hosts; the submission rows are
|
||
structural counts and are intentionally stable.
|
||
|
||
| Benchmark | Result |
|
||
|---|---|
|
||
| `bench_scene_cache_hit_vs_rebuild` | cold 25.9 µs / warm 53 ns (100 parts) — **488×** |
|
||
| `bench_scene_cache_scaling` | 10/50/100/500 parts: cold scales linearly, warm stays ~150 ns |
|
||
| `bench_mesh_cache_hit_cost` | 11.1 µs/frame, **110 ns per part** (100 parts, all hits) |
|
||
| `bench_size_parametric_vs_mesh_derived` | parametric 2 ns vs mesh-derived 886 ns — **443×** |
|
||
| `bench_pick_broadphase_mesh_bounds_vs_size` | 63.3 µs vs 131.0 µs per frame — **2.1× saved** |
|
||
| `bench_pick_broadphase_world_aabb_recompute_vs_cache` | **230 µs → 39 µs per pick at 500 parts — 5.9×.** Transforming 8 corners per part per pick, versus a `PlacedHash` lookup. Hover picking runs on mouse-move, so this is per-frame. Cached in `SceneCache::world_aabb_for` (Phase 3.2). |
|
||
| `bench_parts_script_regeneration_per_drag_frame` | **42 µs (50 parts) / 170 µs (200) / 425 µs (500)** per drag frame, before the editor's own `set_text`. `script_dirty` was set on every motion event; it is now set once on commit (Phase 3.6). |
|
||
| `bench_param_hash_cost_per_frame` | Box **50 ns/node** (25 µs/frame at 500) — not worth a cache. Extruded 64-gon **987 ns/node**; that is the vertex data, and bulk-hashing the slice measured no faster. Phase 3.3 left undone deliberately. |
|
||
| `bench_parallel_threshold_warm_vs_cold_cache` | 200 nodes: cold **8.80 ms seq / 6.21 ms par (1.42×)**, warm **5.81 / 5.73 (1.01×)**. The 32-node proxy is imprecise but harmless; Phase 3.8 rewrite not worth it. |
|
||
| `bench_glb_export_with_and_without_cache` | cold 1.42 ms / warm 1.35 ms (50 parts) |
|
||
| `bench_parallel_vs_sequential_export` | sequential 4.83 ms / parallel 3.58 ms (100 parts) — 1.3× |
|
||
| `bench_command_execute_overhead` | **0.8 µs per execute, 0.5 µs per undo** (1000 commands over a 1000-node scene). Was 254 µs / 246 µs: `CadCommandCtx::new` eagerly built a scene snapshot no production command reads. Now built on first `scene()` call — roughly **300×**. |
|
||
| `bench_gpu_upload_mesh_source` | boxes: rebuild 0.27 µs/part vs warm cache 0.178 µs — **2×**; extruded 24-gons: 1.11 µs vs 0.516 µs — **2×**. Cold cache 0.79 µs/part is **slower** than rebuilding (hash + insert) |
|
||
| `bench_delete_invalidation_clear_vs_evict` | deleting 1 of 100 parts: clear+rebuild 52.9 µs vs evict+reuse 14.5 µs — **4×** |
|
||
| `bench_viewport_snapshot_sync_per_frame` | deep copies 21.2 µs/frame + forced rebuilds 77.2 µs/frame = **98.5 µs/frame**; generation-checked 0.051 µs — **1930×** (100 parts) |
|
||
|
||
### Frame submission budget
|
||
|
||
`bench_frame_submission_budget`, added by Phase 0 of
|
||
`REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md` and extended through Phase 5.
|
||
Every other benchmark in this file measures CPU work *behind* the
|
||
renderer; none measures a live GPU frame, so the submission report stays
|
||
as an explicit structural count rather than pretending to be frame time.
|
||
|
||
Counts, not timings — counts do not flake under CI load. 1920×1080,
|
||
parts on a lattice across a 200 m site (deliberately larger than the
|
||
viewport at either zoom: a scene that fits on screen is the case where
|
||
culling can do nothing, and reporting only that would flatter it).
|
||
|
||
**2D, after Phase 1 culling, Phase 4 stroke batching, and Phase 5
|
||
sub-pixel LOD.** "Before" is the Phase 0 baseline: every part in the
|
||
document submitted, one `stroke()` per part and one per grid line.
|
||
|
||
| Zoom | Parts | Visible | Full outlines | LOD points | Part calls | Tessellations | Before |
|
||
|---|---|---:|---:|---:|---:|---:|---:|
|
||
| 5 m (170 grid lines) | 100 | 1 | 1 | 0 | 1 stroke | **3** (2 grid + 1 parts) | 270 |
|
||
| | 500 | 4 | 4 | 0 | 2 strokes | **4** | 670 |
|
||
| | 2,000 | 12 | 12 | 0 | 2 strokes | **4** | 2,170 |
|
||
| 200 m (270 grid lines) | 100 | 100 | 1 | 99 | 1 stroke + 1 fill | **4** | 370 |
|
||
| | 500 | 500 | 1 | 499 | 1 stroke + 1 fill | **4** | 770 |
|
||
| | 2,000 | 2,000 | 1 | 1,999 | 1 stroke + 1 fill | **4** | 2,270 |
|
||
| 500 m (170 grid lines) | 100 | 100 | 1 | 99 | 1 stroke + 1 fill | **4** | 270 |
|
||
| | 500 | 500 | 1 | 499 | 1 stroke + 1 fill | **4** | 670 |
|
||
| | 2,000 | 2,000 | 1 | 1,999 | 1 stroke + 1 fill | **4** | 2,170 |
|
||
|
||
The benchmark models one material plus one selected part. Selection is
|
||
intentionally kept at full detail, so the first visible part remains an
|
||
outline while the ordinary parts cross the LOD threshold. At 5 m, a 1 m
|
||
part is comfortably resolvable and every survivor stays full detail. At
|
||
200 m, the same part is only about 2.7 logical pixels across; at 500 m it
|
||
is about 1.1 pixels, so the ordinary survivors become 2×2 logical-pixel
|
||
filled markers.
|
||
|
||
Phase 5 does **not** claim another draw-call win. The 2D scene remains one
|
||
`DrawVector::end()` submission, and the one material still makes one
|
||
outline stroke plus one marker fill. The win is geometry volume: a filled
|
||
marker is bounded and does not carry a full rectangle/circle outline for
|
||
a part whose shape cannot be resolved. `FrameBudget` reports full
|
||
outlines and LOD markers separately so this change cannot be mistaken for
|
||
another tessellation reduction.
|
||
|
||
**3D, frustum culling (Phase 1), instanced batching (Phase 3), and
|
||
bounding-box LOD (Phase 5)**, 60° fov, camera orbiting the site centre.
|
||
"Shapes" is how many distinct full-detail solids the visible parts use —
|
||
6 for a catalogue-built model, all distinct for the worst case. LOD boxes
|
||
are all instances of one shared unit-cube proxy batch:
|
||
|
||
| Camera | Parts | Visible | Full meshes | LOD boxes | Draw calls, 6 shapes | Draw calls, all distinct | Before |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|
|
||
| 20 m | 100 | 39 | 39 | 0 | **6** | 39 | 100 |
|
||
| | 500 | 198 | 198 | 0 | **6** | 198 | 500 |
|
||
| | 2,000 | 733 | 733 | 0 | **6** | 733 | 2,000 |
|
||
| 80 m | 100 | 76 | 76 | 0 | **6** | 76 | 100 |
|
||
| | 500 | 386 | 386 | 0 | **6** | 386 | 500 |
|
||
| | 2,000 | 1,459 | 1,459 | 0 | **6** | 1,459 | 2,000 |
|
||
| 400 m | 100 | 100 | 20 | 80 | **7** | 21 | 100 |
|
||
| | 500 | 500 | 63 | 437 | **7** | 64 | 500 |
|
||
| | 2,000 | 2,000 | 380 | 1,620 | **7** | 381 | 2,000 |
|
||
|
||
The bottom row is the one to look at. At 400 m the whole site is on
|
||
screen, so culling removes nothing. LOD replaces 1,620 unresolvable
|
||
meshes with instances of one proxy, and the six full-shape groups plus
|
||
that proxy draw in **7 calls instead of 2,000**. The "all distinct" column
|
||
is the same scene with every full-detail part a different size; 380 remain
|
||
full and the 1,620 tiny parts still share the proxy, so the count is 381.
|
||
The selected/hovered exemption is not present in this synthetic 3-D
|
||
benchmark; the policy tests pin that decorated parts remain full detail.
|
||
|
||
Where culling does work it is decisive on its own: drafting at a 5 m
|
||
zoom over a 200 m site, the 2,000-part scene queues 12 outlines instead
|
||
of 2,000. Where it cannot help — everything on screen — batching still
|
||
holds the call count at 4.
|
||
|
||
Two different costs, deliberately not summed:
|
||
|
||
- **Tessellations** are CPU calls into `tessellate_path_stroke`. The
|
||
whole 2D scene is **one** draw call from `DrawVector::end()` no matter
|
||
how many parts, so this — not draw calls — was the number that scaled
|
||
there. Since Phase 4 it scales with colour groups instead.
|
||
- **Draw calls** are GPU submissions. 3D used to be one per uploaded,
|
||
visible part; since Phase 3 full-detail parts use one call per distinct
|
||
shape, with the parts riding inside as instance rows. Since Phase 5,
|
||
all tiny parts use one additional shared proxy batch.
|
||
|
||
The grid line count is unchanged by culling, as it should be: it was
|
||
already virtualised, and stays at 170–270 lines across the measured zoom
|
||
range because the step adapts. Phase 4 changed what those lines cost —
|
||
two `stroke()` calls rather than one each — and Phase 5 changes only the
|
||
representation of tiny parts. The tests `culling_shows_up_as_fewer_part_tessellations`,
|
||
`the_grid_costs_two_strokes_at_any_zoom`, and the LOD policy tests pin
|
||
those separate costs.
|
||
|
||
Phase 2 (sharing geometry between parts of the same shape, keyed by a
|
||
new `ShapeHash` — **not** `ParamHash`, which includes the node id) and
|
||
Phase 3 (instancing) brought the draw-call column down in the zoomed-out
|
||
case where culling cannot; Phase 4 did the same for the 2-D tessellation
|
||
column. Phase 5 addresses the remaining geometry volume in both paths:
|
||
the 2,000-part rows at 200 m and 500 m keep one selected outline but
|
||
replace 1,999 unresolvable outlines with bounded point markers, while the
|
||
400 m 3-D row replaces 1,620 unresolvable meshes with one shared proxy
|
||
batch. A real GPU/window visual check is still required before treating
|
||
the structural counts as frame-time results.
|
||
|
||
### GPU geometry buffers, keyed by shape
|
||
|
||
`bench_geometry_buffers_shared_by_shape`, added by Phase 2. Counts, not
|
||
timings. The number it replaces was never measured because it was never
|
||
in doubt: **one buffer per part**, however many of those parts were the
|
||
same column.
|
||
|
||
| Model | Parts | Buffers | Ratio |
|
||
|---|---|---|---|
|
||
| 200 identical walls (`build_house_scene(200)`) | 200 | **1** | 200× |
|
||
| 420-part repetitive model (columns, walls in three lengths, two opening types) | 420 | **6** | 70× |
|
||
| 420 parts, every one a different size | 420 | **420** | 1× |
|
||
|
||
Read the last row too. Sharing is a property of the model, not of the
|
||
code: a scene where every part differs gets nothing from this phase.
|
||
That is the honest ceiling, and it is why the draw-call column is
|
||
untouched here — Phase 3 (instancing) is what collapses draw calls, and
|
||
it needs the shared buffers this phase produces.
|
||
|
||
The middle row is the one to plan against. Real architectural models are
|
||
catalogue-built: a few column types, a few wall lengths, a door and a
|
||
window repeated everywhere. 70× fewer GPU buffers, and 70× fewer mesh
|
||
uploads on the first frame after a load.
|
||
|
||
### Note: two benchmarks had drifted off the caching path
|
||
|
||
If you ran these between the Phase 5.1 generation-tracked store and the
|
||
commit that added this note, `bench_scene_cache_hit_vs_rebuild` reported
|
||
**1.2×** against the **488×** in the table above, and
|
||
`bench_scene_cache_scaling` reported **1× at every part count**. That
|
||
looked exactly like a catastrophic cache regression. It was not.
|
||
|
||
Both benchmarks called `SceneCache::scene(&[CadNode])`, which is
|
||
documented as *always rebuilding* — it takes a bare slice, so it has no
|
||
generation to compare and cannot cache. The caching entry point the
|
||
editor uses is `SceneCache::scene_for(&PartsStore)`. When the store
|
||
landed, these two benchmarks were not moved with it, so their "warm"
|
||
sample was a second full rebuild and the printed speedup was allocator
|
||
noise.
|
||
|
||
Repointed at `scene_for`, they reproduce the table: cold 24.6 µs / warm
|
||
**48 ns**, **512×**, and the warm read stays flat at ~55 ns from 10 parts
|
||
to 500. `bench_scene_cache_hit_vs_rebuild` now asserts `Arc::ptr_eq` on
|
||
the two samples, so it fails loudly rather than quietly measuring two
|
||
rebuilds if it is ever pointed at a non-caching path again.
|
||
|
||
`SceneCache::scene()` is not the bug and was left alone: its docstring
|
||
says what it does, and nine tests legitimately use it for exactly the
|
||
case it describes.
|
||
|
||
### Note on `bench_gpu_upload_mesh_source`
|
||
|
||
Phase 5.3 routed the GPU upload path through `MeshCache` instead of
|
||
calling `build_solid()` directly. The measured speed win is **2× on a warm
|
||
cache and negative on a cold one** — a first-ever build now pays a
|
||
`ParamHash` and a map insert on top of the meshing.
|
||
|
||
That trade is worth taking, but not primarily for speed:
|
||
|
||
- `ensure_part_geometry` only builds for parts with no uploaded buffer, so
|
||
the cold case is a scene load or an undo/redo, not a frame in a loop.
|
||
- The warm case is every part that was already exported, hovered or
|
||
previously uploaded — which, during editing, is most of them.
|
||
- The real gain is that there is now **one** node-to-triangles path. Two
|
||
independent `match` arms over the same enum were free to drift, and
|
||
nothing would have caught it: the preview would simply have disagreed
|
||
with the export.
|
||
|
||
Do not "optimise" this back by special-casing cheap primitives. The 0.5 µs
|
||
is not worth reintroducing the second pipeline.
|
||
|
||
### Note on `bench_viewport_snapshot_sync_per_frame`
|
||
|
||
This benchmark measures the *primitives* the sync path is built from --
|
||
`Vec<CadNode>` clones and `SceneCache::scene()` rebuilds -- not the widget
|
||
method itself, which needs a live `Cx` and three real viewports. Its
|
||
printed numbers therefore do **not** move when the sync path is fixed;
|
||
they quantify what one frame of syncing costs and what the same frame
|
||
costs when a generation check is allowed to short-circuit it.
|
||
|
||
Read it as the size of the prize, not as a regression gate. The Phase 5.2
|
||
change is verified by behaviour tests
|
||
(`replace_all_bumps_the_generation_so_destinations_can_compare`,
|
||
`mesh_cache_self_invalidates_on_parameter_and_transform_edits`,
|
||
`geometry_is_retained_for_surviving_ids_only`), each negative-tested.
|
||
|
||
---
|
||
|
||
## Reading these
|
||
|
||
**The scene cache is doing its job.** 488× on a warm hit, and the scaling
|
||
run shows the warm path is O(1) while the cold path is linear. No work
|
||
needed here.
|
||
|
||
**`MeshCache` hits are not free.** 186 ns per part is `ParamHash::from_node`
|
||
being recomputed on every lookup — it hashes the solid's every field, the
|
||
transform and the material, even when the entry is present and valid. At
|
||
100 parts that is ~19 µs per hover frame. Not the dominant cost today, but
|
||
it is the next thing to fix if picking needs to get faster: memoise the
|
||
hash on `CadNode` and invalidate it on mutation.
|
||
|
||
**`CadNode::size()` is a trap.** 2 ns for a parametric solid, 886 ns for a
|
||
CSG or extruded one, because those have no closed-form extent and the
|
||
function meshes the solid to derive a bounding box (added in Phase 1.7 to
|
||
fix unpickable parts). It reads like a cheap field access at the call site.
|
||
Anything calling it in a loop should take bounds from a cached mesh
|
||
instead — which is exactly what `pick_part` now does.
|
||
|
||
**Export is already fast enough.** 1.4 ms for 50 parts, and the mesh cache
|
||
barely moves it because JSON/binary serialisation dominates, not meshing.
|
||
The parallel path wins only 1.3× at 100 parts. Neither is worth tuning
|
||
before the UI hot paths.
|
||
|
||
---
|
||
|
||
## What Phase 3 changed
|
||
|
||
| Change | Effect |
|
||
|---|---|
|
||
| `pick_part` takes AABB bounds from the mesh it already holds instead of calling `size()` | **2.1×** less broad-phase work per hover frame; also *more correct* — see below |
|
||
| Hover picking throttled by `HOVER_PICK_MIN_MOVE_PX` (3 px) | Sub-pixel jitter no longer triggers a full ray cast against every triangle |
|
||
| Removed `cx.redraw_all()` from the part-drag `MouseMove` handler | No full-widget-tree relayout per motion event during a drag |
|
||
|
||
The AABB change also fixed a latent correctness bug. `size()` returns a
|
||
*symmetric* extent about the origin, but an extruded polygon grows along
|
||
+Y from its base plane — so the old broad-phase box was in the wrong
|
||
place and could reject a ray that actually hits. `pick_bounds_tests::
|
||
mesh_bounds_are_asymmetric_where_size_is_not` pins this.
|
||
|
||
---
|
||
|
||
## Not done, and why
|
||
|
||
**Memoising `ParamHash` on `CadNode`** (~186 ns/part/frame). It needs a
|
||
cache field on the node plus invalidation on every mutation path, which is
|
||
the same ownership problem Phase 5.1 solves properly with a generation
|
||
counter. Doing it now would add a second thing to keep in sync by hand —
|
||
exactly the pattern this codebase already suffers from.
|
||
|
||
**A BVH or spatial index for picking.** Current cost is linear in part
|
||
count with a cheap early-out. At the 10–100 parts these models actually
|
||
contain, a BVH would cost more to build than it saves. Revisit if scenes
|
||
grow an order of magnitude.
|
||
|
||
**The other 81 `cx.redraw_all()` calls.** Most are on discrete user
|
||
actions (keypress, button, tool change) where a full redraw is once-per-
|
||
gesture and harmless. Only the per-motion ones were worth removing, and
|
||
`orbit_3d_by_screen_delta` / `pan_3d` were already correct.
|
||
|
||
|
||
## `bench_command_execute_overhead` — resolved
|
||
|
||
Recorded here because the number moved 300× and the reason is worth
|
||
keeping.
|
||
|
||
`CadCommandCtx::new` eagerly built a scene snapshot:
|
||
|
||
```rust
|
||
let built_scene = scene_cache.scene_for(parts); // O(nodes)
|
||
```
|
||
|
||
`CommandContext::scene()` needs one, but **no production command reads
|
||
it** — `MoveNode`, `ResizeNode`, `RotateNode`, `YawNode`, `ModifyNode`,
|
||
`CreateNode` and `DeleteNode` all address a node by id. Every command
|
||
mutation bumps the `PartsStore` generation, so the next context
|
||
construction was a guaranteed cache miss that cloned every node and
|
||
re-registered every material. `move_selected` issues one command per
|
||
selected part per frame, so a drag was O(commands × nodes).
|
||
|
||
The snapshot is now built on the first `scene()` call and memoised, and
|
||
dropped by every mutating method (paired with the existing
|
||
`mark_dirty()` calls, so a new mutator cannot silently miss it).
|
||
|
||
The eager build was also a latent **correctness** bug: it captured the
|
||
scene *before* the command ran, so a command that mutated and then read
|
||
`scene()` saw its own edit missing. Pinned by
|
||
`a_lazily_built_scene_reflects_edits_made_earlier_in_the_same_context`.
|
||
|
||
Two measurement notes:
|
||
|
||
- The old 0.4 µs figure predated `PartsStore`; the generation counter
|
||
turned a previously accidental cache hit into a guaranteed miss. The
|
||
regression was real, not a measurement artefact.
|
||
- The benchmark itself called `ctx.scene()` each iteration to read the
|
||
start position, which no production path does — `move_selected` and
|
||
`finish_part_drag` both read `p.pos()` off the parts list. Fixing the
|
||
lazy build alone moved undo 228 → 0.5 µs but left execute at 231 µs,
|
||
because the benchmark was timing its own scene read. It now mirrors
|
||
the production callers.
|
||
|
||
---
|
||
|
||
## Engine baselines (doc + spreadsheet)
|
||
|
||
Added in PLAN_PERF_OPTIMIZATION.md Phase 0. Release-only, `#[ignore]`-d
|
||
benches that **print, never assert** wall-clock numbers (CI-load flakes).
|
||
Host: x86-64, macOS, `--release`, `target-cpu` not set. Reproduce:
|
||
|
||
```bash
|
||
cargo test --release -p spreadsheet-engine --lib -- \
|
||
--ignored --nocapture --test-threads=1 bench_
|
||
cargo test --release -p doc-engine --lib -- \
|
||
--ignored --nocapture --test-threads=1 bench_
|
||
```
|
||
|
||
| Benchmark | Baseline (this host) | What it measures |
|
||
|---|---|---|
|
||
| `bench_grid_set_get_10k` | 1.65 µs per get_raw+display | HashMap-backed grid read hot loop |
|
||
| `bench_formula_chain_recalc_1k` | 22.6 ms per full recalc | 1000-cell dependency-chain rebuild+eval |
|
||
| `bench_column_sum_10k` | 91.8 ms per recalc | SUM(A1:A10000) — aggregate hot loop |
|
||
| `bench_workbook_roundtrip_10k` | 97.7 ms per roundtrip | serialize+deserialize 10k cells |
|
||
| `bench_crdt_materialize_2k` | **29.9 ms per materialize** | whole-op-log → projection, 2000 chars |
|
||
| `bench_crdt_save_wire_2k` | **9.9 ms per to_json+from_json** | full-doc save wire, 2000 chars |
|
||
| `bench_insert_char_1k` | **3.65 ms per keystroke** | insert atom + re-materialize per key |
|
||
|
||
The doc rows confirm the Phase 6 targets are real: typing is ~3.65 ms/key
|
||
(materialize dominates), and each save rewrites the whole op log.
|