nigig-org/REVIEWS/CAD_DRAWCALL_STRATEGY_ANALYSIS.md
andodeki cbce585eb4
Some checks failed
nigig-build (CAD) / supply-chain (push) Has been cancelled
nigig-build (CAD) / cad-module (push) Has been cancelled
nigig-build (CAD) / full-crate-check (push) Has been cancelled
nigig-build (CAD) / cad-engine-coverage (push) Has been cancelled
nigig-build (CAD) / doc-workspace-coverage (push) Has been cancelled
nigig-build (CAD) / cad-widget-coverage (push) Has been cancelled
repo hygiene / hygiene (push) Has been cancelled
perf(cad): complete 2D and 3D LOD -- Phase 5
The first Phase 5 tranche correctly reduced sub-pixel 2D outlines, but it stopped short of the plan's full level-of-detail phase. Complete the phase by applying the same measured three-pixel policy to the 3D path.

Visible 3D parts below the projected AABB threshold now use one shared unit-cube geometry with an axis-aligned world-bounds transform. Full-detail shape batches, proxy instance rows and proxy draw calls are reported separately. Wireframe and hidden-line overlays use the same policy, and selected or hovered parts remain full detail.

Extend the host-testable projection and proxy policy tests, frame benchmark, source guards and documentation. Keep the real GPU/window visual check explicit: compilation, policy arithmetic, submission structure and coverage pass here, but this environment cannot execute Makepad's draw submission.
2026-08-26 16:25:45 +00:00

259 lines
12 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Does the CAD viewport have a minimal-drawcall strategy?
Short answer: **the pre-Phase-0 renderer did not, but the CAD module now
uses the transferable parts of the strategy.** Phases 1–5 added viewport
culling, shape-shared 3-D instancing, 2-D colour batching and measured
sub-pixel LOD. The 2-D scene remains one Makepad vector draw call; the
useful work there is reducing tessellation and geometry volume. The 3-D
submission is now one call per visible shape, with instance rows.
This was prompted by a datagrid brief asking for "virtual viewport on
both axes" and "an optimal minimal drawcall strategy". Those are two
distinct techniques. CAD has no cell recycling or row/column index
virtualisation, but continuous-space culling, batching, instancing and
LOD transfer directly.
The first sections below are the historical pre-Phase-0 audit, from
reading `viewport_render.rs` and `viewport.rs`. The current measurements
and phase status live in `REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md` and
`BENCH_BASELINE.md`; the original "not measured" caveat is retained and
marked as historical rather than left to read as the current state.
## 0. Corrections to earlier versions of this note
### 0.1 `stroke()` is not a draw call
The first version of this document called `stroke()` a "tessellation/flush
point" and then reasoned about the 2D path as if each one cost a draw
call — "~670 flush points per frame where four would do". That is wrong,
and it was wrong because I inferred Makepad's batching model from a
comment instead of reading it.
Read now, in `draw/src/shader/draw_vector.rs`:
- `begin()` clears CPU-side accumulation buffers (`acc_verts`, `acc_indices`).
- `stroke()` calls `tessellate_path_stroke(...)` and `append_geometry(...)`.
It tessellates the current path into those buffers. **It issues no
draw call.**
- `end(cx)` is the only place `cx.new_draw_call(&self.draw_vars)` appears
— twice in the whole file, both inside `end()`, for the gradient and
plain paths.
So the entire 2D vector scene — grid, parts, overlays — is **one draw
call**. Makepad's `DrawVector` is already a batching design and the CAD
code uses it correctly in that respect. The guard test's own wording,
"thousands of tessellations per frame", was accurate and literal; I read
"draw call" into it.
### 0.2 `ParamHash` is not a shape key
Section 3 and the suggested order both said identical parts could share
a GPU buffer by re-keying `part_geoms` on `ParamHash`. They cannot:
`ParamHash::from_node` hashes **`node.id` first** (`cad_scene.rs:1871`),
before any geometric parameter. Two identical columns at different
coordinates therefore have different `ParamHash`es, and re-keying on it
would share nothing whatsoever.
I took the type name and its docstring ("keyed on the node's mesh
parameters") for the whole story and did not read the body. The same
assumption appeared independently in a second optimisation review, which
is the sort of coincidence that argues for a test rather than a note:
`param_hash_is_not_a_shape_key_it_includes_the_node_id` in
`cad_scene.rs` now pins it.
The id is not a bug in `ParamHash` — its two users, `MeshCache` and
`part_geoms`, are keyed by `NodeId` and only ask "is this cached entry
still valid for *this* node", where the id is a constant. Sharing
geometry needs a *second* hash, `ShapeHash`: the same body without the
id. `REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md` Phase 2 carries it.
What survives the correction, and what changes:
| Claim | Status |
|---|---|
| No viewport culling anywhere | **Stands.** Still the main finding. |
| 3D issues one draw call per part | **Stands.** That is where draw-call multiplication is real. |
| 2D costs ~670 draw calls a frame | **Wrong.** One draw call; the per-item cost is CPU tessellation. |
| Batching the 2D loops is worthwhile | **Weaker, still true.** It cuts tessellation-call overhead, not draw calls. |
The rest of this document is the corrected version.
## 1. Historical baseline: batching was present, used once, guarded by a test
`draw_vector` is an immediate-mode path builder. `stroke()` **tessellates**
the current path into an accumulation buffer; `end()` turns the whole
accumulation into one draw call. So the cost of calling `stroke()` per
item is CPU tessellation and per-call setup, not GPU submission. The
codebase demonstrably knows this. From the axis grid:
```rust
fn queue_dashed_line(&mut self, x1: f32, y1: f32, x2: f32, y2: f32) {
while pos < dist {
self.draw_vector.move_to(...);
self.draw_vector.line_to(...); // queue only
pos = seg_end + gap;
}
} // no stroke() here
```
…and a test in `viewport.rs` enforces it:
> `"the grid lines should share exactly one stroke"`
>
> `"draw_dashed_line strokes per dash, which is thousands of
> tessellations per frame"`
That is precisely the minimal-drawcall discipline the datagrid brief
asks for. It is applied to `draw_axis_grid_2d` and nowhere else.
### Where it is not applied
`draw_2d_vector_scene` has 13 `stroke()` sites, three of them **inside
loops**:
```rust
for i in start_x..=end_x { // base grid, vertical lines
self.draw_vector.move_to(p1);
self.draw_vector.line_to(p2);
self.draw_vector.stroke(if is_major { 1.6 } else { 0.55 }); // per line
}
// …same again for horizontals…
for part in read_parts(&doc).iter() { // every part, every frame
self.draw_vector.set_color(col);
self.draw_vector.rect(...);
self.draw_vector.stroke(1.8); // per part
}
```
The grid targets ~18 px spacing, so a 1920×1080 viewport is roughly
107 vertical + 60 horizontal ≈ **167 tessellation calls for the grid
alone**, plus **one per part**. At 500 parts that is ~670 calls into
`tessellate_path_stroke` per frame where the file's own proven idiom
would give roughly four.
To be clear about the magnitude after the correction above: this is CPU
work — tessellator setup, two `std::mem::take`s and an `append_geometry`
per call — not 670 draw calls. The same total segment count still has to
be tessellated either way. The win is per-call overhead, and it is worth
having, but it is a smaller prize than culling or 3D instancing.
The base grid is the easier of the two: it only alternates between two
stroke widths, so it is two batches, not 167. The parts loop alternates
colour per part (selected / hovered / own colour), so it needs grouping
by colour or a per-instance colour attribute.
### Historical 3-D path: one draw call per part
```rust
for part in read_parts(&doc).iter() {
if let Some(geom) = self.part_geoms.get(&part.id.raw())… {
self.draw_mesh.transform = part_model_matrix_cadnode(part);
self.draw_mesh.color = …;
self.draw_mesh.draw(cx, geom.geometry_id()); // one call, per part
}
}
```
Each part carries its own geometry buffer and its own transform/colour
uniforms, so N parts is N draw calls. No instancing, no batching, no
sorting by state.
## 2. Historical baseline: virtual viewport present for the grid, absent for content
The grid **is** virtualised, and well:
```rust
let world_left = self.pan_2d.x - half_w * 1.2; // visible bounds + 20%
let start_x = ((world_left - gx) / step).floor() as i32;
let end_x = ((world_right - gx) / step).ceil() as i32;
for i in start_x..=end_x { … }
```
Only visible grid lines are emitted, and the step adapts to zoom with a
1/2/5 nice-number progression. That is the datagrid's "only materialise
what is on screen", done properly.
In the historical pre-Phase-1 tree the parts were not culled at all.
`grep -niE "cull|frustum|offscreen|in_view"` over the whole 2,461-line
renderer returned **nothing**. Every part was projected and submitted
every frame whether or not it was on screen, in both the 2D and 3D paths.
Phase 1 now replaces this with the shared 2-D rect and 3-D frustum tests;
see the current plan and benchmark rather than treating this paragraph as
an assertion about the present tree.
### The part that stings
The broad-phase structure needed for culling already exists:
```rust
pub fn world_aabb_for(&self, node: &CadNode, build: impl FnOnce(&CadNode) -> WorldAabb) -> WorldAabb
```
It is cached by `PlacedHash`, and `BENCH_BASELINE.md` records it at
**4.72× faster than recomputing** (94 µs → 20 µs at 500 parts). Its only
caller is `CadViewport::pick_part` — the *picking* path, which runs on
mouse-move and is additionally throttled by `HOVER_PICK_MIN_MOVE_PX`.
So the cheap visibility test exists, is cached, is benchmarked, and is
wired to the path that runs occasionally rather than the path that runs
every frame.
## 3. Which datagrid techniques transfer
| Technique | Transfers? | How it maps to CAD |
|---|---|---|
| Virtual viewport | **Yes, directly** | Test each part's cached world AABB against the view rect (2D) or frustum (3D) before submitting. Phase 1 wires the shared predicates into the draw loops and `frame_budget`. |
| Minimal draw calls | **Partly — 2D is already one call** | `DrawVector` batches to a single draw call at `end()`. Phase 4 reduces tessellation calls via queued colour groups; Phase 3 is the real GPU draw-call win in 3D. |
| Instancing | **Yes, and it is the big one** | Phase 2 keys local geometry by `ShapeHash`, and Phase 3 sends repeated transforms/colours as one instanced draw per visible shape. `ParamHash` is intentionally not the shape key — see correction 0.2 above. |
| Proper clipping | Partly | The datagrid needs nested clip rects per cell; CAD has one viewport. Parts are emitted in screen coordinates and left to the GPU to clip, while culling removes work before clipping. |
| Level of detail | **Yes, measured and implemented in 2-D and 3-D** | `lod.rs` uses a shared three-pixel threshold: ordinary sub-pixel 2-D parts become bounded filled markers and 3-D parts become shared bounding-box proxy instances; selected/hovered parts stay full detail. |
### What does not transfer
- **Widget recycling.** Datagrid cells host child widgets and need a
pool. CAD parts are geometry, not widgets; there is nothing to recycle.
- **Index-range virtualisation.** A grid virtualises by row/column
index, which is O(1) to compute. CAD is continuous space, so the
equivalent needs a spatial test. At the scale this app targets a
linear pass over cached AABBs is fine — it is a few microseconds for
500 parts. A BVH or grid hash only earns its complexity somewhere
north of ~50k parts, and nothing here suggests that is the target.
## 4. What was not measured in the original note
The original audit had no frame-submission benchmark. Its structural
claims were read off the loops, and the consequence — whether the old
path dropped frames — was unmeasured. That was corrected by Phase 0:
`profile_benchmarks.rs::bench_frame_submission_budget` now reports the
visible geometry, tessellation groups and 3-D shape batches, and the
same decision functions drive both the renderer and `frame_budget`.
Phase 5 extends that report with full outlines versus sub-pixel markers
and full 3-D meshes versus bounding-box proxy instances.
The structural count is still not a live GPU timing: there is no `Cx`,
window or GPU in the host coverage harness, and the 3-D instanced/proxy
submission still needs the documented human visual check. The measured
CPU-side cache results remain useful — the mesh cache, scene cache and
AABB cache are independently benchmarked — but they must not be sold as
frame-time measurements.
## 5. Suggested order, cheapest first — current status
1. **Cull parts against the view rect.** Done in Phase 1, using the
shared cached-AABB/frustum predicates.
2. **Key `part_geoms` by shape and instance repeated parts.** Done in
Phases 2–3. The key is `ShapeHash`, **not** `ParamHash` — correction
0.2 above.
3. **Add a frame-submission benchmark.** Done in Phase 0 and extended
through Phase 5; it reports counts without pretending they are GPU
timings.
4. **Batch the base grid and part outlines.** Done in Phase 4, with
marker fills included in the Phase 5 budget.
5. **Use level of detail for sub-pixel parts.** Done for both paths in
Phase 5: 2-D uses bounded filled markers and 3-D uses one shared
bounding-box proxy batch. Live GPU visual verification remains the
release check, not an implementation blocker.
Phased details and the measured rows are in
`REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md` and `BENCH_BASELINE.md`.