Some checks failed
nigig-build (CAD) / supply-chain (push) Has been cancelled
nigig-build (CAD) / cad-module (push) Has been cancelled
nigig-build (CAD) / full-crate-check (push) Has been cancelled
nigig-build (CAD) / cad-engine-coverage (push) Has been cancelled
nigig-build (CAD) / doc-workspace-coverage (push) Has been cancelled
nigig-build (CAD) / cad-widget-coverage (push) Has been cancelled
repo hygiene / hygiene (push) Has been cancelled
The first Phase 5 tranche correctly reduced sub-pixel 2D outlines, but it stopped short of the plan's full level-of-detail phase. Complete the phase by applying the same measured three-pixel policy to the 3D path. Visible 3D parts below the projected AABB threshold now use one shared unit-cube geometry with an axis-aligned world-bounds transform. Full-detail shape batches, proxy instance rows and proxy draw calls are reported separately. Wireframe and hidden-line overlays use the same policy, and selected or hovered parts remain full detail. Extend the host-testable projection and proxy policy tests, frame benchmark, source guards and documentation. Keep the real GPU/window visual check explicit: compilation, policy arithmetic, submission structure and coverage pass here, but this environment cannot execute Makepad's draw submission.
259 lines
12 KiB
Markdown
259 lines
12 KiB
Markdown
# Does the CAD viewport have a minimal-drawcall strategy?
|
||
|
||
Short answer: **the pre-Phase-0 renderer did not, but the CAD module now
|
||
uses the transferable parts of the strategy.** Phases 1–5 added viewport
|
||
culling, shape-shared 3-D instancing, 2-D colour batching and measured
|
||
sub-pixel LOD. The 2-D scene remains one Makepad vector draw call; the
|
||
useful work there is reducing tessellation and geometry volume. The 3-D
|
||
submission is now one call per visible shape, with instance rows.
|
||
|
||
This was prompted by a datagrid brief asking for "virtual viewport on
|
||
both axes" and "an optimal minimal drawcall strategy". Those are two
|
||
distinct techniques. CAD has no cell recycling or row/column index
|
||
virtualisation, but continuous-space culling, batching, instancing and
|
||
LOD transfer directly.
|
||
|
||
The first sections below are the historical pre-Phase-0 audit, from
|
||
reading `viewport_render.rs` and `viewport.rs`. The current measurements
|
||
and phase status live in `REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md` and
|
||
`BENCH_BASELINE.md`; the original "not measured" caveat is retained and
|
||
marked as historical rather than left to read as the current state.
|
||
|
||
## 0. Corrections to earlier versions of this note
|
||
|
||
### 0.1 `stroke()` is not a draw call
|
||
|
||
The first version of this document called `stroke()` a "tessellation/flush
|
||
point" and then reasoned about the 2D path as if each one cost a draw
|
||
call — "~670 flush points per frame where four would do". That is wrong,
|
||
and it was wrong because I inferred Makepad's batching model from a
|
||
comment instead of reading it.
|
||
|
||
Read now, in `draw/src/shader/draw_vector.rs`:
|
||
|
||
- `begin()` clears CPU-side accumulation buffers (`acc_verts`, `acc_indices`).
|
||
- `stroke()` calls `tessellate_path_stroke(...)` and `append_geometry(...)`.
|
||
It tessellates the current path into those buffers. **It issues no
|
||
draw call.**
|
||
- `end(cx)` is the only place `cx.new_draw_call(&self.draw_vars)` appears
|
||
— twice in the whole file, both inside `end()`, for the gradient and
|
||
plain paths.
|
||
|
||
So the entire 2D vector scene — grid, parts, overlays — is **one draw
|
||
call**. Makepad's `DrawVector` is already a batching design and the CAD
|
||
code uses it correctly in that respect. The guard test's own wording,
|
||
"thousands of tessellations per frame", was accurate and literal; I read
|
||
"draw call" into it.
|
||
|
||
### 0.2 `ParamHash` is not a shape key
|
||
|
||
Section 3 and the suggested order both said identical parts could share
|
||
a GPU buffer by re-keying `part_geoms` on `ParamHash`. They cannot:
|
||
`ParamHash::from_node` hashes **`node.id` first** (`cad_scene.rs:1871`),
|
||
before any geometric parameter. Two identical columns at different
|
||
coordinates therefore have different `ParamHash`es, and re-keying on it
|
||
would share nothing whatsoever.
|
||
|
||
I took the type name and its docstring ("keyed on the node's mesh
|
||
parameters") for the whole story and did not read the body. The same
|
||
assumption appeared independently in a second optimisation review, which
|
||
is the sort of coincidence that argues for a test rather than a note:
|
||
`param_hash_is_not_a_shape_key_it_includes_the_node_id` in
|
||
`cad_scene.rs` now pins it.
|
||
|
||
The id is not a bug in `ParamHash` — its two users, `MeshCache` and
|
||
`part_geoms`, are keyed by `NodeId` and only ask "is this cached entry
|
||
still valid for *this* node", where the id is a constant. Sharing
|
||
geometry needs a *second* hash, `ShapeHash`: the same body without the
|
||
id. `REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md` Phase 2 carries it.
|
||
|
||
What survives the correction, and what changes:
|
||
|
||
| Claim | Status |
|
||
|---|---|
|
||
| No viewport culling anywhere | **Stands.** Still the main finding. |
|
||
| 3D issues one draw call per part | **Stands.** That is where draw-call multiplication is real. |
|
||
| 2D costs ~670 draw calls a frame | **Wrong.** One draw call; the per-item cost is CPU tessellation. |
|
||
| Batching the 2D loops is worthwhile | **Weaker, still true.** It cuts tessellation-call overhead, not draw calls. |
|
||
|
||
The rest of this document is the corrected version.
|
||
|
||
## 1. Historical baseline: batching was present, used once, guarded by a test
|
||
|
||
`draw_vector` is an immediate-mode path builder. `stroke()` **tessellates**
|
||
the current path into an accumulation buffer; `end()` turns the whole
|
||
accumulation into one draw call. So the cost of calling `stroke()` per
|
||
item is CPU tessellation and per-call setup, not GPU submission. The
|
||
codebase demonstrably knows this. From the axis grid:
|
||
|
||
```rust
|
||
fn queue_dashed_line(&mut self, x1: f32, y1: f32, x2: f32, y2: f32) {
|
||
while pos < dist {
|
||
self.draw_vector.move_to(...);
|
||
self.draw_vector.line_to(...); // queue only
|
||
pos = seg_end + gap;
|
||
}
|
||
} // no stroke() here
|
||
```
|
||
|
||
…and a test in `viewport.rs` enforces it:
|
||
|
||
> `"the grid lines should share exactly one stroke"`
|
||
>
|
||
> `"draw_dashed_line strokes per dash, which is thousands of
|
||
> tessellations per frame"`
|
||
|
||
That is precisely the minimal-drawcall discipline the datagrid brief
|
||
asks for. It is applied to `draw_axis_grid_2d` and nowhere else.
|
||
|
||
### Where it is not applied
|
||
|
||
`draw_2d_vector_scene` has 13 `stroke()` sites, three of them **inside
|
||
loops**:
|
||
|
||
```rust
|
||
for i in start_x..=end_x { // base grid, vertical lines
|
||
self.draw_vector.move_to(p1);
|
||
self.draw_vector.line_to(p2);
|
||
self.draw_vector.stroke(if is_major { 1.6 } else { 0.55 }); // per line
|
||
}
|
||
// …same again for horizontals…
|
||
|
||
for part in read_parts(&doc).iter() { // every part, every frame
|
||
self.draw_vector.set_color(col);
|
||
self.draw_vector.rect(...);
|
||
self.draw_vector.stroke(1.8); // per part
|
||
}
|
||
```
|
||
|
||
The grid targets ~18 px spacing, so a 1920×1080 viewport is roughly
|
||
107 vertical + 60 horizontal ≈ **167 tessellation calls for the grid
|
||
alone**, plus **one per part**. At 500 parts that is ~670 calls into
|
||
`tessellate_path_stroke` per frame where the file's own proven idiom
|
||
would give roughly four.
|
||
|
||
To be clear about the magnitude after the correction above: this is CPU
|
||
work — tessellator setup, two `std::mem::take`s and an `append_geometry`
|
||
per call — not 670 draw calls. The same total segment count still has to
|
||
be tessellated either way. The win is per-call overhead, and it is worth
|
||
having, but it is a smaller prize than culling or 3D instancing.
|
||
|
||
The base grid is the easier of the two: it only alternates between two
|
||
stroke widths, so it is two batches, not 167. The parts loop alternates
|
||
colour per part (selected / hovered / own colour), so it needs grouping
|
||
by colour or a per-instance colour attribute.
|
||
|
||
### Historical 3-D path: one draw call per part
|
||
|
||
```rust
|
||
for part in read_parts(&doc).iter() {
|
||
if let Some(geom) = self.part_geoms.get(&part.id.raw())… {
|
||
self.draw_mesh.transform = part_model_matrix_cadnode(part);
|
||
self.draw_mesh.color = …;
|
||
self.draw_mesh.draw(cx, geom.geometry_id()); // one call, per part
|
||
}
|
||
}
|
||
```
|
||
|
||
Each part carries its own geometry buffer and its own transform/colour
|
||
uniforms, so N parts is N draw calls. No instancing, no batching, no
|
||
sorting by state.
|
||
|
||
## 2. Historical baseline: virtual viewport present for the grid, absent for content
|
||
|
||
The grid **is** virtualised, and well:
|
||
|
||
```rust
|
||
let world_left = self.pan_2d.x - half_w * 1.2; // visible bounds + 20%
|
||
let start_x = ((world_left - gx) / step).floor() as i32;
|
||
let end_x = ((world_right - gx) / step).ceil() as i32;
|
||
for i in start_x..=end_x { … }
|
||
```
|
||
|
||
Only visible grid lines are emitted, and the step adapts to zoom with a
|
||
1/2/5 nice-number progression. That is the datagrid's "only materialise
|
||
what is on screen", done properly.
|
||
|
||
In the historical pre-Phase-1 tree the parts were not culled at all.
|
||
`grep -niE "cull|frustum|offscreen|in_view"` over the whole 2,461-line
|
||
renderer returned **nothing**. Every part was projected and submitted
|
||
every frame whether or not it was on screen, in both the 2D and 3D paths.
|
||
Phase 1 now replaces this with the shared 2-D rect and 3-D frustum tests;
|
||
see the current plan and benchmark rather than treating this paragraph as
|
||
an assertion about the present tree.
|
||
|
||
### The part that stings
|
||
|
||
The broad-phase structure needed for culling already exists:
|
||
|
||
```rust
|
||
pub fn world_aabb_for(&self, node: &CadNode, build: impl FnOnce(&CadNode) -> WorldAabb) -> WorldAabb
|
||
```
|
||
|
||
It is cached by `PlacedHash`, and `BENCH_BASELINE.md` records it at
|
||
**4.72× faster than recomputing** (94 µs → 20 µs at 500 parts). Its only
|
||
caller is `CadViewport::pick_part` — the *picking* path, which runs on
|
||
mouse-move and is additionally throttled by `HOVER_PICK_MIN_MOVE_PX`.
|
||
|
||
So the cheap visibility test exists, is cached, is benchmarked, and is
|
||
wired to the path that runs occasionally rather than the path that runs
|
||
every frame.
|
||
|
||
## 3. Which datagrid techniques transfer
|
||
|
||
| Technique | Transfers? | How it maps to CAD |
|
||
|---|---|---|
|
||
| Virtual viewport | **Yes, directly** | Test each part's cached world AABB against the view rect (2D) or frustum (3D) before submitting. Phase 1 wires the shared predicates into the draw loops and `frame_budget`. |
|
||
| Minimal draw calls | **Partly — 2D is already one call** | `DrawVector` batches to a single draw call at `end()`. Phase 4 reduces tessellation calls via queued colour groups; Phase 3 is the real GPU draw-call win in 3D. |
|
||
| Instancing | **Yes, and it is the big one** | Phase 2 keys local geometry by `ShapeHash`, and Phase 3 sends repeated transforms/colours as one instanced draw per visible shape. `ParamHash` is intentionally not the shape key — see correction 0.2 above. |
|
||
| Proper clipping | Partly | The datagrid needs nested clip rects per cell; CAD has one viewport. Parts are emitted in screen coordinates and left to the GPU to clip, while culling removes work before clipping. |
|
||
| Level of detail | **Yes, measured and implemented in 2-D and 3-D** | `lod.rs` uses a shared three-pixel threshold: ordinary sub-pixel 2-D parts become bounded filled markers and 3-D parts become shared bounding-box proxy instances; selected/hovered parts stay full detail. |
|
||
|
||
### What does not transfer
|
||
|
||
- **Widget recycling.** Datagrid cells host child widgets and need a
|
||
pool. CAD parts are geometry, not widgets; there is nothing to recycle.
|
||
- **Index-range virtualisation.** A grid virtualises by row/column
|
||
index, which is O(1) to compute. CAD is continuous space, so the
|
||
equivalent needs a spatial test. At the scale this app targets a
|
||
linear pass over cached AABBs is fine — it is a few microseconds for
|
||
500 parts. A BVH or grid hash only earns its complexity somewhere
|
||
north of ~50k parts, and nothing here suggests that is the target.
|
||
|
||
## 4. What was not measured in the original note
|
||
|
||
The original audit had no frame-submission benchmark. Its structural
|
||
claims were read off the loops, and the consequence — whether the old
|
||
path dropped frames — was unmeasured. That was corrected by Phase 0:
|
||
`profile_benchmarks.rs::bench_frame_submission_budget` now reports the
|
||
visible geometry, tessellation groups and 3-D shape batches, and the
|
||
same decision functions drive both the renderer and `frame_budget`.
|
||
Phase 5 extends that report with full outlines versus sub-pixel markers
|
||
and full 3-D meshes versus bounding-box proxy instances.
|
||
|
||
The structural count is still not a live GPU timing: there is no `Cx`,
|
||
window or GPU in the host coverage harness, and the 3-D instanced/proxy
|
||
submission still needs the documented human visual check. The measured
|
||
CPU-side cache results remain useful — the mesh cache, scene cache and
|
||
AABB cache are independently benchmarked — but they must not be sold as
|
||
frame-time measurements.
|
||
|
||
## 5. Suggested order, cheapest first — current status
|
||
|
||
1. **Cull parts against the view rect.** Done in Phase 1, using the
|
||
shared cached-AABB/frustum predicates.
|
||
2. **Key `part_geoms` by shape and instance repeated parts.** Done in
|
||
Phases 2–3. The key is `ShapeHash`, **not** `ParamHash` — correction
|
||
0.2 above.
|
||
3. **Add a frame-submission benchmark.** Done in Phase 0 and extended
|
||
through Phase 5; it reports counts without pretending they are GPU
|
||
timings.
|
||
4. **Batch the base grid and part outlines.** Done in Phase 4, with
|
||
marker fills included in the Phase 5 budget.
|
||
5. **Use level of detail for sub-pixel parts.** Done for both paths in
|
||
Phase 5: 2-D uses bounded filled markers and 3-D uses one shared
|
||
bounding-box proxy batch. Live GPU visual verification remains the
|
||
release check, not an implementation blocker.
|
||
|
||
Phased details and the measured rows are in
|
||
`REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md` and `BENCH_BASELINE.md`.
|