nigig-org/REVIEWS/CAD_DRAWCALL_STRATEGY_ANALYSIS.md
andodeki cbce585eb4
Some checks failed
nigig-build (CAD) / supply-chain (push) Has been cancelled
nigig-build (CAD) / cad-module (push) Has been cancelled
nigig-build (CAD) / full-crate-check (push) Has been cancelled
nigig-build (CAD) / cad-engine-coverage (push) Has been cancelled
nigig-build (CAD) / doc-workspace-coverage (push) Has been cancelled
nigig-build (CAD) / cad-widget-coverage (push) Has been cancelled
repo hygiene / hygiene (push) Has been cancelled
perf(cad): complete 2D and 3D LOD -- Phase 5
The first Phase 5 tranche correctly reduced sub-pixel 2D outlines, but it stopped short of the plan's full level-of-detail phase. Complete the phase by applying the same measured three-pixel policy to the 3D path.

Visible 3D parts below the projected AABB threshold now use one shared unit-cube geometry with an axis-aligned world-bounds transform. Full-detail shape batches, proxy instance rows and proxy draw calls are reported separately. Wireframe and hidden-line overlays use the same policy, and selected or hovered parts remain full detail.

Extend the host-testable projection and proxy policy tests, frame benchmark, source guards and documentation. Keep the real GPU/window visual check explicit: compilation, policy arithmetic, submission structure and coverage pass here, but this environment cannot execute Makepad's draw submission.
2026-08-26 16:25:45 +00:00

12 KiB
Raw Permalink Blame History

Does the CAD viewport have a minimal-drawcall strategy?

Short answer: the pre-Phase-0 renderer did not, but the CAD module now uses the transferable parts of the strategy. Phases 1–5 added viewport culling, shape-shared 3-D instancing, 2-D colour batching and measured sub-pixel LOD. The 2-D scene remains one Makepad vector draw call; the useful work there is reducing tessellation and geometry volume. The 3-D submission is now one call per visible shape, with instance rows.

This was prompted by a datagrid brief asking for "virtual viewport on both axes" and "an optimal minimal drawcall strategy". Those are two distinct techniques. CAD has no cell recycling or row/column index virtualisation, but continuous-space culling, batching, instancing and LOD transfer directly.

The first sections below are the historical pre-Phase-0 audit, from reading viewport_render.rs and viewport.rs. The current measurements and phase status live in REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md and BENCH_BASELINE.md; the original "not measured" caveat is retained and marked as historical rather than left to read as the current state.

0. Corrections to earlier versions of this note

0.1 stroke() is not a draw call

The first version of this document called stroke() a "tessellation/flush point" and then reasoned about the 2D path as if each one cost a draw call — "~670 flush points per frame where four would do". That is wrong, and it was wrong because I inferred Makepad's batching model from a comment instead of reading it.

Read now, in draw/src/shader/draw_vector.rs:

  • begin() clears CPU-side accumulation buffers (acc_verts, acc_indices).
  • stroke() calls tessellate_path_stroke(...) and append_geometry(...). It tessellates the current path into those buffers. It issues no draw call.
  • end(cx) is the only place cx.new_draw_call(&self.draw_vars) appears — twice in the whole file, both inside end(), for the gradient and plain paths.

So the entire 2D vector scene — grid, parts, overlays — is one draw call. Makepad's DrawVector is already a batching design and the CAD code uses it correctly in that respect. The guard test's own wording, "thousands of tessellations per frame", was accurate and literal; I read "draw call" into it.

0.2 ParamHash is not a shape key

Section 3 and the suggested order both said identical parts could share a GPU buffer by re-keying part_geoms on ParamHash. They cannot: ParamHash::from_node hashes node.id first (cad_scene.rs:1871), before any geometric parameter. Two identical columns at different coordinates therefore have different ParamHashes, and re-keying on it would share nothing whatsoever.

I took the type name and its docstring ("keyed on the node's mesh parameters") for the whole story and did not read the body. The same assumption appeared independently in a second optimisation review, which is the sort of coincidence that argues for a test rather than a note: param_hash_is_not_a_shape_key_it_includes_the_node_id in cad_scene.rs now pins it.

The id is not a bug in ParamHash — its two users, MeshCache and part_geoms, are keyed by NodeId and only ask "is this cached entry still valid for this node", where the id is a constant. Sharing geometry needs a second hash, ShapeHash: the same body without the id. REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md Phase 2 carries it.

What survives the correction, and what changes:

Claim Status
No viewport culling anywhere Stands. Still the main finding.
3D issues one draw call per part Stands. That is where draw-call multiplication is real.
2D costs ~670 draw calls a frame Wrong. One draw call; the per-item cost is CPU tessellation.
Batching the 2D loops is worthwhile Weaker, still true. It cuts tessellation-call overhead, not draw calls.

The rest of this document is the corrected version.

1. Historical baseline: batching was present, used once, guarded by a test

draw_vector is an immediate-mode path builder. stroke() tessellates the current path into an accumulation buffer; end() turns the whole accumulation into one draw call. So the cost of calling stroke() per item is CPU tessellation and per-call setup, not GPU submission. The codebase demonstrably knows this. From the axis grid:

fn queue_dashed_line(&mut self, x1: f32, y1: f32, x2: f32, y2: f32) {
    while pos < dist {
        self.draw_vector.move_to(...);
        self.draw_vector.line_to(...);      // queue only
        pos = seg_end + gap;
    }
}                                            // no stroke() here

…and a test in viewport.rs enforces it:

"the grid lines should share exactly one stroke"

"draw_dashed_line strokes per dash, which is thousands of tessellations per frame"

That is precisely the minimal-drawcall discipline the datagrid brief asks for. It is applied to draw_axis_grid_2d and nowhere else.

Where it is not applied

draw_2d_vector_scene has 13 stroke() sites, three of them inside loops:

for i in start_x..=end_x {              // base grid, vertical lines
    self.draw_vector.move_to(p1);
    self.draw_vector.line_to(p2);
    self.draw_vector.stroke(if is_major { 1.6 } else { 0.55 });   // per line
}
// …same again for horizontals…

for part in read_parts(&doc).iter() {   // every part, every frame
    self.draw_vector.set_color(col);
    self.draw_vector.rect(...);
    self.draw_vector.stroke(1.8);                                  // per part
}

The grid targets ~18 px spacing, so a 1920×1080 viewport is roughly 107 vertical + 60 horizontal ≈ 167 tessellation calls for the grid alone, plus one per part. At 500 parts that is ~670 calls into tessellate_path_stroke per frame where the file's own proven idiom would give roughly four.

To be clear about the magnitude after the correction above: this is CPU work — tessellator setup, two std::mem::takes and an append_geometry per call — not 670 draw calls. The same total segment count still has to be tessellated either way. The win is per-call overhead, and it is worth having, but it is a smaller prize than culling or 3D instancing.

The base grid is the easier of the two: it only alternates between two stroke widths, so it is two batches, not 167. The parts loop alternates colour per part (selected / hovered / own colour), so it needs grouping by colour or a per-instance colour attribute.

Historical 3-D path: one draw call per part

for part in read_parts(&doc).iter() {
    if let Some(geom) = self.part_geoms.get(&part.id.raw())… {
        self.draw_mesh.transform = part_model_matrix_cadnode(part);
        self.draw_mesh.color = …;
        self.draw_mesh.draw(cx, geom.geometry_id());   // one call, per part
    }
}

Each part carries its own geometry buffer and its own transform/colour uniforms, so N parts is N draw calls. No instancing, no batching, no sorting by state.

2. Historical baseline: virtual viewport present for the grid, absent for content

The grid is virtualised, and well:

let world_left  = self.pan_2d.x - half_w * 1.2;   // visible bounds + 20%
let start_x = ((world_left - gx) / step).floor() as i32;
let end_x   = ((world_right - gx) / step).ceil()  as i32;
for i in start_x..=end_x { … }

Only visible grid lines are emitted, and the step adapts to zoom with a 1/2/5 nice-number progression. That is the datagrid's "only materialise what is on screen", done properly.

In the historical pre-Phase-1 tree the parts were not culled at all. grep -niE "cull|frustum|offscreen|in_view" over the whole 2,461-line renderer returned nothing. Every part was projected and submitted every frame whether or not it was on screen, in both the 2D and 3D paths. Phase 1 now replaces this with the shared 2-D rect and 3-D frustum tests; see the current plan and benchmark rather than treating this paragraph as an assertion about the present tree.

The part that stings

The broad-phase structure needed for culling already exists:

pub fn world_aabb_for(&self, node: &CadNode, build: impl FnOnce(&CadNode) -> WorldAabb) -> WorldAabb

It is cached by PlacedHash, and BENCH_BASELINE.md records it at 4.72× faster than recomputing (94 µs → 20 µs at 500 parts). Its only caller is CadViewport::pick_part — the picking path, which runs on mouse-move and is additionally throttled by HOVER_PICK_MIN_MOVE_PX.

So the cheap visibility test exists, is cached, is benchmarked, and is wired to the path that runs occasionally rather than the path that runs every frame.

3. Which datagrid techniques transfer

Technique Transfers? How it maps to CAD
Virtual viewport Yes, directly Test each part's cached world AABB against the view rect (2D) or frustum (3D) before submitting. Phase 1 wires the shared predicates into the draw loops and frame_budget.
Minimal draw calls Partly — 2D is already one call DrawVector batches to a single draw call at end(). Phase 4 reduces tessellation calls via queued colour groups; Phase 3 is the real GPU draw-call win in 3D.
Instancing Yes, and it is the big one Phase 2 keys local geometry by ShapeHash, and Phase 3 sends repeated transforms/colours as one instanced draw per visible shape. ParamHash is intentionally not the shape key — see correction 0.2 above.
Proper clipping Partly The datagrid needs nested clip rects per cell; CAD has one viewport. Parts are emitted in screen coordinates and left to the GPU to clip, while culling removes work before clipping.
Level of detail Yes, measured and implemented in 2-D and 3-D lod.rs uses a shared three-pixel threshold: ordinary sub-pixel 2-D parts become bounded filled markers and 3-D parts become shared bounding-box proxy instances; selected/hovered parts stay full detail.

What does not transfer

  • Widget recycling. Datagrid cells host child widgets and need a pool. CAD parts are geometry, not widgets; there is nothing to recycle.
  • Index-range virtualisation. A grid virtualises by row/column index, which is O(1) to compute. CAD is continuous space, so the equivalent needs a spatial test. At the scale this app targets a linear pass over cached AABBs is fine — it is a few microseconds for 500 parts. A BVH or grid hash only earns its complexity somewhere north of ~50k parts, and nothing here suggests that is the target.

4. What was not measured in the original note

The original audit had no frame-submission benchmark. Its structural claims were read off the loops, and the consequence — whether the old path dropped frames — was unmeasured. That was corrected by Phase 0: profile_benchmarks.rs::bench_frame_submission_budget now reports the visible geometry, tessellation groups and 3-D shape batches, and the same decision functions drive both the renderer and frame_budget. Phase 5 extends that report with full outlines versus sub-pixel markers and full 3-D meshes versus bounding-box proxy instances.

The structural count is still not a live GPU timing: there is no Cx, window or GPU in the host coverage harness, and the 3-D instanced/proxy submission still needs the documented human visual check. The measured CPU-side cache results remain useful — the mesh cache, scene cache and AABB cache are independently benchmarked — but they must not be sold as frame-time measurements.

5. Suggested order, cheapest first — current status

  1. Cull parts against the view rect. Done in Phase 1, using the shared cached-AABB/frustum predicates.
  2. Key part_geoms by shape and instance repeated parts. Done in Phases 2–3. The key is ShapeHash, not ParamHash — correction 0.2 above.
  3. Add a frame-submission benchmark. Done in Phase 0 and extended through Phase 5; it reports counts without pretending they are GPU timings.
  4. Batch the base grid and part outlines. Done in Phase 4, with marker fills included in the Phase 5 budget.
  5. Use level of detail for sub-pixel parts. Done for both paths in Phase 5: 2-D uses bounded filled markers and 3-D uses one shared bounding-box proxy batch. Live GPU visual verification remains the release check, not an implementation blocker.

Phased details and the measured rows are in REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md and BENCH_BASELINE.md.