The first Phase 5 tranche correctly reduced sub-pixel 2D outlines, but it stopped short of the plan's full level-of-detail phase. Complete the phase by applying the same measured three-pixel policy to the 3D path. Visible 3D parts below the projected AABB threshold now use one shared unit-cube geometry with an axis-aligned world-bounds transform. Full-detail shape batches, proxy instance rows and proxy draw calls are reported separately. Wireframe and hidden-line overlays use the same policy, and selected or hovered parts remain full detail. Extend the host-testable projection and proxy policy tests, frame benchmark, source guards and documentation. Keep the real GPU/window visual check explicit: compilation, policy arithmetic, submission structure and coverage pass here, but this environment cannot execute Makepad's draw submission.
12 KiB
Does the CAD viewport have a minimal-drawcall strategy?
Short answer: the pre-Phase-0 renderer did not, but the CAD module now uses the transferable parts of the strategy. Phases 1–5 added viewport culling, shape-shared 3-D instancing, 2-D colour batching and measured sub-pixel LOD. The 2-D scene remains one Makepad vector draw call; the useful work there is reducing tessellation and geometry volume. The 3-D submission is now one call per visible shape, with instance rows.
This was prompted by a datagrid brief asking for "virtual viewport on both axes" and "an optimal minimal drawcall strategy". Those are two distinct techniques. CAD has no cell recycling or row/column index virtualisation, but continuous-space culling, batching, instancing and LOD transfer directly.
The first sections below are the historical pre-Phase-0 audit, from
reading viewport_render.rs and viewport.rs. The current measurements
and phase status live in REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md and
BENCH_BASELINE.md; the original "not measured" caveat is retained and
marked as historical rather than left to read as the current state.
0. Corrections to earlier versions of this note
0.1 stroke() is not a draw call
The first version of this document called stroke() a "tessellation/flush
point" and then reasoned about the 2D path as if each one cost a draw
call — "~670 flush points per frame where four would do". That is wrong,
and it was wrong because I inferred Makepad's batching model from a
comment instead of reading it.
Read now, in draw/src/shader/draw_vector.rs:
begin()clears CPU-side accumulation buffers (acc_verts,acc_indices).stroke()callstessellate_path_stroke(...)andappend_geometry(...). It tessellates the current path into those buffers. It issues no draw call.end(cx)is the only placecx.new_draw_call(&self.draw_vars)appears — twice in the whole file, both insideend(), for the gradient and plain paths.
So the entire 2D vector scene — grid, parts, overlays — is one draw
call. Makepad's DrawVector is already a batching design and the CAD
code uses it correctly in that respect. The guard test's own wording,
"thousands of tessellations per frame", was accurate and literal; I read
"draw call" into it.
0.2 ParamHash is not a shape key
Section 3 and the suggested order both said identical parts could share
a GPU buffer by re-keying part_geoms on ParamHash. They cannot:
ParamHash::from_node hashes node.id first (cad_scene.rs:1871),
before any geometric parameter. Two identical columns at different
coordinates therefore have different ParamHashes, and re-keying on it
would share nothing whatsoever.
I took the type name and its docstring ("keyed on the node's mesh
parameters") for the whole story and did not read the body. The same
assumption appeared independently in a second optimisation review, which
is the sort of coincidence that argues for a test rather than a note:
param_hash_is_not_a_shape_key_it_includes_the_node_id in
cad_scene.rs now pins it.
The id is not a bug in ParamHash — its two users, MeshCache and
part_geoms, are keyed by NodeId and only ask "is this cached entry
still valid for this node", where the id is a constant. Sharing
geometry needs a second hash, ShapeHash: the same body without the
id. REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md Phase 2 carries it.
What survives the correction, and what changes:
| Claim | Status |
|---|---|
| No viewport culling anywhere | Stands. Still the main finding. |
| 3D issues one draw call per part | Stands. That is where draw-call multiplication is real. |
| 2D costs ~670 draw calls a frame | Wrong. One draw call; the per-item cost is CPU tessellation. |
| Batching the 2D loops is worthwhile | Weaker, still true. It cuts tessellation-call overhead, not draw calls. |
The rest of this document is the corrected version.
1. Historical baseline: batching was present, used once, guarded by a test
draw_vector is an immediate-mode path builder. stroke() tessellates
the current path into an accumulation buffer; end() turns the whole
accumulation into one draw call. So the cost of calling stroke() per
item is CPU tessellation and per-call setup, not GPU submission. The
codebase demonstrably knows this. From the axis grid:
fn queue_dashed_line(&mut self, x1: f32, y1: f32, x2: f32, y2: f32) {
while pos < dist {
self.draw_vector.move_to(...);
self.draw_vector.line_to(...); // queue only
pos = seg_end + gap;
}
} // no stroke() here
…and a test in viewport.rs enforces it:
"the grid lines should share exactly one stroke"
"draw_dashed_line strokes per dash, which is thousands of tessellations per frame"
That is precisely the minimal-drawcall discipline the datagrid brief
asks for. It is applied to draw_axis_grid_2d and nowhere else.
Where it is not applied
draw_2d_vector_scene has 13 stroke() sites, three of them inside
loops:
for i in start_x..=end_x { // base grid, vertical lines
self.draw_vector.move_to(p1);
self.draw_vector.line_to(p2);
self.draw_vector.stroke(if is_major { 1.6 } else { 0.55 }); // per line
}
// …same again for horizontals…
for part in read_parts(&doc).iter() { // every part, every frame
self.draw_vector.set_color(col);
self.draw_vector.rect(...);
self.draw_vector.stroke(1.8); // per part
}
The grid targets ~18 px spacing, so a 1920×1080 viewport is roughly
107 vertical + 60 horizontal ≈ 167 tessellation calls for the grid
alone, plus one per part. At 500 parts that is ~670 calls into
tessellate_path_stroke per frame where the file's own proven idiom
would give roughly four.
To be clear about the magnitude after the correction above: this is CPU
work — tessellator setup, two std::mem::takes and an append_geometry
per call — not 670 draw calls. The same total segment count still has to
be tessellated either way. The win is per-call overhead, and it is worth
having, but it is a smaller prize than culling or 3D instancing.
The base grid is the easier of the two: it only alternates between two stroke widths, so it is two batches, not 167. The parts loop alternates colour per part (selected / hovered / own colour), so it needs grouping by colour or a per-instance colour attribute.
Historical 3-D path: one draw call per part
for part in read_parts(&doc).iter() {
if let Some(geom) = self.part_geoms.get(&part.id.raw())… {
self.draw_mesh.transform = part_model_matrix_cadnode(part);
self.draw_mesh.color = …;
self.draw_mesh.draw(cx, geom.geometry_id()); // one call, per part
}
}
Each part carries its own geometry buffer and its own transform/colour uniforms, so N parts is N draw calls. No instancing, no batching, no sorting by state.
2. Historical baseline: virtual viewport present for the grid, absent for content
The grid is virtualised, and well:
let world_left = self.pan_2d.x - half_w * 1.2; // visible bounds + 20%
let start_x = ((world_left - gx) / step).floor() as i32;
let end_x = ((world_right - gx) / step).ceil() as i32;
for i in start_x..=end_x { … }
Only visible grid lines are emitted, and the step adapts to zoom with a 1/2/5 nice-number progression. That is the datagrid's "only materialise what is on screen", done properly.
In the historical pre-Phase-1 tree the parts were not culled at all.
grep -niE "cull|frustum|offscreen|in_view" over the whole 2,461-line
renderer returned nothing. Every part was projected and submitted
every frame whether or not it was on screen, in both the 2D and 3D paths.
Phase 1 now replaces this with the shared 2-D rect and 3-D frustum tests;
see the current plan and benchmark rather than treating this paragraph as
an assertion about the present tree.
The part that stings
The broad-phase structure needed for culling already exists:
pub fn world_aabb_for(&self, node: &CadNode, build: impl FnOnce(&CadNode) -> WorldAabb) -> WorldAabb
It is cached by PlacedHash, and BENCH_BASELINE.md records it at
4.72× faster than recomputing (94 µs → 20 µs at 500 parts). Its only
caller is CadViewport::pick_part — the picking path, which runs on
mouse-move and is additionally throttled by HOVER_PICK_MIN_MOVE_PX.
So the cheap visibility test exists, is cached, is benchmarked, and is wired to the path that runs occasionally rather than the path that runs every frame.
3. Which datagrid techniques transfer
| Technique | Transfers? | How it maps to CAD |
|---|---|---|
| Virtual viewport | Yes, directly | Test each part's cached world AABB against the view rect (2D) or frustum (3D) before submitting. Phase 1 wires the shared predicates into the draw loops and frame_budget. |
| Minimal draw calls | Partly — 2D is already one call | DrawVector batches to a single draw call at end(). Phase 4 reduces tessellation calls via queued colour groups; Phase 3 is the real GPU draw-call win in 3D. |
| Instancing | Yes, and it is the big one | Phase 2 keys local geometry by ShapeHash, and Phase 3 sends repeated transforms/colours as one instanced draw per visible shape. ParamHash is intentionally not the shape key — see correction 0.2 above. |
| Proper clipping | Partly | The datagrid needs nested clip rects per cell; CAD has one viewport. Parts are emitted in screen coordinates and left to the GPU to clip, while culling removes work before clipping. |
| Level of detail | Yes, measured and implemented in 2-D and 3-D | lod.rs uses a shared three-pixel threshold: ordinary sub-pixel 2-D parts become bounded filled markers and 3-D parts become shared bounding-box proxy instances; selected/hovered parts stay full detail. |
What does not transfer
- Widget recycling. Datagrid cells host child widgets and need a pool. CAD parts are geometry, not widgets; there is nothing to recycle.
- Index-range virtualisation. A grid virtualises by row/column index, which is O(1) to compute. CAD is continuous space, so the equivalent needs a spatial test. At the scale this app targets a linear pass over cached AABBs is fine — it is a few microseconds for 500 parts. A BVH or grid hash only earns its complexity somewhere north of ~50k parts, and nothing here suggests that is the target.
4. What was not measured in the original note
The original audit had no frame-submission benchmark. Its structural
claims were read off the loops, and the consequence — whether the old
path dropped frames — was unmeasured. That was corrected by Phase 0:
profile_benchmarks.rs::bench_frame_submission_budget now reports the
visible geometry, tessellation groups and 3-D shape batches, and the
same decision functions drive both the renderer and frame_budget.
Phase 5 extends that report with full outlines versus sub-pixel markers
and full 3-D meshes versus bounding-box proxy instances.
The structural count is still not a live GPU timing: there is no Cx,
window or GPU in the host coverage harness, and the 3-D instanced/proxy
submission still needs the documented human visual check. The measured
CPU-side cache results remain useful — the mesh cache, scene cache and
AABB cache are independently benchmarked — but they must not be sold as
frame-time measurements.
5. Suggested order, cheapest first — current status
- Cull parts against the view rect. Done in Phase 1, using the shared cached-AABB/frustum predicates.
- Key
part_geomsby shape and instance repeated parts. Done in Phases 2–3. The key isShapeHash, notParamHash— correction 0.2 above. - Add a frame-submission benchmark. Done in Phase 0 and extended through Phase 5; it reports counts without pretending they are GPU timings.
- Batch the base grid and part outlines. Done in Phase 4, with marker fills included in the Phase 5 budget.
- Use level of detail for sub-pixel parts. Done for both paths in Phase 5: 2-D uses bounded filled markers and 3-D uses one shared bounding-box proxy batch. Live GPU visual verification remains the release check, not an implementation blocker.
Phased details and the measured rows are in
REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md and BENCH_BASELINE.md.