# Does the CAD viewport have a minimal-drawcall strategy? Short answer: **the pre-Phase-0 renderer did not, but the CAD module now uses the transferable parts of the strategy.** Phases 1–5 added viewport culling, shape-shared 3-D instancing, 2-D colour batching and measured sub-pixel LOD. The 2-D scene remains one Makepad vector draw call; the useful work there is reducing tessellation and geometry volume. The 3-D submission is now one call per visible shape, with instance rows. This was prompted by a datagrid brief asking for "virtual viewport on both axes" and "an optimal minimal drawcall strategy". Those are two distinct techniques. CAD has no cell recycling or row/column index virtualisation, but continuous-space culling, batching, instancing and LOD transfer directly. The first sections below are the historical pre-Phase-0 audit, from reading `viewport_render.rs` and `viewport.rs`. The current measurements and phase status live in `REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md` and `BENCH_BASELINE.md`; the original "not measured" caveat is retained and marked as historical rather than left to read as the current state. ## 0. Corrections to earlier versions of this note ### 0.1 `stroke()` is not a draw call The first version of this document called `stroke()` a "tessellation/flush point" and then reasoned about the 2D path as if each one cost a draw call — "~670 flush points per frame where four would do". That is wrong, and it was wrong because I inferred Makepad's batching model from a comment instead of reading it. Read now, in `draw/src/shader/draw_vector.rs`: - `begin()` clears CPU-side accumulation buffers (`acc_verts`, `acc_indices`). - `stroke()` calls `tessellate_path_stroke(...)` and `append_geometry(...)`. It tessellates the current path into those buffers. **It issues no draw call.** - `end(cx)` is the only place `cx.new_draw_call(&self.draw_vars)` appears — twice in the whole file, both inside `end()`, for the gradient and plain paths. So the entire 2D vector scene — grid, parts, overlays — is **one draw call**. Makepad's `DrawVector` is already a batching design and the CAD code uses it correctly in that respect. The guard test's own wording, "thousands of tessellations per frame", was accurate and literal; I read "draw call" into it. ### 0.2 `ParamHash` is not a shape key Section 3 and the suggested order both said identical parts could share a GPU buffer by re-keying `part_geoms` on `ParamHash`. They cannot: `ParamHash::from_node` hashes **`node.id` first** (`cad_scene.rs:1871`), before any geometric parameter. Two identical columns at different coordinates therefore have different `ParamHash`es, and re-keying on it would share nothing whatsoever. I took the type name and its docstring ("keyed on the node's mesh parameters") for the whole story and did not read the body. The same assumption appeared independently in a second optimisation review, which is the sort of coincidence that argues for a test rather than a note: `param_hash_is_not_a_shape_key_it_includes_the_node_id` in `cad_scene.rs` now pins it. The id is not a bug in `ParamHash` — its two users, `MeshCache` and `part_geoms`, are keyed by `NodeId` and only ask "is this cached entry still valid for *this* node", where the id is a constant. Sharing geometry needs a *second* hash, `ShapeHash`: the same body without the id. `REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md` Phase 2 carries it. What survives the correction, and what changes: | Claim | Status | |---|---| | No viewport culling anywhere | **Stands.** Still the main finding. | | 3D issues one draw call per part | **Stands.** That is where draw-call multiplication is real. | | 2D costs ~670 draw calls a frame | **Wrong.** One draw call; the per-item cost is CPU tessellation. | | Batching the 2D loops is worthwhile | **Weaker, still true.** It cuts tessellation-call overhead, not draw calls. | The rest of this document is the corrected version. ## 1. Historical baseline: batching was present, used once, guarded by a test `draw_vector` is an immediate-mode path builder. `stroke()` **tessellates** the current path into an accumulation buffer; `end()` turns the whole accumulation into one draw call. So the cost of calling `stroke()` per item is CPU tessellation and per-call setup, not GPU submission. The codebase demonstrably knows this. From the axis grid: ```rust fn queue_dashed_line(&mut self, x1: f32, y1: f32, x2: f32, y2: f32) { while pos < dist { self.draw_vector.move_to(...); self.draw_vector.line_to(...); // queue only pos = seg_end + gap; } } // no stroke() here ``` …and a test in `viewport.rs` enforces it: > `"the grid lines should share exactly one stroke"` > > `"draw_dashed_line strokes per dash, which is thousands of > tessellations per frame"` That is precisely the minimal-drawcall discipline the datagrid brief asks for. It is applied to `draw_axis_grid_2d` and nowhere else. ### Where it is not applied `draw_2d_vector_scene` has 13 `stroke()` sites, three of them **inside loops**: ```rust for i in start_x..=end_x { // base grid, vertical lines self.draw_vector.move_to(p1); self.draw_vector.line_to(p2); self.draw_vector.stroke(if is_major { 1.6 } else { 0.55 }); // per line } // …same again for horizontals… for part in read_parts(&doc).iter() { // every part, every frame self.draw_vector.set_color(col); self.draw_vector.rect(...); self.draw_vector.stroke(1.8); // per part } ``` The grid targets ~18 px spacing, so a 1920×1080 viewport is roughly 107 vertical + 60 horizontal ≈ **167 tessellation calls for the grid alone**, plus **one per part**. At 500 parts that is ~670 calls into `tessellate_path_stroke` per frame where the file's own proven idiom would give roughly four. To be clear about the magnitude after the correction above: this is CPU work — tessellator setup, two `std::mem::take`s and an `append_geometry` per call — not 670 draw calls. The same total segment count still has to be tessellated either way. The win is per-call overhead, and it is worth having, but it is a smaller prize than culling or 3D instancing. The base grid is the easier of the two: it only alternates between two stroke widths, so it is two batches, not 167. The parts loop alternates colour per part (selected / hovered / own colour), so it needs grouping by colour or a per-instance colour attribute. ### Historical 3-D path: one draw call per part ```rust for part in read_parts(&doc).iter() { if let Some(geom) = self.part_geoms.get(&part.id.raw())… { self.draw_mesh.transform = part_model_matrix_cadnode(part); self.draw_mesh.color = …; self.draw_mesh.draw(cx, geom.geometry_id()); // one call, per part } } ``` Each part carries its own geometry buffer and its own transform/colour uniforms, so N parts is N draw calls. No instancing, no batching, no sorting by state. ## 2. Historical baseline: virtual viewport present for the grid, absent for content The grid **is** virtualised, and well: ```rust let world_left = self.pan_2d.x - half_w * 1.2; // visible bounds + 20% let start_x = ((world_left - gx) / step).floor() as i32; let end_x = ((world_right - gx) / step).ceil() as i32; for i in start_x..=end_x { … } ``` Only visible grid lines are emitted, and the step adapts to zoom with a 1/2/5 nice-number progression. That is the datagrid's "only materialise what is on screen", done properly. In the historical pre-Phase-1 tree the parts were not culled at all. `grep -niE "cull|frustum|offscreen|in_view"` over the whole 2,461-line renderer returned **nothing**. Every part was projected and submitted every frame whether or not it was on screen, in both the 2D and 3D paths. Phase 1 now replaces this with the shared 2-D rect and 3-D frustum tests; see the current plan and benchmark rather than treating this paragraph as an assertion about the present tree. ### The part that stings The broad-phase structure needed for culling already exists: ```rust pub fn world_aabb_for(&self, node: &CadNode, build: impl FnOnce(&CadNode) -> WorldAabb) -> WorldAabb ``` It is cached by `PlacedHash`, and `BENCH_BASELINE.md` records it at **4.72× faster than recomputing** (94 µs → 20 µs at 500 parts). Its only caller is `CadViewport::pick_part` — the *picking* path, which runs on mouse-move and is additionally throttled by `HOVER_PICK_MIN_MOVE_PX`. So the cheap visibility test exists, is cached, is benchmarked, and is wired to the path that runs occasionally rather than the path that runs every frame. ## 3. Which datagrid techniques transfer | Technique | Transfers? | How it maps to CAD | |---|---|---| | Virtual viewport | **Yes, directly** | Test each part's cached world AABB against the view rect (2D) or frustum (3D) before submitting. Phase 1 wires the shared predicates into the draw loops and `frame_budget`. | | Minimal draw calls | **Partly — 2D is already one call** | `DrawVector` batches to a single draw call at `end()`. Phase 4 reduces tessellation calls via queued colour groups; Phase 3 is the real GPU draw-call win in 3D. | | Instancing | **Yes, and it is the big one** | Phase 2 keys local geometry by `ShapeHash`, and Phase 3 sends repeated transforms/colours as one instanced draw per visible shape. `ParamHash` is intentionally not the shape key — see correction 0.2 above. | | Proper clipping | Partly | The datagrid needs nested clip rects per cell; CAD has one viewport. Parts are emitted in screen coordinates and left to the GPU to clip, while culling removes work before clipping. | | Level of detail | **Yes, measured and implemented in 2-D and 3-D** | `lod.rs` uses a shared three-pixel threshold: ordinary sub-pixel 2-D parts become bounded filled markers and 3-D parts become shared bounding-box proxy instances; selected/hovered parts stay full detail. | ### What does not transfer - **Widget recycling.** Datagrid cells host child widgets and need a pool. CAD parts are geometry, not widgets; there is nothing to recycle. - **Index-range virtualisation.** A grid virtualises by row/column index, which is O(1) to compute. CAD is continuous space, so the equivalent needs a spatial test. At the scale this app targets a linear pass over cached AABBs is fine — it is a few microseconds for 500 parts. A BVH or grid hash only earns its complexity somewhere north of ~50k parts, and nothing here suggests that is the target. ## 4. What was not measured in the original note The original audit had no frame-submission benchmark. Its structural claims were read off the loops, and the consequence — whether the old path dropped frames — was unmeasured. That was corrected by Phase 0: `profile_benchmarks.rs::bench_frame_submission_budget` now reports the visible geometry, tessellation groups and 3-D shape batches, and the same decision functions drive both the renderer and `frame_budget`. Phase 5 extends that report with full outlines versus sub-pixel markers and full 3-D meshes versus bounding-box proxy instances. The structural count is still not a live GPU timing: there is no `Cx`, window or GPU in the host coverage harness, and the 3-D instanced/proxy submission still needs the documented human visual check. The measured CPU-side cache results remain useful — the mesh cache, scene cache and AABB cache are independently benchmarked — but they must not be sold as frame-time measurements. ## 5. Suggested order, cheapest first — current status 1. **Cull parts against the view rect.** Done in Phase 1, using the shared cached-AABB/frustum predicates. 2. **Key `part_geoms` by shape and instance repeated parts.** Done in Phases 2–3. The key is `ShapeHash`, **not** `ParamHash` — correction 0.2 above. 3. **Add a frame-submission benchmark.** Done in Phase 0 and extended through Phase 5; it reports counts without pretending they are GPU timings. 4. **Batch the base grid and part outlines.** Done in Phase 4, with marker fills included in the Phase 5 budget. 5. **Use level of detail for sub-pixel parts.** Done for both paths in Phase 5: 2-D uses bounded filled markers and 3-D uses one shared bounding-box proxy batch. Live GPU visual verification remains the release check, not an implementation blocker. Phased details and the measured rows are in `REVIEWS/CAD_RENDER_OPTIMISATION_PLAN.md` and `BENCH_BASELINE.md`.