The parallel batch machinery frombe21d627aonly ever existed to chase a default-on multi-threaded win that never came (washer w8 still +54% with the hybrid on). At w1 — the only configuration this opt-in flag is for — parallel_for runs inline, so the serial path gives the identical result for ~340 fewer lines. Wire up the previously-dead dynamic_tree_self_pairs/cross_pairs into a serial collect_batch_candidates (three BVTT self/cross traversals -> canonical (a,b,child) sort -> serial filter into move_results[0]) and delete BatchWork, BatchCtx, batch_drain_*, BatchFilterCtx, batch_filter_*, bvtt_step, dynamic_tree_bvtt_drain/expand, and the batch_frontier/worker_* scratch fields. Determinism preserved exactly: OFF 0x61E35C31/step314 bit-identical, ON 0xBE99C5F7/step313 identical across workers 1/2/4. The debug SET-equality oracle and the determinism_broad_phase_hybrid_across_worker_counts test are unchanged and still pass; zero warnings. PGO: pgo.sh never trained -bp=1, so an off-path-only profile laid the hybrid branch out cold and collapsed the win to ~-5%. Add one -b=8 -bp=1 training run (neutral for the default path — counts merge, the OFF branch stays hot) and retrain. Corrected README numbers to measured values (hybrid-trained profile, paired -bp toggle, washer w1): pair-finding stage 7.6k -> 3.6k ms/1000 (-52%); total ~-17% (~19.0k vs ~22.6k), which beats C (20661) and narrows Rapier's lead from ~23% to ~12% — not the "17.7k / within 5% / -19%"be21d627aclaimed. Still default-off (w8 regresses ~+54%), opt-in single-threaded accelerator for churn-heavy scenes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
335 KiB
335 KiB