makepad/skills/makepad-game-director/references/workflow-evaluations.md
Arena Agent fe21c07d84 skills: add Makepad game skills pack ported from threejs-game-skills
Nine agent skills for building Makepad/Rust games, ported from
majidmanzarpour/threejs-game-skills. Same director-routed workflow and
premium bar; runtime rewritten for this fork's game crates.

- makepad-game-director        entrypoint, routing, continuity, asset probe
- makepad-gameplay-systems     loop, movers vs rigid bodies, input, camera, netplay
- makepad-aaa-graphics-builder lighting, shaders, budget, visual scorecard
- makepad-game-ui-designer     HUD, menus, touch and XR UI
- makepad-debug-profiler       defect bisection and profiling
- makepad-qa-release           verification ladder, evidence, packaging
- makepad-3d-generator         CC0 model search, casts, procedural geometry
- makepad-image-generator      texgen textures, palettes, sky, icons
- makepad-audio-generator      sample bank, mixer, material impacts, 3D audio

Written against the real APIs in libs/game/*, libs/sim and apps/arcade:
the game.* verb table, GameRenderer adaptive quality, the packed 6-float
GameMeshVertex layout, script_mod! splash styling, makepad-test driving,
and the BUDGETS.md numbers. No paid generation API is required - the CC0
library plus seeded makepad-game-gen replaces them.

Includes install.sh (Codex/Claude), validate-skills.sh and a repo-aware
probe_assets.sh; both scripts verified against this checkout.
2026-09-05 21:18:58 +00:00

4.2 KiB

Workflow and evaluations

Phase order

  1. Brief — core loop in one sentence, win/lose condition, input map, camera scale, art direction, target device. Write it into artifacts/game-progress.md before implementing.
  2. Route decision — script (A) or Rust crate (B). See architecture.md. Say which and why.
  3. Playable slice — one level, real input, win and lose reachable. Primitives are acceptable here and only here.
  4. Assets — probe the library, resolve model ids and one cast, confirm scale against the player capsule.
  5. Representative scene — the slice with real assets, real lighting, real camera. Assess this before multiplying levels or enemy variants. Multiplying a scene that is wrong multiplies the wrongness.
  6. Depth — graphics, UI, audio, feel, in that order of visible impact.
  7. Verification — the ladder in makepad-qa-release, scaled to the change.
  8. Report — outcome first, evidence second, remaining risks third.

Phases 4 and 3 overlap deliberately: start asset resolution while the loop is being written.

Scaling to request size

Request Phases Verification
HUD tweak, one constant, copy change 7 cargo check + the one affected test
New mechanic in an existing game 1, 3, 6, 7 crate tests + a makepad-test run of the loop
New level / enemy family 1, 4, 5, 6, 7 above + screenshot evidence
Full game, or "premium"/"AAA"/"less basic" all full ladder, release build, scorecard

A narrow edit to a premium game stays a narrow edit. Do not run the whole pipeline because the codebase is large.

Delegation

Lead plus up to two workers. Rust-specific rules:

  • Split by crate or module, never by lines in one file. Two workers editing arcade_view.rs will produce a merge that does not compile.
  • Every worker runs cargo check -p <their crate> before handing back. A worker that hands back non-compiling code has produced negative value.
  • The lead owns: the root Cargo.toml members list, cross-crate trait/struct signatures, and the final integration build.
  • Good split: worker A on libs/game/gen level generation, worker B on HUD in the app crate, lead on the sim/render loop that joins them.
  • Bad split: two workers on "the graphics".

When delegation tools are absent, work directly.

Review gate

For substantial gameplay, graphics, or animation changes, one focused independent review catches missed defects. Supply raw captures, the diff, and the relevant rubric; ask for concrete defects, not endorsement of the lead's score. Resolve findings without recursive review cycles.

Reviewer prompt shape:

Here is a screenshot set, the frame-time reading, and the diff. Using visual-scorecard.md, list every category scoring below 2 and the specific artefact that caused it. Do not agree with my score; find what I missed.

Self-evaluation before reporting

Answer these honestly. Any "no" is unfinished work, not a caveat to mention in passing.

  • Does cargo check --workspace pass?
  • Do the affected crates' tests pass?
  • Has the game been run, not just compiled?
  • Is the win condition reachable and the lose condition reachable?
  • Does every asset referenced by id actually resolve, or does the scene silently fall back to primitives?
  • Is frame time inside apps/arcade/BUDGETS.md on the target device class?
  • For a premium request: does the scorecard average ≥ 2.3 with nothing below 2?
  • Did anything regress that the change had no business touching?

Failure patterns specific to this stack

  • Compiling is not running. Makepad resolves the splash DSL at runtime; a style or widget error appears on launch, not at cargo check.
  • Silent asset fallback. A wrong model id degrades to a primitive rather than erroring. Assert on model_is_loaded(id) rather than trusting the screenshot to look intentional.
  • Testing through the widget. If a rule can be tested without a Cx, it must be — put it in a Cx-free crate and unit-test it there.
  • Variable-dt simulation. Breaks networking and replay. Fixed TICK_DT always.
  • Second frame pacer. GameRenderer already does adaptive quality; a competing one oscillates.
  • Per-entity script on_tick. Exhausts the shared 500k-instruction pool. Batch.