makepad/libs/ai/models/trellis/README.md
Admin 60615db4ed libs: ai hub/models/cuda, speech, chat_ui
Squash of 54 work commits (Sep 1–12):
  6251f7c  ai-hub: body domain — live pose packets ride the realtime session
  ea50c77  chat_ui: the feed's session gets its profile brief back
  f51b5f3  ai-body: the crate for the native SAM 3D Body port, with its weights reader
  8211ae6  ai-body: the MHR rig and the pose head's parameter decoding, oracle-exact
  9e343a8  ai-body: the DINOv3 ViT-H+/16 backbone, crop and ray conditioning; Metal gains rope-half and affine layer norm
  69d842c  ai-body: the promptable pose decoder and its refinement loop, oracle-matched on Metal
  66e5e2f  ai-hub: SAM 3D Body runs natively — `sam3dbody` on the body domain, oracle-matched end to end
  a634198  ai-hub: the body-native commit carried a peer's in-flight hub hunks; put them back where they were
  9ff44e8  ai-hub: the body-native wiring, this time only the lane's hunks
  6a1c16b  ai-body: third-party notices — what the port is implemented after, and what it is not
  d78411a  ai-body: the per-step work moves to the GPU
  b22259b  ai-body: the context stays on the GPU; only the pose token leaves the loop
  346f31f  ai-body: flash attention for the head-dim-64 blocks
  45b5b98  ai-body: the crop size is a runtime knob, and the loop reports where its time goes
  4be6d19  ai-body: the test modules import the grid constants they still use
  7598346  ai-body: tensor-core GEMMs for the backbone, and the rig's correctives only where they count
  a9ce596  ai-body: the crop warp runs across cores
  8964ba6  ai-body: an FP8 backbone mode, off by default, measured against the oracle
  a2aaa8f  ai-body: the FP8 bias rides a column-broadcast add on the device
  d53c77d  metal: a device-resident ViT stack, and the body backbone rides it
  d006d0a  metal: resident f32 linears keep their weight on the device
  525ba1c  metal: a device-resident two-way decoder layer, and the body decoder rides it
  c9e6d88  ai-body: the hands pass — hand crops, the hand decoder, the hand-mode rig and the wrist fusion
  62dff26  ai-body: the mask prompt — a person's segmentation mask conditions the body pass
  a648cf8  ai-hub: body session options — hands, detect, persons=N
  8c568df  ai-hub: drop the SAM 3D Body reference worker backend
  7ff875a  ai-hub: keep a peer's in-flight beats/notes/local work out of the body commits
  31e5faa  ai-hub: local model runner, licence acknowledgements, a shared install panel; Beat This!, Basic Pitch and the Salamander drum-kit entries
  b94bc58  ai-services: the wire, the app port and the panel state — one conversation, many apps
  2acb798  ai-services: wire v2 — endpoints, receiver-side caps, result disposition
  8ae0ffb  ai-services: the engine core — registry, router and conversation, tested against a scripted model
  2308736  ai-services: the real models behind the engine feature — local through the hub, Claude, and none
  c3f631d  livepipe: one reusable pipe from a camera to a fleet node and back
  ff62db3  ai libs: the runtime env-var cleanup — precision is a per-caller policy, not an environment side channel
  04a94ef  realtime: one service-log line when a live session opens and one when it closes
  0ecb81c  ai models: the model-crates env-var cleanup — 172 research knobs gone, the unset default is the code
  4ca36c1  ai hub + services: the assistant's model comes from wherever it is resident — the fleet chat box, with tools, then the local weights
  432121e  aichat engine + wm: launch, then use — the assistant continues in the same turn once the app it started is on the bus
  7a5bf69  ai-hub registry: the Salamander drumkit samples come from the makepad.nl mirror — the GitHub repo only carries the .sfz files
  102ffc5  ai-services: messages on the bus — a manifest declares topics, the engine subscribes on a tool's behalf or by ToolResult.subscribe, a service publishes Message frames, an idle conversation wakes on a message as an event turn under rate laws; the WM bus forwards the new frames; every app that matches the wire gets its arm
  a837792  hub + flow: a whitespace-only chat completion is retried once and then fails instead of passing as an answer; a flow's model is a fleet model id unless it names a weight file on disk; chat models show under the text domain in /v1/models
  bc6c620  hub + flow: what the chat review found — the in-process route retries an empty completion too, a node says whether its prefill opened thinking so a brief-mode answer is never discarded, a preferred model falls back to normal election when no node has it, discovery keeps looking for the preferred model until patience runs out
  75c3441  hub: the PRO 6000 serves image as well as chat and text
  ad5e98b  hub registry: flux2-dev's VRAM estimate is its measured peak, 30 GB
  c7241e0  hub: a node that evicted every resident releases its cached allocator pool before refusing a load or publishing usable VRAM
  30575f0  flow: route generation by request workload
  1be1e21  ai-hub: gate downloads by disk capacity and recover fleet admission
  df6b394  filesystem_watcher, bounded_http, ai services: live and tool prerequisites
  79ebdb9  ai-hub: add a native Pixal3D image-to-3D backend
  0ba0d74  ai-hub: propagate typed refusals under reject queue policy
  cc6c872  Speed up H3 conditioning and video decoding
  e512059  Fix Qwen vision residency and generated material colors
  2864f68  ai-hub http client: bound every plain TCP connect to 3 s per address
  3d93229  ai: CUDA is a Linux/Windows-only dependency; the hub library defaults to llm + stt

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 13:40:31 +02:00

164 lines
7 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Native TRELLIS.2 and Pixal3D
The `pixal3d` mesh backend runs image conditioning, sparse structure, low/high
resolution shape diffusion, sparse shape/texture decoders and PBR GLB export
through Rust and CUDA. It loads the combined Comfy-Org BF16 flow checkpoint by
component prefix, without extracting four copies. Checkpoints are downloaded
and SHA-256 verified through the hub registry. Production inference does not
start ComfyUI or Python.
The [Image to Pixal3D Flow](../../../flow/recipes/templates/image-to-pixal3d.splash)
connects matting, native mesh generation and the mesh output. Its JSON settings
input controls the reconstruction. The model's tensor stages remain inside the
hub backend so intermediate GPU tensors stay on the GPU worker.
## Request
Submit a mesh request to the hub with PNG input bytes in `input_b64`:
```json
{
"model": "pixal3d",
"input_b64": "<PNG base64>",
"input_content_type": "image/png",
"seed": 42,
"texture": true,
"remesh_resolution": 256,
"decimation_target": 80000,
"texture_size": 1024,
"pixal": {
"resolution": 1024,
"camera_fov": 49.13,
"structure_seed": 56,
"texture_seed": 43,
"shape_steps": 20
}
}
```
`resolution` accepts 1024 or 1536. `camera_fov` is the horizontal field of view
in degrees (1–170); the default is 49.13. The main seed drives shape sampling.
Structure defaults to the main seed and texture to main seed + 1. Shape steps
control the high resolution stage (default 20, range 1–100). Native BiRefNet
segments opaque input; pre-segmented RGBA input skips that pass.
The defaults favor fast asset export: a 256 remesh grid, about 80,000 faces and
a 1024 texture atlas. Shape resolution and output mesh density are separate
controls. `remesh_resolution: 0` retains the existing raw mesh export path.
## Performance implementation
- NAF evaluates attention only at the four bilinear sample pixels consumed by
each active voxel. It avoids the dense 1024 × 1024 × 1024 feature image
(4 GiB at F32), dense unfolded neighborhoods and third-party NATTEN kernels.
- Image cross-attention K/V and up to 512 MiB of fixed projection outputs are
cached across diffusion steps. Negative CFG preserves the projection bias.
- Guide transpose and 2D positional rotation are fused on the GPU. At 1024,
this avoids constructing and uploading two 128 MiB CPU phase tables.
- Device weights reuse the existing native namespace cache; conditioning guide
tensors are released before the corresponding diffusion stage.
This is currently a CUDA BF16 backend. The registry advertises a 24 GiB budget;
1536, larger images/meshes and device workspace behavior can require more.
Actual validation used an RTX PRO 6000 Blackwell with 96 GiB, not a 24 GiB card.
## Relationship to the ComfyUI workflow
The port follows Tencent's Pixal3D projection architecture and the Comfy-Org
combined checkpoint layout. It supports the 1024/1536 cascades and independent
sampling seeds. It is not a bit-for-bit reproduction of a ComfyUI seed:
Makepad uses its native Gaussian RNG and GPU math.
The downloaded workflow also uses features not implemented here: INT8 convrot
and 6 GiB offloading, MoGe automatic FOV, PEC UV unwrapping, and explicit normal
and AO texture baking. This backend uses manual FOV and the existing native
FaithC remesh, decimation, xatlas and color/metallic/roughness baking. Its
current shared decoders use the original Microsoft F16 checkpoints rather
than Comfy-Org's BF16 VAE repack. The downloaded 768 remesh / 700,000-face /
2048-atlas settings are not the fast defaults above (the native remesh grid is
currently capped at 512).
See [third-party notices](THIRD_PARTY_NOTICES.md) for pinned source identities
and licenses. Model and component licenses remain separate; the bundle is not
represented as uniformly MIT. ComfyUI GPL code is not vendored into this crate.
## Reproduce and check
From `libs/ai/hub`, on a CUDA build machine:
```sh
cargo build --release --no-default-features --features mesh --examples
```
Run the standalone release executable `pixal_generate` with:
```text
pixal_generate WEIGHTS INPUT.png OUTPUT.glb 1024 3
```
The optional final argument repeats generation with one loaded backend; each
trial writes `OUTPUT.glb.N.glb`. Trial 0 includes preparation/first use of model
weights. Later trials retain the device weight cache. The timer includes
matting, neural generation, mesh processing and GLB encoding; registry file
verification occurs before the timer. Weight filenames for this harness:
```text
pixal3d_bf16.safetensors
dino_naf.safetensors
ss-decoder.safetensors
shape-decoder.safetensors
texture-decoder.safetensors
native-matte.safetensors
```
All revisions, sizes and hashes come from `libs/ai/hub/registry.json`; the
harness only remaps local cache filenames. Standard hub operation uses the
ordinary registry cache paths.
Numerical checks:
```sh
# From libs/ai:
cargo test --release -p makepad-ai-trellis --lib
# Run native checkpoint-driven encoder fixture, then the independent oracle:
pixal_naf_check dino_naf.safetensors guide.f32
python pixal_guide_oracle.py dino_naf.safetensors guide.f32
```
The Python tests live in `tests/` and need PyTorch/safetensors. The complete
encoder fixture checks both cached and uncached native paths against ordinary
PyTorch convolution, group normalization, pooling, RoPE, neighborhood
attention and grid sampling. F16 tensor-core operands account for its looser
tolerance than the separate F32 sampling kernel test.
`tests/pixal_naf_oracle.py` accepts a shared library built from
`../../cuda/kernels/pixal.cu` with `nvcc -shared -O3` and the target GPU's
architecture. On Windows export `makepad_cuda_pixal_naf_sample_f32` from the
DLL. `--benchmark` measures sparse sampling only, excluding guide encoding and
the rest of the model.
## Measured run (2026-09-06)
RTX PRO 6000 Blackwell 96 GiB, CUDA 13.2, release build, Comfy's
`viking_wolf_rune_axe.png`, the 1024 request above:
| Trial | Through texture decode | Complete PBR GLB |
| --- | ---: | ---: |
| First use in a fresh process | 27.38 s | 47.76 s |
| Cached model weights, second call | 10.08 s | 32.40 s |
| Cached model weights, third call | 10.13 s | 30.49 s |
The final inspected GLB has 79,158 faces, valid UVs and two embedded 1024 PNG
textures. The 1536 cascade also completed in 68.33 s with the same fast export
settings, before the fused GPU phase change. Mesh cleanup and xatlas account
for most of the warm end-to-end time and vary between runs. These are native
measurements, not a speedup claim against an end-to-end ComfyUI baseline.
The separate F32 sparse NAF sampling oracle differed from dense PyTorch by at
most 1.1e-6, with roughly 7.3 ms sampling for 30,000 voxels. That kernel timing
excludes guide encoding. The complete checkpoint-driven encoder/NAF fixture
had maximum absolute error 0.000561 against F32 PyTorch; its native encoded
feature reuse and recomputation paths agreed exactly. A larger projection
cache and a persistent 1 GiB image encoding cache did not show a reliable
end-to-end benefit and were not retained in the production path.