Squash of 54 work commits (Sep 1–12):6251f7cai-hub: body domain — live pose packets ride the realtime sessionea50c77chat_ui: the feed's session gets its profile brief backf51b5f3ai-body: the crate for the native SAM 3D Body port, with its weights reader8211ae6ai-body: the MHR rig and the pose head's parameter decoding, oracle-exact9e343a8ai-body: the DINOv3 ViT-H+/16 backbone, crop and ray conditioning; Metal gains rope-half and affine layer norm69d842cai-body: the promptable pose decoder and its refinement loop, oracle-matched on Metal66e5e2fai-hub: SAM 3D Body runs natively — `sam3dbody` on the body domain, oracle-matched end to enda634198ai-hub: the body-native commit carried a peer's in-flight hub hunks; put them back where they were9ff44e8ai-hub: the body-native wiring, this time only the lane's hunks6a1c16bai-body: third-party notices — what the port is implemented after, and what it is notd78411aai-body: the per-step work moves to the GPUb22259bai-body: the context stays on the GPU; only the pose token leaves the loop346f31fai-body: flash attention for the head-dim-64 blocks45b5b98ai-body: the crop size is a runtime knob, and the loop reports where its time goes4be6d19ai-body: the test modules import the grid constants they still use7598346ai-body: tensor-core GEMMs for the backbone, and the rig's correctives only where they counta9ce596ai-body: the crop warp runs across cores8964ba6ai-body: an FP8 backbone mode, off by default, measured against the oraclea2aaa8fai-body: the FP8 bias rides a column-broadcast add on the deviced53c77dmetal: a device-resident ViT stack, and the body backbone rides itd006d0ametal: resident f32 linears keep their weight on the device525ba1cmetal: a device-resident two-way decoder layer, and the body decoder rides itc9e6d88ai-body: the hands pass — hand crops, the hand decoder, the hand-mode rig and the wrist fusion62dff26ai-body: the mask prompt — a person's segmentation mask conditions the body passa648cf8ai-hub: body session options — hands, detect, persons=N8c568dfai-hub: drop the SAM 3D Body reference worker backend7ff875aai-hub: keep a peer's in-flight beats/notes/local work out of the body commits31e5faaai-hub: local model runner, licence acknowledgements, a shared install panel; Beat This!, Basic Pitch and the Salamander drum-kit entriesb94bc58ai-services: the wire, the app port and the panel state — one conversation, many apps2acb798ai-services: wire v2 — endpoints, receiver-side caps, result disposition8ae0ffbai-services: the engine core — registry, router and conversation, tested against a scripted model2308736ai-services: the real models behind the engine feature — local through the hub, Claude, and nonec3f631dlivepipe: one reusable pipe from a camera to a fleet node and backff62db3ai libs: the runtime env-var cleanup — precision is a per-caller policy, not an environment side channel04a94efrealtime: one service-log line when a live session opens and one when it closes0ecb81cai models: the model-crates env-var cleanup — 172 research knobs gone, the unset default is the code4ca36c1ai hub + services: the assistant's model comes from wherever it is resident — the fleet chat box, with tools, then the local weights432121eaichat engine + wm: launch, then use — the assistant continues in the same turn once the app it started is on the bus7a5bf69ai-hub registry: the Salamander drumkit samples come from the makepad.nl mirror — the GitHub repo only carries the .sfz files102ffc5ai-services: messages on the bus — a manifest declares topics, the engine subscribes on a tool's behalf or by ToolResult.subscribe, a service publishes Message frames, an idle conversation wakes on a message as an event turn under rate laws; the WM bus forwards the new frames; every app that matches the wire gets its arma837792hub + flow: a whitespace-only chat completion is retried once and then fails instead of passing as an answer; a flow's model is a fleet model id unless it names a weight file on disk; chat models show under the text domain in /v1/modelsbc6c620hub + flow: what the chat review found — the in-process route retries an empty completion too, a node says whether its prefill opened thinking so a brief-mode answer is never discarded, a preferred model falls back to normal election when no node has it, discovery keeps looking for the preferred model until patience runs out75c3441hub: the PRO 6000 serves image as well as chat and textad5e98bhub registry: flux2-dev's VRAM estimate is its measured peak, 30 GBc7241e0hub: a node that evicted every resident releases its cached allocator pool before refusing a load or publishing usable VRAM30575f0flow: route generation by request workload1be1e21ai-hub: gate downloads by disk capacity and recover fleet admission df6b394 filesystem_watcher, bounded_http, ai services: live and tool prerequisites 79ebdb9 ai-hub: add a native Pixal3D image-to-3D backend 0ba0d74 ai-hub: propagate typed refusals under reject queue policy cc6c872 Speed up H3 conditioning and video decoding e512059 Fix Qwen vision residency and generated material colors 2864f68 ai-hub http client: bound every plain TCP connect to 3 s per address 3d93229 ai: CUDA is a Linux/Windows-only dependency; the hub library defaults to llm + stt Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
164 lines
7 KiB
Markdown
164 lines
7 KiB
Markdown
# Native TRELLIS.2 and Pixal3D
|
||
|
||
The `pixal3d` mesh backend runs image conditioning, sparse structure, low/high
|
||
resolution shape diffusion, sparse shape/texture decoders and PBR GLB export
|
||
through Rust and CUDA. It loads the combined Comfy-Org BF16 flow checkpoint by
|
||
component prefix, without extracting four copies. Checkpoints are downloaded
|
||
and SHA-256 verified through the hub registry. Production inference does not
|
||
start ComfyUI or Python.
|
||
|
||
The [Image to Pixal3D Flow](../../../flow/recipes/templates/image-to-pixal3d.splash)
|
||
connects matting, native mesh generation and the mesh output. Its JSON settings
|
||
input controls the reconstruction. The model's tensor stages remain inside the
|
||
hub backend so intermediate GPU tensors stay on the GPU worker.
|
||
|
||
## Request
|
||
|
||
Submit a mesh request to the hub with PNG input bytes in `input_b64`:
|
||
|
||
```json
|
||
{
|
||
"model": "pixal3d",
|
||
"input_b64": "<PNG base64>",
|
||
"input_content_type": "image/png",
|
||
"seed": 42,
|
||
"texture": true,
|
||
"remesh_resolution": 256,
|
||
"decimation_target": 80000,
|
||
"texture_size": 1024,
|
||
"pixal": {
|
||
"resolution": 1024,
|
||
"camera_fov": 49.13,
|
||
"structure_seed": 56,
|
||
"texture_seed": 43,
|
||
"shape_steps": 20
|
||
}
|
||
}
|
||
```
|
||
|
||
`resolution` accepts 1024 or 1536. `camera_fov` is the horizontal field of view
|
||
in degrees (1–170); the default is 49.13. The main seed drives shape sampling.
|
||
Structure defaults to the main seed and texture to main seed + 1. Shape steps
|
||
control the high resolution stage (default 20, range 1–100). Native BiRefNet
|
||
segments opaque input; pre-segmented RGBA input skips that pass.
|
||
|
||
The defaults favor fast asset export: a 256 remesh grid, about 80,000 faces and
|
||
a 1024 texture atlas. Shape resolution and output mesh density are separate
|
||
controls. `remesh_resolution: 0` retains the existing raw mesh export path.
|
||
|
||
## Performance implementation
|
||
|
||
- NAF evaluates attention only at the four bilinear sample pixels consumed by
|
||
each active voxel. It avoids the dense 1024 × 1024 × 1024 feature image
|
||
(4 GiB at F32), dense unfolded neighborhoods and third-party NATTEN kernels.
|
||
- Image cross-attention K/V and up to 512 MiB of fixed projection outputs are
|
||
cached across diffusion steps. Negative CFG preserves the projection bias.
|
||
- Guide transpose and 2D positional rotation are fused on the GPU. At 1024,
|
||
this avoids constructing and uploading two 128 MiB CPU phase tables.
|
||
- Device weights reuse the existing native namespace cache; conditioning guide
|
||
tensors are released before the corresponding diffusion stage.
|
||
|
||
This is currently a CUDA BF16 backend. The registry advertises a 24 GiB budget;
|
||
1536, larger images/meshes and device workspace behavior can require more.
|
||
Actual validation used an RTX PRO 6000 Blackwell with 96 GiB, not a 24 GiB card.
|
||
|
||
## Relationship to the ComfyUI workflow
|
||
|
||
The port follows Tencent's Pixal3D projection architecture and the Comfy-Org
|
||
combined checkpoint layout. It supports the 1024/1536 cascades and independent
|
||
sampling seeds. It is not a bit-for-bit reproduction of a ComfyUI seed:
|
||
Makepad uses its native Gaussian RNG and GPU math.
|
||
|
||
The downloaded workflow also uses features not implemented here: INT8 convrot
|
||
and 6 GiB offloading, MoGe automatic FOV, PEC UV unwrapping, and explicit normal
|
||
and AO texture baking. This backend uses manual FOV and the existing native
|
||
FaithC remesh, decimation, xatlas and color/metallic/roughness baking. Its
|
||
current shared decoders use the original Microsoft F16 checkpoints rather
|
||
than Comfy-Org's BF16 VAE repack. The downloaded 768 remesh / 700,000-face /
|
||
2048-atlas settings are not the fast defaults above (the native remesh grid is
|
||
currently capped at 512).
|
||
|
||
See [third-party notices](THIRD_PARTY_NOTICES.md) for pinned source identities
|
||
and licenses. Model and component licenses remain separate; the bundle is not
|
||
represented as uniformly MIT. ComfyUI GPL code is not vendored into this crate.
|
||
|
||
## Reproduce and check
|
||
|
||
From `libs/ai/hub`, on a CUDA build machine:
|
||
|
||
```sh
|
||
cargo build --release --no-default-features --features mesh --examples
|
||
```
|
||
|
||
Run the standalone release executable `pixal_generate` with:
|
||
|
||
```text
|
||
pixal_generate WEIGHTS INPUT.png OUTPUT.glb 1024 3
|
||
```
|
||
|
||
The optional final argument repeats generation with one loaded backend; each
|
||
trial writes `OUTPUT.glb.N.glb`. Trial 0 includes preparation/first use of model
|
||
weights. Later trials retain the device weight cache. The timer includes
|
||
matting, neural generation, mesh processing and GLB encoding; registry file
|
||
verification occurs before the timer. Weight filenames for this harness:
|
||
|
||
```text
|
||
pixal3d_bf16.safetensors
|
||
dino_naf.safetensors
|
||
ss-decoder.safetensors
|
||
shape-decoder.safetensors
|
||
texture-decoder.safetensors
|
||
native-matte.safetensors
|
||
```
|
||
|
||
All revisions, sizes and hashes come from `libs/ai/hub/registry.json`; the
|
||
harness only remaps local cache filenames. Standard hub operation uses the
|
||
ordinary registry cache paths.
|
||
|
||
Numerical checks:
|
||
|
||
```sh
|
||
# From libs/ai:
|
||
cargo test --release -p makepad-ai-trellis --lib
|
||
|
||
# Run native checkpoint-driven encoder fixture, then the independent oracle:
|
||
pixal_naf_check dino_naf.safetensors guide.f32
|
||
python pixal_guide_oracle.py dino_naf.safetensors guide.f32
|
||
```
|
||
|
||
The Python tests live in `tests/` and need PyTorch/safetensors. The complete
|
||
encoder fixture checks both cached and uncached native paths against ordinary
|
||
PyTorch convolution, group normalization, pooling, RoPE, neighborhood
|
||
attention and grid sampling. F16 tensor-core operands account for its looser
|
||
tolerance than the separate F32 sampling kernel test.
|
||
|
||
`tests/pixal_naf_oracle.py` accepts a shared library built from
|
||
`../../cuda/kernels/pixal.cu` with `nvcc -shared -O3` and the target GPU's
|
||
architecture. On Windows export `makepad_cuda_pixal_naf_sample_f32` from the
|
||
DLL. `--benchmark` measures sparse sampling only, excluding guide encoding and
|
||
the rest of the model.
|
||
|
||
## Measured run (2026-09-06)
|
||
|
||
RTX PRO 6000 Blackwell 96 GiB, CUDA 13.2, release build, Comfy's
|
||
`viking_wolf_rune_axe.png`, the 1024 request above:
|
||
|
||
| Trial | Through texture decode | Complete PBR GLB |
|
||
| --- | ---: | ---: |
|
||
| First use in a fresh process | 27.38 s | 47.76 s |
|
||
| Cached model weights, second call | 10.08 s | 32.40 s |
|
||
| Cached model weights, third call | 10.13 s | 30.49 s |
|
||
|
||
The final inspected GLB has 79,158 faces, valid UVs and two embedded 1024 PNG
|
||
textures. The 1536 cascade also completed in 68.33 s with the same fast export
|
||
settings, before the fused GPU phase change. Mesh cleanup and xatlas account
|
||
for most of the warm end-to-end time and vary between runs. These are native
|
||
measurements, not a speedup claim against an end-to-end ComfyUI baseline.
|
||
|
||
The separate F32 sparse NAF sampling oracle differed from dense PyTorch by at
|
||
most 1.1e-6, with roughly 7.3 ms sampling for 30,000 voxels. That kernel timing
|
||
excludes guide encoding. The complete checkpoint-driven encoder/NAF fixture
|
||
had maximum absolute error 0.000561 against F32 PyTorch; its native encoded
|
||
feature reuse and recomputation paths agreed exactly. A larger projection
|
||
cache and a persistent 1 GiB image encoding cache did not show a reliable
|
||
end-to-end benefit and were not retained in the production path.
|