makepad/libs/ai/models/trellis/README.md
Admin 60615db4ed libs: ai hub/models/cuda, speech, chat_ui
Squash of 54 work commits (Sep 1–12):
  6251f7c  ai-hub: body domain — live pose packets ride the realtime session
  ea50c77  chat_ui: the feed's session gets its profile brief back
  f51b5f3  ai-body: the crate for the native SAM 3D Body port, with its weights reader
  8211ae6  ai-body: the MHR rig and the pose head's parameter decoding, oracle-exact
  9e343a8  ai-body: the DINOv3 ViT-H+/16 backbone, crop and ray conditioning; Metal gains rope-half and affine layer norm
  69d842c  ai-body: the promptable pose decoder and its refinement loop, oracle-matched on Metal
  66e5e2f  ai-hub: SAM 3D Body runs natively — `sam3dbody` on the body domain, oracle-matched end to end
  a634198  ai-hub: the body-native commit carried a peer's in-flight hub hunks; put them back where they were
  9ff44e8  ai-hub: the body-native wiring, this time only the lane's hunks
  6a1c16b  ai-body: third-party notices — what the port is implemented after, and what it is not
  d78411a  ai-body: the per-step work moves to the GPU
  b22259b  ai-body: the context stays on the GPU; only the pose token leaves the loop
  346f31f  ai-body: flash attention for the head-dim-64 blocks
  45b5b98  ai-body: the crop size is a runtime knob, and the loop reports where its time goes
  4be6d19  ai-body: the test modules import the grid constants they still use
  7598346  ai-body: tensor-core GEMMs for the backbone, and the rig's correctives only where they count
  a9ce596  ai-body: the crop warp runs across cores
  8964ba6  ai-body: an FP8 backbone mode, off by default, measured against the oracle
  a2aaa8f  ai-body: the FP8 bias rides a column-broadcast add on the device
  d53c77d  metal: a device-resident ViT stack, and the body backbone rides it
  d006d0a  metal: resident f32 linears keep their weight on the device
  525ba1c  metal: a device-resident two-way decoder layer, and the body decoder rides it
  c9e6d88  ai-body: the hands pass — hand crops, the hand decoder, the hand-mode rig and the wrist fusion
  62dff26  ai-body: the mask prompt — a person's segmentation mask conditions the body pass
  a648cf8  ai-hub: body session options — hands, detect, persons=N
  8c568df  ai-hub: drop the SAM 3D Body reference worker backend
  7ff875a  ai-hub: keep a peer's in-flight beats/notes/local work out of the body commits
  31e5faa  ai-hub: local model runner, licence acknowledgements, a shared install panel; Beat This!, Basic Pitch and the Salamander drum-kit entries
  b94bc58  ai-services: the wire, the app port and the panel state — one conversation, many apps
  2acb798  ai-services: wire v2 — endpoints, receiver-side caps, result disposition
  8ae0ffb  ai-services: the engine core — registry, router and conversation, tested against a scripted model
  2308736  ai-services: the real models behind the engine feature — local through the hub, Claude, and none
  c3f631d  livepipe: one reusable pipe from a camera to a fleet node and back
  ff62db3  ai libs: the runtime env-var cleanup — precision is a per-caller policy, not an environment side channel
  04a94ef  realtime: one service-log line when a live session opens and one when it closes
  0ecb81c  ai models: the model-crates env-var cleanup — 172 research knobs gone, the unset default is the code
  4ca36c1  ai hub + services: the assistant's model comes from wherever it is resident — the fleet chat box, with tools, then the local weights
  432121e  aichat engine + wm: launch, then use — the assistant continues in the same turn once the app it started is on the bus
  7a5bf69  ai-hub registry: the Salamander drumkit samples come from the makepad.nl mirror — the GitHub repo only carries the .sfz files
  102ffc5  ai-services: messages on the bus — a manifest declares topics, the engine subscribes on a tool's behalf or by ToolResult.subscribe, a service publishes Message frames, an idle conversation wakes on a message as an event turn under rate laws; the WM bus forwards the new frames; every app that matches the wire gets its arm
  a837792  hub + flow: a whitespace-only chat completion is retried once and then fails instead of passing as an answer; a flow's model is a fleet model id unless it names a weight file on disk; chat models show under the text domain in /v1/models
  bc6c620  hub + flow: what the chat review found — the in-process route retries an empty completion too, a node says whether its prefill opened thinking so a brief-mode answer is never discarded, a preferred model falls back to normal election when no node has it, discovery keeps looking for the preferred model until patience runs out
  75c3441  hub: the PRO 6000 serves image as well as chat and text
  ad5e98b  hub registry: flux2-dev's VRAM estimate is its measured peak, 30 GB
  c7241e0  hub: a node that evicted every resident releases its cached allocator pool before refusing a load or publishing usable VRAM
  30575f0  flow: route generation by request workload
  1be1e21  ai-hub: gate downloads by disk capacity and recover fleet admission
  df6b394  filesystem_watcher, bounded_http, ai services: live and tool prerequisites
  79ebdb9  ai-hub: add a native Pixal3D image-to-3D backend
  0ba0d74  ai-hub: propagate typed refusals under reject queue policy
  cc6c872  Speed up H3 conditioning and video decoding
  e512059  Fix Qwen vision residency and generated material colors
  2864f68  ai-hub http client: bound every plain TCP connect to 3 s per address
  3d93229  ai: CUDA is a Linux/Windows-only dependency; the hub library defaults to llm + stt

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 13:40:31 +02:00

7 KiB
Raw Permalink Blame History

Native TRELLIS.2 and Pixal3D

The pixal3d mesh backend runs image conditioning, sparse structure, low/high resolution shape diffusion, sparse shape/texture decoders and PBR GLB export through Rust and CUDA. It loads the combined Comfy-Org BF16 flow checkpoint by component prefix, without extracting four copies. Checkpoints are downloaded and SHA-256 verified through the hub registry. Production inference does not start ComfyUI or Python.

The Image to Pixal3D Flow connects matting, native mesh generation and the mesh output. Its JSON settings input controls the reconstruction. The model's tensor stages remain inside the hub backend so intermediate GPU tensors stay on the GPU worker.

Request

Submit a mesh request to the hub with PNG input bytes in input_b64:

{
  "model": "pixal3d",
  "input_b64": "<PNG base64>",
  "input_content_type": "image/png",
  "seed": 42,
  "texture": true,
  "remesh_resolution": 256,
  "decimation_target": 80000,
  "texture_size": 1024,
  "pixal": {
    "resolution": 1024,
    "camera_fov": 49.13,
    "structure_seed": 56,
    "texture_seed": 43,
    "shape_steps": 20
  }
}

resolution accepts 1024 or 1536. camera_fov is the horizontal field of view in degrees (1–170); the default is 49.13. The main seed drives shape sampling. Structure defaults to the main seed and texture to main seed + 1. Shape steps control the high resolution stage (default 20, range 1–100). Native BiRefNet segments opaque input; pre-segmented RGBA input skips that pass.

The defaults favor fast asset export: a 256 remesh grid, about 80,000 faces and a 1024 texture atlas. Shape resolution and output mesh density are separate controls. remesh_resolution: 0 retains the existing raw mesh export path.

Performance implementation

  • NAF evaluates attention only at the four bilinear sample pixels consumed by each active voxel. It avoids the dense 1024 × 1024 × 1024 feature image (4 GiB at F32), dense unfolded neighborhoods and third-party NATTEN kernels.
  • Image cross-attention K/V and up to 512 MiB of fixed projection outputs are cached across diffusion steps. Negative CFG preserves the projection bias.
  • Guide transpose and 2D positional rotation are fused on the GPU. At 1024, this avoids constructing and uploading two 128 MiB CPU phase tables.
  • Device weights reuse the existing native namespace cache; conditioning guide tensors are released before the corresponding diffusion stage.

This is currently a CUDA BF16 backend. The registry advertises a 24 GiB budget; 1536, larger images/meshes and device workspace behavior can require more. Actual validation used an RTX PRO 6000 Blackwell with 96 GiB, not a 24 GiB card.

Relationship to the ComfyUI workflow

The port follows Tencent's Pixal3D projection architecture and the Comfy-Org combined checkpoint layout. It supports the 1024/1536 cascades and independent sampling seeds. It is not a bit-for-bit reproduction of a ComfyUI seed: Makepad uses its native Gaussian RNG and GPU math.

The downloaded workflow also uses features not implemented here: INT8 convrot and 6 GiB offloading, MoGe automatic FOV, PEC UV unwrapping, and explicit normal and AO texture baking. This backend uses manual FOV and the existing native FaithC remesh, decimation, xatlas and color/metallic/roughness baking. Its current shared decoders use the original Microsoft F16 checkpoints rather than Comfy-Org's BF16 VAE repack. The downloaded 768 remesh / 700,000-face / 2048-atlas settings are not the fast defaults above (the native remesh grid is currently capped at 512).

See third-party notices for pinned source identities and licenses. Model and component licenses remain separate; the bundle is not represented as uniformly MIT. ComfyUI GPL code is not vendored into this crate.

Reproduce and check

From libs/ai/hub, on a CUDA build machine:

cargo build --release --no-default-features --features mesh --examples

Run the standalone release executable pixal_generate with:

pixal_generate WEIGHTS INPUT.png OUTPUT.glb 1024 3

The optional final argument repeats generation with one loaded backend; each trial writes OUTPUT.glb.N.glb. Trial 0 includes preparation/first use of model weights. Later trials retain the device weight cache. The timer includes matting, neural generation, mesh processing and GLB encoding; registry file verification occurs before the timer. Weight filenames for this harness:

pixal3d_bf16.safetensors
dino_naf.safetensors
ss-decoder.safetensors
shape-decoder.safetensors
texture-decoder.safetensors
native-matte.safetensors

All revisions, sizes and hashes come from libs/ai/hub/registry.json; the harness only remaps local cache filenames. Standard hub operation uses the ordinary registry cache paths.

Numerical checks:

# From libs/ai:
cargo test --release -p makepad-ai-trellis --lib

# Run native checkpoint-driven encoder fixture, then the independent oracle:
pixal_naf_check dino_naf.safetensors guide.f32
python pixal_guide_oracle.py dino_naf.safetensors guide.f32

The Python tests live in tests/ and need PyTorch/safetensors. The complete encoder fixture checks both cached and uncached native paths against ordinary PyTorch convolution, group normalization, pooling, RoPE, neighborhood attention and grid sampling. F16 tensor-core operands account for its looser tolerance than the separate F32 sampling kernel test.

tests/pixal_naf_oracle.py accepts a shared library built from ../../cuda/kernels/pixal.cu with nvcc -shared -O3 and the target GPU's architecture. On Windows export makepad_cuda_pixal_naf_sample_f32 from the DLL. --benchmark measures sparse sampling only, excluding guide encoding and the rest of the model.

Measured run (2026-09-06)

RTX PRO 6000 Blackwell 96 GiB, CUDA 13.2, release build, Comfy's viking_wolf_rune_axe.png, the 1024 request above:

Trial Through texture decode Complete PBR GLB
First use in a fresh process 27.38 s 47.76 s
Cached model weights, second call 10.08 s 32.40 s
Cached model weights, third call 10.13 s 30.49 s

The final inspected GLB has 79,158 faces, valid UVs and two embedded 1024 PNG textures. The 1536 cascade also completed in 68.33 s with the same fast export settings, before the fused GPU phase change. Mesh cleanup and xatlas account for most of the warm end-to-end time and vary between runs. These are native measurements, not a speedup claim against an end-to-end ComfyUI baseline.

The separate F32 sparse NAF sampling oracle differed from dense PyTorch by at most 1.1e-6, with roughly 7.3 ms sampling for 30,000 voxels. That kernel timing excludes guide encoding. The complete checkpoint-driven encoder/NAF fixture had maximum absolute error 0.000561 against F32 PyTorch; its native encoded feature reuse and recomputation paths agreed exactly. A larger projection cache and a persistent 1 GiB image encoding cache did not show a reliable end-to-end benefit and were not retained in the production path.