The model code was spread across eight crates that had grown into each other:
ggml and cuda and mlx each owned part of a tensor runtime, llama and tts and
voice2 each owned part of a model, and libs/diffusion owned everything else.
They are now one tree with an explicit shape:
libs/ai/cuda — kernels and launch surface
libs/ai/metal — Metal shaders and the shim
libs/ai/llm — the language-model runtime (sessions, lanes, contexts,
the CUDA and Metal executors, the compiled Metal path)
libs/ai/models/ — common, flux, h3, music, paint, speech, stems, vision
libs/diffusion is not deleted but demoted: what remains is the VALIDATOR
crate — several dozen `*_validate.rs` oracles that check a native
implementation against a reference, which is where they belong now that the
implementations live next door.
The functional work inside the move is mostly in the LLM runtime: N lanes that
draft while one verify batch serves all of them, per-slot prefill over a shared
folded attention arena, speculation that survives batching, and a scheduler
that reports rather than publishes. And in the CUDA build: a machine without
usable CUDA must still LINK (and say so), the default kernel arch is the
building machine's GPU, `NO_CUDA` forces the stub even where the toolkit
exists, and kernels compile in parallel with progress.
libs/video_flow is new here: classical optical flow estimation and the `mkfl`
motion-field payload — a flow field measured from a clip without a model,
which is what drives free-rate bounce-looping playback and the uprez/tween
enhance pipe.
33 lines
1.6 KiB
Rust
33 lines
1.6 KiB
Rust
use std::env;
|
|
|
|
// Whether `cuda_exec/real.rs` compiles at all.
|
|
//
|
|
// The dispatcher calls both the CUDA runtime (`cudaMalloc`, `cudaFree`, ...)
|
|
// and the `mkllm_*` kernels, and BOTH arrive from makepad-ai-cuda: its nvcc
|
|
// build produces the kernel objects, and its build script emits the
|
|
// `cargo:rustc-link-lib` lines for cudart/cuBLAS. So makepad-ai-cuda's answer
|
|
// is the only honest one, and cargo hands it to us — we are a direct
|
|
// dependent — through the `links = "makepad_ai_cuda"` handshake as
|
|
// DEP_MAKEPAD_AI_CUDA_KERNELS. libs/voice, libs/ai/metal and
|
|
// libs/ai/models/common already gate on exactly this.
|
|
//
|
|
// Probing for nvcc here instead, which is what this script used to do, is how
|
|
// a machine WITH the CUDA toolkit still failed to link: nvcc existed, so
|
|
// real.rs compiled and referenced `cudaFree`, while makepad-ai-cuda had
|
|
// dropped out of its kernel build (MAKEPAD_GGML_NO_CUDA, an nvcc failure, no
|
|
// MSVC lib.exe, a toolkit whose lib dir it could not find) and emitted no
|
|
// CUDA link directives at all. Two build scripts answering the same question
|
|
// separately can only ever agree by luck.
|
|
fn main() {
|
|
println!("cargo:rustc-check-cfg=cfg(makepad_llama_cuda_kernels)");
|
|
if env::var("DEP_MAKEPAD_AI_CUDA_KERNELS").as_deref() != Ok("1") {
|
|
return;
|
|
}
|
|
// The arch makepad-ai-cuda actually compiled for, so real.rs's
|
|
// SASS-compatibility check tests the kernels that exist rather than a
|
|
// second guess at them.
|
|
if let Ok(arch) = env::var("DEP_MAKEPAD_AI_CUDA_ARCH") {
|
|
println!("cargo:rustc-env=MAKEPAD_LLAMA_CUDA_ARCH={arch}");
|
|
}
|
|
println!("cargo:rustc-cfg=makepad_llama_cuda_kernels");
|
|
}
|