Squashed from work: - asset-ai: FastH3 4-step fast video backend; clip keyframes on the wire - asset-ui: loop video chains — text→image→video that ends where it began - h3: safetensors -> pruned-Q4_K GGUF quantizer for the 24GB DiT tiers - h3_quant_gguf verify: row-error gates calibrated to the measured Q4_K floor - asset-ai realtime: the feedback loop — the source anchors, the drifted frame inits - asset-ai realtime: a feedback loop survives a resize and travels by default - asset-ai realtime: the feedback loop frees itself from the feed handshake and pauses for its listener - asset-ai realtime: the outbound encode leaves the loop's critical path - asset-ai ocr: the ocr domain — Chandra 2 at page resolution, and the tower goes planner-owned - llm slots: a lane can hold an image span — embedding prefill and a rope cursor of its own - vision tower on CUDA: the encode leg gets its two missing kernels - llm/ocr: one M-RoPE grid encoder for both image paths, and a livelock made an error - vision tower on CUDA: the f16 GEMM keeps the precision it was throwing away - live: a feed that moves box takes its trip with it — one seed image - vision tower on CUDA: the tiled attention becomes bit-exact, and tensor cores go - llm prefill on CUDA: the MMA attention kernel gets the tile a 4-to-1 model needs - asset-ai ocr: the CUDA encode lane joins the integration — vision-parity sits beside run's three arms, and the kernels - Merge branch 'ocr-perf-integration' into work - asset-ai: the live anchor can follow the trip, and text leaves the 5090 - asset-ai: the camera moves the world, and the world starts still - asset-import: the EA strategy classics, in the one 2D contract - rtsmap: one seeded generator for tiled strategy maps - asset-ui: one card for the strategy classics, with a pack dropdown - asset-ai: music3 reference-audio path, ocr/h3 backends, registry - asset: mp4 sample index for range-streaming, chat tools, import profiles - cnc: tiberium is twelve growth frames, not twelve empty variants - platform: native file and save dialogs, in-house on all three desktops - chat: the scan holds out for a lane home - chat: a full home queues you — take the free lane - chat: the preload has a percentage, and the boundless cap stops showing - llm cuda: the 32x2 attention tile — even GQA ratios stay on MMA - sa3 gets a bake path: the sfx model's tables precomputed by a diffusion-side bin - sqlite_query: anti-join regression test - td import: HARV's second frame block is its harvesting cycle, not a turret - asset-ui: sprite enhancement runs on the 32B dev DiT — distillation, not the prompt, was the ceiling - ai-hub: makepad-asset-ai becomes makepad-ai-hub at libs/ai/hub, the chat pane becomes makepad-chat-ui, the service bin - asset-ui: test health fixtures grow the realtime field they were born without - ai-hub: one home at ~/.makepad — weights/ run/ cache/ logs/, the service cache migrates from ai_content by a single re - ai-hub: subprocess workers die with the node — process groups everywhere, PDEATHSIG on linux, one KILL_ON_JOB_CLOSE Jo - ai-hub: the hub object — AiHub::in_process, pipes vocabulary, and the local LLM engine generalized out of mpfiles (aic - strict-json: the dependency-free JSON module gets its own crate; asset-client re-exports it so nothing downstream move - ai-hub: the machine layer — node entries, the 0600 machine token, and the residency election that IS the lock (aicore - ai-hub: MPHUB1 — the fabric beacon only dedicated nodes can send (aicore §4) - ai-hub: job leases — work lives only while it is renewed (aicore §8) - asset-creator: the pipeline library is born — specs, the deps gate, and the derived-state law (aicore §9) - ai-hub: RAM residency facts — the CPU-side twin of residency.rs (aicore §3) - ai-hub: ETA placement primitives — relative GPU throughput, the four-term estimate, and an observable breakdown (aicor - ai-hub: leases go live on the wire — origin fields on submit, /job/<id>/keepalive, /bye, and the reaper that cancels w - ai-hub: the chat providers move in — fleet qwen, openai, grok, claude/codex/grok CLIs, the responses driver, and the w - asset-creator: the engine — one pipeline run against the hub, deps-gated, spliced, cancellable, resumable-by-construct - ai-hub: the machine node mode — --machine binds loopback, registers in ~/.makepad/run, and exits on its own once idle - asset-creator: makepad-creator-run — the detached client for runs that must outlive a window (aicore §9) - ai-hub: a native Claude Messages-API provider — API-key or Claude Code OAuth, bounded SSE streaming, injected tools (a - route + converse: off makepad_ai — the Agent seam moves to converse, route's cloud dispatcher rides the hub's Claude p - asset-creator: the preset tables move in — fifteen chain-policy constants shared by every creator app (aicore §9 / P6) - makepad_ai is deleted — every backend is a hub pipe, the agent seam lives with its consumers (aicore §14, decided 2026 - ai-hub: loads hold the machine residency election — set_model_state claims on Loaded and publishes the service port (a - ai-hub: chats run the machine election — route to a serving holder, wait on a loading one, claim and publish when open - ai-hub: pick_for_domain_eta — ETA-ranked placement over the shared hard-filter core (aicore §6 / P4) - asset-creator: the engine picks a provider per stage at dispatch time — a chain's later stages see fresh fleet state ( - ai-hub: the fabric secret gates the service HTTP surface — bearer on everything but /health and the ticketed peer path - vj: DREAM runs execute in the app — pipelines.rs becomes the run it used to watch (aicore §9 / F1) - asset-creator: the runner — generate one thing and put it in the catalog, one implementation for every surface (aicore - chat-ui: the session runs in the app — no broker anywhere on the chat path (aicore P8 / F5) - asset-store: assets.query is a first-class query endpoint — the bounded SQL surface outlives the broker (aicore P8 / F - asset-creator: CreatorTools — the chat tool pack for a store that only stores (aicore §9 / P8) - asset-store: the shrink — the store stores (aicore P7) - importer + asset-server host: the coordination era ends (aicore P7) - store config purge + asset-ui goes fleet-direct; the derive protocol gets its route proof (aicore P7) - client + chat dispatcher: the dead wire comes out (aicore P7/P8) - ai-hub: 0.3.0 — the health version says which era a node runs - ai-hub: the default fleet is 'gen' — apps hear the LAN without env plumbing - ai-hub: the preload note percents the prefill, not the job bar - ai-hub: conversations keep their KV — the wire mirror, the lane identity, the in-turn dynamic context (aicore §7) - ai-hub: an open-think model is thinking from its first token - libs: the zero-warning sweep — stitch casts say what they mean, xatlas keeps upstream's surface quietly - zero-warning sweep, round two — the first full-workspace pass - zero-warning sweep, round three — the model lanes and the deep examples - zero-warning sweep, round four — the last stragglers - zero-warning sweep, round five — vj and chat-ui - zero-warning sweep, round six — three cascades Co-authored-by: Claude <info@makepad.nl>
378 lines
13 KiB
Python
378 lines
13 KiB
Python
# music3_oracle_dump.py - standalone MiniMax-Music3 stage dumps for the native
|
|
# port. Does NOT talk to the live :8123 service. Loads the official
|
|
# ModularPipeline from a local cache dir (same view rewrite as music3_worker.py)
|
|
# and writes numpy + wav + meta.json for a short seeded clip.
|
|
#
|
|
# Usage (169):
|
|
# C:\ai\music3venv\Scripts\python.exe music3_oracle_dump.py \
|
|
# --model-dir C:\ai\asset_node_cache\music\MiniMax-Music3 \
|
|
# --out-dir C:\ai\music3_oracle\pine_5s_seed7 \
|
|
# --seconds 5 --seed 7
|
|
import argparse
|
|
import json
|
|
import os
|
|
import shutil
|
|
import sys
|
|
import time
|
|
import wave
|
|
|
|
os.environ.setdefault("HF_HUB_OFFLINE", "1")
|
|
os.environ.setdefault("TRANSFORMERS_OFFLINE", "1")
|
|
os.environ.setdefault("HF_HUB_DISABLE_TELEMETRY", "1")
|
|
|
|
# Official fixture (docs example lyrics/caption, shortened duration).
|
|
DEFAULT_PROMPT = (
|
|
"Genre: acoustic pop. BPM: 96. Key: C major. Warm and intimate, "
|
|
"building gently into the chorus. Vocals: soft female lead, close and "
|
|
"breathy, light stacked harmonies in the chorus. Arrangement: "
|
|
"fingerpicked guitar and soft piano; brushed drums and upright bass "
|
|
"enter in the chorus."
|
|
)
|
|
DEFAULT_LYRICS = (
|
|
"[verse]\n"
|
|
"Morning light filtering through the pine\n"
|
|
"Every quiet street is yours and mine\n"
|
|
"[chorus]\n"
|
|
"Softly the world begins to breathe"
|
|
)
|
|
|
|
|
|
def build_local_view(model_dir, view_dir):
|
|
for root, _dirs, files in os.walk(model_dir):
|
|
rel = os.path.relpath(root, model_dir)
|
|
dst_root = os.path.join(view_dir, rel) if rel != "." else view_dir
|
|
os.makedirs(dst_root, exist_ok=True)
|
|
for name in files:
|
|
src = os.path.join(root, name)
|
|
dst = os.path.join(dst_root, name)
|
|
if os.path.exists(dst) and os.path.getsize(dst) == os.path.getsize(src):
|
|
continue
|
|
if os.path.exists(dst):
|
|
os.remove(dst)
|
|
if name == "modular_model_index.json":
|
|
with open(src, "r", encoding="utf-8") as handle:
|
|
index = json.load(handle)
|
|
for value in index.values():
|
|
if isinstance(value, list) and len(value) == 3 and isinstance(value[2], dict):
|
|
if "pretrained_model_name_or_path" in value[2]:
|
|
value[2]["pretrained_model_name_or_path"] = view_dir
|
|
with open(dst, "w", encoding="utf-8") as handle:
|
|
json.dump(index, handle, indent=1)
|
|
continue
|
|
try:
|
|
os.link(src, dst)
|
|
except OSError:
|
|
shutil.copyfile(src, dst)
|
|
|
|
|
|
def to_numpy(value):
|
|
import numpy as np
|
|
import torch
|
|
|
|
if isinstance(value, torch.Tensor):
|
|
return value.detach().to(torch.float32).cpu().numpy()
|
|
return np.asarray(value)
|
|
|
|
|
|
def write_wav_i16(path, data, sample_rate):
|
|
import numpy as np
|
|
|
|
pcm = (np.clip(data.T, -1.0, 1.0) * 32767.0).astype("<i2")
|
|
with wave.open(path, "wb") as handle:
|
|
handle.setnchannels(pcm.shape[1])
|
|
handle.setsampwidth(2)
|
|
handle.setframerate(int(sample_rate))
|
|
handle.writeframes(pcm.tobytes())
|
|
|
|
|
|
def main():
|
|
parser = argparse.ArgumentParser()
|
|
parser.add_argument("--model-dir", required=True)
|
|
parser.add_argument("--out-dir", required=True)
|
|
parser.add_argument("--view-dir", default=None)
|
|
parser.add_argument("--prompt", default=DEFAULT_PROMPT)
|
|
parser.add_argument("--lyrics", default=DEFAULT_LYRICS)
|
|
parser.add_argument("--seconds", type=float, default=5.0)
|
|
parser.add_argument("--seed", type=int, default=7)
|
|
parser.add_argument("--steps", type=int, default=30)
|
|
args = parser.parse_args()
|
|
|
|
out_dir = os.path.abspath(args.out_dir)
|
|
os.makedirs(out_dir, exist_ok=True)
|
|
view_dir = args.view_dir or os.path.join(out_dir, "_model_view")
|
|
stats = {}
|
|
t0 = time.time()
|
|
|
|
def save(name, value):
|
|
import numpy as np
|
|
|
|
arr = np.ascontiguousarray(to_numpy(value))
|
|
np.save(os.path.join(out_dir, name + ".npy"), arr)
|
|
stats[name] = {
|
|
"shape": list(arr.shape),
|
|
"dtype": str(arr.dtype),
|
|
"absmax": float(np.abs(arr).max()) if arr.size else 0.0,
|
|
"mean": float(arr.astype("float64").mean()) if arr.size else 0.0,
|
|
"std": float(arr.astype("float64").std()) if arr.size else 0.0,
|
|
}
|
|
print("saved", name, arr.shape, "absmax", stats[name]["absmax"], flush=True)
|
|
|
|
print("build-view", args.model_dir, "->", view_dir, flush=True)
|
|
build_local_view(args.model_dir, view_dir)
|
|
|
|
print("load-libs", flush=True)
|
|
import numpy as np
|
|
import torch
|
|
from diffusers import ModularPipeline
|
|
from diffusers.modular_pipelines.minimax_music3.encoders import (
|
|
_AUDIO_CFG_TOKEN_ID,
|
|
_AUDIO_CODE_OFFSET,
|
|
_AUDIO_END_TOKEN_ID,
|
|
_AR_CFG_SCALE,
|
|
_AR_CFG_TOP_K,
|
|
_AR_SAMPLING_TOP_K,
|
|
_clean_caption,
|
|
_normalize_lyrics,
|
|
)
|
|
|
|
print("load-components", flush=True)
|
|
pipe = ModularPipeline.from_pretrained(view_dir)
|
|
pipe.load_components(dtype=torch.bfloat16)
|
|
print("to-gpu", flush=True)
|
|
pipe.to("cuda")
|
|
print(
|
|
"ready sr=%s fr=%s hop=%s codebooks=%s vram=%.1fGiB"
|
|
% (
|
|
pipe.sampling_rate,
|
|
pipe.frame_rate,
|
|
pipe.latent_hop_length,
|
|
pipe.num_codebooks,
|
|
torch.cuda.max_memory_allocated() / (1024**3),
|
|
),
|
|
flush=True,
|
|
)
|
|
|
|
assembled = (
|
|
"<|im_start|><|caption_start|>"
|
|
+ _clean_caption(args.prompt)
|
|
+ "<|caption_end|><|lyrics_start|>"
|
|
+ _normalize_lyrics(args.lyrics)
|
|
+ "<|lyrics_end|><|im_end|><|audio_start|>"
|
|
)
|
|
with open(os.path.join(out_dir, "assembled_prompt.txt"), "w", encoding="utf-8") as handle:
|
|
handle.write(assembled)
|
|
tok = pipe.tokenizer
|
|
text_ids = tok(assembled, return_tensors="pt")["input_ids"]
|
|
uncond = text_ids.clone()
|
|
uncond[:, 1:-2] = _AUDIO_CFG_TOKEN_ID
|
|
text_ids_pair = torch.cat((text_ids, uncond), dim=0)
|
|
save("text_ids", text_ids_pair.cpu().numpy().astype(np.int64))
|
|
specials = {
|
|
name: int(tok.convert_tokens_to_ids(name))
|
|
for name in (
|
|
"<|im_start|>",
|
|
"<|im_end|>",
|
|
"<|caption_start|>",
|
|
"<|caption_end|>",
|
|
"<|lyrics_start|>",
|
|
"<|lyrics_end|>",
|
|
"<|audio_start|>",
|
|
"<|endoftext|>",
|
|
)
|
|
}
|
|
|
|
taps = {
|
|
"lm_calls": 0,
|
|
"dit_calls": 0,
|
|
"rvq_calls": 0,
|
|
"cond_calls": 0,
|
|
"vocoder_calls": 0,
|
|
"semantic_codes": [],
|
|
"rvq_codes": [],
|
|
}
|
|
|
|
lm_model = pipe.language_model.model
|
|
orig_lm = lm_model.forward
|
|
|
|
def lm_forward(*a, **kw):
|
|
out = orig_lm(*a, **kw)
|
|
taps["lm_calls"] += 1
|
|
if taps["lm_calls"] == 1:
|
|
save("lm_prefill_last_hidden", out.last_hidden_state)
|
|
last = out.last_hidden_state[:, -1]
|
|
logits = pipe.language_model.lm_head(last).float()
|
|
save("lm_prefill_logits", logits)
|
|
return out
|
|
|
|
lm_model.forward = lm_forward
|
|
|
|
orig_dit = pipe.transformer.forward
|
|
|
|
def dit_forward(*a, **kw):
|
|
out = orig_dit(*a, **kw)
|
|
taps["dit_calls"] += 1
|
|
sample = out[0] if isinstance(out, (tuple, list)) else getattr(out, "sample", out)
|
|
if taps["dit_calls"] == 1:
|
|
hidden = kw.get("hidden_states", a[0] if a else None)
|
|
timestep = kw.get("timestep")
|
|
cond = kw.get("encoder_hidden_states")
|
|
if hidden is not None:
|
|
save("dit_step0_x", hidden)
|
|
if timestep is not None:
|
|
save("dit_step0_t", timestep)
|
|
if cond is not None:
|
|
save("dit_step0_cond", cond)
|
|
save("dit_step0_v_cond", sample)
|
|
elif taps["dit_calls"] == 2:
|
|
save("dit_step0_v_uncond", sample)
|
|
return out
|
|
|
|
pipe.transformer.forward = dit_forward
|
|
|
|
orig_rvq = pipe.rvq_depth_decoder.forward
|
|
|
|
def rvq_forward(*a, **kw):
|
|
out = orig_rvq(*a, **kw)
|
|
taps["rvq_calls"] += 1
|
|
if taps["rvq_calls"] == 1:
|
|
hidden = a[0] if a else kw.get("inputs_embeds")
|
|
if hidden is not None:
|
|
save("rvq_step0_in", hidden)
|
|
save("rvq_step0_out", out)
|
|
return out
|
|
|
|
pipe.rvq_depth_decoder.forward = rvq_forward
|
|
|
|
orig_cond = pipe.condition_encoder.forward
|
|
|
|
def cond_forward(*a, **kw):
|
|
out = orig_cond(*a, **kw)
|
|
taps["cond_calls"] += 1
|
|
if taps["cond_calls"] == 1:
|
|
hidden = a[0] if a else kw.get("hidden_states")
|
|
if hidden is not None:
|
|
save("cond_enc_in", hidden)
|
|
save("cond_enc_out", out)
|
|
return out
|
|
|
|
pipe.condition_encoder.forward = cond_forward
|
|
|
|
orig_voc = pipe.vocoder.forward
|
|
|
|
def voc_forward(*a, **kw):
|
|
out = orig_voc(*a, **kw)
|
|
taps["vocoder_calls"] += 1
|
|
if taps["vocoder_calls"] == 1:
|
|
hidden = a[0] if a else kw.get("latents")
|
|
if hidden is not None:
|
|
save("vocoder_in", hidden)
|
|
save("vocoder_out", out)
|
|
return out
|
|
|
|
pipe.vocoder.forward = voc_forward
|
|
|
|
# Capture sampled semantic codes from lm_head-guided first token of each
|
|
# inner-model decode after prefill by wrapping the official sampler.
|
|
import diffusers.modular_pipelines.minimax_music3.encoders as enc
|
|
|
|
orig_sample = enc._sample_top_k
|
|
sample_n = {"n": 0}
|
|
|
|
def sample_top_k(logits, generator):
|
|
token = orig_sample(logits, generator)
|
|
sample_n["n"] += 1
|
|
# Global LM samples once per frame (after CFG), then RVQ samples 7
|
|
# residual codes. Record the first 8 sample values of frame 0 plus
|
|
# every global-LM sample as semantic_codes when the vocab is the LM.
|
|
if logits.shape[-1] == pipe.language_model.config.vocab_size:
|
|
taps["semantic_codes"].append(int(token.reshape(-1)[0].item()))
|
|
elif logits.shape[-1] == pipe.audio_vocab_size:
|
|
taps["rvq_codes"].append(int(token.reshape(-1)[0].item()))
|
|
if sample_n["n"] == 1:
|
|
save("first_sample_logits", logits)
|
|
return token
|
|
|
|
enc._sample_top_k = sample_top_k
|
|
|
|
print("generate seconds=%s seed=%s steps=%s" % (args.seconds, args.seed, args.steps), flush=True)
|
|
gen_t0 = time.time()
|
|
generator = torch.Generator("cuda").manual_seed(int(args.seed))
|
|
audio = pipe(
|
|
prompt=args.prompt,
|
|
lyrics=args.lyrics,
|
|
audio_duration=float(args.seconds),
|
|
num_inference_steps=int(args.steps),
|
|
generator=generator,
|
|
output="audios",
|
|
)[0]
|
|
gen_s = time.time() - gen_t0
|
|
|
|
data = to_numpy(audio)
|
|
if data.ndim == 3:
|
|
data = data[0]
|
|
if data.ndim == 1:
|
|
data = data[None, :]
|
|
if data.shape[0] > data.shape[1]:
|
|
data = data.T
|
|
save("audio", data)
|
|
wav_path = os.path.join(out_dir, "song.wav")
|
|
write_wav_i16(wav_path, data, int(pipe.sampling_rate))
|
|
|
|
if taps["semantic_codes"]:
|
|
save("semantic_codes", np.asarray(taps["semantic_codes"], dtype=np.int64))
|
|
if taps["rvq_codes"]:
|
|
codes = np.asarray(taps["rvq_codes"], dtype=np.int64)
|
|
save("rvq_codes_flat", codes)
|
|
n_cb = int(pipe.num_codebooks) - 1
|
|
if n_cb > 0 and codes.size % n_cb == 0:
|
|
save("rvq_codes", codes.reshape(-1, n_cb))
|
|
|
|
meta = {
|
|
"fixture": "pine_5s_seed7",
|
|
"prompt": args.prompt,
|
|
"lyrics": args.lyrics,
|
|
"assembled_prompt": assembled,
|
|
"seconds": float(args.seconds),
|
|
"seed": int(args.seed),
|
|
"num_inference_steps": int(args.steps),
|
|
"sample_rate": int(pipe.sampling_rate),
|
|
"frame_rate": float(pipe.frame_rate),
|
|
"latent_hop_length": int(pipe.latent_hop_length),
|
|
"num_codebooks": int(pipe.num_codebooks),
|
|
"audio_vocab_size": int(pipe.audio_vocab_size),
|
|
"num_channels_latents": int(pipe.num_channels_latents),
|
|
"special_tokens": specials,
|
|
"audio_end_token_id": int(_AUDIO_END_TOKEN_ID),
|
|
"audio_cfg_token_id": int(_AUDIO_CFG_TOKEN_ID),
|
|
"audio_code_offset": int(_AUDIO_CODE_OFFSET),
|
|
"ar_cfg_scale": float(_AR_CFG_SCALE),
|
|
"ar_cfg_top_k": int(_AR_CFG_TOP_K),
|
|
"ar_sampling_top_k": int(_AR_SAMPLING_TOP_K),
|
|
"lm_vocab_size": int(pipe.language_model.config.vocab_size),
|
|
"lm_hidden_size": int(pipe.language_model.config.hidden_size),
|
|
"lm_num_layers": int(pipe.language_model.config.num_hidden_layers),
|
|
"counts": {
|
|
"lm_calls": taps["lm_calls"],
|
|
"dit_calls": taps["dit_calls"],
|
|
"rvq_calls": taps["rvq_calls"],
|
|
"cond_calls": taps["cond_calls"],
|
|
"vocoder_calls": taps["vocoder_calls"],
|
|
"semantic_codes": len(taps["semantic_codes"]),
|
|
"rvq_codes": len(taps["rvq_codes"]),
|
|
},
|
|
"wall_s": {
|
|
"total": time.time() - t0,
|
|
"generate": gen_s,
|
|
},
|
|
"peak_vram_bytes": int(torch.cuda.max_memory_allocated()),
|
|
"stats": stats,
|
|
"wav": wav_path,
|
|
}
|
|
with open(os.path.join(out_dir, "meta.json"), "w", encoding="utf-8") as handle:
|
|
json.dump(meta, handle, indent=2)
|
|
print("done", json.dumps({k: meta[k] for k in ("counts", "wall_s", "peak_vram_bytes", "wav")}), flush=True)
|
|
return 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
sys.exit(main())
|