makepad/libs/ai/models/speech/tools/build_lexicon.py
Admin 7f59912916 libs/ai: one AI stack, replacing libs/ggml, llama, mlx, cuda, tts, voice2 and pbr_paint
The model code was spread across eight crates that had grown into each other:
ggml and cuda and mlx each owned part of a tensor runtime, llama and tts and
voice2 each owned part of a model, and libs/diffusion owned everything else.
They are now one tree with an explicit shape:

  libs/ai/cuda     — kernels and launch surface
  libs/ai/metal    — Metal shaders and the shim
  libs/ai/llm      — the language-model runtime (sessions, lanes, contexts,
                     the CUDA and Metal executors, the compiled Metal path)
  libs/ai/models/  — common, flux, h3, music, paint, speech, stems, vision

libs/diffusion is not deleted but demoted: what remains is the VALIDATOR
crate — several dozen `*_validate.rs` oracles that check a native
implementation against a reference, which is where they belong now that the
implementations live next door.

The functional work inside the move is mostly in the LLM runtime: N lanes that
draft while one verify batch serves all of them, per-slot prefill over a shared
folded attention arena, speculation that survives batching, and a scheduler
that reports rather than publishes. And in the CUDA build: a machine without
usable CUDA must still LINK (and say so), the default kernel arch is the
building machine's GPU, `NO_CUDA` forces the stub even where the toolkit
exists, and kernels compile in parallel with progress.

libs/video_flow is new here: classical optical flow estimation and the `mkfl`
motion-field payload — a flow field measured from a clip without a model,
which is what drives free-rate bounce-looping playback and the uprez/tween
enhance pipe.
2026-08-23 01:34:35 +02:00

75 lines
2 KiB
Python

#!/usr/bin/env python3
"""Pack misaki's IPA lexicon into a binary the Rust G2P can binary-search.
misaki's `us_gold.json` already speaks Kokoro's own symbol set, so no ARPAbet
conversion is needed. Values are either a string or a POS-keyed dict; we take
`DEFAULT`.
python3 build_lexicon.py us_gold.json ../data/us_lexicon.bin
Format (little endian), entries sorted by key bytes:
magic b"MKLEX\\0\\0\\0" 8 bytes
count u32
index count * { key_off u32, val_off u32 } offsets into blob
blob NUL-terminated utf8 strings
"""
import json
import struct
import sys
MAGIC = b"MKLEX\0\0\0"
def pronunciation(value):
if isinstance(value, str):
return value
if isinstance(value, dict):
chosen = value.get("DEFAULT")
if isinstance(chosen, str):
return chosen
for candidate in value.values():
if isinstance(candidate, str):
return candidate
return None
def main():
src, dst = sys.argv[1], sys.argv[2]
raw = json.load(open(src))
entries = []
for word, value in raw.items():
ipa = pronunciation(value)
if not word or not ipa:
continue
entries.append((word.encode(), ipa.encode()))
entries.sort(key=lambda kv: kv[0])
blob = bytearray()
index = []
seen = {}
for key, val in entries:
# Dedupe identical strings; plenty of words share a pronunciation.
for text in (key, val):
if text not in seen:
seen[text] = len(blob)
blob += text + b"\0"
index.append((seen[key], seen[val]))
with open(dst, "wb") as out:
out.write(MAGIC)
out.write(struct.pack("<I", len(index)))
for key_off, val_off in index:
out.write(struct.pack("<II", key_off, val_off))
out.write(blob)
size = len(MAGIC) + 4 + 8 * len(index) + len(blob)
print(f"{src} -> {dst}")
print(f" entries : {len(index)}")
print(f" bytes : {size/1048576:.2f} MB")
if __name__ == "__main__":
main()