Idle listening now runs PLAIN capture (no VPIO, zero ducking — music
untouched); the app arms the voice-processing unit only while its own
TTS is audibly playing (+0.8s tail), which is the only window where echo
cancellation matters. Options changes rebuild running input units
(AudioUnitAccess.last_input_options); WindowVoiceInput/VoiceWave gain
set_echo_cancellation; route app polls speech playback at 4Hz.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- bridge-dz baked for the entire Netherlands (17026 tiles, 29MB):
AHN-refined over the 8 Amsterdam sheet pairs, solver-only elsewhere.
Baked as 4 north-south strips with 2-tile overlap — the full-bbox
global solve peaked past 60GB RSS — merged with the new mbtiles-merge
subcommand (later input wins on overlaps, block-major write order).
Route app bridge_dz path ams -> nl.
- audio_unit: the new VoiceInput (VoiceProcessingIO) kind takes the input
setup/handler paths (set_input_handler panicked and aborted the app);
VPIO capture forced mono on every platform.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- ggml: read-only mmap module (unix-gated) + two-region Context (mapped
weights / dirty caches) + segmented Metal buffer binding; llama loads
GGUF weights as file-backed clean pages (jetsam-exempt) with owned-arena
fallback (MAKEPAD_LLAMA_NO_MMAP=1). Route app dirty footprint 12GB -> 4-7GB;
model load becomes lazy page-in; A/B byte-identical on 4B + 9B.
- llama: fix graph-cache keying corruption — a graph keyed wider than the
KV cache corrupted attention for any prefill batch >= 2 (flash op reads
permute-node dims baked at build; view reconfigure never reached the
kernel; masks were written cache-narrow). Masks now always fill the full
graph key width and graphs key by 1024-buckets; reconfigure path removed.
Verified byte-exact vs per-length reference across batch 1/2/8/64/512,
short+long prompts, mmap on/off, plus a two-session concurrency probe.
- voice: passive VoiceWaves no longer register the global audio-input
callback (the invisible caption-bar wave stole mic audio — last
registrant wins — and spawned duplicate whisper workers); whisper back
to F16 default (voice Metal library has no quantized kernels; q5_0
failed every GPU matmul); raw transcripts render immediately.
- route: kokoro TTS voice output (speaker toggle; streams reply sentences,
announces nav maneuvers + arrival) with barge-in — voice activity on the
mic stops playback instantly; dispatcher context 8k -> 32k (hybrid KV is
12/48 layers, ~48KB/token); window caption bar suppressed under studio.
- llama-generate: --max-context/--prefill-batch-size + state fingerprints;
new llama_concurrent_probe bin.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>