Apps ask the hub for a recognizer or a voice and get one; where it runs is the hub's decision. AiHub::start_stt / start_tts return poll-driven sessions shaped like the chat session. The Auto ladder is Whisper/Kokoro in this process (weights present, machine election), on the machine node over loopback, on a LAN node, else the OS engine; SpeechReach::Local is the "don't reach out" knob. Audio always comes back as PCM: the app owns the device. Three layers: - makepad-ai-speech is the whole speech model family, engines only. libs/voice (Whisper + Silero VAD) folds in as the `whisper` and `vad` modules next to kokoro and indextts, each a cargo feature; the Apple bridges and the Speaker/VoiceTranscriber selection leave it. - makepad-system-speech (new) is the OS speech services as blocking fns: Apple SpeechAnalyzer/AVSpeechSynthesizer via Swift, Windows.Media.Speech* on the vendored bindings, Android SpeechRecognizer/TextToSpeech through MakepadSpeech.java (API 26 floor), espeak-ng on Linux. It models the two STT shapes honestly: PCM in (Whisper, Apple) versus an engine that owns the microphone (Android, Windows), with capabilities the caller reads. - the hub grows speech sessions, in-process Whisper/Kokoro workers with the residency election, a `whisper` wire backend (stt domain, registry entry pinned to ggerganov/whisper.cpp) so a Mac can serve a Quest, and a `language` field on the generate request. Consumers: the Window voice input runs on an STT session and switches to engine-mic mode when the recognizer owns the microphone; converse's SpeechOutput is a lazily started TTS session plus a pump thread; route drops its private speech copy for converse; vj's lyrics fallback and the alignment bakes call the engines directly. Verified here: speech-roundtrip through the real sessions (Apple voice in, in-process Whisper on Metal out, 4.3% WER); system-speech-test TTS->STT verbatim; hub/converse/system-speech unit tests; msvc, aarch64-android and linux-gnu cross-checks; Java against android-34. Windows, Android and Linux bridges are compile-checked only. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
30 lines
1.4 KiB
TOML
30 lines
1.4 KiB
TOML
[package]
|
|
name = "makepad-converse"
|
|
version = "0.1.0"
|
|
edition = "2021"
|
|
description = "Conversational voice pipeline for Makepad apps: streamed agent replies spoken as they arrive, with a pluggable transcript filter between speech-to-text and the backend LLM."
|
|
license = "MIT OR Apache-2.0"
|
|
|
|
[features]
|
|
# Speech synthesis through the ai-hub (Kokoro or the OS voice). Default ON so existing consumers are unchanged,
|
|
# but separable: the model plus its embedded pronunciation lexicon is ~10 MB of
|
|
# binary and ~327 MB resident, which a build that can never speak should not
|
|
# carry. With it off, SpeechOutput still exists and every non-audio path
|
|
# behaves identically — the worker just drains its queue in silence.
|
|
default = ["tts"]
|
|
tts = ["dep:makepad-ai-hub"]
|
|
# The local filtering LLM (QwenFilter) on makepad-ai-llm; heavy, so opt-in.
|
|
local-llm = ["dep:makepad-ai-llm"]
|
|
|
|
[dependencies]
|
|
makepad-widgets = { path = "../../widgets" }
|
|
# The hub picks the voice: Kokoro in-process / on the machine node / on a LAN
|
|
# node, else the OS synthesizer (makepad-system-speech). No default features:
|
|
# the image/video/mesh backends are not this crate's business.
|
|
makepad-ai-hub = { path = "../ai/hub", default-features = false, features = ["tts", "speech"], optional = true }
|
|
makepad-ai-llm = { path = "../ai/llm", optional = true }
|
|
|
|
[[bin]]
|
|
name = "filter-repl"
|
|
path = "src/bin/filter_repl.rs"
|
|
required-features = ["local-llm"]
|