makepad/libs/chat_ui/tests/live_feed.rs
Admin e37c263b9c ai backbone: the hub era — makepad_ai deleted, every backend is a hub pipe; machine residency elections, job leases, ETA placement; the store only stores; creator pipelines run in the app (aicore)
Squashed from work:
- asset-ai: FastH3 4-step fast video backend; clip keyframes on the wire
- asset-ui: loop video chains — text→image→video that ends where it began
- h3: safetensors -> pruned-Q4_K GGUF quantizer for the 24GB DiT tiers
- h3_quant_gguf verify: row-error gates calibrated to the measured Q4_K floor
- asset-ai realtime: the feedback loop — the source anchors, the drifted frame inits
- asset-ai realtime: a feedback loop survives a resize and travels by default
- asset-ai realtime: the feedback loop frees itself from the feed handshake and pauses for its listener
- asset-ai realtime: the outbound encode leaves the loop's critical path
- asset-ai ocr: the ocr domain — Chandra 2 at page resolution, and the tower goes planner-owned
- llm slots: a lane can hold an image span — embedding prefill and a rope cursor of its own
- vision tower on CUDA: the encode leg gets its two missing kernels
- llm/ocr: one M-RoPE grid encoder for both image paths, and a livelock made an error
- vision tower on CUDA: the f16 GEMM keeps the precision it was throwing away
- live: a feed that moves box takes its trip with it — one seed image
- vision tower on CUDA: the tiled attention becomes bit-exact, and tensor cores go
- llm prefill on CUDA: the MMA attention kernel gets the tile a 4-to-1 model needs
- asset-ai ocr: the CUDA encode lane joins the integration — vision-parity sits beside run's three arms, and the kernels
- Merge branch 'ocr-perf-integration' into work
- asset-ai: the live anchor can follow the trip, and text leaves the 5090
- asset-ai: the camera moves the world, and the world starts still
- asset-import: the EA strategy classics, in the one 2D contract
- rtsmap: one seeded generator for tiled strategy maps
- asset-ui: one card for the strategy classics, with a pack dropdown
- asset-ai: music3 reference-audio path, ocr/h3 backends, registry
- asset: mp4 sample index for range-streaming, chat tools, import profiles
- cnc: tiberium is twelve growth frames, not twelve empty variants
- platform: native file and save dialogs, in-house on all three desktops
- chat: the scan holds out for a lane home
- chat: a full home queues you — take the free lane
- chat: the preload has a percentage, and the boundless cap stops showing
- llm cuda: the 32x2 attention tile — even GQA ratios stay on MMA
- sa3 gets a bake path: the sfx model's tables precomputed by a diffusion-side bin
- sqlite_query: anti-join regression test
- td import: HARV's second frame block is its harvesting cycle, not a turret
- asset-ui: sprite enhancement runs on the 32B dev DiT — distillation, not the prompt, was the ceiling
- ai-hub: makepad-asset-ai becomes makepad-ai-hub at libs/ai/hub, the chat pane becomes makepad-chat-ui, the service bin
- asset-ui: test health fixtures grow the realtime field they were born without
- ai-hub: one home at ~/.makepad — weights/ run/ cache/ logs/, the service cache migrates from ai_content by a single re
- ai-hub: subprocess workers die with the node — process groups everywhere, PDEATHSIG on linux, one KILL_ON_JOB_CLOSE Jo
- ai-hub: the hub object — AiHub::in_process, pipes vocabulary, and the local LLM engine generalized out of mpfiles (aic
- strict-json: the dependency-free JSON module gets its own crate; asset-client re-exports it so nothing downstream move
- ai-hub: the machine layer — node entries, the 0600 machine token, and the residency election that IS the lock (aicore
- ai-hub: MPHUB1 — the fabric beacon only dedicated nodes can send (aicore §4)
- ai-hub: job leases — work lives only while it is renewed (aicore §8)
- asset-creator: the pipeline library is born — specs, the deps gate, and the derived-state law (aicore §9)
- ai-hub: RAM residency facts — the CPU-side twin of residency.rs (aicore §3)
- ai-hub: ETA placement primitives — relative GPU throughput, the four-term estimate, and an observable breakdown (aicor
- ai-hub: leases go live on the wire — origin fields on submit, /job/<id>/keepalive, /bye, and the reaper that cancels w
- ai-hub: the chat providers move in — fleet qwen, openai, grok, claude/codex/grok CLIs, the responses driver, and the w
- asset-creator: the engine — one pipeline run against the hub, deps-gated, spliced, cancellable, resumable-by-construct
- ai-hub: the machine node mode — --machine binds loopback, registers in ~/.makepad/run, and exits on its own once idle
- asset-creator: makepad-creator-run — the detached client for runs that must outlive a window (aicore §9)
- ai-hub: a native Claude Messages-API provider — API-key or Claude Code OAuth, bounded SSE streaming, injected tools (a
- route + converse: off makepad_ai — the Agent seam moves to converse, route's cloud dispatcher rides the hub's Claude p
- asset-creator: the preset tables move in — fifteen chain-policy constants shared by every creator app (aicore §9 / P6)
- makepad_ai is deleted — every backend is a hub pipe, the agent seam lives with its consumers (aicore §14, decided 2026
- ai-hub: loads hold the machine residency election — set_model_state claims on Loaded and publishes the service port (a
- ai-hub: chats run the machine election — route to a serving holder, wait on a loading one, claim and publish when open
- ai-hub: pick_for_domain_eta — ETA-ranked placement over the shared hard-filter core (aicore §6 / P4)
- asset-creator: the engine picks a provider per stage at dispatch time — a chain's later stages see fresh fleet state (
- ai-hub: the fabric secret gates the service HTTP surface — bearer on everything but /health and the ticketed peer path
- vj: DREAM runs execute in the app — pipelines.rs becomes the run it used to watch (aicore §9 / F1)
- asset-creator: the runner — generate one thing and put it in the catalog, one implementation for every surface (aicore
- chat-ui: the session runs in the app — no broker anywhere on the chat path (aicore P8 / F5)
- asset-store: assets.query is a first-class query endpoint — the bounded SQL surface outlives the broker (aicore P8 / F
- asset-creator: CreatorTools — the chat tool pack for a store that only stores (aicore §9 / P8)
- asset-store: the shrink — the store stores (aicore P7)
- importer + asset-server host: the coordination era ends (aicore P7)
- store config purge + asset-ui goes fleet-direct; the derive protocol gets its route proof (aicore P7)
- client + chat dispatcher: the dead wire comes out (aicore P7/P8)
- ai-hub: 0.3.0 — the health version says which era a node runs
- ai-hub: the default fleet is 'gen' — apps hear the LAN without env plumbing
- ai-hub: the preload note percents the prefill, not the job bar
- ai-hub: conversations keep their KV — the wire mirror, the lane identity, the in-turn dynamic context (aicore §7)
- ai-hub: an open-think model is thinking from its first token
- libs: the zero-warning sweep — stitch casts say what they mean, xatlas keeps upstream's surface quietly
- zero-warning sweep, round two — the first full-workspace pass
- zero-warning sweep, round three — the model lanes and the deep examples
- zero-warning sweep, round four — the last stragglers
- zero-warning sweep, round five — vj and chat-ui
- zero-warning sweep, round six — three cascades

Co-authored-by: Claude <info@makepad.nl>
2026-09-01 16:46:31 +02:00

136 lines
5.4 KiB
Rust

//! The try-it harness: one REAL turn against a running Asset Server and a
//! real fleet box, with no window in the way.
//!
//! Ignored by default (it needs a live store and a live GPU). Run it with:
//!
//! ```text
//! MAKEPAD_LIVE_STORE=127.0.0.1:55463:55464 \
//! MAKEPAD_LIVE_TOKEN_FILE=local/asset-ui/asset-server/admin-token \
//! cargo test -p makepad-asset-chat-ui --test live_feed -- --ignored --nocapture
//! ```
//!
//! It proves the things a screenshot cannot: that the reply STREAMS, that
//! the rate meter reads the serving box's own token counts, and that
//! Escape's cancel really ends a turn in flight.
use makepad_chat_ui::feed::{ChatFeed, FeedConfig, NoClientTools};
use makepad_chat_ui::transcript::{ChatData, ChatRole, CHAT};
use makepad_asset_client::ApiEndpoints;
use std::net::SocketAddr;
use std::time::{Duration, Instant};
fn live_config(cache: &str) -> Option<FeedConfig> {
let spec = std::env::var("MAKEPAD_LIVE_STORE").ok()?;
let parts: Vec<&str> = spec.split(':').collect();
let [ip, control, data] = parts.as_slice() else {
panic!("MAKEPAD_LIVE_STORE must be ip:control:data");
};
let control: SocketAddr = format!("{ip}:{control}").parse().expect("control addr");
let data: SocketAddr = format!("{ip}:{data}").parse().expect("data addr");
let token = std::env::var("MAKEPAD_LIVE_TOKEN_FILE")
.ok()
.and_then(|p| std::fs::read_to_string(p).ok())
.map(|t| t.trim().to_string());
Some(FeedConfig::new(
ApiEndpoints { control, data },
token,
std::env::temp_dir().join(format!("mp_chat_ui_live_{}_{cache}", std::process::id())),
"gen",
"gen",
))
}
fn wait_until(what: &str, secs: u64, mut done: impl FnMut() -> bool) {
let deadline = Instant::now() + Duration::from_secs(secs);
while Instant::now() < deadline {
if done() {
return;
}
// The feed reports every failure as a system line; waiting out the
// full timeout on one is a slow way to read a message we already
// have.
if let Ok(data) = CHAT.read() {
if let Some(bad) = data.messages.iter().find(|m| m.role == ChatRole::System) {
panic!("the feed refused the turn: {}", bad.text);
}
}
std::thread::sleep(Duration::from_millis(50));
}
let data = CHAT.read().unwrap();
panic!(
"timed out waiting for {what}; activity={:?} streaming={} messages={:?}",
data.activity,
data.is_streaming,
data.messages.iter().map(|m| (m.role, m.text.clone())).collect::<Vec<_>>()
);
}
#[test]
#[ignore = "needs a live Asset Server and fleet box (MAKEPAD_LIVE_STORE)"]
fn a_live_turn_streams_and_reports_a_real_rate() {
let Some(cfg) = live_config("turn") else {
eprintln!("MAKEPAD_LIVE_STORE unset — nothing to talk to");
return;
};
ChatData::clear();
let feed = ChatFeed::start(cfg, Box::new(NoClientTools));
// Long enough to be MEASURED: a one-delta reply lands inside a single
// sample and has no honest rate to report.
feed.send("count from 1 to 40, one number per line, nothing else.".into(), Vec::new());
// The live readout exists WHILE it streams; that is the whole point of
// the meter (a number that only appears at the end teaches nothing).
// A cold 27B box spends real time on the first session: the provider
// probe, the assembled context, then prefill.
let mut live: Option<String> = None;
wait_until("the turn to land", 240, || {
if let Some(rate) = ChatData::live_rate_label() {
live = Some(rate);
}
!ChatData::is_streaming() && ChatData::item_count() > 1
});
println!("live rate while streaming: {live:?}");
assert!(live.is_some(), "the meter must read a rate DURING the reply");
let data = CHAT.read().unwrap();
let reply = data
.messages
.iter()
.rev()
.find(|m| m.role == ChatRole::Assistant)
.expect("an assistant reply");
println!("reply: {}", reply.text);
println!("meta: {:?}", reply.meta);
assert!(!reply.text.trim().is_empty());
let meta = reply.meta.as_deref().expect("a rate footnote on the landed reply");
assert!(meta.contains("tok/s"), "{meta}");
assert!(
!meta.starts_with('~'),
"a real serving box counts its own tokens — an estimate means the \
serving block never arrived: {meta}"
);
drop(data);
ChatData::clear();
}
#[test]
#[ignore = "needs a live Asset Server and fleet box (MAKEPAD_LIVE_STORE)"]
fn escape_ends_a_live_turn_in_flight() {
let Some(cfg) = live_config("cancel") else {
eprintln!("MAKEPAD_LIVE_STORE unset — nothing to talk to");
return;
};
ChatData::clear();
let feed = ChatFeed::start(cfg, Box::new(NoClientTools));
feed.send("write a long detailed essay about rust ownership.".into(), Vec::new());
// A cold 27B box spends real time on the first session: the provider
// probe, the assembled context, then prefill.
wait_until("the first token", 240, || {
CHAT.read().map(|d| !d.streaming_text.is_empty()).unwrap_or(false)
});
feed.cancel();
// Cancel is the broker's route, not a local flag: the turn has to STOP,
// and it has to stop soon enough to be worth pressing.
wait_until("the cancelled turn to stop", 30, || !ChatData::is_streaming());
println!("cancelled after: {:?}", ChatData::item_count());
ChatData::clear();
}