makepad/tools/qwen38/tok_parity.ps1
Admin 384d0e031c tools: the Builder replaces makepad_loader, the web server moves to makepad/webserver, fleet scripts, docs and the workspace members
tools/makepad_builder replaces tools/makepad_loader: one build target
shared across app builds, workspace package selection, checkout
progress on the public Git API, detached built apps with a completion
state, waits for Windows security scans, manual retry after compiler
locks, dedicated-folder installer checks, catalog and runtime fixes.
tools/web_server and its scripts leave for github.com/makepad/webserver.
Arch USB clone/restore scripts, the qwen38 box scripts and the G-belt
serial test join tools/. docs/agents records the agent workflow and the
remote-control handoff protocol; AGENTS.md forbids vendored sources and
bulk imports. Cargo.toml lists apps/wm-dyn, libs/code_language,
libs/search, libs/tar, libs/loader_bundle and tools/makepad_builder,
and drops the two removed crates.

Squashed from work:
- Share Builder target across Makepad app builds
- Fix Builder workspace package selection
- Align Builder checkout progress with public Git API
- Detach built apps and show completion state
- Wait for Windows security scans
- Offer manual retry after Windows compiler locks
- docs: the agent workflow of record and the remote-control handoff protocol
- builder: dedicated-folder installer checks, catalog and runtime fixes; Windows job objects hold c_void handles
- tools: Arch USB clone/restore scripts, the qwen38 box scripts, and the G-belt serial test
- tools: the web server moves to makepad/webserver
- AGENTS.md: no vendored sources or bulk imports in the tree

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 12:17:59 +02:00

40 lines
1.8 KiB
PowerShell

# Tokenizer parity: HF AutoTokenizer (pinned snapshot) vs llama-tokenize on the produced GGUF
$ErrorActionPreference = 'Continue'
$dir = 'C:\ai\models\qwen38'
$tf = "$dir\tok_test.txt"
$lines = @(
'Hello, world! The quick brown fox jumps over 12345 lazy dogs.',
'Streaming JSON {"key": [1, 2.5, null]} plus code: fn main() { println!("hi"); }',
'Unicode: caf'+[char]0x00E9+' na'+[char]0x00EF+'ve '+[char]0x4E2D+[char]0x6587+[char]0x6D4B+[char]0x8BD5+' '+[char]0xD83D+[char]0xDE80+[char]0xD83C+[char]0xDF0D+' end',
'<|im_start|>plain text with special-looking tokens<|im_end|> and <think> tags'
)
[System.IO.File]::WriteAllText($tf, ($lines -join "`n"), (New-Object System.Text.UTF8Encoding($false)))
$py = @'
import json, subprocess, sys
dir_ = r"C:\ai\models\qwen38"
tf = dir_ + r"\tok_test.txt"
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained(dir_ + r"\hf")
text = open(tf, encoding="utf-8").read()
hf_ids = tok.encode(text, add_special_tokens=False)
out = subprocess.run([r"C:\ai\qwen38\bin\llama-tokenize.exe", "-m", dir_ + r"\Qwen3.8-27B-Q4_K_M.gguf",
"-f", tf, "--ids", "--no-bos"], capture_output=True, text=True)
raw = out.stdout.strip()
try:
ll_ids = json.loads(raw)
except Exception:
ll_ids = [int(x) for x in raw.replace("[", " ").replace("]", " ").replace(",", " ").split()]
print("hf_n", len(hf_ids), "llama_n", len(ll_ids))
if hf_ids == ll_ids:
print("TOKENIZER_PARITY PASS")
else:
print("TOKENIZER_PARITY FAIL")
for i, (a, b) in enumerate(zip(hf_ids, ll_ids)):
if a != b:
print("first_diff_at", i, "hf", hf_ids[max(0,i-3):i+3], "llama", ll_ids[max(0,i-3):i+3])
break
if out.stderr:
print("stderr_tail", out.stderr[-400:])
'@
Set-Content -Path $dir\tok_parity.py -Value $py -Encoding UTF8
python $dir\tok_parity.py