A 7M-parameter dispatcher that knows when to say no.
Pure MFL. 8 MB. Your CPU.
$ curl localhost:8097/v1/route -d '{"state":"remind me to stretch every hour"}' {"action":"tool_call","route":"set_reminder","confidence":1.0,"latency_ms":6} $ curl localhost:8097/v1/route -d '{"state":"write me a 2000-word essay","execute":true}' {"action":"delegate","reason":"route_escalate","confidence":0.96, "via":"qwen3-1.7b","content":"**The Roman Empire: A Legacy of…"}
Low confidence → delegate. Escalation head fires → delegate. Head and trunk disagree → delegate. A dispatcher that guesses wrong confidently is worse than none — so it doesn't.
Linear heads read the frozen trunk's hidden state in one forward pass: route (16 intents), noul (should-I-escalate), score (complexity tier). No generation needed for a decision.
head_studio.py turns a config of route→example-phrases into a trained, evaluated .head — per-customer intents without fine-tuning the trunk. Demo helpdesk config: 100% eval acc, ECE 0.0002.
ANVIL_KEYS gates /v1/* behind Bearer auth; ANVIL_TENANTS maps keys to heads. Two customers, one trunk, different routers — under the model lock.
Every decision is logged; calibrate.py replays labeled traffic into per-bucket empirical ECE. On served traffic: ≥0.99 confidence → 98% right; <0.9 → coin flip, gated to delegate.
Tokenizer training, pretraining, fine-tune, int8 export, OpenAI-compatible CPU server — all machin/MFL, compiled to C. The appliance is a static binary plus an 8 MB model.
ANVIL_DELEGATE_URL points at any OpenAI-compatible upstream. Abstained requests forward after the model lock releases — a slow upstream never stalls the router.
# 7.5 MB, static binary + model + heads + systemd unit tar xzf mtlm-router-m7router3s384-linux-amd64.tar.gz cd mtlm-router-m7router3s384-linux-amd64 ./start.sh # → /v1/route on :8097