Docs · Modes and capacity

Source docs/MODES-AND-CAPACITY.md in the ZoeyOS source. Last verified 2026-10-08. State current, fixed against the running house on 2026-10-08. The capacity numbers below were measured 2026-09-05 on the earlier engines and have not been re-measured since the reply lane moved to vLLM 0.30. The crew lane runs without speculative decoding, the voice is Chatterbox Turbo, and a heavy-render mode was added 2026-09-29.

Modes and capacity

Every number here came off the box. Nothing is a spec-sheet figure or a rule of thumb.

The box

Boardmini-ITX
CPUAMD Ryzen 9 9950X, 16 cores / 32 threads
RAM89 GiB usable
GPUNVIDIA RTX PRO 6000 Blackwell Workstation, 97,887 MiB
GPU power600 W limit; 90 W idle with both engines resident

One card. One desktop-sized case. Everything below runs on it at the same time: the engines, the voice, the face and presence rails.

How to read the page numbers

Legal prose was tokenised against the checkpoint's own tokeniser: the GPL-3 at 1.35 tokens a word, Apache-2.0 at 1.47, a mixed licence corpus at 1.37. Working figure: 1.4 tokens per word. Pages assume 500 words for a dense single-spaced contract page, 250 for a double-spaced pleading page, about 225 for a transcript page, and reserve 10,000 tokens for the system prompt, the tool schemas and the answer.

Standard: the daily house, and ordinary work

Two Qwen3.8-27B-NVFP4 engines, fp8 key-value cache, 8 seats each.

windowKV poolconcurrencydense pages
Zoey :8000131,07214 GiB / 363,488 tokens2.77x~173
Crew :8005102,40017.76 GiB / 443,301 tokens4.33x~132

Alongside them: the resident voice and the face and presence rails. Responsiveness with four concurrent callers: Zoey 207 ms to first token, 0.78 s median latency, 143 tokens a second aggregate; crew 75 ms, 0.67 s, 166 tokens a second.

Beast: when the document will not fit

One engine takes the whole card. The ops lane parks; Zoey's seats move to the other engine and she shares it with the crew. She keeps talking and hearing; the mode refuses to start if the voice is unhealthy, and reserves room for it.

Window262,144, the checkpoint's native ceiling
KV pool57.61 GiB = 1,653,294 tokens
Seats5, with 6.31x reported concurrency at a full window
Dense pages~360 per caller
Load3 min 26 s in; 2 min 56 s back to Standard

Proven, not projected: a 200,048-token brief with the clause buried at 50 % depth answered correctly in 107 s; five such briefs at once (1,000,290 tokens in flight) all correct in 8 min 14 s. Zero CUDA errors across the run.

The honest caveat. Ingest runs about 1,900 to 2,000 tokens a second whether you send one matter or five. The five prefills serialise. Concurrency buys queue depth and unattended batch, not speed. Five seats means five matters can be held at a full window and answered without babysitting. It does not mean five times faster, and it should never be quoted that way.

Render: the image engine takes the card

The ops lane parks, the render engine starts, the crew lane stays so the board continues. Zoey limps: the ear stays up; her voice depends on the recipe. Leaving Render is a cold boot of the ops lane, about two minutes. Recipes are sized from 10 GB to 36 GB. A separate heavy-render mode (2026-09-29) takes the crew off the card on purpose for the high-VRAM recipes.

All OS Off: an empty card

Parks both engines, the voice and the render engine. The ear and presence stay on the CPU. The house goes mute. Not a host power-off.

Said plainly

Standard handles a matter up to roughly 170 dense pages, and four callers can each hold a window that size at once. If the document is bigger than that, Beast takes you to about 360 dense pages: five readers, five matters, answered in one unattended pass.

Or in other units: about 800 transcript pages in Beast, about 384 in Standard.