Docs · Modes and capacity
docs/MODES-AND-CAPACITY.md in the ZoeyOS source. Last verified 2026-10-08. State current, fixed against the running house on 2026-10-08. The capacity numbers below were measured 2026-09-05 on the earlier engines and have not been re-measured since the reply lane moved to vLLM 0.30. The crew lane runs without speculative decoding, the voice is Chatterbox Turbo, and a heavy-render mode was added 2026-09-29.Modes and capacity
Every number here came off the box. Nothing is a spec-sheet figure or a rule of thumb.
The box
| Board | mini-ITX |
|---|---|
| CPU | AMD Ryzen 9 9950X, 16 cores / 32 threads |
| RAM | 89 GiB usable |
| GPU | NVIDIA RTX PRO 6000 Blackwell Workstation, 97,887 MiB |
| GPU power | 600 W limit; 90 W idle with both engines resident |
One card. One desktop-sized case. Everything below runs on it at the same time: the engines, the voice, the face and presence rails.
How to read the page numbers
Legal prose was tokenised against the checkpoint's own tokeniser: the GPL-3 at 1.35 tokens a word, Apache-2.0 at 1.47, a mixed licence corpus at 1.37. Working figure: 1.4 tokens per word. Pages assume 500 words for a dense single-spaced contract page, 250 for a double-spaced pleading page, about 225 for a transcript page, and reserve 10,000 tokens for the system prompt, the tool schemas and the answer.
Standard: the daily house, and ordinary work
Two Qwen3.8-27B-NVFP4 engines, fp8 key-value cache, 8 seats each.
| window | KV pool | concurrency | dense pages | |
| Zoey :8000 | 131,072 | 14 GiB / 363,488 tokens | 2.77x | ~173 |
| Crew :8005 | 102,400 | 17.76 GiB / 443,301 tokens | 4.33x | ~132 |
Alongside them: the resident voice and the face and presence rails. Responsiveness with four concurrent callers: Zoey 207 ms to first token, 0.78 s median latency, 143 tokens a second aggregate; crew 75 ms, 0.67 s, 166 tokens a second.
Beast: when the document will not fit
One engine takes the whole card. The ops lane parks; Zoey's seats move to the other engine and she shares it with the crew. She keeps talking and hearing; the mode refuses to start if the voice is unhealthy, and reserves room for it.
| Window | 262,144, the checkpoint's native ceiling |
|---|---|
| KV pool | 57.61 GiB = 1,653,294 tokens |
| Seats | 5, with 6.31x reported concurrency at a full window |
| Dense pages | ~360 per caller |
| Load | 3 min 26 s in; 2 min 56 s back to Standard |
Proven, not projected: a 200,048-token brief with the clause buried at 50 % depth answered correctly in 107 s; five such briefs at once (1,000,290 tokens in flight) all correct in 8 min 14 s. Zero CUDA errors across the run.
The honest caveat. Ingest runs about 1,900 to 2,000 tokens a second whether you send one matter or five. The five prefills serialise. Concurrency buys queue depth and unattended batch, not speed. Five seats means five matters can be held at a full window and answered without babysitting. It does not mean five times faster, and it should never be quoted that way.
Render: the image engine takes the card
The ops lane parks, the render engine starts, the crew lane stays so the board continues. Zoey limps: the ear stays up; her voice depends on the recipe. Leaving Render is a cold boot of the ops lane, about two minutes. Recipes are sized from 10 GB to 36 GB. A separate heavy-render mode (2026-09-29) takes the crew off the card on purpose for the high-VRAM recipes.
All OS Off: an empty card
Parks both engines, the voice and the render engine. The ear and presence stay on the CPU. The house goes mute. Not a host power-off.
Said plainly
Standard handles a matter up to roughly 170 dense pages, and four callers can each hold a window that size at once. If the document is bigger than that, Beast takes you to about 360 dense pages: five readers, five matters, answered in one unattended pass.
Or in other units: about 800 transcript pages in Beast, about 384 in Standard.