Docs · Speech to speech
Source
docs/STT-TTS-ARCHITECTURE-20261007.md in the ZoeyOS source. Last verified 2026-10-07. State current.Speech to speech: the ear, the bridge and the voice
As they were running on 2026-10-07, read from the services themselves. The team's own notes were checked and not used as the source.
The path
room microphone (reSpeaker beam array), raw capture
│
▼
ear rail :8190 voice activity (TEN VAD), the WeSpeaker speaker stamp
│
▼
recogniser :8096 Audio8 ark-asr 0.6B, int8 ONNX, on the CPU, its own environment
│
▼
ear bridge :8194 the turn, the quiet wait, the speak; the reply brain is Zoey's gateway → ops :8000
│
▼
voice :7789 Chatterbox Turbo, 24 kHz PCM, watermark on
│
▼
PipeWire the speakers
announcer, separate: Audio8's TTS runtime :8024. Not the conversation voice.
Health at measurement, idle room: the ear on the beam microphone, TEN VAD at 0.5 (Silero installed beside it at 0.35, not the live switch), WeSpeaker as the voice encoder. The bridge waits two seconds of quiet before she speaks and flushes fast at 0.3 s.
The units
| voice | Chatterbox Turbo on :7789. Replaced CosyVoice the morning of 2026-10-07; the contract (health, stream, abort) did not move, so the bridge kept up without a restart. |
|---|---|
| recogniser | the ear worker on :8096, its own Python environment |
| ear rail | capture, VAD, stamp, HTTP to the recogniser, on :8190 |
| ear bridge | hears a turn, asks Zoey, speaks through :7789; abort on :8194 |
| announcer | Audio8's TTS runtime on :8024, for house lines, never her voice |
| ops | the reply brain: vLLM 0.30.0, Qwen3.8-27B-NVFP4, :8000 |
| crew | the other lane: vLLM 0.27.1, the same weight family, :8005 |
Versions read from the environments
| Chatterbox | chatterbox-tts 0.1.7, MIT, Resemble AI; Python 3.14.7, torch 2.14.1, transformers 5.2.0 |
|---|---|
| Perth watermark | resemble-perth 1.0.1, on |
| TEN VAD | 1.0.6.8, Apache 2.0 with additional conditions; one pitch file carries LPCNet's BSD code |
| Silero VAD | 6.2.1, installed, not the live switch |
| WeSpeaker | 0.0.1, Apache-2.0 |
| ear worker | Python 3.12.13, numpy 1.26.4 |
| PipeWire | 1.6.2 |
Two rules the ear keeps
- The ear skips the house's own voice before recognising. A line from a bench, or the announcer on air, is dropped before the recogniser sees it. A recogniser restart is an ear restart: the ear hears nothing while it is down.
- Length is coached, not capped. The spoken-length cap was reverted 2026-10-06; the running bridge carries no cap.
What this page does not decide
Loudness, the mouth tiles and the sheet lead live with the mouth instruments. This page is the ear, the bridge and the voice as they were running on 2026-10-07.