hal0 exposes two voice directions through OpenAI-compatible endpoints:
- Speech-to-text (
POST /v1/audio/transcriptions) — served from the
single canonicalsttslot, which can run either engine: Moonshine
(ONNX) on CPU, or whisper-v3:turbo on the XDNA NPU, co-loaded with the
chat model in one FLM process (the “NPU trio”: chat + STT + embed).
Choosing a device swaps which engine thesttslot runs — there’s no
separate slot to configure. - Text-to-speech (
POST /v1/audio/speech) — served from the single
canonicalttsslot, which can run either engine: Kokoro-82M ONNX
on CPU, or Qwen3-TTS on GPU (ROCm, native gfx1151). Choosing a device
swaps which provider thettsslot runs — there’s no separate slot to
configure.
Both are children of the voice capability: voice.stt maps to the
stt slot and voice.tts maps to the tts slot. Both children are
now device-keyed engine switches with the same shape. Before this
redesign, voice.stt could only ever run backend=npu — any host
without the XDNA NPU had zero speech-to-text. Reinstating Moonshine as a
CPU-side STT engine and keying stt on device the same way tts
already was closed that gap (see ADR-0001, “Reinstate Moonshine as the
CPU STT engine”, docs/adr/0001-moonshine-cpu-stt-reinstatement.md in
the repository — an internal design record, not published to these
docs).
Enable the voice capability
hal0 capabilities set voice stt --model <your-stt-model>
hal0 capabilities set voice tts --model <your-tts-model>
Or over the API, which accepts any subset of { backend, provider, model, enabled }:
Speech-to-text (CPU, Moonshine)
curl -X POST http://localhost:8080/api/capabilities/voice/stt \
-H 'Content-Type: application/json' \
-d '{"enabled": true, "backend": "cpu", "provider": "moonshine"}'
Speech-to-text (NPU, FLM)
curl -X POST http://localhost:8080/api/capabilities/voice/stt \
-H 'Content-Type: application/json' \
-d '{"enabled": true, "backend": "npu", "provider": "flm"}'
Text-to-speech (CPU, Kokoro)
curl -X POST http://localhost:8080/api/capabilities/voice/tts \
-H 'Content-Type: application/json' \
-d '{"enabled": true, "backend": "cpu", "provider": "kokoro"}'
Text-to-speech (GPU, Qwen3-TTS)
curl -X POST http://localhost:8080/api/capabilities/voice/tts \
-H 'Content-Type: application/json' \
-d '{"enabled": true, "backend": "gpu-rocm", "provider": "qwen3tts"}'
The voice.stt engine switch
| Device | Provider | Runs on |
|---|---|---|
cpu |
moonshine |
CPU (ONNX) |
npu |
flm |
XDNA NPU, coresident with chat + embed in the FLM trio |
There is deliberately no GPU row — hal0 ships no GPU STT engine. This is a
two-engine, two-device special case, not a general provider selector: if
a second CPU (or GPU) STT engine ever lands, device alone stops being
enough to disambiguate and provider must become a first-class part of
capability selection — this future-work trigger is recorded in ADR-0001
(docs/adr/0001-moonshine-cpu-stt-reinstatement.md, repo-only).
The voice.tts engine switch
| Device | Provider | Runs on |
|---|---|---|
cpu |
kokoro |
CPU (ONNX) |
gpu-rocm |
qwen3tts |
Strix Halo iGPU (ROCm, native gfx1151) |
Selecting a device rewrites which provider the tts (or stt) slot
loads — the slot itself doesn’t move. Kokoro, Qwen3-TTS, and Moonshine
weights are all operator-staged (not pulled through the hal0 registry;
whisper-v3:turbo on the FLM/NPU path resolves through FastFlowLM’s own
model handling instead). Moonshine’s weights preflight at slot spawn — a
missing or empty bundle fails loudly, by name (slot.weights_missing),
rather than starting a container that 500s on the first request. The
tts slot additionally mounts a writable cache directory for
kernel/tokenizer caches.
Staging the Moonshine weights
The bundle is a directory the loader reads directly — it resolves
encoder_model.ort (or .onnx) plus decoder_model_merged.* and the
tokenizer inside whatever directory it is handed. Two consequences:
- Stage a non-streaming variant. The
base-enbundle has the file
names the loader wants. The streaming bundles (small-streaming-en,
medium-streaming-en) ship a different file set (encoder.ort,
frontend.ort,adapter.ort) belonging to the Moonshine streaming SDK,
which this image does not consume. Pointing the slot at one is rejected
at spawn withslot.weights_missingnaming the streaming case. - Point
--model_pathanywhere sensible. hal0 walks the tree to find
the directory holding the encoder, so both…/base-enand
…/base-en/quantized/base-enwork. If the resolved path is outside the
model store root, the slot mounts it read-only at the same absolute path
(plus the symlink target, if any) so the container can read it.
Verify staging before enabling the capability:
hal0 doctor all --json | jq '.[] | select(.key == "stt-weights")'
A pass names the staged path; a fail names exactly what is missing.
How the CPU STT path works (Moonshine)
Enabling voice.stt with backend=cpu spawns the stt slot’s own
container from the hal0-toolbox-moonshine image (MoonshineProvider),
distinct from the NPU path below — it’s a standalone process, not a
co-loaded FLM role. The container wraps an OpenAI-compatible FastAPI
server (POST /v1/audio/transcriptions, GET /v1/models, GET /health).
It also exposes WS /v1/audio/stream for live PCM16 (16 kHz mono)
streaming transcription — this endpoint is in-container only and is
not routed by the hal0 dispatcher; it isn’t part of the public /v1/*
surface. hal0 binds 0.0.0.0:8080 with no built-in auth by design
(LAN-trust posture); voice endpoints accept file uploads on that
unauthenticated bind.
How the NPU STT path works
Enabling voice.stt with backend=npu does not spawn a standalone
process — it drives the FLM trio, one flm serve anchor process serving
chat, transcription, and embeddings together. The orchestrator toggles
the anchor’s [npu].asr flag and writes a type=transcription slot
record for dispatch gating. The anchor is not auto-restarted — the
response carries pending_reload: true, and the change takes effect once
you reload the FLM anchor yourself (hal0 slot restart <anchor-slot>).
NPU transcription requires the FLM chat anchor to already be loaded.
Transcribe audio (speech to text)
curl -X POST http://localhost:8080/v1/audio/transcriptions \
-F 'file=@recording.wav' \
-F 'model=<your-stt-model>'
The model form field is required; omitting it returns 400
(request.missing_model).
Synthesize speech (text to speech)
curl -X POST http://localhost:8080/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{
"model": "<your-tts-model>",
"input": "Hello from hal0.",
"voice": "<voice-id>"
}' \
--output speech.wav
model is required. voice, speed, and response_format are
optional — see TTS request defaults below.
TTS request defaults
The tts slot can carry persisted request defaults that
/v1/audio/speech seeds into a request whenever the body omits the
matching field — an explicit value in the request always wins:
default_voice— voice id injected when the request omitsvoice.default_speed— playback speed (0.25–4.0) injected when the
request omitsspeed.default_response_format—mp3/wav/opus/flac/pcm,
injected when the request omitsresponse_format.
These are set from the dashboard’s Settings → Voice panel and take effect
immediately — no container restart. The dashboard’s voice picker reads a
live voice list from GET /api/slots/tts/voices; if the tts slot is
cold or unreachable, it falls back to a built-in seed list.
See also
- Manage slots — the
sttandttsslots. - Choose models — model-owned launch
defaults now apply to voice models the same as chat models.