For headless servers, dedicated GPUs, or "I want one command" deployments. The docker image bundles the backend; the UI is served over HTTP and you open it in a normal browser.
Official images: ghcr.io/debpalash/omnivoice-studio
and palashdeb/omnivoice-studio on Docker Hub — same images, same tags.
Image ↔ version mapping
Tag What you get :latestRolling preview — latest commit on main, at or ahead of the last release. This is the preview channel; pin:stablefor production.:stableMost recent versioned release (updated on every v*git tag):0.5.1Exact release version :0.5Latest patch within the 0.5 minor :mainAlias of the same rolling mainbuild as:latest:sha-xxxxxxxSpecific commit (produced by manual workflow dispatch) :rocmAMD GPU (ROCm) build of the rolling preview — the ROCm analogue of :latest:stable-rocm,:0.5.1-rocm,:0.5-rocm,:sha-xxxxxxx-rocmROCm builds of the corresponding CUDA tags above Versioning rule: preview builds always come from
mainand never version-sort below:stable— upgrades flow naturally.Note on the update-channel toggle: The update-channel UI (Settings → About → Update channel) is part of the Tauri desktop app's built-in auto-updater. It does not apply to the Docker image — the Docker image is the headless web-server build. To update your Docker deployment, pull the new image tag and recreate the container (
docker compose pull && docker compose up -d).
Docker's NAT prevents the backend from proving that a browser is on the host,
so server-mode settings and diagnostics require an administrator API key even
when the published port is loopback-only. Generate one before using any Studio
profile or docker run command below:
export OMNIVOICE_API_KEY="$(python3 -c 'import secrets; print(secrets.token_urlsafe(32))')"Keep that shell open until the container starts. The web UI asks for this key and exchanges it for a short-lived browser session; it does not persist the master key.
docker pull ghcr.io/debpalash/omnivoice-studio:latest
docker run -d --name omnivoice \
-p 127.0.0.1:3900:3900 \
-e OMNIVOICE_API_KEY="$OMNIVOICE_API_KEY" \
-v omnivoice-data:/app/omnivoice_data \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/debpalash/omnivoice-studio:latestDocker Hub mirror: the same images are published to
palashdeb/omnivoice-studioon Docker Hub with identical tags — swap the image forpalashdeb/omnivoice-studio:latestif you prefer Docker Hub. Tag semantics (:latest= rolling main preview,:stable/:X.Y.Z= releases) are the same on both registries.
Open http://localhost:3900. The first run downloads
~2.4 GB of model weights — follow docker logs -f omnivoice to watch.
docker run -d --name omnivoice --gpus all \
-p 127.0.0.1:3900:3900 \
-e OMNIVOICE_API_KEY="$OMNIVOICE_API_KEY" \
-v omnivoice-data:/app/omnivoice_data \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/debpalash/omnivoice-studio:latestGPU mode requires the NVIDIA Container Toolkit on the host.
AMD GPUs use the dedicated :rocm image variant — the default (CUDA)
image runs CPU-only on AMD hardware. The ROCm userspace ships inside the
image; the host only needs the amdgpu kernel driver. Pass the GPU through
as plain device nodes (no container toolkit needed):
docker run -d --name omnivoice \
--device /dev/kfd --device /dev/dri \
-p 127.0.0.1:3900:3900 \
-e OMNIVOICE_API_KEY="$OMNIVOICE_API_KEY" \
-v omnivoice-data:/app/omnivoice_data \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/debpalash/omnivoice-studio:rocmWSL exposes AMD compute through /dev/dxg, not native Linux's /dev/kfd and
/dev/dri. First install ROCm and librocdxg in the WSL distribution and
confirm the host-side rocminfo lists the GPU. Then use the WSL-specific
bridge flags from AMD's librocdxg container contract:
docker run -d --name omnivoice \
--device /dev/dxg \
-v /usr/lib/wsl/lib/libdxcore.so:/usr/lib/libdxcore.so \
-v /opt/rocm/lib/librocdxg.so:/usr/lib/librocdxg.so \
-v /opt/rocm/share/rocdxg/dids.conf:/usr/share/rocdxg/dids.conf \
-e HSA_ENABLE_DXG_DETECTION=1 \
--cap-add SYS_PTRACE \
--security-opt seccomp=unconfined \
--ipc=host --shm-size 8G \
-p 127.0.0.1:3900:3900 \
-e OMNIVOICE_API_KEY="$OMNIVOICE_API_KEY" \
-v omnivoice-data:/app/omnivoice_data \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/debpalash/omnivoice-studio:rocmThe image currently uses ROCm 7.2.x, so HSA_ENABLE_DXG_DETECTION=1 is
required; AMD removed that requirement only in ROCk 7.13. The ptrace and
unconfined-seccomp flags weaken container isolation, so keep the published port
on 127.0.0.1 and do not run untrusted workloads in this container. See AMD's
librocdxg WSL container instructions
for the driver/runtime compatibility matrix.
The same flags work with Podman (podman run --device /dev/kfd --device /dev/dri …); in a Quadlet unit that's two AddDevice= lines:
# ~/.config/containers/systemd/omnivoice.container
[Container]
Image=ghcr.io/debpalash/omnivoice-studio:rocm
AddDevice=/dev/kfd
AddDevice=/dev/dri
PublishPort=127.0.0.1:3900:3900
Volume=omnivoice-data:/app/omnivoice_data
Environment=OMNIVOICE_API_KEY=replace-with-a-long-random-keyRelease pins exist too: :stable-rocm, :0.5.1-rocm, :0.5-rocm mirror
the CUDA tags exactly.
Consumer cards and APUs (RX 6000/7000, Strix Point/Halo): the backend auto-sets
HSA_OVERRIDE_GFX_VERSIONwhen — and only when — your card's GFX ID is missing from the shipped ROCm build's architecture list, so try without any override first. Overriding a natively-supported GPU (gfx1151 on ROCm 7.x, for example) only forces it onto foreign kernels. If the GPU still isn't used, force it explicitly with-e HSA_OVERRIDE_GFX_VERSION=11.0.0(user-set on the container — it is deliberately not baked into the image, because the right value depends on your card); a value you set is always respected as-is.Rootless / non-root hosts: if
/dev/kfdis group-owned, the container user needs those groups too — add--group-addfor your host'srenderandvideoGIDs (getent group render video).
Verify the container sees the GPU:
docker exec <container> python3 -c \
"import torch; ok = torch.cuda.is_available(); print(ok, torch.cuda.get_device_name(0) if ok else 'unavailable')"Use omnivoice for the docker run examples above. Docker Compose names the
ROCm container omnivoice-studio-rocm (CPU: omnivoice-studio, NVIDIA:
omnivoice-studio-gpu); docker compose ps shows the exact active name.
(ROCm-built PyTorch reports through torch.cuda.* — True plus your card's
name means torch can see the GPU.) That check alone isn't proof the app is
using it: Settings → Performance & Device shows the device VoiceStudio
actually resolved.
Model Catalogue → Engines should report both omnivoice and
omnivoice-subprocess as accelerated on ROCm, rather than a CPU-fallback
warning.
If it reads cpu while the command above prints True, the backend log line
starting Falling back to CPU: names the architecture mismatch it hit.
The image installs and launches VoiceStudio through that same python3
interpreter. To verify this invariant on an older or custom image, compare
docker exec <container> python3 -c "import sys, torch; print(sys.executable, torch.version.hip)" with docker exec <container> sh -c 'tr "\\0" " " </proc/1/cmdline'; PID 1 must begin with python3 -m uvicorn.
If the command prints False, run Settings → About → Run self-check;
the GPU row says why. Native Linux has three common answers:
| What it says | What to do |
|---|---|
/dev/kfd is not present |
The container was started without --device /dev/kfd --device /dev/dri, or the host's amdgpu driver isn't loaded. |
this process cannot open it |
A group problem. Run ls -l /dev/kfd /dev/dri/render* on the host, and pass those GIDs with --group-add. The numbers differ between machines — a --group-add 39 copied from someone else's command grants nothing. |
no GPU was enumerated |
The device nodes are fine and the runtime still found nothing — usually a card newer than the image's ROCm. Check rocminfo on the host, and see the HSA_OVERRIDE_GFX_VERSION note above. |
On WSL, the self-check instead distinguishes a missing /dev/dxg permission,
the pre-7.13 HSA_ENABLE_DXG_DETECTION opt-in, and incomplete ROCDXG runtime
mounts.
# Generate this once in the shell that runs Compose.
export OMNIVOICE_API_KEY="$(python3 -c 'import secrets; print(secrets.token_urlsafe(32))')"
# CPU
docker compose -f deploy/docker-compose.yml --profile cpu up -d
# NVIDIA GPU
docker compose -f deploy/docker-compose.yml --profile gpu up -d
# AMD GPU (ROCm)
docker compose -f deploy/docker-compose.yml --profile rocm up -dThe docker-compose.yml shipped in deploy/ defaults to 127.0.0.1:3900
on the host. The backend inside the container binds to 0.0.0.0 so the
host port mapping can forward — the host-side 127.0.0.1 binding is what
enforces loopback-only.
To lend a headless GPU to VoiceStudio running on another machine, generate a join code on that control plane and start one of the worker profiles:
# NVIDIA
OMNIVOICE_WORKER_TOKEN='ovw_…' docker compose \
-f deploy/docker-compose.yml --profile worker-gpu up -d
# AMD / ROCm
OMNIVOICE_WORKER_TOKEN='ovw_…' docker compose \
-f deploy/docker-compose.yml --profile worker-rocm up -dThese profiles publish no HTTP port and require no browser UI. The join code
must advertise a LAN or private-overlay address the container can reach, not
the control plane's 127.0.0.1. Enrollment state persists in a dedicated
volume, so the container reconnects after a restart even though the join code
is single-use. Container health becomes green only after the control plane
accepts that registration; a missing or invalid token stays unhealthy instead
of reporting the generic web backend as ready. See Remote GPU
workers for enrollment, approval, routing, and security
details.
To expose VoiceStudio on your LAN (e.g. you're running it on a homelab box and opening the UI from a laptop), change the host port mapping:
# deploy/docker-compose.yml
services:
omnivoice:
ports:
- "0.0.0.0:3900:3900" # ← was 127.0.0.1:3900:3900The VoiceStudio frontend defaults to the same origin the page was served
from, so opening the UI from http://<lan-ip>:3900 Just Works for both the
page load and the API/media requests it makes afterwards.
If you front the app with a reverse proxy and the API and UI land on
different origins, pin the API base explicitly. Use OMNIVOICE_PUBLIC_API_BASE
— a runtime env var the backend injects into the page, so it works with the
prebuilt image via docker run -e (the older VITE_OMNIVOICE_API is inlined at
build time and cannot be set on a prebuilt image):
docker run -e OMNIVOICE_API_KEY="$OMNIVOICE_API_KEY" \
-e OMNIVOICE_PUBLIC_API_BASE=https://api.your-host.example \
-p 0.0.0.0:3900:3900 \
ghcr.io/debpalash/omnivoice-studio:latest
OMNIVOICE_PUBLIC_API_BASEmust be a plainhttp(s)://…URL; anything else is ignored and the app falls back to same-origin. If you build from source you may instead bakeVITE_OMNIVOICE_APIat build time, but the runtime var above is simpler and image-agnostic.
Security: Loopback-only publishing is the safe default. On a trusted LAN, set a long random
OMNIVOICE_API_KEYwithdocker run -eor Compose; the browser will prompt for it. The optional six-digit share PIN permits casual consumption access but does not authorize administration or dictation. On any untrusted network, plain HTTP is not safe for the API key or session cookie. Keep the backend on an encrypted private overlay such as Tailscale/ZeroTier; do not expose it directly to the public internet. See API authentication for the complete access model.
Two paths are worth persisting across container restarts:
| Mount | Purpose | Why |
|---|---|---|
omnivoice_data:/app/omnivoice_data |
Project DB, user voices, settings | Survives upgrade; encrypted HF token lives here |
~/.cache/huggingface:/root/.cache/huggingface |
HF model cache | Re-using your host's cache saves ~2.4 GB of re-downloads |
- Container reports 0.2.7 but image is tagged 0.3.x: This was a workflow bug
(fixes #249, #251) — the
:latesttag was not being updated on release tag pushes. Pull the image again after the fix is merged:docker pull ghcr.io/debpalash/omnivoice-studio:latest. The running version is now shown in Settings → About → Version (read live from the backend), so the web UI no longer displays a dash in Docker. - Checking which version is running:
docker exec <container> python3 -c "import importlib.metadata; print(importlib.metadata.version('omnivoice'))", or hit the/healthendpoint — it returns{"status": "ok", "device": ..., "version": "0.3.x"}. Use the container name listed bydocker compose ps(oromnivoicefor thedocker runexamples). - Watching startup: the port answers within about a second of container
start, but heavy initialization (PyTorch, API routes, database migration)
continues in the background. During that window
/healthreturns 503 with the current step, andGET /startup/progressreturns the full step-by-step ledger (status, currentstep/label, per-step states) — useful when a start seems slow and you want to see where it actually is. The DockerHEALTHCHECKflips healthy only once/healthis 200. - "Loopback origin required" errors (and a blank version): the desktop
build restricts the
/system/*and/api/settings/*routes to a loopback origin, but Docker's NAT makes every request look non-loopback, so the gate used to 403 the whole admin UI (issue #261). The image now ships withOMNIVOICE_SERVER_MODE=1, which relaxes that gate for the headless deployment. Admin mutations still requireOMNIVOICE_API_KEY; all commands above pass it into the container, and the UI prompts for it on first use. Exposure is governed by your-pport mapping (keep the127.0.0.1:prefix to stay local) plus authentication. If you front the container with your own auth proxy on loopback, setOMNIVOICE_SERVER_MODE=0to re-enable the strict gate. - Media-preview 404 in LAN mode: see the LAN access section
above — the
window.location.hostfix shipped in v0.3. - GPU not detected (NVIDIA): verify
docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu22.04 nvidia-smisucceeds first. - GPU not detected (AMD): make sure you pulled the
:rocmtag (the default image is CUDA-only) and passed--device /dev/kfd --device /dev/dri. Check the container sees the card withdocker exec omnivoice rocminfo | grep -i gfx. On consumer cards, run without anyHSA_OVERRIDE_GFX_VERSIONfirst — the backend sets it itself when your card needs it, and overriding a natively-supported GPU only forces it onto foreign kernels. See Pull and run (AMD GPU / ROCm) above for when to set one by hand. - More entries: docs/install/troubleshooting.md.