Self-hosted speech-to-text and speaker diarization with an OpenAI-compatible transcription API. Built to keep audio private.
- NVIDIA Parakeet-TDT-0.6b-v2 for transcription
- NVIDIA streaming Sortformer v2 for speaker diarization
- Batch and realtime WebSocket APIs
- Multi-GPU load balancing via nginx
least_conn
Most transcription workflows send audio to a third-party API. That means the audio leaves your infrastructure, lands on someone else's servers, and you have limited visibility into how it is handled, retained, or logged. For medical, legal, and sensitive environments, that tradeoff is often unacceptable.
This repository is a complete stack that runs entirely on your own hardware. It pairs NeMo ASR with streaming Sortformer diarization, merges them into speaker-attributed turns, and provides separate batch and realtime transcription APIs. The audio-handling design is fully inspectable, and the privacy properties are enforced by architecture.
- Inference is local. Audio and transcripts move only between the
containers in this stack. Image builds pull Python packages at build
time. On first startup, each worker downloads its model from Hugging
Face and caches it in a Docker volume. After the caches are populated,
HF_HUB_OFFLINE=1makes Hub calls fail locally. - Audio and transcripts are never written to persistent storage.
/tmpis a tmpfs (a RAM-only filesystem that is wiped when the container stops) in every service. All audio data lives in tmpfs and is discarded when the request completes or the container stops. Transcripts exist only in the HTTP response and in memory for the life of the request: no database, cache, or file. nginx streams request bodies and never spills buffered responses to temporary files, and every container runs on a read-only root filesystem, so nothing can write outside tmpfs and the model-cache volumes. - Temp files are removed on every exit path. Paths are bound before
the first write, so a failure mid-write is still cleaned up. Cleanup
runs in
finallyblocks that tolerate an already-deleted file. The failure paths are covered by tests. - Useful logs without recording request content. nginx emits a bounded JSON access record containing request ID, client address, method, URI without query parameters, status, byte count and timings. It does not log bodies, transcripts, filenames, query strings, authorization headers or arbitrary headers. The orchestrator keeps Uvicorn access logs enabled; GPU-worker Uvicorn access logs are disabled while structured application logs remain on. A validated request ID correlates nginx, orchestrator, ASR and Sortformer records. Docker log retention is capped.
┌──────────────────────────────┐
client │ nginx :8080 (least_conn LB) │
│ └──────┬───────────┬───────────┘
│ POST /v1/audio/… │ │
├─────────────────────────┤ │
│ ▼ ▼
│ ┌─────────────┐ ┌──────────────┐
│ │ parakeet- │ │ diarize- │ 1 of each
│ │ gpuN :8000 │ │ gpuN :8000 │ per GPU
│ │ (batch ASR) │ │ (Sortformer) │
│ └─────────────┘ └──────────────┘
│ ▲ ▲
│ POST /…/diarized ┌─────┴───────────┴─────┐
├──────────────────▶│ orchestrator :8000 │ ASR ∥ diarize,
│ │ (align + merge) │ then word-level merge
│ └───────────────────────┘
│ WS /v1/realtime ┌───────────────────────┐
└──────────────────▶│ parakeet-rt-gpuN │ streaming partial/final
│ (WebSocket ASR) │ transcript events
└───────────────────────┘
- One batch ASR worker + one diarization worker per GPU (plus one realtime worker per GPU if enabled).
- nginx
least_connspreads requests across GPUs; each request runs on exactly one GPU. - The orchestrator fans a request out to ASR and diarization in parallel, then assigns each transcribed word to the nearest speaker segment and merges consecutive same-speaker words into turns.
| Role | Model | License |
|---|---|---|
| Transcription | nvidia/parakeet-tdt-0.6b-v2 | CC-BY-4.0 |
| Diarization | nvidia/diar_streaming_sortformer_4spk-v2 | CC-BY-4.0 |
Models are downloaded directly from Hugging Face on first start and cached in a Docker volume.
- NVIDIA GPU or GPUs with a CUDA 12.x driver.
- Docker with the NVIDIA Container Toolkit
The stack supports one or more NVIDIA GPUs. The Compose generator emits one worker set per selected GPU.
# 1. Build the worker image (used by ASR, diarize, and realtime workers)
docker build -t parakeet-asr -f asr/Dockerfile asr/
# 2. Optional local env file — the stack runs without one
cp .env.example .env
# 3. (Optional) The committed docker-compose.yml covers 1 GPU with realtime
# enabled. For a different GPU count or limits, regenerate:
NUM_GPUS=<your-gpu-count> ENABLE_REALTIME=1 python3 generate_compose.py > docker-compose.yml
# 4. Launch (first start downloads models; wait for healthy)
docker compose up -d
docker compose ps
# 5. Test
curl http://localhost:8080/health
curl -X POST http://localhost:8080/v1/audio/transcriptions -F "file=@sample.mp3"The API is published on 127.0.0.1:8080 by default.
There is no authentication and no TLS. Anyone who can reach the port can
transcribe audio. The generated compose therefore binds to loopback; put a
TLS-terminating, authenticating gateway in front before setting BIND_ADDR to
anything else.
Services run non-root with read-only filesystems, dropped capabilities, bounded resources, and tmpfs-backed working storage. See SECURITY.md for the full threat model and how to report a vulnerability.
Access and structured logs go to Docker stdout/stderr so an operator can send them to a log collector or SIEM.
The Compose generator is configured by environment variables: NUM_GPUS,
ENABLE_REALTIME, BIND_ADDR, HOST_PORT, MAX_UPLOAD_MB, TMPFS_SIZE,
WORKER_MEM_LIMIT, MAX_AUDIO_SECONDS, MAX_REALTIME_CONNECTIONS,
MAX_CONCURRENT_REQUESTS, and MAX_SESSION_SECONDS.
HF_HUB_OFFLINE and an optional read-only HF_TOKEN are read by Docker Compose
from the shell or .env at launch. Set HF_HUB_OFFLINE=1 after the caches are populated.
No service declares env_file, so those two settings are the only part of a
local .env that reaches a container.
Each setting is documented in generate_compose.py or .env.example.
All endpoints are served through nginx on port 8080. The batch transcription endpoint is OpenAI-compatible; diarization, combined transcription, and realtime streaming are custom extensions.
OpenAI-compatible. file (multipart) plus optional response_format
(json | verbose_json). verbose_json adds word-level timestamps:
curl -X POST http://localhost:8080/v1/audio/transcriptions \
-F "file=@audio.mp3" -F "response_format=verbose_json"{
"text": "…", "duration": 3612.4, "realtime_factor": 1445.0,
"processing_time": 2.5,
"words": [{"word": "Hello", "start": 0.12, "end": 0.31}, …]
}Timing values are illustrative, not benchmark results.
file plus optional response_format=verbose_json (adds per-speaker
talk-time stats). num/min/max_speakers are accepted for API
compatibility but ignored (Sortformer auto-detects).
{
"segments": [{"speaker": "SPEAKER_00", "start": 0.08, "end": 4.31}, …],
"speakers": ["SPEAKER_00", "SPEAKER_01"], "num_speakers": 2
}Transcribes and diarizes in parallel, then merges into speaker-attributed turns:
curl -X POST http://localhost:8080/v1/audio/transcriptions/diarized \
-F "file=@audio.mp3"{
"text": "[SPEAKER_00] Hello, how are you today?\n[SPEAKER_01] …",
"turns": [{"speaker": "SPEAKER_00", "start": 0.12, "end": 2.84, "text": "…"}, …],
"words": [{"word": "Hello", "start": 0.12, "end": 0.31, "speaker": "SPEAKER_00"}, …]
}Each WebSocket connection is pinned to one worker and carries one session. Client → server messages:
| Type | Payload |
|---|---|
session.start |
{"type": "session.start", "language": "en", "partials": true} |
audio.append |
{"type": "audio.append", "audioBase64": "<PCM16 mono>", "sampleRate": 16000} |
session.stop |
{"type": "session.stop"} |
Server → client: session.started, transcript.partial (unstable tail),
transcript.final (committed text with word timestamps), turn.end
(silence detected), session.stopped, and error. Inference runs every 500 ms
on the uncommitted buffer; text stable for 2 cycles is committed as final.
sampleRate must be between 8000 and 192000. A frame may contain at most two
seconds of decoded audio, and pending audio is bounded. Session duration and
per-worker connection limits are configurable. Diarization is not run on the
realtime path.
Every HTTP response and WebSocket connection receives an X-Request-ID. A
lowercase 32-hex value supplied by a trusted upstream is preserved; other
values are replaced. The orchestrator forwards the same ID to both GPU legs.
asr/ Parakeet batch API, shared Dockerfile, and resolved dependency lock
diarize/ Sortformer streaming diarization API
realtime/ Parakeet streaming WebSocket API
orchestrator/ Transcription+diarization service and resolved dependency lock
nginx/ Generated load-balancer config — edit generate_compose.py, not this
scripts/ Batch-processing helpers (these WRITE TRANSCRIPTS to disk)
tests/ Unit tests — no GPU required
.github/ CI: tests, dependency audit, compose generation, secret scan
generate_compose.py Emits docker-compose.yml + nginx/nginx.conf for N GPUs
Worker API scripts are bind-mounted into containers, so code changes only need a container restart. Container shutdown allows in-flight work to finish during the configured grace period; Docker sends SIGKILL only if the process does not exit before that period expires. Working audio remains on tmpfs throughout shutdown.
pip install -r tests/requirements.txt
pytest tests/ -q- Sortformer config: workers run the model card's "very high latency"
streaming preset (
chunk_len=340,right_context=40,fifo_len=40,update_period=300,spkcache_len=188), the most accurate setting, with bounded memory per 27.2 s chunk. - Long audio: batch transcription chunks audio at 120 s with a 2 s
overlap so words aren't cut at the boundary; a small amount of speech
may repeat around chunk boundaries. Word timestamps are only decoded when
verbose_jsonis requested. Their cost depends on the pinned NeMo/Torch build. Diarization streams, so VRAM stays flat on hour-plus files. - Health: each worker exposes its own
/healthwith model status.GET :8080/healthis an nginx liveness check only. Usedocker compose psfor worker health.
Code: Copyright 2026 palmER, licensed under Apache-2.0. See NOTICE for attribution that must be preserved in redistributions.
Models are downloaded from Hugging Face and are licensed by NVIDIA under CC-BY-4.0: Parakeet-TDT-0.6b-v2, diar_streaming_sortformer_4spk-v2.
This is the transcription stack behind palmER, the AI scribe and MDM documentation platform built by emergency medicine physicians. The design rationale is in We open-sourced our transcription stack.