Skip to content

About

Self-hosted speech-to-text and speaker diarization with an OpenAI-compatible API. NVIDIA Parakeet + streaming Sortformer, built to keep audio private.

Topics

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Self-Hosted ASR with Speaker Diarization

Self-hosted speech-to-text and speaker diarization with an OpenAI-compatible transcription API. Built to keep audio private.

  • NVIDIA Parakeet-TDT-0.6b-v2 for transcription
  • NVIDIA streaming Sortformer v2 for speaker diarization
  • Batch and realtime WebSocket APIs
  • Multi-GPU load balancing via nginx least_conn

Why this exists

Most transcription workflows send audio to a third-party API. That means the audio leaves your infrastructure, lands on someone else's servers, and you have limited visibility into how it is handled, retained, or logged. For medical, legal, and sensitive environments, that tradeoff is often unacceptable.

This repository is a complete stack that runs entirely on your own hardware. It pairs NeMo ASR with streaming Sortformer diarization, merges them into speaker-attributed turns, and provides separate batch and realtime transcription APIs. The audio-handling design is fully inspectable, and the privacy properties are enforced by architecture.

How audio is handled

  • Inference is local. Audio and transcripts move only between the containers in this stack. Image builds pull Python packages at build time. On first startup, each worker downloads its model from Hugging Face and caches it in a Docker volume. After the caches are populated, HF_HUB_OFFLINE=1 makes Hub calls fail locally.
  • Audio and transcripts are never written to persistent storage. /tmp is a tmpfs (a RAM-only filesystem that is wiped when the container stops) in every service. All audio data lives in tmpfs and is discarded when the request completes or the container stops. Transcripts exist only in the HTTP response and in memory for the life of the request: no database, cache, or file. nginx streams request bodies and never spills buffered responses to temporary files, and every container runs on a read-only root filesystem, so nothing can write outside tmpfs and the model-cache volumes.
  • Temp files are removed on every exit path. Paths are bound before the first write, so a failure mid-write is still cleaned up. Cleanup runs in finally blocks that tolerate an already-deleted file. The failure paths are covered by tests.
  • Useful logs without recording request content. nginx emits a bounded JSON access record containing request ID, client address, method, URI without query parameters, status, byte count and timings. It does not log bodies, transcripts, filenames, query strings, authorization headers or arbitrary headers. The orchestrator keeps Uvicorn access logs enabled; GPU-worker Uvicorn access logs are disabled while structured application logs remain on. A validated request ID correlates nginx, orchestrator, ASR and Sortformer records. Docker log retention is capped.

Architecture

                        ┌──────────────────────────────┐
   client               │  nginx :8080 (least_conn LB) │
     │                  └──────┬───────────┬───────────┘
     │ POST /v1/audio/…        │           │
     ├─────────────────────────┤           │
     │                         ▼           ▼
     │              ┌─────────────┐  ┌──────────────┐
     │              │ parakeet-   │  │ diarize-     │   1 of each
     │              │ gpuN :8000  │  │ gpuN :8000   │   per GPU
     │              │ (batch ASR) │  │ (Sortformer) │
     │              └─────────────┘  └──────────────┘
     │                         ▲           ▲
     │ POST /…/diarized  ┌─────┴───────────┴─────┐
     ├──────────────────▶│ orchestrator :8000    │  ASR ∥ diarize,
     │                   │ (align + merge)       │  then word-level merge
     │                   └───────────────────────┘
     │ WS /v1/realtime   ┌───────────────────────┐
     └──────────────────▶│ parakeet-rt-gpuN      │  streaming partial/final
                         │ (WebSocket ASR)       │  transcript events
                         └───────────────────────┘
  • One batch ASR worker + one diarization worker per GPU (plus one realtime worker per GPU if enabled).
  • nginx least_conn spreads requests across GPUs; each request runs on exactly one GPU.
  • The orchestrator fans a request out to ASR and diarization in parallel, then assigns each transcribed word to the nearest speaker segment and merges consecutive same-speaker words into turns.

Models

Role Model License
Transcription nvidia/parakeet-tdt-0.6b-v2 CC-BY-4.0
Diarization nvidia/diar_streaming_sortformer_4spk-v2 CC-BY-4.0

Models are downloaded directly from Hugging Face on first start and cached in a Docker volume.

Requirements

The stack supports one or more NVIDIA GPUs. The Compose generator emits one worker set per selected GPU.

Quickstart

# 1. Build the worker image (used by ASR, diarize, and realtime workers)
docker build -t parakeet-asr -f asr/Dockerfile asr/

# 2. Optional local env file — the stack runs without one
cp .env.example .env

# 3. (Optional) The committed docker-compose.yml covers 1 GPU with realtime
#    enabled. For a different GPU count or limits, regenerate:
NUM_GPUS=<your-gpu-count> ENABLE_REALTIME=1 python3 generate_compose.py > docker-compose.yml

# 4. Launch (first start downloads models; wait for healthy)
docker compose up -d
docker compose ps

# 5. Test
curl http://localhost:8080/health
curl -X POST http://localhost:8080/v1/audio/transcriptions -F "file=@sample.mp3"

The API is published on 127.0.0.1:8080 by default.

Security

There is no authentication and no TLS. Anyone who can reach the port can transcribe audio. The generated compose therefore binds to loopback; put a TLS-terminating, authenticating gateway in front before setting BIND_ADDR to anything else.

Services run non-root with read-only filesystems, dropped capabilities, bounded resources, and tmpfs-backed working storage. See SECURITY.md for the full threat model and how to report a vulnerability.

Access and structured logs go to Docker stdout/stderr so an operator can send them to a log collector or SIEM.

Configuration

The Compose generator is configured by environment variables: NUM_GPUS, ENABLE_REALTIME, BIND_ADDR, HOST_PORT, MAX_UPLOAD_MB, TMPFS_SIZE, WORKER_MEM_LIMIT, MAX_AUDIO_SECONDS, MAX_REALTIME_CONNECTIONS, MAX_CONCURRENT_REQUESTS, and MAX_SESSION_SECONDS. HF_HUB_OFFLINE and an optional read-only HF_TOKEN are read by Docker Compose from the shell or .env at launch. Set HF_HUB_OFFLINE=1 after the caches are populated. No service declares env_file, so those two settings are the only part of a local .env that reaches a container. Each setting is documented in generate_compose.py or .env.example.

API

All endpoints are served through nginx on port 8080. The batch transcription endpoint is OpenAI-compatible; diarization, combined transcription, and realtime streaming are custom extensions.

POST /v1/audio/transcriptions — batch transcription

OpenAI-compatible. file (multipart) plus optional response_format (json | verbose_json). verbose_json adds word-level timestamps:

curl -X POST http://localhost:8080/v1/audio/transcriptions \
  -F "file=@audio.mp3" -F "response_format=verbose_json"
{
  "text": "…", "duration": 3612.4, "realtime_factor": 1445.0,
  "processing_time": 2.5,
  "words": [{"word": "Hello", "start": 0.12, "end": 0.31}, …]
}

Timing values are illustrative, not benchmark results.

POST /v1/audio/diarizations — speaker diarization

file plus optional response_format=verbose_json (adds per-speaker talk-time stats). num/min/max_speakers are accepted for API compatibility but ignored (Sortformer auto-detects).

{
  "segments": [{"speaker": "SPEAKER_00", "start": 0.08, "end": 4.31}, …],
  "speakers": ["SPEAKER_00", "SPEAKER_01"], "num_speakers": 2
}

POST /v1/audio/transcriptions/diarized — combined

Transcribes and diarizes in parallel, then merges into speaker-attributed turns:

curl -X POST http://localhost:8080/v1/audio/transcriptions/diarized \
  -F "file=@audio.mp3"
{
  "text": "[SPEAKER_00] Hello, how are you today?\n[SPEAKER_01] …",
  "turns": [{"speaker": "SPEAKER_00", "start": 0.12, "end": 2.84, "text": "…"}, …],
  "words": [{"word": "Hello", "start": 0.12, "end": 0.31, "speaker": "SPEAKER_00"}, …]
}

WS /v1/realtime — streaming transcription

Each WebSocket connection is pinned to one worker and carries one session. Client → server messages:

Type Payload
session.start {"type": "session.start", "language": "en", "partials": true}
audio.append {"type": "audio.append", "audioBase64": "<PCM16 mono>", "sampleRate": 16000}
session.stop {"type": "session.stop"}

Server → client: session.started, transcript.partial (unstable tail), transcript.final (committed text with word timestamps), turn.end (silence detected), session.stopped, and error. Inference runs every 500 ms on the uncommitted buffer; text stable for 2 cycles is committed as final. sampleRate must be between 8000 and 192000. A frame may contain at most two seconds of decoded audio, and pending audio is bounded. Session duration and per-worker connection limits are configurable. Diarization is not run on the realtime path.

Every HTTP response and WebSocket connection receives an X-Request-ID. A lowercase 32-hex value supplied by a trusted upstream is preserved; other values are replaced. The orchestrator forwards the same ID to both GPU legs.

Repo layout

asr/            Parakeet batch API, shared Dockerfile, and resolved dependency lock
diarize/        Sortformer streaming diarization API
realtime/       Parakeet streaming WebSocket API
orchestrator/   Transcription+diarization service and resolved dependency lock
nginx/          Generated load-balancer config — edit generate_compose.py, not this
scripts/        Batch-processing helpers (these WRITE TRANSCRIPTS to disk)
tests/          Unit tests — no GPU required
.github/        CI: tests, dependency audit, compose generation, secret scan
generate_compose.py   Emits docker-compose.yml + nginx/nginx.conf for N GPUs

Worker API scripts are bind-mounted into containers, so code changes only need a container restart. Container shutdown allows in-flight work to finish during the configured grace period; Docker sends SIGKILL only if the process does not exit before that period expires. Working audio remains on tmpfs throughout shutdown.

Testing

pip install -r tests/requirements.txt
pytest tests/ -q

Implementation notes

  • Sortformer config: workers run the model card's "very high latency" streaming preset (chunk_len=340, right_context=40, fifo_len=40, update_period=300, spkcache_len=188), the most accurate setting, with bounded memory per 27.2 s chunk.
  • Long audio: batch transcription chunks audio at 120 s with a 2 s overlap so words aren't cut at the boundary; a small amount of speech may repeat around chunk boundaries. Word timestamps are only decoded when verbose_json is requested. Their cost depends on the pinned NeMo/Torch build. Diarization streams, so VRAM stays flat on hour-plus files.
  • Health: each worker exposes its own /health with model status. GET :8080/health is an nginx liveness check only. Use docker compose ps for worker health.

License

Code: Copyright 2026 palmER, licensed under Apache-2.0. See NOTICE for attribution that must be preserved in redistributions.

Models are downloaded from Hugging Face and are licensed by NVIDIA under CC-BY-4.0: Parakeet-TDT-0.6b-v2, diar_streaming_sortformer_4spk-v2.

About

This is the transcription stack behind palmER, the AI scribe and MDM documentation platform built by emergency medicine physicians. The design rationale is in We open-sourced our transcription stack.

About

Self-hosted speech-to-text and speaker diarization with an OpenAI-compatible API. NVIDIA Parakeet + streaming Sortformer, built to keep audio private.

Topics

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Used by

Contributors

Languages