A small, reliable Python pipeline that turns raw audio/video into clean, structured training data for speech models. It acquires media, normalises the audio, runs speaker diarization ("who spoke when") and optional emotion/style tagging ("how it was spoken"), and writes RTTM / JSON / CSV outputs — designed to run unattended over many files.
source (URL or file)
│ yt-dlp FFmpeg
▼ ▼
download ───────────▶ extract 16 kHz mono WAV
│
pyannote / mock ▼
┌────────────────────────────────┐
│ diarization → who spoke when │
└───────────────┬─────────────────┘
│ segments
HF / mock ▼ (optional)
┌────────────────────────────────┐
│ emotion tagging → how spoken │
└───────────────┬─────────────────┘
▼
outputs: <id>.rttm <id>.json summary.csv
git clone <your-repo-url>
cd speech-data-pipeline
# create + activate a virtual environment
python -m venv .venv
source .venv/bin/activate # macOS/Linux
# .venv\Scripts\Activate.ps1 # Windows PowerShell
pip install -r requirements-dev.txt
pytest -vThe tests synthesise their own audio and use the mock backends, so this works with no HuggingFace token, no model download, and no audio files. (Three end-to-end tests also exercise FFmpeg; without FFmpeg installed they skip cleanly rather than fail.)
- Acquisition with
yt-dlp, plus support for local media files. - Audio normalisation with FFmpeg to 16 kHz mono 16-bit WAV (what diarization models expect).
- Speaker diarization via
pyannote.audio, behind a small interface so the backend is swappable. Amockbackend lets the whole pipeline (and the tests) run with no GPU, no token, and no model download. - Emotion & style tagging (optional) via a HuggingFace audio-classification
model, behind the same interface pattern. The default model is public (no token),
and a
mockbackend keeps it offline-testable. - Structured outputs: NIST RTTM, a rich JSON annotation per file, and a flat CSV summary.
- Built for unattended runs:
- Idempotent — a manifest records finished items; re-running skips them, so a crashed batch resumes without redoing work.
- Failure isolation — one bad input is logged and skipped, never killing the batch.
- Retries with backoff on downloads, and structured logging to console and file.
speech-data-pipeline/
├── run.py # CLI entry point
├── requirements.txt # runtime dependencies
├── requirements-dev.txt # tests + tooling
├── pytest.ini # test config (puts project root on sys.path)
├── urls.example.txt # input format: one URL/path per line
├── .env.example # documents HUGGINGFACE_TOKEN (no value committed)
├── SECURITY.md
├── .github/ # CI (tests, bandit, gitleaks) + Dependabot
├── src/
│ ├── acquire.py # yt-dlp download + FFmpeg extraction (Stage 1)
│ ├── diarize.py # Diarizer interface, pyannote + mock backends (Stage 2)
│ ├── emotion.py # EmotionTagger interface, HF + mock backends (Stage 2b)
│ ├── outputs.py # RTTM / JSON / CSV writers (Stage 3)
│ ├── pipeline.py # orchestration: batch, manifest/idempotency, isolation
│ └── utils.py # logging, stable IDs, filename helpers
└── tests/
└── test_pipeline.py # offline tests via the mock backends
- Python 3.10+
- FFmpeg on your PATH — required for the acquisition stage (
ffmpeg -versionto check).- macOS:
brew install ffmpeg - Linux:
sudo apt install ffmpeg - Windows:
winget install Gyan.FFmpeg(reopen the terminal afterwards so PATH updates)
- macOS:
pip install -r requirements.txtThis installs torch, pyannote.audio, and transformers — a large download.
Then provide a HuggingFace token via an untracked .env (never commit it):
cp .env.example .env # macOS/Linux (Windows: copy .env.example .env)
# then edit .env and set HUGGINGFACE_TOKEN=hf_...pyannote/speaker-diarization-3.1 is gated. While logged in to HuggingFace, accept
the terms on both model pages (access is granted instantly):
- https://huggingface.co/pyannote/speaker-diarization-3.1
- https://huggingface.co/pyannote/segmentation-3.0
The emotion model (superb/wav2vec2-base-superb-er) is public — nothing to accept.
A GPU is used automatically if available; otherwise everything runs on CPU.
Offline first — no token or model needed:
# add a local file or URL to urls.txt (one per line), then:
python run.py --input urls.txt --backend mock --emotion mockReal diarization, with emotion tagging:
python run.py --input urls.txt --backend pyannote --emotion hf --verboseUseful flags: --emotion {none,hf,mock} (tagging backend), --retries N
(download attempts), --force (reprocess items already in the manifest),
--verbose (debug logging), --work DIR (scratch dir).
For each input with id <id>:
outputs/rttm/<id>.rttm— one line per speech turn, NIST RTTM:SPEAKER <id> 1 <start> <duration> <NA> <NA> <speaker> <NA> <NA>outputs/json/<id>.json— full record: source, title, duration, speaker count, and the list of{start, end, duration, speaker}segments. When emotion tagging is on, each segment also carriesemotionandemotion_score, plus a top-levelemotion_model.outputs/summary.csv— one row per file (id, title, duration, #speakers, #segments, emotion model, dominant emotion).outputs/manifest.json/run_report.json— idempotency state and per-run status.outputs/logs/pipeline.log— full run log.
- Swappable model backends. Diarization and emotion tagging each sit behind a
tiny interface with a real backend and a
mockbackend. Adding the emotion stage required no change to the diarizer, and the mocks make the whole pipeline testable with no GPU/model/network. - In-memory audio decoding. Audio is read with
soundfileand passed to pyannote as an in-memory{waveform, sample_rate}tensor rather than a file path. This is pyannote's recommended path and sidesteps the optionaltorchcodecdecoder, which is brittle to install on some platforms. - Reliability over throughput. The orchestrator favours resumability (manifest), isolation (one failure never stops the batch), and observability (structured logs) so it can run unattended on large inputs.
pytest -vThe suite synthesises a short WAV with the standard library and runs the pipeline through the mock backends, so it needs no network, model, or audio libraries. It covers segment generation, all three output formats, idempotent re-runs, and failure isolation. The end-to-end tests use FFmpeg; if FFmpeg isn't installed they are skipped with a clear reason rather than failing.
- Parallel processing of the batch with a worker pool.
- Quality filtering — drop segments below a duration/SNR threshold before export.
- Speaking-rate / pitch features as additional per-segment annotations.
- No secrets in the repo — the only secret (a HuggingFace token) is read from
the environment or an untracked
.env. See SECURITY.md. - Secret scanning with
gitleaksin CI and as a pre-commit hook. - Static analysis with
banditin CI; Dependabot for weekly dependency updates. - Pre-commit hooks block large files, private keys, and merge markers:
pip install pre-commit && pre-commit install - Runtime artefacts (downloads, WAVs,
outputs/), virtualenvs, caches, and model weights are gitignored so they never bloat the repo.
Only download content you have the right to use. Configure your own sources in the input file before running.