Skip to content

About

Python pipeline that turns raw audio/video into structured speech-training data: acquisition, speaker diarization, and emotion tagging, with RTTM/JSON/CSV outputs.

Topics

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Speech Data Pipeline — Diarization & Audio Annotation

A small, reliable Python pipeline that turns raw audio/video into clean, structured training data for speech models. It acquires media, normalises the audio, runs speaker diarization ("who spoke when") and optional emotion/style tagging ("how it was spoken"), and writes RTTM / JSON / CSV outputs — designed to run unattended over many files.

 source (URL or file)
        │  yt-dlp                 FFmpeg
        ▼                          ▼
   download  ───────────▶  extract 16 kHz mono WAV
                                    │
                pyannote / mock     ▼
            ┌────────────────────────────────┐
            │  diarization → who spoke when    │
            └───────────────┬─────────────────┘
                            │  segments
              HF / mock      ▼  (optional)
            ┌────────────────────────────────┐
            │  emotion tagging → how spoken    │
            └───────────────┬─────────────────┘
                            ▼
        outputs:  <id>.rttm   <id>.json   summary.csv

Quickstart (no token, no models, ~1 minute)

git clone <your-repo-url>
cd speech-data-pipeline

# create + activate a virtual environment
python -m venv .venv
source .venv/bin/activate                 # macOS/Linux
# .venv\Scripts\Activate.ps1              # Windows PowerShell

pip install -r requirements-dev.txt
pytest -v

The tests synthesise their own audio and use the mock backends, so this works with no HuggingFace token, no model download, and no audio files. (Three end-to-end tests also exercise FFmpeg; without FFmpeg installed they skip cleanly rather than fail.)

Features

  • Acquisition with yt-dlp, plus support for local media files.
  • Audio normalisation with FFmpeg to 16 kHz mono 16-bit WAV (what diarization models expect).
  • Speaker diarization via pyannote.audio, behind a small interface so the backend is swappable. A mock backend lets the whole pipeline (and the tests) run with no GPU, no token, and no model download.
  • Emotion & style tagging (optional) via a HuggingFace audio-classification model, behind the same interface pattern. The default model is public (no token), and a mock backend keeps it offline-testable.
  • Structured outputs: NIST RTTM, a rich JSON annotation per file, and a flat CSV summary.
  • Built for unattended runs:
    • Idempotent — a manifest records finished items; re-running skips them, so a crashed batch resumes without redoing work.
    • Failure isolation — one bad input is logged and skipped, never killing the batch.
    • Retries with backoff on downloads, and structured logging to console and file.

Project structure

speech-data-pipeline/
├── run.py                  # CLI entry point
├── requirements.txt        # runtime dependencies
├── requirements-dev.txt    # tests + tooling
├── pytest.ini              # test config (puts project root on sys.path)
├── urls.example.txt        # input format: one URL/path per line
├── .env.example            # documents HUGGINGFACE_TOKEN (no value committed)
├── SECURITY.md
├── .github/                # CI (tests, bandit, gitleaks) + Dependabot
├── src/
│   ├── acquire.py          # yt-dlp download + FFmpeg extraction (Stage 1)
│   ├── diarize.py          # Diarizer interface, pyannote + mock backends (Stage 2)
│   ├── emotion.py          # EmotionTagger interface, HF + mock backends (Stage 2b)
│   ├── outputs.py          # RTTM / JSON / CSV writers (Stage 3)
│   ├── pipeline.py         # orchestration: batch, manifest/idempotency, isolation
│   └── utils.py            # logging, stable IDs, filename helpers
└── tests/
    └── test_pipeline.py    # offline tests via the mock backends

Prerequisites

  • Python 3.10+
  • FFmpeg on your PATH — required for the acquisition stage (ffmpeg -version to check).
    • macOS: brew install ffmpeg
    • Linux: sudo apt install ffmpeg
    • Windows: winget install Gyan.FFmpeg (reopen the terminal afterwards so PATH updates)

Setup for real models

pip install -r requirements.txt

This installs torch, pyannote.audio, and transformers — a large download.

Then provide a HuggingFace token via an untracked .env (never commit it):

cp .env.example .env       # macOS/Linux   (Windows: copy .env.example .env)
# then edit .env and set HUGGINGFACE_TOKEN=hf_...

Gated diarization model

pyannote/speaker-diarization-3.1 is gated. While logged in to HuggingFace, accept the terms on both model pages (access is granted instantly):

The emotion model (superb/wav2vec2-base-superb-er) is public — nothing to accept. A GPU is used automatically if available; otherwise everything runs on CPU.

Usage

Offline first — no token or model needed:

# add a local file or URL to urls.txt (one per line), then:
python run.py --input urls.txt --backend mock --emotion mock

Real diarization, with emotion tagging:

python run.py --input urls.txt --backend pyannote --emotion hf --verbose

Useful flags: --emotion {none,hf,mock} (tagging backend), --retries N (download attempts), --force (reprocess items already in the manifest), --verbose (debug logging), --work DIR (scratch dir).

Output formats

For each input with id <id>:

  • outputs/rttm/<id>.rttm — one line per speech turn, NIST RTTM:
    SPEAKER <id> 1 <start> <duration> <NA> <NA> <speaker> <NA> <NA>
    
  • outputs/json/<id>.json — full record: source, title, duration, speaker count, and the list of {start, end, duration, speaker} segments. When emotion tagging is on, each segment also carries emotion and emotion_score, plus a top-level emotion_model.
  • outputs/summary.csv — one row per file (id, title, duration, #speakers, #segments, emotion model, dominant emotion).
  • outputs/manifest.json / run_report.json — idempotency state and per-run status.
  • outputs/logs/pipeline.log — full run log.

Design notes

  • Swappable model backends. Diarization and emotion tagging each sit behind a tiny interface with a real backend and a mock backend. Adding the emotion stage required no change to the diarizer, and the mocks make the whole pipeline testable with no GPU/model/network.
  • In-memory audio decoding. Audio is read with soundfile and passed to pyannote as an in-memory {waveform, sample_rate} tensor rather than a file path. This is pyannote's recommended path and sidesteps the optional torchcodec decoder, which is brittle to install on some platforms.
  • Reliability over throughput. The orchestrator favours resumability (manifest), isolation (one failure never stops the batch), and observability (structured logs) so it can run unattended on large inputs.

Testing

pytest -v

The suite synthesises a short WAV with the standard library and runs the pipeline through the mock backends, so it needs no network, model, or audio libraries. It covers segment generation, all three output formats, idempotent re-runs, and failure isolation. The end-to-end tests use FFmpeg; if FFmpeg isn't installed they are skipped with a clear reason rather than failing.

Roadmap / possible extensions

  • Parallel processing of the batch with a worker pool.
  • Quality filtering — drop segments below a duration/SNR threshold before export.
  • Speaking-rate / pitch features as additional per-segment annotations.

Security & repo hygiene

  • No secrets in the repo — the only secret (a HuggingFace token) is read from the environment or an untracked .env. See SECURITY.md.
  • Secret scanning with gitleaks in CI and as a pre-commit hook.
  • Static analysis with bandit in CI; Dependabot for weekly dependency updates.
  • Pre-commit hooks block large files, private keys, and merge markers:
    pip install pre-commit && pre-commit install
  • Runtime artefacts (downloads, WAVs, outputs/), virtualenvs, caches, and model weights are gitignored so they never bloat the repo.

Notes

Only download content you have the right to use. Configure your own sources in the input file before running.

About

Python pipeline that turns raw audio/video into structured speech-training data: acquisition, speaker diarization, and emotion tagging, with RTTM/JSON/CSV outputs.

Topics

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages