Skip to content

feat(transcripts): normalize SRT and WebVTT caption text deterministically #17

Description

@EloyEMC

Problem / requested outcome

TubeMaster currently performs only narrow SRT cleanup. Caption payloads can still contain line-ending differences, structural SRT/WebVTT metadata, markup, excess whitespace, and repeated adjacent cue lines. Phase 4 needs a deterministic normalization boundary before future chunking or indexing.

Scope

Add one pure provider-local caption normalizer, based on issue #16, without changing public transcript contracts:

  • Normalize CRLF/CR line endings to LF.
  • Remove SRT indexes, SRT/WebVTT timestamps, WEBVTT headers, and supported cue metadata.
  • Strip basic caption markup and VTT cue tags without becoming a rendering engine.
  • Normalize whitespace per caption line and remove empty output lines.
  • Deduplicate only adjacent identical normalized cue lines, preserving order and non-adjacent repetition.
  • Decode equivalent string, Buffer, and ArrayBuffer payloads identically.
  • Tolerate malformed structural lines without throwing.
  • Return existing no-captions behavior when normalized text is empty.

Acceptance criteria

  • SRT input with CRLF, indexes, timestamps, blank lines, and markup produces clean deterministic text.
  • WebVTT input with header, timestamps, cue tags, and metadata produces the same clean text boundary.
  • Equivalent string, Buffer, and ArrayBuffer payloads produce equivalent normalized text.
  • Empty cue blocks do not create blank output.
  • Adjacent duplicate cue lines collapse once; non-adjacent repeated lines remain.
  • Malformed timestamp/index lines never throw and have deterministic tested behavior.
  • Available output remains { status: "available", text, language? }.
  • Existing unavailable/unsupported results and diagnostics remain unchanged.
  • API, CLI, and MCP envelopes remain unchanged.
  • Tests cover normalization and existing transcript regressions.

Non-goals

  • New transcript providers or fallback changes.
  • Language detection, language translation, or language normalization.
  • Chunking, persistence, indexing, embeddings, vector databases, semantic search, or topic extraction.
  • API/CLI/MCP surface changes.
  • Global duplicate removal or sentence reconstruction.
  • OAuth, quota, analytics, playlist, reporting, or metadata changes.

Planned validation

  • Focused transcript provider/service/API/CLI/MCP tests.
  • npm test.
  • npx tsc --noEmit.
  • npm run lint.
  • git diff --check.

Dependency: issue #16 (feat(transcripts): deterministically select YouTube caption tracks, commit cab2140).
Roadmap: Phase 4 Transcriptions.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions