Problem / requested outcome
TubeMaster currently performs only narrow SRT cleanup. Caption payloads can still contain line-ending differences, structural SRT/WebVTT metadata, markup, excess whitespace, and repeated adjacent cue lines. Phase 4 needs a deterministic normalization boundary before future chunking or indexing.
Scope
Add one pure provider-local caption normalizer, based on issue #16, without changing public transcript contracts:
- Normalize CRLF/CR line endings to LF.
- Remove SRT indexes, SRT/WebVTT timestamps, WEBVTT headers, and supported cue metadata.
- Strip basic caption markup and VTT cue tags without becoming a rendering engine.
- Normalize whitespace per caption line and remove empty output lines.
- Deduplicate only adjacent identical normalized cue lines, preserving order and non-adjacent repetition.
- Decode equivalent string, Buffer, and ArrayBuffer payloads identically.
- Tolerate malformed structural lines without throwing.
- Return existing
no-captions behavior when normalized text is empty.
Acceptance criteria
- SRT input with CRLF, indexes, timestamps, blank lines, and markup produces clean deterministic text.
- WebVTT input with header, timestamps, cue tags, and metadata produces the same clean text boundary.
- Equivalent string, Buffer, and ArrayBuffer payloads produce equivalent normalized text.
- Empty cue blocks do not create blank output.
- Adjacent duplicate cue lines collapse once; non-adjacent repeated lines remain.
- Malformed timestamp/index lines never throw and have deterministic tested behavior.
- Available output remains
{ status: "available", text, language? }.
- Existing unavailable/unsupported results and diagnostics remain unchanged.
- API, CLI, and MCP envelopes remain unchanged.
- Tests cover normalization and existing transcript regressions.
Non-goals
- New transcript providers or fallback changes.
- Language detection, language translation, or language normalization.
- Chunking, persistence, indexing, embeddings, vector databases, semantic search, or topic extraction.
- API/CLI/MCP surface changes.
- Global duplicate removal or sentence reconstruction.
- OAuth, quota, analytics, playlist, reporting, or metadata changes.
Planned validation
- Focused transcript provider/service/API/CLI/MCP tests.
npm test.
npx tsc --noEmit.
npm run lint.
git diff --check.
Dependency: issue #16 (feat(transcripts): deterministically select YouTube caption tracks, commit cab2140).
Roadmap: Phase 4 Transcriptions.
Problem / requested outcome
TubeMaster currently performs only narrow SRT cleanup. Caption payloads can still contain line-ending differences, structural SRT/WebVTT metadata, markup, excess whitespace, and repeated adjacent cue lines. Phase 4 needs a deterministic normalization boundary before future chunking or indexing.
Scope
Add one pure provider-local caption normalizer, based on issue #16, without changing public transcript contracts:
no-captionsbehavior when normalized text is empty.Acceptance criteria
{ status: "available", text, language? }.Non-goals
Planned validation
npm test.npx tsc --noEmit.npm run lint.git diff --check.Dependency: issue #16 (
feat(transcripts): deterministically select YouTube caption tracks, commitcab2140).Roadmap: Phase 4 Transcriptions.