Skip to content

Improve track matching and add comprehensive job progress tracking - #9

Merged
Stensel8 merged 16 commits into
mainfrom
claude/number-matching-download-progress-q6ypw2
Sep 25, 2026
Merged

Stensel8 merged 16 commits into
mainfrom
claude/number-matching-download-progress-q6ypw2

Conversation

@Stensel8

@Stensel8 Stensel8 commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Track Matching Improvements

  • Enhanced version tag detection: Expanded VERSION_TAGS from a simple tuple to a dictionary with regex patterns, supporting more variations like "sped up", "slowed", "re-recorded", and alternative spellings (e.g., "rmx" for "remix", "acapella" variants)
  • Better title normalization: Improved _comparable() function to handle abbreviations ("Pt." → "Part", "Vol." → "Volume", "n" → "and") and curly apostrophes
  • Smarter search queries: New search_queries() function generates progressively looser search queries (specific to loose) without duplicates, improving the chances of finding tracks
  • Improved similarity scoring: Enhanced similarity() function with better handling of word order and spacing
  • Better artist parsing: Added _ARTIST_JOINS regex to properly split artist names on various separators (commas, ampersands, "x", "vs", "feat", "with")

Progress Tracking Architecture

  • New Step dataclass: Replaces the old callback signature with a structured progress report containing phase, text, done/total counts, current track, match result, and closest candidate for unmatched tracks
  • Phase-based workflow: Introduced Phase type and PHASES mapping to track distinct steps: "read", "match", "check", "add"
  • ETA calculation: Added remaining() function to estimate time left based on current pace
  • Job lifecycle tracking: Enhanced Job class to track all phases, timestamps, and near-misses (unmatched tracks with their closest candidates)

Web Interface Enhancements

  • Extracted job display: Moved job progress UI to reusable _job.html template
  • Rich progress display: New app.js shows:
    • Step indicators (read → match → check → add)
    • Real-time progress bar with percentage
    • Time elapsed and estimated time remaining
    • Count of found vs. not found tracks
    • Current track being processed
    • List of unmatched tracks with closest candidates and match scores
    • Download links for results and unmatched CSV
  • Better error handling: Improved error messages and status display
  • Animated progress bar: CSS animation for indeterminate progress (when total is unknown)

Matching (refs #8):
- A track without artists (a CSV of titles only, or a search result
  without artist data) was always refused: the unknown artist scored
  0.5, below the 0.6 cut-off. It is now judged on title and length.
- A search the service refuses (400/404) no longer stops the other
  searches for that track.
- Tidal's bulk ISRC lookup now follows links.next. One ISRC is often on
  a single, an album and a compilation, so 20 codes can fill more than
  one page; the codes on later pages fell back to a text search.
- Four searches, from specific to loose (as written, without feat. and
  remaster noise, without accents and punctuation, title only). The
  best hit over all of them wins; a convincing one ends the search.
- Artists written together or apart ("Macklemore & Ryan Lewis",
  "Nicky Jam x J Balvin"), "The", "Ke$ha", compact spellings (ACDC).
- Titles: en/em dashes, curly apostrophes, & vs and, Pt. vs Part,
  word order, main title in brackets, more remaster noise.
- More version tags (sped up, slowed, a cappella, re-recordings such as
  "Taylor's Version"), looked for only in the version part of a title,
  so the studio "Live Wire" no longer matches "Live Wire (Live)".
- A title of only brackets, like "(Untitled)", no longer shrinks to ""
  and matches every other such title.
- Ties go to the closest full title (the right remixer), then the same
  album (the original release over a compilation).
- Tracks that are not found are listed with the closest candidate and
  its score.

Progress:
- Every run reports its steps: reading the source, finding the tracks,
  checking what the playlist already has, adding (per request batch),
  with totals where the service gives them (Spotify's liked count).
- CLI: one live line per step with a bar, counts, found / not found,
  time left and the current track; export gets it too, and -q.
- Web: export is now a background job like import and transfer, so the
  tab no longer just spins; the CSV downloads when it is ready. The
  panel shows the steps, a bar, time taken and left, found / not found,
  the tracks not found so far with their closest candidate, and the
  browser tab title shows the percentage. Buttons are off while a job
  runs. The page script moved to static/app.js.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
Copilot AI lite review requested due to automatic review settings September 25, 2026 20:11
@Stensel8 Stensel8 linked an issue Sep 25, 2026 that may be closed by this pull request
@Stensel8 Stensel8 self-assigned this Sep 25, 2026

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Job-state synchronization and malformed export request handling must be fixed before approval.

Review effort: Lite
Findings: None

What changed in this PR

Improves cross-service track matching and adds phase-based progress tracking for CLI and web background jobs.

Changes:

  • Expands normalization, search fallback, scoring, and near-miss reporting.
  • Adds structured job phases, ETA tracking, downloads, and live progress UI.
  • Updates providers, CLI output, documentation, packaging, and tests.

Review findings include two moderate issues: synchronize job state updates with JSON serialization, and validate export request bodies before accessing .get. README wording should also clarify fuzzy artist matching.

File Description
tests/​test_web.py Tests background jobs and web progress.
tests/​test_tidal.py Tests paginated ISRC lookups.
tests/​test_sync.py Tests phases, progress, and near misses.
tests/​test_spotify.py Tests liked-track counts and lookups.
tests/​test_matching.py Tests expanded matching behavior.
tests/​test_cli.py Tests CLI progress output.
README.md Documents matching and progress behavior.
pyproject.toml Packages JavaScript assets.
musicsync/​web/​views.py Adds background export and download routes.
musicsync/​web/​templates/​dashboard.html Uses reusable job controls.
musicsync/​web/​templates/​base.html Loads shared JavaScript.
musicsync/​web/​templates/​account.html Adds asynchronous export actions.
musicsync/​web/​templates/​_job.html Defines reusable job-progress markup.
musicsync/​web/​static/​style.css Styles progress and job displays.
musicsync/​web/​static/​app.js Renders live job progress.
musicsync/​web/​jobs.py Tracks job lifecycle and progress state.
musicsync/​sync.py Adds phases and structured progress callbacks.
musicsync/​providers/​tidal.py Adds paginated lookups and batch sizing.
musicsync/​providers/​spotify.py Adds liked-track totals and batch sizing.
musicsync/​providers/​base.py Adds progressive searches and provider lookup flow.
musicsync/​matching.py Enhances normalization, comparison, and scoring.
musicsync/​cli.py Implements terminal-aware progress reporting.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Stensel8 and others added 15 commits September 25, 2026 22:28
Clarify virtual environment activation for Unix, fish, and PowerShell users, and add a config file screenshot to the setup documentation.
Playlists that Music-Sync creates now say "Imported with Music-Sync:
https://github.com/Stensel8/Music-Sync". Existing playlists are left
as they are.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
The old path drew half a diamond and a shifted one over each other.
The logo is four diamonds: three in a row and one below the middle.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
… numbers

Every track that is not found now comes with a short reason instead of
a bare percentage, in the CLI and on the web page:
- "not on Tidal": nothing by this artist came up;
- "only other songs on Tidal": the artist is there, this song is not
  (the candidate is no longer shown: it said nothing);
- "only another version": live, remix and the like, shown as closest;
- "score too low for a match (0.75, needs 0.80)": close, not sure enough.

Matching:
- Artist names must be 0.7 alike instead of 0.6, so "Roy Blair" no
  longer passes for "Radio Blazers". "and" now joins artists too, so
  "Tom Petty and the Heartbreakers" matches "Tom Petty".
- Titles with other numbers are other songs: "Part 1" and "Part 2",
  "Song 5" and "Song 55" scored above 0.9. Roman numerals II-IX count as
  numbers; a number on one side only (a year, "Op. 67") does not matter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
- A service without a client ID now leads to /<service>/setup instead of
  an error page: a link to the developer dashboard, the redirect URI and
  the settings file path with copy buttons, what to tick, a snippet to
  fill in, and a "Done" button that logs in once the ID is there (or
  says it is not in the file yet). The login page, the dashboard and the
  account and login routes all go there.
- Fix: a login stored before the client ID was taken out of the settings
  counted as "Connected", and clicking it gave the error page. Without a
  client ID a service now shows as not set up.
- The CLI error for a missing client ID names the developer dashboard.
- A logo, beamed notes whose beam is an arrow, as favicon, in the top
  bar and on the login page.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
Added usage examples for music-sync commands in README.
The logo and one sentence at the top, then installation, getting
started (music-sync web walks you through the setup) and usage in the
browser and on the command line, with a table of the options. The setup
steps with their screenshots, how matching works (shortened), the CSV
format and the rest follow below. The config screenshot moved next to
the other images in docs/images.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
…redits

From spotify_to_tidal: before anything is looked up, the source tracks
are matched against the tracks already in the target playlist, by id,
ISRC or title and artist. Those need no ISRC lookup or search, so a
second run of the same transfer asks the service almost nothing. The
"check" step therefore comes before "match" now.

The credits name and link every project Music-Sync grew out of:
csv2tidal (Nugman, roland.behme), RZetko's gist, spotify_to_tidal (Tim
Rae and contributors), python-tidal (Thomas Amland, tehkillerbee,
morguldir), the first Music-Sync, and daisyUI for the colours.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
Roland Behme and Nugman are the same person; the credits now say so,
in one short paragraph instead of a list.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
From a review of this branch for reuse, simplification, efficiency and
depth; no feature is lost and every test still passes.

- Cache the pure text functions of the matcher: a wanted track was
  folded and normalised again for every candidate. About 6x less CPU
  per searched track, with the same results.
- Batch "add to playlist" once, in import_tracks; the providers no
  longer split the same list a second time.
- Step carries the running found count, so the CLI and the web page no
  longer count on their own. The CLI progress line keeps less state.
- Miss.why builds "reason; closest: ..." once, for the CLI and the page
  (the job JSON now sends it as one string).
- Drop an except that could not be reached (find already handles a
  refused search), the unmatched shims, a helper used twice, and the
  duplicate CSV writing in the web views.
- Tests: FakeProvider records searches, a transfer() helper.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
`music-sync transfer spotify tidal --sync-favorites` adds your liked
songs to the target's favourites instead of to a playlist, skipping
what is already there; the flag has the name spotify_to_tidal uses.
`import --to-favorites` does the same for a CSV, and the web page has
an "or to favourites" box for import and transfer.

- Provider.add_favorite_tracks (after python-tidal's favorites.add_track):
  Spotify PUT /me/library with up to 40 URIs (the endpoint since
  February 2026, as spotipy uses it), TIDAL POST
  /userCollectionTracks/me/relationships/items with up to 20 (from the
  TIDAL OpenAPI spec).
- New scopes: user-library-modify (Spotify) and collection.write
  (TIDAL). A login from before gets a 403; Music-Sync then says to log
  in again instead of showing the bare error.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
csv2tidal (and RZetko's gist before it) read a CSV of `artist,album`
and put each album in your Tidal favourites; that was lost on the way
to Music-Sync. `music-sync import tidal albums.csv --albums` does it
again, for Tidal and Spotify, and the web page has an "albums" box.

- Provider.search_albums and add_favorite_albums (after python-tidal's
  favorites.add_album). Tidal reuses the track code with a kind:
  searchResults include=albums, /albums?filter[id]=&include=artists,
  and /userCollectionAlbums/me/relationships/items (from the TIDAL
  OpenAPI spec). Spotify searches type=album and saves album URIs with
  the same PUT /me/library as tracks.
- Albums are matched with the same scoring as tracks; one not found
  says why ("only other albums on Tidal").
- import_tracks and import_albums share the dedupe and the batched
  adding.
- A CSV with an album column and no title reads the album as the title.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
spotify_to_tidal looks for the album first and then for the track in
its tracklist. Music-Sync now does that too, as the last search for a
track that the ISRC and the text searches did not find convincingly:
search the album with its first artist, take the best album (0.8 or
more), and score the tracks on it like any other search result.

- Provider.album_tracks: Tidal reads the album's item ids and fetches
  those tracks the way it already does (/albums/{id}/relationships/items,
  then /tracks); Spotify reads /albums/{id}/tracks.
- In find() the album is just one more search, so it shares the error
  handling and stops as soon as a search convinces.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
Each form is now a grid of labelled fields with the button bottom right.
Tracks or albums, and playlist or favourites, are segmented choices
instead of checkboxes and loose selects, and the fields that do not
apply are hidden (CSS :has). The forms are read with FormData, so the
field names are the request keys.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
TIDAL's current OpenAPI spec (1.10.133) allows 50 items per POST to
/playlists/{id}/relationships/items and /userCollection*/relationships/items,
up from 20. Lookups by id or ISRC are still capped at 20.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MNJmFQ2fzEzpGMMVteondW
@Stensel8
Stensel8 merged commit 1667117 into main Sep 25, 2026
5 of 6 checks passed
@Stensel8
Stensel8 deleted the claude/number-matching-download-progress-q6ypw2 branch September 25, 2026 21:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

task: improve number matching algo

3 participants