Compare commits

...
Sign in to create a new pull request.

3 commits

Author SHA1 Message Date
Шурупов Илья Викторович
83f782d125 Remove hardcoded credentials; read them from the environment
- Env.py now reads SPOTIFY_CLIENT_ID/SPOTIFY_CLIENT_SECRET (and optional
  path overrides) from the environment instead of hardcoding them.
- Add .env.example and ignore .env / .venv.
- Rewrite README with setup, usage and layout (drops the leaked secret).

Note: secrets are also purged from prior history in this push.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 00:39:59 +03:00
Шурупов Илья Викторович
36058c318a spotify: resilient fetching, progress bars, venv launcher
- fetch() uses a retrying session (connection resets, 5xx, 429 Retry-After)
  with backoff, so a transient drop mid-fetch no longer aborts the run.
- tqdm progress bars on all paginated fetches (driven by paging `total`).
- Flow.fetch_spotify(): drop stray `self` from the @staticmethod.
- gui.sh runs the venv's streamlit from the script dir; add requirements.txt.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 00:39:59 +03:00
Шурупов Илья Викторович
16b55b0674 spotify: update fetching to current Web API
- OAuth: URL-encode authorize params (scope spaces were unencoded),
  bind loopback redirect server to 127.0.0.1 to match the redirect_uri
  per Spotify's 2025 loopback rules, capture refresh_token.
- Search: URL-encode the query and drop the bogus track=1 param.
- Pagination: follow the paging object's `next` field instead of manual
  offset/total bookkeeping (Spotify's recommended approach).
- Fix invalid `raise "string"` in save_image_from_url.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 00:03:10 +03:00
10 changed files with 212 additions and 78 deletions

9
.env.example Normal file
View file

@ -0,0 +1,9 @@
# Copy to `.env` and fill in, then `source .env` before running.
# Get these from https://developer.spotify.com/dashboard (your app's settings).
# Register `http://127.0.0.1:8888` as a Redirect URI in that app.
export SPOTIFY_CLIENT_ID=your_client_id_here
export SPOTIFY_CLIENT_SECRET=your_client_secret_here
# Optional overrides:
# export MUSIC_LIBRARY_PATH=~/Music
# export WORK_DIR=./work_dir/new

2
.gitignore vendored
View file

@ -2,6 +2,8 @@ prefetched*
__pycache__
.idea
secret
.env
.venv
work_dir
old
tmp

14
Env.py
View file

@ -1,7 +1,13 @@
import os
from pathlib import Path
local_library_path = Path("/home/auser/Music/")
work_dir = Path("./work_dir/new")
# Local config. Paths can be overridden via environment variables.
local_library_path = Path(os.environ.get("MUSIC_LIBRARY_PATH", "~/Music")).expanduser()
work_dir = Path(os.environ.get("WORK_DIR", "./work_dir/new"))
client_id = '***REDACTED***'
client_secret = "***REDACTED***"
# Spotify app credentials. NEVER hardcode these here — set them in the
# environment (see .env.example) so they don't end up in version control.
# export SPOTIFY_CLIENT_ID=...
# export SPOTIFY_CLIENT_SECRET=...
client_id = os.environ.get("SPOTIFY_CLIENT_ID")
client_secret = os.environ.get("SPOTIFY_CLIENT_SECRET")

View file

@ -26,7 +26,7 @@ class Flow:
itunes_library : Library = None
@staticmethod
def fetch_spotify(self):
def fetch_spotify():
fetch_all(client_id, client_secret, work_dir)
@staticmethod

103
README.md
View file

@ -1,3 +1,100 @@
# spotifyFetcher
loads all data through web api<br>
client secret - ***REDACTED***
# MusicIndexer
Pull your Spotify library and listening stats via the Spotify Web API, index a
local music library (audio tags + iTunes XML), and match the two so you can see
which tracks you own, what's missing, and your top artists/tracks/playlists — in
a Streamlit GUI.
## Features
- **Spotify backend** — OAuth (Authorization Code) login, then fetch the user
profile, liked tracks, all playlists and their tracks, long-term top tracks
and top artists, and playlist/profile cover art. Resilient paginated fetching
(follows `next`, retries dropped connections, honours `429` rate limits) with
progress bars.
- **Local backend** — index a folder of audio files via `mutagen` tags.
- **iTunes backend** — parse an iTunes Library XML export.
- **Matching** — map local tracks to Spotify tracks (fuzzy matcher; ISRC/MBID
matching planned).
- **Stats & lyrics** — generate listening stats and fetch lyrics.
- **GUI** — browse libraries as artist → album → track tables.
## Setup
Requires Python 3.13+.
```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```
### Credentials
Create an app at https://developer.spotify.com/dashboard and add
`http://127.0.0.1:8888` as a **Redirect URI**. Credentials are read from the
environment — never commit them.
```bash
cp .env.example .env
# edit .env with your client id/secret
source .env
```
`Env.py` reads `SPOTIFY_CLIENT_ID` / `SPOTIFY_CLIENT_SECRET` (plus optional
`MUSIC_LIBRARY_PATH` and `WORK_DIR`) from the environment.
## Usage
Edit `main.py` to uncomment the steps you want, then run:
```bash
python main.py
```
The flow steps (see `Flow.py`) include:
| Step | What it does |
|------|--------------|
| `fetch_spotify()` | OAuth login + download all Spotify data into `work_dir/spotify/` |
| `fetch_local()` | index local audio files into `work_dir/local/` |
| `parse_spotify_library()` / `parse_itunes_library()` / `parse_local_library()` | build normalized libraries |
| `load_libraries()` / `save_libraries()` | (de)serialize libraries to `work_dir` |
| `map_local_to_spotify()` | match local tracks against the Spotify library |
| `fetch_and_save_lyrics()` | fetch lyrics for tracks |
| `log_stats()` | generate stats |
On first `fetch_spotify()` a browser opens for the Spotify login; a local server
on `127.0.0.1:8888` catches the redirect and exchanges the code for a token.
### GUI
```bash
./gui.sh # runs: .venv/bin/streamlit run Gui.py
```
## Layout
```
main.py entry point (CLI flow)
Gui.py / gui.sh Streamlit GUI
Flow.py orchestrates fetch / parse / map / stats
Env.py config (reads credentials from the environment)
src/
Library.py data model (Library / Artist / Album / Track)
LibrarySerializer.py load/save libraries
backends/
spotify/ Web API client, OAuth, parser
local/ local-file indexer, parser, playlist generator
itunes/ iTunes XML parser
navidrome/ Navidrome backend
mappers/ local↔Spotify matching (fuzzy / search)
lyrics/ lyrics fetching
work_dir/ fetched data and serialized libraries (gitignored)
```
## Notes
- Data is cached as JSON under `work_dir/` (gitignored); re-running reuses it.
- The deprecated Spotify endpoints (audio-features, recommendations, etc.) are
not used — only currently-supported endpoints.

6
gui.sh
View file

@ -1 +1,5 @@
streamlit run Gui.py
#!/usr/bin/env bash
# Run from the script's directory using the project virtualenv,
# so it works regardless of whether the venv is activated / on PATH.
cd "$(dirname "$0")"
exec .venv/bin/streamlit run Gui.py "$@"

12
main.py
View file

@ -17,17 +17,17 @@ def test1():
def main():
flow = Flow.Flow()
# flow.fetch_spotify()
flow.fetch_spotify()
# flow.fetch_local()
flow.parse_spotify_library()
flow.parse_itunes_library()
flow.parse_local_library()
#flow.parse_itunes_library()
#flow.parse_local_library()
flow.load_libraries()
#flow.load_libraries()
flow.map_local_to_spotify()
flow.save_libraries()
#flow.map_local_to_spotify()
#flow.save_libraries()
# flow.log_stats()

9
requirements.txt Normal file
View file

@ -0,0 +1,9 @@
streamlit==1.58.0
streamlit-aggrid==1.2.1.post2
pandas==3.0.3
Pillow==12.2.0
mutagen==1.47.0
rapidfuzz==3.14.5
requests==2.34.2
tqdm==4.68.2
xmltodict==1.0.4

View file

@ -11,6 +11,7 @@ import os.path
client_id = None
client_secret = None
token = None
refresh_token = None
def retrieve_web_api_token(authorization_code):
@ -44,7 +45,10 @@ def retrieve_web_api_token(authorization_code):
print(token_response.json())
raise ValueError
token = token_response.json()['access_token']
global refresh_token
response_json = token_response.json()
token = response_json['access_token']
refresh_token = response_json.get('refresh_token')
class RequestHandler(BaseHTTPRequestHandler):
@ -73,7 +77,7 @@ class RequestHandler(BaseHTTPRequestHandler):
def run_redirect_server():
httpd = HTTPServer(('localhost', 8888), RequestHandler)
httpd = HTTPServer(('127.0.0.1', 8888), RequestHandler)
print('Starting HTTP redirection server')
httpd.serve_forever()
@ -88,8 +92,9 @@ def open_authenticate_page():
"scope": "user-top-read playlist-read-collaborative user-read-email user-library-read user-read-private playlist-read-private",
}
auth_url = f"{authorize_url}?client_id={params['client_id']}&response_type={params['response_type']}&redirect_uri={params['redirect_uri']}&scope={params['scope']}"
auth_url = f"{authorize_url}?{urllib.parse.urlencode(params)}"
webbrowser.open(auth_url)
print(f"If the browser did not open, visit:\n{auth_url}")
def authenticate(id, secret):

View file

@ -1,6 +1,10 @@
import requests
import urllib.parse
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
from PIL import Image
from io import BytesIO
from tqdm import tqdm
from src.backends.spotify.SpotifyAuthenticator import authenticate
@ -8,96 +12,94 @@ from src.helpers import *
FETCH_STEP = 50
token = ''
REQUEST_TIMEOUT = 30
def _build_session():
"""Session that retries transient failures (dropped connections, 5xx) and
honours Spotify's 429 Retry-After, so a single reset mid-fetch doesn't
abort the whole run."""
session = requests.Session()
retry = Retry(
total=6,
connect=6,
read=6,
backoff_factor=1.5, # sleeps ~0, 1.5, 3, 6, 12, 24s between attempts
status_forcelist=(429, 500, 502, 503, 504),
allowed_methods=frozenset({"GET", "POST"}),
respect_retry_after_header=True,
raise_on_status=False,
)
adapter = HTTPAdapter(max_retries=retry)
session.mount("https://", adapter)
session.mount("http://", adapter)
return session
_session = _build_session()
def fetch(endpoint, method='GET', body=None):
url = f'https://api.spotify.com/{endpoint}'
url = endpoint if endpoint.startswith('http') else f'https://api.spotify.com/{endpoint}'
headers = {
'Authorization': f'Bearer {token}',
}
response = requests.request(method, url, headers=headers, json=body)
response = _session.request(method, url, headers=headers, json=body, timeout=REQUEST_TIMEOUT)
response.raise_for_status()
return response.json()
def fetch_paginated(endpoint, desc="Fetching", leave=True):
"""Collect every item of a paging object by following its `next` field,
which is Spotify's recommended way to page (more robust than tracking
offset/total manually). Shows a tqdm progress bar driven by the paging
object's `total`."""
items = []
page = fetch(endpoint)
with tqdm(total=page.get('total'), desc=desc, unit="item", leave=leave) as pbar:
while True:
batch = page.get('items', [])
items += batch
pbar.update(len(batch))
next_url = page.get('next')
if not next_url:
break
page = fetch(next_url)
return items
def find_song(song_pattern):
return fetch(f"v1/search/?type=track&track=1&q={song_pattern}")['tracks']['items']
query = urllib.parse.quote(song_pattern)
return fetch(f"v1/search?type=track&limit=10&q={query}")['tracks']['items']
def fetch_playlists(user_id):
print("Fetching playlists")
endpoint = f"v1/me/playlists"
pl_total = fetch(f'{endpoint}?limit=1')['total']
pl_count = 0
playlists = []
while pl_count < pl_total:
res = fetch(f'{endpoint}?limit={FETCH_STEP}&offset={pl_count}')
playlists += res['items']
pl_count += FETCH_STEP
return playlists
return fetch_paginated(f"v1/me/playlists?limit={FETCH_STEP}", desc="Fetching playlists")
def fetch_playlist_items(playlist_id):
endpoint = f"v1/playlists/{playlist_id}/tracks"
items = fetch(f'{endpoint}?fields=total')['total']
loaded = 0
tracks = []
while loaded < items:
res = fetch(f'{endpoint}?limit={FETCH_STEP}&offset={loaded}')
tracks += res['items']
loaded += FETCH_STEP
return tracks
return fetch_paginated(
f"v1/playlists/{playlist_id}/tracks?limit={FETCH_STEP}",
desc="Fetching playlist tracks",
leave=False,
)
def fetch_tracks(user_id):
print("Fetching tracks")
endpoint = f"v1/me/tracks"
total = fetch(f"{endpoint}?limit=1&market=ES")['total']
current = 0
tracks = []
while total > current:
res = fetch(f"{endpoint}?limit={FETCH_STEP}&offset={current}&market=ES")
tracks += res['items']
current += FETCH_STEP
tracks = fetch_paginated(f"v1/me/tracks?limit={FETCH_STEP}&market=ES", desc="Fetching liked tracks")
save_data(tracks, "tracks")
return tracks
def fetch_tracks_top():
print("Fetching top tracks")
total = fetch(f'v1/me/top/tracks?fields=total')['total']
current = 0
tracks = []
while current < total:
res = fetch(f'v1/me/top/tracks?limit={FETCH_STEP}&offset={current}&time_range=long_term')
tracks += res['items']
current += FETCH_STEP
tracks = fetch_paginated(f"v1/me/top/tracks?limit={FETCH_STEP}&time_range=long_term", desc="Fetching top tracks")
save_data(tracks, "top_tracks")
return tracks
def fetch_artists_top():
print("Fetching top artists")
total = fetch(f'v1/me/top/artists?fields=total')['total']
current = 0
artists = []
while current < total:
res = fetch(f'v1/me/top/artists?limit={FETCH_STEP}&offset={current}&time_range=long_term')
artists += res['items']
current += FETCH_STEP
artists = fetch_paginated(f"v1/me/top/artists?limit={FETCH_STEP}&time_range=long_term", desc="Fetching top artists")
save_data(artists, "top_artists")
return artists
@ -140,7 +142,7 @@ def save_image_from_url(url, name, image_dir="playlist_covers"):
response = requests.get(url)
if response.status_code != 200:
raise "Cannot fetch the playlist cover"
raise RuntimeError("Cannot fetch the playlist cover")
image = Image.open(BytesIO(response.content))