10 Commits

Author SHA1 Message Date
andrew 55a059b6da 0.6.14: Fix Bandcamp sync silently failing to import new purchases
sync-bandcamp.sh set a hardcoded PATH at startup that left out /opt/venv/bin,
which is where beet lives in this image. Every time the Bandcamp sync had a
new purchase to import, its call into import-track.sh would run beet import
with that broken PATH, fail with "command not found", and import-track.sh
would just log the exit code and carry on, so sync-bandcamp.sh still reported
success. The purchase sat on disk, never entered the beets library, and never
showed up in Navidrome. This went unnoticed for weeks because most daily
syncs have nothing new to import, so the broken code path rarely ran.

Fixed the same copy-pasted PATH line in upgrade-mp3-to-flac.sh (which also
calls beet directly) and notify-telegram.sh (harmless there, but fixed for
consistency).

Also moved the post-import duplicate cleanup (replace-with-better.sh and
dedup-library.sh) out of sync-bandcamp.sh and into import-track.sh itself, so
every import gets the same cleanup, not just Bandcamp purchases. A manual
import or SMB drop that happens to match something already in the library no
longer leaves a duplicate copy sitting there until someone runs the dedup
review by hand.
2026-07-23 10:09:18 -06:00
andrew 068e7e9534 0.6.13: Fix watermark maintenance jobs, honest buy-link write reporting, Qobuz 403 diagnosis
scrub-watermark-text.py crashed on every run (missing import os). strip-watermark-art.py
was starved every Sunday by lock contention from strip-mb-tags' growing runtime -- moved
it to 01:00 so the rest of the weekly chain has room. enrich-buy-url.py's summary line
counted matches found, not successful writes, so a run where every metaflac write failed
(root-owned files) still reported "N tagged". pipeline-status.sh now tells a 403 from
Qobuz (Akamai/CDN block) apart from a 401 (actually expired token) -- re-exporting the
token does nothing for the former. Also fixes the dedup review table's "Caught by" badge
overflowing into the Keep column's format pill.
2026-07-22 13:09:34 -06:00
andrew e07e931c45 0.6.12: Redesign the dedup review table for faster, more accurate scanning
The keep/delete columns were raw paths in a fixed-width cell: ellipsis-
truncated, full text only on hover. Slow to scan (mousing over every
row to read the song name) and, once "fixed" with a tag-derived title/
artist header in a first pass of this change, actively worse for the
thing dedup review actually needs -- confirming two files are really
the same recording. Tags can be wrong; the filename on disk can't.

New layout: each candidate gets a title row (song title, large,
unclipped -- parsed from beets' known singleton path template,
Artist/Album/Title.ext) followed by a compact detail row. The detail
row shows a color-coded format pill (FLAC/MP3/etc, matching the same
quality judgment dedup-library.sh's rank_file() already makes) and the
full relative file path, always rendered as visible text -- never
hover-only -- so a real difference in filename, album, or folder is
still plainly visible even when tags line up. Shared artist collapses
onto the title row so it isn't repeated per side; a colored left rail
marks which side survives without having to read the words.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 10:15:27 -06:00
andrew 6114e6dc7a 0.6.11: Auto-apply exact-tag format upgrades; self-heal stale pending rows
Two gaps left obvious cases stuck in the manual dedup queue:

1. Format-upgrade pairs found by Pass 2/3 (case-insensitive/normalized
   tag match) still required manual confirm even though the pass
   itself already proved same song via identical tags, and rank_file()
   already guarantees FLAC beats any other format regardless of
   size/dirty-name. A directory-casing difference (e.g. "All About
   This Ep" vs "...EP") meant these never matched the narrower
   same-directory numbered-twin rule from 0.6.10. New is_format_upgrade()
   auto-deletes any Pass 2/3 pair whose extensions differ, gated on
   pass_name so Pass 4 (cross-album fuzzy -- different masters, DJ-mix
   edits, genuinely ambiguous) is untouched and still requires a human.

2. A pending row can go stale without ever being confirmed -- some
   other action (a later auto-apply, upgrade-mp3-to-flac, a manual fix)
   already resolved one side of the pair -- and nothing pruned it, so
   an already-fixed duplicate kept surfacing in the queue indefinitely.
   scan() now sweeps and auto-resolves any pending candidate whose
   keep_path or delete_path no longer exists before persisting new
   ones, so the queue reflects current reality on every run instead of
   accumulating dead entries.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 09:50:55 -06:00
andrew f05f076464 0.6.10: Auto-delete identical numbered-sibling duplicates, no review needed
Every dedup pass previously funneled through the same manual-confirm
queue, including the obviously-safe case: two files in the same
directory with the same name and extension, differing only by the
".N" collision suffix sldl/beets inserts when a filename collides.
That's not a fuzzy match or a different version -- it's the same
download landing twice -- so it doesn't need a human in the loop the
way case-insensitive tag matches or cross-album fuzzy matches do
(different masters, DJ-mix versions, etc., which still require
confirmation).

dedup-library.sh gains --auto-apply-numbered: an is_numbered_twin()
filename check (same dir, same name modulo the numeric infix, same
extension) that deletes matching pairs unconditionally, regardless of
which pass found them -- Pass 1 only fires when the canonical name is
a beets ghost; the common case where both copies are already tracked
in beets defers to Pass 2's tag-based grouping instead, and still gets
the same free pass here. When the auto-deleted pair's survivor is
itself the numbered-named file, it's renamed back to canonical so the
library doesn't accumulate ".1."/".2." names for tracks that no longer
have a duplicate.

dedup_review_service.scan() (both the manual "Scan now" button and the
nightly scheduled job) now passes this flag and records auto-applied
deletions as pre-confirmed DedupCandidate rows for audit visibility --
they never appear as pending review items. Tag-based and cross-album
fuzzy passes are unaffected and still require manual confirmation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 09:34:16 -06:00
andrew 62251d764e 0.6.9: Finish the Stage 4 path revisit in the three remaining scripts
The 2026-07-08 path rewrite changed beets' displayed paths from the
transitional /music/... mount to real /data/music/Library/... paths.
dedup-library.sh was fixed in 0.6.8; three more scripts still assumed
the old prefix, each failing silently:

- replace-with-better.sh resolved every library match to a doubled
  nonexistent path, logged [stale-db], skipped the in-place replace,
  and then imported the staged FLAC as a NEW track via the
  leftover-import step. Net effect: the mp3-to-flac upgrade flow
  produced flac+mp3 twin pairs in the library (surfacing in the dedup
  queue) instead of replacing the MP3 in place.

- fix-track-metadata.py queried beets only by the legacy /music/...
  form, matched nothing, and silently skipped beet update/move after
  retagging.

- export-laptop-playlists.py filtered out every library track (no
  displayed path starts with /music/ anymore), producing empty
  exports.

All three now use the real displayed path and keep the /music/... form
only as a legacy fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 11:51:55 -06:00
andrew d756aab7fc 0.6.8: Pin sldl skip index + fix dedup scan blinded by the path rewrite
Two bugs conspired to fill the library with .1.flac duplicates on the
night of 2026-07-15/16:

1. sldl derives its skip-index folder from the input: the Spotify
   playlist display name before 0.6.2, the CSV filename after. The
   switch stranded every playlist's download history in the old
   emoji-named subfolder, so sldl started from an empty index and
   re-downloaded every playlist in full. Fix: pin index-path in
   _template.conf, seed the pinned file from all stranded indexes
   (merge-sldl-indexes.py, called from run-playlist.sh whenever the
   pinned index is missing or older than a stranded one). Recovered
   rows keep the CSV-era identity fields but are marked downloaded, so
   tracks that failed last night yet exist in the library since months
   ago are not fetched again.

2. The safety net that should have flagged the flood, dedup-library.sh,
   has silently reported 0 groups since the 2026-07-08 path rewrite:
   host_path() still assumed /music/... display paths and prepended the
   library dir to already-correct /data/music/Library/... paths, so
   every candidate failed the -f check. Fix: pass real paths through
   unchanged, keep the /music prefix swap only for legacy entries.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 09:23:53 -06:00
andrew 24738d815f 0.6.7: Drop slskd from the digest health check
slskd is not part of the pipeline -- sldl is its own Soulseek client and
nothing in alembic calls slskd's API; the digest's reachability ping was its
only mention. It keeps running on the host purely as the user's own
file-sharing presence, so checking it here just risked a permanent bogus
warning line for something alembic doesn't depend on.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:50:20 -06:00
andrew 900bbff8aa 0.6.6: Digest format final: plain reflowing rows instead of <pre> columns
Iterating on the 0.6.5 HTML digest on a real phone: <pre> column layout
forces fixed-width lines that wrap into unreadable fragments ("54 tracks
in / M3U"), and one <pre> per section rendered as a stack of copy-button
widgets. Every row is now plain proportional text -- colored dot, bold
label, "-- value" -- which reflows cleanly at any screen width; breakdowns
(added-today per playlist, library by format) become bullet lists. Header
banner (brand line, dot tally, all-clear/N-issues blockquote) unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:44:26 -06:00
andrew be33664be2 0.6.5: HTML Telegram digest + fix silently-empty beet sections + watermark-art resilience
The daily digest is now formatted with Telegram HTML: bold section headers
with <pre>-aligned columns, colored circle emoji as status dots, a branded
header line, and a top blockquote callout (all clear vs N issues). All free
text is &/</> escaped so a stray angle bracket in a log snippet can't break
the parse and swallow the whole message. notify-telegram.sh gains an --html
flag (parse_mode=HTML) used by the digest send; plain-text callers are
unchanged. Truncation is now tag-safe in HTML mode: cut on a line boundary,
re-close an open <pre>, and use an escaped marker -- the old literal
"...<truncated>" tail would itself have 400'd the send.

Two real bugs surfaced while testing:

- pipeline-status.sh pins its own PATH, which lacked /opt/venv/bin where
  beet lives, so every beet probe ("added today", "Library by format", the
  mp3-now count) has been silently empty behind 2>/dev/null since the cron
  migration. PATH now includes the venv; both sections show real numbers.

- strip-watermark-art.py aborted its entire weekly run when metaflac stalled
  on ONE file (seen today: a healthy 190KB cover took >15s under disk
  contention, TimeoutExpired killed the job). Per-file timeout is now 60s
  and a timeout skips that file with a warning instead of failing the run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:30:58 -06:00
21 changed files with 831 additions and 193 deletions
+3 -3
View File
@@ -73,10 +73,10 @@ Optional, add these later if you want them:
Pull the prebuilt image onto your Docker host:
```bash
docker pull git.kretzer.club/andrew/alembic:0.6.4
docker pull git.kretzer.club/andrew/alembic:0.6.14
```
That is the whole install. You do not need to download the source or build anything. The `0.6.4` is the version; you can pin to it so nothing changes under you, or use `latest` to always get the newest.
That is the whole install. You do not need to download the source or build anything. The `0.6.14` is the version; you can pin to it so nothing changes under you, or use `latest` to always get the newest.
(If you would rather build it yourself from source, you can, but you do not need to.)
@@ -136,7 +136,7 @@ Create a file called `docker-compose.yml` on your server (put it wherever you ke
```yaml
services:
alembic:
image: git.kretzer.club/andrew/alembic:0.6.4
image: git.kretzer.club/andrew/alembic:0.6.14
container_name: alembic
ports:
- "8420:8420"
+68 -5
View File
@@ -1,3 +1,6 @@
import os
import re
from fastapi import APIRouter, BackgroundTasks, Depends, Request
from fastapi.responses import RedirectResponse
from fastapi.templating import Jinja2Templates
@@ -30,16 +33,76 @@ def _display_path(path: str) -> str:
return path[len(_LIBRARY_PREFIX):] if path.startswith(_LIBRARY_PREFIX) else path
# Compilation-style import paths carry a "$track " prefix on the filename
# (see pipeline/configs/beets config.yaml paths: comp/albumtype_soundtrack);
# singleton paths (the vast majority of this library) don't. Strip it either
# way so the title reads clean.
_TRACK_PREFIX_RE = re.compile(r"^\d{1,3}[\s.\-]+")
# rank_file() in dedup-library.sh always ranks FLAC above every other format
# and WAV below MP3 (a download glitch, never desirable) -- mirror that
# judgment in the format pill's color so it reads as "this one's the keeper"
# at a glance, not just a bare extension string.
_FORMAT_BADGE = {"flac": "badge-success", "wav": "badge-warning"}
def _split_display(path: str) -> dict:
"""Break a library display path into (artist, album, title, ext).
Beets' singleton path template is always
%the{$albumartist}/$album/$title (pipeline/configs/beets config.yaml,
paths:) -- every track in this library lands at exactly that shape, so
the last two segments are reliably album/filename. A path that doesn't
fit (an unexpected root-level file) degrades to showing the raw display
path as the title instead of guessing at structure that isn't there.
"""
display = _display_path(path)
parts = display.split("/")
if len(parts) >= 2:
artist, album, filename = parts[0], parts[-2], parts[-1]
title, ext = os.path.splitext(filename)
title = _TRACK_PREFIX_RE.sub("", title)
else:
artist, album = "", ""
title, ext = os.path.splitext(display)
return {
"path": path,
"display": display,
"artist": artist,
"album": album,
"title": title,
"ext": ext.lstrip(".").lower(),
}
def _file_size(path: str) -> int | None:
try:
return os.path.getsize(path)
except OSError:
return None
def _to_row(c) -> dict:
keep = _split_display(c.keep_path)
delete = _split_display(c.delete_path)
keep["size_bytes"] = _file_size(c.keep_path)
keep["badge"] = _FORMAT_BADGE.get(keep["ext"], "badge-muted")
delete["size_bytes"] = c.delete_size_bytes if c.delete_size_bytes is not None else _file_size(c.delete_path)
delete["badge"] = _FORMAT_BADGE.get(delete["ext"], "badge-muted")
return {
"id": c.id,
"pass_name": c.pass_name,
"pass_label": _pass_label(c.pass_name),
"keep_path": c.keep_path,
"delete_path": c.delete_path,
"keep_display": _display_path(c.keep_path),
"delete_display": _display_path(c.delete_path),
"delete_size_bytes": c.delete_size_bytes,
"keep": keep,
"delete": delete,
# Same-tag passes (numbered_sibling/case_insensitive/normalized) share
# artist/title almost always; cross_album_fuzzy and fuzzy_audio can
# legitimately differ (different credit, different album entirely) --
# the template only collapses shared context onto one line when it's
# actually shared, otherwise shows both sides in full.
"same_title": keep["title"].lower() == delete["title"].lower(),
"same_artist": keep["artist"].lower() == delete["artist"].lower(),
"same_album": keep["album"].lower() == delete["album"].lower(),
}
+83 -10
View File
@@ -27,6 +27,38 @@ def _script_for_pass(pass_name: str) -> str:
return _FUZZY_SCRIPT if pass_name == _FUZZY_PASS else _SCRIPT
def _prune_stale_pending(db) -> int:
"""Mark pending candidates whose delete_path or keep_path no longer
exists as resolved, instead of leaving them to surface forever.
A pending row can go stale without ever being confirmed through the UI --
e.g. a later scan's numbered-twin/format-upgrade auto-apply already
removed one side, or upgrade-mp3-to-flac replaced it, or someone deleted
it by hand. confirm_and_apply() already does this "already_gone" check
for candidates a user explicitly acts on; this generalizes it to run at
the start of every scan, so the pending list reflects current reality
instead of stale groups that something else already resolved."""
now = time.time()
rows = db.execute(
select(DedupCandidate).where(
DedupCandidate.applied == False, # noqa: E712
DedupCandidate.confirmed == False, # noqa: E712
DedupCandidate.ignored == False, # noqa: E712
)
).scalars().all()
pruned = 0
for c in rows:
if not Path(c.delete_path).exists() or not Path(c.keep_path).exists():
c.confirmed = True
c.confirmed_by = "auto:stale"
c.confirmed_at = now
c.applied = True
pruned += 1
if pruned:
db.commit()
return pruned
def _is_pair_ignored(db, path_a: str, path_b: str) -> bool:
"""True if this path pair was ever marked 'keep both', regardless of
which path was on the keep/delete side that time -- a later scan can
@@ -58,7 +90,9 @@ def _is_pair_already_pending(db, keep_path: str, delete_path: str) -> bool:
return db.execute(query).first() is not None
async def _run_scan(job_key: str, script_name: str, triggered_by: str) -> DedupRun | None:
async def _run_scan(
job_key: str, script_name: str, triggered_by: str, extra_args: tuple[str, ...] = ()
) -> DedupRun | None:
"""Returns None (persisting nothing) if the job never actually ran --
e.g. skipped_lock because something else was using the pipeline lock
at that moment. Same lesson as genre_review_service.run(): recording a
@@ -67,7 +101,7 @@ async def _run_scan(job_key: str, script_name: str, triggered_by: str) -> DedupR
pending-candidates list isn't scoped to "latest run only"."""
script = str(settings.pipeline_dir / "lib" / script_name)
job_run, output = await pipeline_runner.run_job_capture(
job_key, [script, "--json"], triggered_by=triggered_by, timeout=_JOB_TIMEOUT_SECONDS
job_key, [script, "--json", *extra_args], triggered_by=triggered_by, timeout=_JOB_TIMEOUT_SECONDS
)
if job_run.status != "success":
return None
@@ -76,6 +110,8 @@ async def _run_scan(job_key: str, script_name: str, triggered_by: str) -> DedupR
db = SessionLocal()
try:
_prune_stale_pending(db)
now = time.time()
dedup_run = DedupRun(
started_at=job_run.started_at,
finished_at=job_run.finished_at,
@@ -91,7 +127,30 @@ async def _run_scan(job_key: str, script_name: str, triggered_by: str) -> DedupR
skipped_ignored = 0
skipped_duplicate = 0
auto_applied = 0
for c in candidates:
if c.get("auto_applied"):
# dedup-library.sh already deleted this pair unconditionally
# (identical numbered-sibling twin: same dir/name/ext, no
# tag/fuzzy ambiguity involved) -- record it pre-applied for
# audit history; it never shows up as a pending review item.
db.add(
DedupCandidate(
dedup_run_id=dedup_run.id,
pass_name=c.get("pass", "unknown"),
keep_path=c["keep_path"],
delete_path=c["delete_path"],
delete_id=c.get("delete_id"),
delete_size_bytes=c.get("delete_size_bytes"),
confirmed=True,
confirmed_by="auto:numbered_twin",
confirmed_at=now,
applied=True,
)
)
auto_applied += 1
db.flush()
continue
if _is_pair_ignored(db, c["keep_path"], c["delete_path"]):
skipped_ignored += 1
continue
@@ -116,8 +175,11 @@ async def _run_scan(job_key: str, script_name: str, triggered_by: str) -> DedupR
db.commit()
db.refresh(dedup_run)
skipped_total = skipped_ignored + skipped_duplicate
if skipped_total:
dedup_run.kept = (dedup_run.kept or 0) + skipped_total
if skipped_total or auto_applied:
if skipped_total:
dedup_run.kept = (dedup_run.kept or 0) + skipped_total
if auto_applied:
dedup_run.deleted = auto_applied
db.commit()
db.refresh(dedup_run)
return dedup_run
@@ -126,12 +188,23 @@ async def _run_scan(job_key: str, script_name: str, triggered_by: str) -> DedupR
async def scan(triggered_by: str = "manual") -> DedupRun | None:
"""Dry-run dedup-library.sh --json, persist every candidate deletion
into a fresh dedup_runs/dedup_candidates pair. Never deletes anything
-- the scheduled maintenance:dedup job also only ever calls this (no
--apply), matching the false-negative-biased dedup preference; actual
deletion only ever happens through confirm_and_apply() below."""
return await _run_scan("dedup:scan", _SCRIPT, triggered_by)
"""Dry-run dedup-library.sh --json (plus --auto-apply-numbered), persist
every candidate deletion into a fresh dedup_runs/dedup_candidates pair.
Two kinds of match need no human judgment and get deleted unconditionally
by dedup-library.sh itself (see is_numbered_twin/is_format_upgrade there):
identical numbered-sibling twins (same dir, same name, same extension,
differing only by the ".N" collision suffix), and format upgrades within
an EXACT tag match from Pass 2/3 (case-insensitive/normalized already
proved same song via identical tags, so a differing extension there is
just "better format vs. worse"). Everything else -- including Pass 4's
cross-album fuzzy matches, where a wrong auto-delete could take out a
genuinely different version (a different album pressing, a DJ-mix edit,
etc.) -- stays a dry-run candidate awaiting manual confirm_and_apply()
below; that's the false-negative-biased preference for anything with real
ambiguity. Auto-applied deletions are reported here already applied,
purely for audit visibility."""
return await _run_scan("dedup:scan", _SCRIPT, triggered_by, extra_args=("--auto-apply-numbered",))
async def scan_fuzzy(triggered_by: str = "manual") -> DedupRun | None:
+12 -1
View File
@@ -155,8 +155,19 @@ MAINTENANCE_JOBS: dict[str, tuple[dict, callable]] = {
_log_rotation,
),
# ==== Weekly (Sunday) ====
# strip_mb_tags runs first here at 01:00 -- not 08:30 like the rest of
# this chain -- because its runtime isn't stable: 10m19s on 2026-07-12,
# 39.6min on 2026-07-19 (MusicBrainz lookup latency scales with library
# size and isn't under our control). At 08:30 that variance repeatedly
# starved every job behind it: strip_watermark_art waited the full
# SCHEDULED_LOCK_WAIT_SECONDS and gave up every week from 07-12 onward,
# and on 07-19 the starvation cascaded all the way through genre:run,
# normalize_casing, beets_update_sync, dedup:scan, and both gen_*_playlist
# jobs. 01:00 sits in the dead zone before the first playlist sync (05:00)
# and well after log_rotation (00:00), so even a run several times slower
# than 07-19's is guaranteed to release the lock long before 08:30.
"maintenance:strip_mb_tags": (
dict(minute=30, hour=8, day_of_week="sun"),
dict(minute=0, hour=1, day_of_week="sun"),
_lib("maintenance:strip_mb_tags", "strip-mb-tags.sh"),
),
"maintenance:strip_watermark_art": (
+61 -5
View File
@@ -475,11 +475,60 @@ tbody tr:hover { background: rgba(167, 139, 250, 0.055); }
tbody td.actions-cell { display: flex; gap: 0.5rem; flex-wrap: wrap; }
.dedup-table { table-layout: fixed; }
.dedup-table .path-cell {
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
max-width: 0; /* forces the cell to respect the colgroup width instead of the content's natural width */
/* Each candidate is its own <tbody> (title row + detail row) so the pair
highlights together on hover, and so :hover can bubble from either row
up to the shared group without JS. That means every detail row is now
a tbody's last-child, which would otherwise strip the separator between
one candidate and the next (the generic `tbody tr:last-child` rule above
assumed one shared tbody) -- restore it here, and only drop it for the
actual last group in the table. */
.dedup-table .dedup-row-group:hover td { background: rgba(167, 139, 250, 0.06); }
.dedup-table .dedup-detail-row td { border-bottom: 1px solid var(--border); }
.dedup-table .dedup-row-group:last-of-type .dedup-detail-row td { border-bottom: none; }
.dedup-title-row td { padding: 0.75rem 1.1rem 0.2rem; border-bottom: none; }
.dedup-detail-row td { padding-top: 0.15rem; }
.dedup-title {
font-family: var(--font-display);
font-size: 1.02rem;
font-weight: 700;
color: var(--fg);
line-height: 1.35;
overflow-wrap: anywhere;
}
.dedup-title-alt { font-weight: 400; color: var(--muted); }
.dedup-context { font-size: 0.8rem; margin-top: 0.15rem; }
/* A thin colored rail on each side echoes the keep/delete verdict itself --
green for the copy that survives, muted-pink for the one on the chopping
block -- so which side is which reads before you've even read the words. */
.dedup-side {
display: flex;
align-items: center;
gap: 0.55rem;
padding-left: 0.65rem;
border-left: 2px solid transparent;
min-height: 1.6rem;
}
.dedup-side-keep { border-left-color: var(--status-success); }
.dedup-side-delete { border-left-color: var(--status-danger); }
.dedup-side-size { font-size: 0.78rem; white-space: nowrap; flex-shrink: 0; }
/* The actual filename/path -- always visible, never truncated. Tags can be
wrong; this is the ground truth for judging whether two files are really
the same recording, so it can't be hover-only the way it was in the first
pass of this redesign. */
.dedup-side-path {
padding-left: 0.65rem;
margin-top: 0.2rem;
font-size: 0.78rem;
color: var(--muted-2);
overflow-wrap: anywhere;
line-height: 1.4;
}
.empty-row td { padding: 2.5rem 1rem; text-align: center; }
@@ -509,6 +558,13 @@ tbody td.actions-cell { display: flex; gap: 0.5rem; flex-wrap: wrap; }
.badge-pulse::before { animation: pulse-dot 1.4s ease-in-out infinite; }
@keyframes pulse-dot { 0%, 100% { opacity: 1; } 50% { opacity: 0.35; } }
/* The dedup "Caught by" column is only 10.75rem wide -- "Acoustically
Similar" at the default badge size overflowed past it and rendered on
top of the keep-side format pill. Smaller size/padding/tracking keeps it
inside the column instead of shrinking the column itself, which would
just crowd the Keep/Delete path columns next to it. */
.badge-caughtby { font-size: 0.6rem; padding: 0.24rem 0.55rem; letter-spacing: 0.02em; }
/* ---- stacked status bar (playlist track reconciliation) ---- */
.status-bar {
+75 -46
View File
@@ -1,5 +1,65 @@
{% extends "base.html" %}
{% block title %}Dedup — alembic{% endblock %}
{% macro candidate_rows(c, mode) %}
<tbody class="dedup-row-group">
<tr class="dedup-title-row">
<td colspan="{{ 5 if mode == 'pending' else 4 }}">
{% if c.same_title %}
<div class="dedup-title">{{ c.keep.title }}</div>
{% else %}
<div class="dedup-title">{{ c.keep.title }} <span class="dedup-title-alt">/ {{ c.delete.title }}</span></div>
{% endif %}
{% if c.same_artist and c.same_album and c.keep.album %}
<div class="dedup-context muted">{{ c.keep.artist }} · {{ c.keep.album }}</div>
{% elif c.same_artist %}
<div class="dedup-context muted">{{ c.keep.artist }}</div>
{% else %}
<div class="dedup-context muted">{{ c.keep.artist }} / {{ c.delete.artist }}</div>
{% endif %}
</td>
</tr>
<tr class="dedup-detail-row">
{% if mode == 'pending' %}
<td><input type="checkbox" name="candidate_id" value="{{ c.id }}" form="bulk-delete-form"></td>
{% endif %}
<td><span class="badge badge-caughtby {{ 'badge-warning' if c.pass_label == 'Acoustically Similar' else 'badge-info' }}" title="{{ c.pass_name }}">{{ c.pass_label }}</span></td>
<td>
<div class="dedup-side dedup-side-keep">
<span class="badge {{ c.keep.badge }}">{{ c.keep.ext or '?' }}</span>
{% if c.keep.size_bytes %}<span class="muted dedup-side-size">{{ (c.keep.size_bytes / 1024 / 1024) | round(1) }} MB</span>{% endif %}
</div>
<div class="dedup-side-path mono muted">{{ c.keep.display }}</div>
</td>
<td>
<div class="dedup-side dedup-side-delete">
<span class="badge {{ c.delete.badge }}">{{ c.delete.ext or '?' }}</span>
{% if c.delete.size_bytes %}<span class="muted dedup-side-size">{{ (c.delete.size_bytes / 1024 / 1024) | round(1) }} MB</span>{% endif %}
</div>
<div class="dedup-side-path mono muted">{{ c.delete.display }}</div>
</td>
<td class="actions-cell">
{% if mode == 'pending' %}
<form method="post" action="/dedup/confirm" class="inline"
onsubmit="return confirm('Permanently delete this file from disk?\n{{ c.delete.display }}')">
<input type="hidden" name="candidate_id" value="{{ c.id }}">
<button type="submit" class="btn btn-sm btn-danger">Delete</button>
</form>
<form method="post" action="/dedup/ignore" class="inline">
<input type="hidden" name="candidate_id" value="{{ c.id }}">
<button type="submit" class="btn btn-sm btn-ghost" title="Keep both copies and never flag this pair again">Keep both</button>
</form>
{% else %}
<form method="post" action="/dedup/unignore" class="inline">
<input type="hidden" name="candidate_id" value="{{ c.id }}">
<button type="submit" class="btn btn-sm btn-ghost">Un-ignore</button>
</form>
{% endif %}
</td>
</tr>
</tbody>
{% endmacro %}
{% block content %}
<div class="page-header">
<div>
@@ -18,48 +78,30 @@
{% endif %}
<h2>Pending candidates</h2>
<p class="muted" style="margin-top:-0.5rem;">Paths are relative to the library root. Hover a name for the full path.</p>
<p class="muted" style="margin-top:-0.5rem;">Song title up top so you can scan straight down the list; the actual file path is always shown underneath each side so you can check they're really the same file before confirming.</p>
<form id="bulk-delete-form" method="post" action="/dedup/confirm"
onsubmit="return confirm('Permanently delete every selected file from disk? This cannot be undone.')"></form>
<div class="table-wrap scroll">
<table class="dedup-table">
<colgroup>
<col style="width:2.2rem"><col style="width:9.5rem"><col style="width:28%">
<col style="width:28%"><col style="width:5rem"><col style="width:9rem">
<col style="width:2.2rem"><col style="width:10.75rem"><col style="width:37.5%">
<col style="width:37.5%"><col style="width:10rem">
</colgroup>
<thead>
<tr><th></th><th>Caught by</th><th>Keep</th><th>Delete</th><th>Size</th><th>Action</th></tr>
<tr><th></th><th>Caught by</th><th>Keep</th><th>Delete</th><th>Action</th></tr>
</thead>
{% for c in candidates %}
{{ candidate_rows(c, 'pending') }}
{% else %}
<tbody>
{% set pass_badge = {"File Naming": "badge-info", "Acoustically Similar": "badge-warning"} %}
{% for c in candidates %}
<tr>
<td><input type="checkbox" name="candidate_id" value="{{ c.id }}" form="bulk-delete-form"></td>
<td><span class="badge {{ pass_badge.get(c.pass_label, 'badge-info') }}" title="{{ c.pass_name }}">{{ c.pass_label }}</span></td>
<td class="muted mono path-cell" title="{{ c.keep_path }}">{{ c.keep_display }}</td>
<td class="mono path-cell" title="{{ c.delete_path }}">{{ c.delete_display }}</td>
<td class="muted">{{ (c.delete_size_bytes / 1024 / 1024) | round(1) if c.delete_size_bytes else '?' }} MB</td>
<td class="actions-cell">
<form method="post" action="/dedup/confirm" class="inline"
onsubmit="return confirm('Permanently delete this file from disk?\n{{ c.delete_display }}')">
<input type="hidden" name="candidate_id" value="{{ c.id }}">
<button type="submit" class="btn btn-sm btn-danger">Delete</button>
</form>
<form method="post" action="/dedup/ignore" class="inline">
<input type="hidden" name="candidate_id" value="{{ c.id }}">
<button type="submit" class="btn btn-sm btn-ghost" title="Keep both copies and never flag this pair again">Keep both</button>
</form>
</td>
</tr>
{% else %}
<tr class="empty-row"><td colspan="6">
<tr class="empty-row"><td colspan="5">
<div class="empty-state">
<span class="sparkle" style="position:relative; display:inline-block; margin-bottom:0.5rem;"></span>
<p class="muted" style="margin:0;">No pending candidates. Run a scan.</p>
</div>
</td></tr>
{% endfor %}
</tbody>
{% endfor %}
</table>
</div>
{% if candidates %}
@@ -72,28 +114,15 @@
<div class="table-wrap scroll">
<table class="dedup-table">
<colgroup>
<col style="width:9.5rem"><col style="width:33%">
<col style="width:33%"><col style="width:5rem"><col style="width:7rem">
<col style="width:10.75rem"><col style="width:39.25%">
<col style="width:39.25%"><col style="width:7rem">
</colgroup>
<thead>
<tr><th>Caught by</th><th>Keep</th><th>Delete</th><th>Size</th><th></th></tr>
<tr><th>Caught by</th><th>Keep</th><th>Delete</th><th></th></tr>
</thead>
<tbody>
{% for c in ignored %}
<tr>
<td><span class="badge badge-muted" title="{{ c.pass_name }}">{{ c.pass_label }}</span></td>
<td class="muted mono path-cell" title="{{ c.keep_path }}">{{ c.keep_display }}</td>
<td class="muted mono path-cell" title="{{ c.delete_path }}">{{ c.delete_display }}</td>
<td class="muted">{{ (c.delete_size_bytes / 1024 / 1024) | round(1) if c.delete_size_bytes else '?' }} MB</td>
<td>
<form method="post" action="/dedup/unignore" class="inline">
<input type="hidden" name="candidate_id" value="{{ c.id }}">
<button type="submit" class="btn btn-sm btn-ghost">Un-ignore</button>
</form>
</td>
</tr>
{% endfor %}
</tbody>
{% for c in ignored %}
{{ candidate_rows(c, 'ignored') }}
{% endfor %}
</table>
</div>
{% endif %}
+24
View File
@@ -242,6 +242,30 @@ BEETS_EXIT=0
beet import -q -s "$DEST_DIR" >> "$LOG" 2>&1 || BEETS_EXIT=$?
log "beets import finished with exit code $BEETS_EXIT"
# ==== Safety net: clean up stragglers + exact duplicates ====
# With duplicate_action=keep, beets always imports rather than rejecting a
# real conflict (see beets/config.yaml) -- any file whose (albumartist,
# album, title) already exists in the library gets imported anyway as a
# `.1.ext` sibling. That's true whether this run came from the web UI, an
# SMB drop, or sync-bandcamp.sh, so the cleanup has to run here rather than
# per-caller (a 2026-07-23 incident: a Bandcamp sync's own call into this
# script failed and silently stranded ~150 already-owned files in import-me/
# for two weeks; a later unrelated manual import swept them back in and
# duplicated them, and nothing had deduped since).
log "Running replace-with-better safety pass on /downloads stragglers"
if ${PIPELINE_DIR:-/app/pipeline}/lib/replace-with-better.sh --apply >> "$LOG" 2>&1; then
log "replace-with-better OK"
else
log "WARN: replace-with-better.sh exited non-zero (exit $?)"
fi
log "Running dedup pass (exact-match duplicates only; FLAC > MP3, then largest file)"
if ${PIPELINE_DIR:-/app/pipeline}/lib/dedup-library.sh --apply >> "$LOG" 2>&1; then
log "dedup OK"
else
log "WARN: dedup-library.sh exited non-zero (exit $?)"
fi
# ==== Regenerate M3U if playlist was specified ====
if [[ -n "$PLAYLIST_NAME" ]]; then
log "Regenerating M3U for $PLAYLIST_NAME"
+15
View File
@@ -96,6 +96,21 @@ else
log "Fetched $TRACK_COUNT track(s) from Spotify"
fi
# ==== Seed the pinned sldl skip index if it's missing or stranded ====
# The conf pins index-path to $DROPBOX/_index.csv (see _template.conf: sldl's
# default index location is derived from the input, so an input change strands
# the old index and the whole playlist re-downloads -- that's what produced
# the 2026-07-16 duplicate flood). If the pinned file is missing, or an index
# in one of sldl's old input-named subfolders is newer (i.e. sldl last wrote
# somewhere else), fold them all into the pinned location first.
INDEX_FILE="$DROPBOX/_index.csv"
if [[ ! -s "$INDEX_FILE" ]] || \
[[ -n "$(find "$DROPBOX" -mindepth 2 -maxdepth 2 -name _index.csv -newer "$INDEX_FILE" -print -quit 2>/dev/null)" ]]; then
log "Seeding pinned sldl index at $INDEX_FILE from prior indexes"
python3 "${PIPELINE_DIR:-/app/pipeline}/lib/merge-sldl-indexes.py" "$DROPBOX" "$INDEX_FILE" >> "$LOG" 2>&1 \
|| log "WARNING: index merge failed -- sldl may re-download tracks the library already has"
fi
# ==== Run sldl (vendored binary, subprocess of this container) ====
log "Running sldl for: $PLAYLIST_NAME"
+12
View File
@@ -31,6 +31,18 @@ path = SLDL_DROPBOX_ROOT/PLAYLIST_NAME
playlist-path = SLDL_DROPBOX_ROOT/PLAYLIST_NAME/_sldl.m3u8
write-playlist = true
# ==== Skip index (pinned!) ====
# sldl's already-downloaded index defaults to {path}/{playlist-name}/_index.csv
# where {playlist-name} is derived from the INPUT: the Spotify playlist's
# display name for spotify input, the CSV filename for csv input. beets moves
# every download out of the dropbox, so this index is the ONLY thing standing
# between a nightly run and re-downloading the whole playlist. When 0.6.2
# switched input from Spotify URL to CSV, the derived name changed, sldl
# started a fresh empty index in a new subfolder, and every playlist
# re-downloaded in full on 2026-07-16. Pinning the path here decouples the
# index from input naming so that can never happen again.
index-path = SLDL_DROPBOX_ROOT/PLAYLIST_NAME/_index.csv
# ==== Naming ====
name-format = {artist} - {title}
+113 -12
View File
@@ -34,11 +34,26 @@ set -euo pipefail
APPLY=0
JSON_MODE=0
ONLY_PATHS_FILE=""
AUTO_APPLY_NUMBERED=0
while [[ $# -gt 0 ]]; do
case "$1" in
--apply) APPLY=1; shift ;;
--json) JSON_MODE=1; shift ;;
--only-paths) ONLY_PATHS_FILE="$2"; shift 2 ;;
# Delete two kinds of match unconditionally, independent of
# --apply/--only-paths, while everything else stays dry-run/log-only:
# - numbered-sibling twins (see is_numbered_twin): same dir, same name,
# same extension, just the ".N" collision suffix -- the same download
# landing twice.
# - format upgrades within an exact tag match (see is_format_upgrade):
# Pass 2/3 already proved same song via identical tags; a different
# extension there just means "better format vs. worse", never a
# different version.
# Excludes Pass 4 (cross-album fuzzy) entirely -- matching across
# different albums/directories is exactly where a different master or
# DJ-mix edit can share tags without being interchangeable, so it always
# needs a human. See dedup_review_service.scan().
--auto-apply-numbered) AUTO_APPLY_NUMBERED=1; shift ;;
*) shift ;;
esac
done
@@ -78,11 +93,63 @@ fi
# Counters: tracked via log grep at the end (avoids bash subshell pitfalls).
# Functions write KEEP/DELETE lines with consistent prefixes; summary greps them.
# Convert container path (/music/...) to host path (${MUSIC_DATA_DIR:-/data/music}/Library/...)
host_path() { echo "${MUSIC_DATA_DIR:-/data/music}/Library${1#/music}"; }
# Resolve a beets-displayed path to a real path in this container. Since the
# 2026-07-08 path rewrite, beets displays ${MUSIC_DATA_DIR:-/data/music}/Library/...
# paths that are directly usable — pass those through UNCHANGED. Only legacy
# /music/... display paths (the pre-rewrite transitional mount) still need the
# prefix swap. Blindly prepending the library dir to an already-correct path
# (the old behavior) produced /data/music/Library/data/music/Library/... —
# nonexistent, so every candidate failed the -f check and every pass reported
# 0 groups from 2026-07-08 until this fix.
host_path() {
local p="$1"
if [[ "$p" == /music/* ]]; then
echo "${MUSIC_DATA_DIR:-/data/music}/Library${p#/music}"
else
echo "$p"
fi
}
SEP=$'\x1c' # ASCII file separator — safe with any music metadata
# True if two paths are the SAME directory, SAME extension, and identical
# basenames except for a numeric ".N" infix sldl/beets inserts on a filename
# collision (e.g. "Title.flac" vs "Title.1.flac"). This is deliberately
# narrower than any tag-based pass: no fuzzy matching, no cross-album
# ambiguity, just the literal same download landing twice. Used to
# auto-delete regardless of which pass found the pair -- Pass 1 only fires
# when the canonical name is a beets ghost; the very common case where BOTH
# copies are already tracked in beets (so Pass 1 defers) surfaces instead via
# Pass 2/3's tag-based grouping, and still deserves the same free pass.
is_numbered_twin() {
local a="$1" b="$2"
[[ "${a%/*}" == "${b%/*}" ]] || return 1
local ba="${a##*/}" bb="${b##*/}"
local ext="${ba##*.}"
[[ "${ext,,}" == "${bb##*.}" ]] || return 1
[[ "$(echo "$ba" | sed -E "s/\.[0-9]+\.${ext}\$/.${ext}/")" \
== "$(echo "$bb" | sed -E "s/\.[0-9]+\.${ext}\$/.${ext}/")" ]]
}
# True if this pair is a same-song format upgrade found by an EXACT tag match
# (Pass 2 case-insensitive or Pass 3 normalized -- both require identical
# albumartist+album+title after case/punctuation folding, so "same song" is
# already proven by the pass itself) where the two files simply have
# different extensions. rank_file() always ranks FLAC ahead of any other
# format regardless of size/dirty-name/etc, so within a tag-exact group a
# cross-extension pair unconditionally means "the format-preferred copy vs.
# a lesser one" -- no judgment call. Deliberately excludes Pass 4
# (cross_album_fuzzy): that pass matches by artist+title-base ACROSS
# different albums/directories, which is exactly where a different master,
# DJ-mix edit, or reissue can share tags without being an interchangeable
# copy -- those still need a human to look at album/context before deleting
# either side.
is_format_upgrade() {
local pass="$1" keep="$2" delete="$3"
[[ "$pass" == "case_insensitive" || "$pass" == "normalized" ]] || return 1
[[ "${keep##*.}" != "${delete##*.}" ]]
}
# Rank a file: lower = better. FLAC > MP3 > other (WAV/etc.); clean filename
# > sync-conflict artifact (-DESKTOP-, (1), .1.); extended/club mix > radio
# edit at same format; then larger wins.
@@ -163,7 +230,7 @@ process_group() {
esac
local first=1
local keep_path=""
local keep_path="" keep_id="" any_auto=0
while IFS="$SEP" read -r _sk id hp; do
[[ -z "$hp" ]] && continue
local size_bytes; size_bytes=$(stat -c '%s' "$hp" 2>/dev/null || echo 0)
@@ -171,14 +238,31 @@ process_group() {
if [[ $first -eq 1 ]]; then
log " KEEP ($size_h) $hp"
keep_path="$hp"
keep_id="$id"
first=0
else
log " DELETE ($size_h) $hp"
json_emit "{\"pass\":\"${pass_name}\",\"keep_path\":\"$(json_escape "$keep_path")\",\"delete_path\":\"$(json_escape "$hp")\",\"delete_size_bytes\":${size_bytes}}"
# Auto-apply-numbered deletes numbered_sibling matches unconditionally,
# bypassing --only-paths (that gate exists only for the manual confirm
# flow, which never runs alongside this flag). All other passes stay
# dry-run/log-only here regardless.
local do_delete=0 auto_applied=0
if [[ $APPLY -eq 1 ]]; then
do_delete=1
elif [[ $AUTO_APPLY_NUMBERED -eq 1 ]] \
&& { is_numbered_twin "$keep_path" "$hp" || is_format_upgrade "$pass_name" "$keep_path" "$hp"; }; then
do_delete=1
auto_applied=1
any_auto=1
fi
local auto_json="false"; [[ $auto_applied -eq 1 ]] && auto_json="true"
json_emit "{\"pass\":\"${pass_name}\",\"keep_path\":\"$(json_escape "$keep_path")\",\"delete_path\":\"$(json_escape "$hp")\",\"delete_size_bytes\":${size_bytes},\"auto_applied\":${auto_json}}"
if [[ $do_delete -eq 1 ]]; then
# With --only-paths given, only delete entries the caller explicitly
# confirmed; without it, delete everything (original behavior).
if [[ -n "$ONLY_PATHS_FILE" && -z "${ONLY_PATHS[$hp]:-}" ]]; then
if [[ $auto_applied -eq 0 && -n "$ONLY_PATHS_FILE" && -z "${ONLY_PATHS[$hp]:-}" ]]; then
log " (skipped — not in --only-paths confirm list)"
elif [[ -n "$id" ]]; then
# In beets — beet remove -d removes from DB + disk. Best-effort: a
@@ -192,6 +276,18 @@ process_group() {
fi
fi
done < <(echo "$ranked" | grep -v '^$' | sort -n)
# An auto-applied numbered-twin deletion can leave the numbered-named file
# as the sole survivor (it won on size despite the dirty-filename penalty).
# Rename it back to the canonical name so the library doesn't accumulate
# ".1."/".2." filenames for tracks that no longer have a duplicate. Only
# auto_applied triggers this (never a manual --only-paths confirm, which
# may leave the sibling undeleted if the user didn't confirm it -- renaming
# then would collide); only when it's actually in beets (a ghost survivor
# has nothing to move); only when the name still has the dirty infix.
if [[ $any_auto -eq 1 && -n "$keep_id" && "$keep_path" =~ \.[0-9]+\.[A-Za-z0-9]+$ ]]; then
beet move "id:${keep_id}" >> "$LOG" 2>&1 || log " WARN: beet move id:${keep_id} failed (rename survivor)"
fi
return 0 # explicit: the loop's last command status must not leak out under set -e
}
@@ -202,9 +298,10 @@ process_group() {
log ""
log "[$(date -Iseconds)] === Pass 1: numbered siblings ==="
# Dump id + container path from beets once. beets normalizes $path for
# display (always /music/-prefixed) even though the DB stores a mix of
# absolute and relative — so the display paths are safe to compare/convert,
# Dump id + display path from beets once. beets normalizes $path for display
# (real /data/music/Library/... paths since the 2026-07-08 rewrite; /music/...
# on any legacy entry) even though the DB stores a mix of absolute and
# relative — so the display paths are safe to compare/convert via host_path(),
# and the ids are what we hand to beet remove/move.
beet ls -f "\$id${SEP}\$path" 2>/dev/null > /tmp/beets-id-paths.txt
cut -d"$SEP" -f2- /tmp/beets-id-paths.txt > /tmp/beets-all-paths.txt
@@ -231,18 +328,22 @@ while IFS="$SEP" read -r numbered_id numbered_cp; do
if [[ $score_canonical -le $score_numbered ]]; then
# Canonical wins: delete numbered (in beets), keep canonical (ghost)
process_group "P1 numbered-sibling: ${numbered_cp##/music/}" \
process_group "P1 numbered-sibling: ${numbered_cp##*/Library/}" \
"$host_canonical" "${numbered_id}${SEP}${numbered_cp}"
if [[ $APPLY -eq 1 ]]; then
if [[ $APPLY -eq 1 || $AUTO_APPLY_NUMBERED -eq 1 ]]; then
# canonical is now a ghost file; import it into beets
beet import -q -s "${canonical_cp}" >> "$LOG" 2>&1 || log " WARN: beet import failed for ${canonical_cp}"
fi
else
# Numbered wins (canonical is ghost): rm the ghost, then beet move to rename
# the numbered file to the canonical path (id query — see process_group).
process_group "P1 numbered-sibling: ${numbered_cp##/music/}" \
# Safe to reach with just AUTO_APPLY_NUMBERED: this loop only ever builds
# genuine numbered-sibling pairs (see is_numbered_twin), and process_group
# above already deleted the ghost canonical before this runs, so the move
# target is guaranteed clear -- no rename collision.
process_group "P1 numbered-sibling: ${numbered_cp##*/Library/}" \
"${numbered_id}${SEP}${numbered_cp}" "$host_canonical"
if [[ $APPLY -eq 1 ]]; then
if [[ $APPLY -eq 1 || $AUTO_APPLY_NUMBERED -eq 1 ]]; then
beet move "id:${numbered_id}" >> "$LOG" 2>&1 || log " WARN: beet move id:${numbered_id} failed"
fi
fi
+12 -4
View File
@@ -458,7 +458,7 @@ def main():
+ (f" upgrade-from={sorted(upgrade_from)}" if upgrade_from else ""))
print(f"[enrich] {len(flacs)} FLAC files in library\n")
looked = found = upgraded = 0
looked = found = upgraded = write_failed = 0
by_source = {s: 0 for s in cascade}
touched = []
for p in flacs:
@@ -519,17 +519,25 @@ def main():
if set_flac_tag(sp, BUY_URL_TAG, url):
touched.append(sp)
else:
print(" ! metaflac write failed")
write_failed += 1
print(" ! metaflac write failed (permissions? disk full?) -- NOT tagged")
else:
touched.append(sp)
if found % 50 == 0:
print(f" ...{found} links from {looked} lookups so far", flush=True)
print(f"\n[enrich] {looked} lookups, {found} {'tagged' if args.apply else 'would-tag'}"
# touched is only appended to on an actual successful write (apply mode)
# or a would-tag match (dry run) -- len(touched) is what really landed on
# disk. `found` counts matches regardless of write outcome, so reporting
# `found` here as "tagged" lied about full success on 2026-07-22, when
# every write in the run failed (root-owned files under explo/) but the
# summary line still read "51 tagged".
print(f"\n[enrich] {looked} lookups, {len(touched)} {'tagged' if args.apply else 'would-tag'}"
f" ({', '.join(f'{s}={by_source[s]}' for s in cascade)})"
f" no-match={looked - found}"
+ (f" upgraded={upgraded}" if upgrade_from else ""))
+ (f" upgraded={upgraded}" if upgrade_from else "")
+ (f" WRITE-FAILED={write_failed}" if write_failed else ""))
if args.apply and touched and az_key and args.azuracast_base:
print("[enrich] telling AzuraCast to reprocess touched files...")
+11 -7
View File
@@ -57,11 +57,12 @@ OUT_DIR = f"{_MUSIC_DATA_DIR}/playlists-laptop"
# Absolute Windows path to the laptop's library root.
LIBRARY_LAPTOP_ROOT = r"C:\Music\Library"
# Beets stores paths with this prefix (the container-side mount of the
# library); strip it to get the path relative to the library root. This is
# still "/music/" through the migration's Stage 0-3 transitional mount;
# revisit at Stage 4 once beets' directory: becomes MUSIC_DATA_DIR/Library.
BEETS_LIBRARY_PREFIX = "/music/"
# Prefixes beets may display library paths under, tried in order: the real
# library dir (everything since the 2026-07-08 path rewrite) and the legacy
# transitional mount (any stale pre-rewrite entry). Strip whichever matches
# to get the path relative to the library root; when neither matches the
# track is outside the library and is skipped.
BEETS_LIBRARY_PREFIXES = (f"{_MUSIC_DATA_DIR}/Library/", "/music/")
M3U_EXT = ".m3u"
# Older formats from earlier iterations of this script — swept by reconcile.
@@ -157,9 +158,12 @@ def build_beets_index() -> dict[tuple[str, str], list[tuple[str, str, str, str]]
if len(parts) != 4:
continue
aa, alb, ti, path = parts
if not path.startswith(BEETS_LIBRARY_PREFIX):
rel = next(
(path[len(p):] for p in BEETS_LIBRARY_PREFIXES if path.startswith(p)),
None,
)
if rel is None:
continue
rel = path[len(BEETS_LIBRARY_PREFIX):]
idx.setdefault((_normkey(alb), _normkey(ti)), []).append((aa, alb, ti, rel))
return idx
+17 -11
View File
@@ -49,8 +49,10 @@ _ALEMBIC_CONFIG_DIR = os.environ.get("ALEMBIC_CONFIG_DIR", "/config")
_MUSIC_DATA_DIR = os.environ.get("MUSIC_DATA_DIR", "/data/music")
LIBRARY = f"{_MUSIC_DATA_DIR}/Library"
# Still "/music" through the migration's Stage 0-3 transitional beets mount;
# revisit at Stage 4 once beets' directory: becomes MUSIC_DATA_DIR/Library.
# Legacy prefix from the migration's transitional beets mount. Since the
# 2026-07-08 path rewrite beets displays real LIBRARY paths, so this only
# matters for accepting pasted /music/... paths as input and as a fallback
# beets query form for any stale pre-rewrite DB entry.
CONTAINER_PREFIX = "/music"
SPOTIFY_ENV = f"{_ALEMBIC_CONFIG_DIR}/pipeline/_spotify.env"
SPOTIFY_GENRE = f"{os.environ.get('PIPELINE_DIR', '/app/pipeline')}/lib/spotify-genre.py"
@@ -316,15 +318,19 @@ def resolve_input_path(arg):
def beets_id_for(host_path):
cp = host_to_container(host_path)
r = subprocess.run(
["beet", "ls", "-f", "$id|||$path", f"path:{cp}"],
capture_output=True, text=True, check=True)
for line in r.stdout.splitlines():
if "|||" in line:
id_str, p = line.split("|||", 1)
if p == cp:
return int(id_str)
# Real path first (how beets displays everything since the 2026-07-08
# path rewrite), legacy /music/... form second for any stale entry.
# Querying only the legacy form (the old behavior) matched nothing after
# the rewrite, so retag-from-url silently skipped beet update/move.
for cp in (host_path, host_to_container(host_path)):
r = subprocess.run(
["beet", "ls", "-f", "$id|||$path", f"path:{cp}"],
capture_output=True, text=True, check=True)
for line in r.stdout.splitlines():
if "|||" in line:
id_str, p = line.split("|||", 1)
if p == cp:
return int(id_str)
return None
+121
View File
@@ -0,0 +1,121 @@
#!/usr/bin/env python3
"""Fold every sldl skip index under a playlist's dropbox into one pinned file.
sldl names its per-playlist index folder after the *input*: the Spotify
playlist's display name for spotify input, the CSV filename stem for csv
input. So when 0.6.2 switched input from Spotify URL to CSV, sldl started a
fresh empty index in a new subfolder, saw no download history, and
re-downloaded every playlist in full (2026-07-16). The rendered confs now pin
index-path to <dropbox>/<playlist>/_index.csv; this script seeds that pinned
file from all the indexes sldl left behind (run-playlist.sh calls it whenever
the pinned file is missing or older than a stranded one).
Usage: merge-sldl-indexes.py <playlist_dropbox_dir> <output_index_csv>
Merge rules, per (artist, title) lowercased:
- the row from the newest index file wins;
- EXCEPT when that row is a failure (state 2) and any older index has the
track as downloaded (state 1 or 3): then the newest row's identity
(artist/album/title/length -- matching what the current input will
present; Spotify length rounding drifted between the old extractor and
our CSV) is kept but marked state 3 (already downloaded), so sldl does
not re-fetch a track the library already holds;
- rows only present in older indexes are kept as-is (tracks since removed
from the playlist; harmless, and they keep their history if re-added).
"""
import csv
import sys
from pathlib import Path
HEADER = ["filepath", "artist", "album", "title", "length", "tracktype", "state", "failurereason"]
DOWNLOADED_STATES = {"1", "3"} # 1 = downloaded this run, 3 = found in index previously
FAILED_STATE = "2"
def read_rows(path: Path) -> list[dict]:
rows = []
with open(path, newline="", encoding="utf-8") as fh:
reader = csv.DictReader(fh)
for row in reader:
if row.get("artist") is None or row.get("title") is None:
continue
rows.append({k: (row.get(k) or "") for k in HEADER})
return rows
def norm_key(row: dict) -> tuple:
# Primary artist only: sldl's old Spotify extractor recorded just the
# first artist, while spotify-playlist-csv.py joins all of them with
# ", " -- keying on the full string would miss every multi-artist track
# when recovering history across the two index generations. First
# comma-segment matches both forms (and both sides of a comma-in-name
# artist like "Tyler, The Creator" truncate identically).
return (row["artist"].split(",")[0].strip().lower(), row["title"].strip().lower())
def main() -> int:
if len(sys.argv) != 3:
print(__doc__, file=sys.stderr)
return 2
dropbox = Path(sys.argv[1])
out_path = Path(sys.argv[2])
if not dropbox.is_dir():
print(f"[merge-sldl-indexes] not a directory: {dropbox}", file=sys.stderr)
return 2
# Every index at the dropbox root or one level down (sldl's input-named
# subfolders), including the pinned output itself if it already exists --
# newest first, so the most recent record of each track wins.
candidates = sorted(
set(dropbox.glob("_index.csv")) | set(dropbox.glob("*/_index.csv")),
key=lambda p: p.stat().st_mtime,
reverse=True,
)
if not candidates:
print(f"[merge-sldl-indexes] no _index.csv found under {dropbox}; nothing to seed")
return 0
merged: dict[tuple, dict] = {}
recovered = 0
for path in candidates:
try:
rows = read_rows(path)
except (OSError, csv.Error) as exc:
print(f"[merge-sldl-indexes] skipping unreadable {path}: {exc}", file=sys.stderr)
continue
print(f"[merge-sldl-indexes] {path}: {len(rows)} rows")
for row in rows:
key = norm_key(row)
kept = merged.get(key)
if kept is None:
merged[key] = row
elif kept["state"] == FAILED_STATE and row["state"] in DOWNLOADED_STATES:
# Newest attempt failed but an older index proves we already
# have this track: keep the newest identity fields, take the
# old filepath (informational only; skip-mode index never
# checks the file on disk), and mark it downloaded.
kept["filepath"] = row["filepath"]
kept["state"] = "3"
kept["failurereason"] = "0"
recovered += 1
out_path.parent.mkdir(parents=True, exist_ok=True)
tmp = out_path.with_suffix(".csv.tmp")
with open(tmp, "w", newline="", encoding="utf-8") as fh:
writer = csv.DictWriter(fh, fieldnames=HEADER)
writer.writeheader()
writer.writerows(merged.values())
tmp.replace(out_path)
downloaded = sum(1 for r in merged.values() if r["state"] in DOWNLOADED_STATES)
print(
f"[merge-sldl-indexes] wrote {out_path}: {len(merged)} tracks "
f"({downloaded} downloaded, {recovered} recovered from older indexes)"
)
return 0
if __name__ == "__main__":
sys.exit(main())
+40 -10
View File
@@ -5,13 +5,30 @@
# Usage:
# echo "single line message" | notify-telegram.sh
# notify-telegram.sh "single line message"
# notify-telegram.sh < ${ALEMBIC_CONFIG_DIR:-/config}/logs/STATUS.log
# notify-telegram.sh --html < ${ALEMBIC_CONFIG_DIR:-/config}/logs/STATUS.log
#
# Telegram messages are capped at 4096 chars. Anything longer is truncated
# with a "...<truncated>" tail. Returns exit 0 on send-OK, non-zero otherwise.
# --html sends with parse_mode=HTML for messages authored as Telegram-HTML
# (pipeline-status.sh's digest: <b>/<i>/<code>/<pre>/<blockquote>). The
# CALLER is responsible for escaping &, <, > in any free text; this script
# only guarantees the truncation below can't cut a tag in half. Without the
# flag, messages go as plain text exactly as before.
#
# Telegram messages are capped at 4096 chars. Anything longer is truncated;
# in HTML mode the cut lands on a line boundary and re-closes an open <pre>
# so the truncated message still parses (a mid-tag cut, or a bare "<" in the
# tail marker, makes the Bot API reject the ENTIRE message with a 400).
# Returns exit 0 on send-OK, non-zero otherwise.
set -euo pipefail
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
# /opt/venv/bin leads PATH for consistency with the other pipeline scripts,
# even though this one only shells out to curl.
PATH=/opt/venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
PARSE_MODE=""
if [[ "${1:-}" == "--html" ]]; then
PARSE_MODE="HTML"
shift
fi
CONFIG="${ALEMBIC_CONFIG_DIR:-/config}/pipeline/telegram/notify.env"
if [[ ! -f "$CONFIG" ]]; then
@@ -30,24 +47,37 @@ else
MSG=$(cat)
fi
# Telegram caps at 4096 chars (UTF-8 codepoints). Trim conservatively at 3900.
# Use python for proper UTF-8 length handling.
# Telegram caps at 4096 chars. Trim conservatively at 3900. Use python for
# proper UTF-8 length handling. Plain mode appends a literal marker; HTML
# mode cuts at the last newline inside the budget (our tags never span
# lines except <pre> blocks) and re-closes an unbalanced <pre>.
MSG=$(python3 -c "
import sys
s = sys.argv[1]
html = sys.argv[2] == 'HTML'
if len(s) > 3900:
s = s[:3900] + '\n...<truncated>'
s = s[:3900]
if html:
cut = s.rfind('\n')
if cut > 0:
s = s[:cut]
if s.count('<pre>') > s.count('</pre>'):
s += '</pre>'
s += '\n<i>… truncated</i>'
else:
s += '\n...(truncated)'
print(s, end='')
" "$MSG")
" "$MSG" "${PARSE_MODE:-plain}")
# Send via Bot API. Disable web-page preview and use plain text (no parse_mode)
# so log content with special chars doesn't get interpreted as Markdown.
# Send via Bot API. Preview disabled; parse_mode only when requested so log
# content with special chars can't be misread as markup in plain sends.
# `|| true`: a curl timeout/network error must fall through to the explicit
# ok-check below (which reports "send failed" and exits 2), not abort here.
resp=$(curl -s --max-time 15 -X POST \
"https://api.telegram.org/bot${TG_BOT_TOKEN}/sendMessage" \
--data-urlencode "chat_id=${TG_CHAT_ID}" \
--data-urlencode "text=${MSG}" \
${PARSE_MODE:+--data-urlencode "parse_mode=${PARSE_MODE}"} \
--data-urlencode "disable_web_page_preview=true" || true)
ok=$(echo "$resp" | python3 -c "import sys,json;print(json.load(sys.stdin).get('ok',False))" 2>/dev/null || echo False)
+125 -47
View File
@@ -4,9 +4,24 @@
# Writes a one-screen summary to ${ALEMBIC_CONFIG_DIR:-/config}/logs/STATUS.log (overwritten daily)
# and ships it to Telegram via notify-telegram.sh.
#
# Formatting note: the digest uses Telegram HTML entities (<b>, <i>, <code>,
# <blockquote>) — the brand's chrome/dot-badge language translated into what
# Telegram can actually render (no color, no custom font, so status "dots"
# become colored circle emoji). Deliberately NO <pre>/monospace column layout:
# a code block forces fixed-width columns that wrap into unreadable garbage on
# a phone ("54 tracks in / M3U"). Instead every row is plain reflowing text —
# a bold label + " — value" — so it wraps cleanly at any screen width. This
# REQUIRES notify-telegram.sh to be called with --html (parse_mode=HTML) —
# sent as plain text the tags would show up literally. All free-text going
# through mark_ok/mark_warn/mark_skip/section is escaped (esc()) for &, <, >
# so a stray angle bracket in a log line can't break the HTML parse and
# swallow the whole message.
#
# Reports on:
# - Sibling service reachability (slskd, navidrome) via HTTP, not docker ps —
# alembic has no Docker socket access
# - Sibling service reachability (navidrome) via HTTP, not docker ps —
# alembic has no Docker socket access. slskd is deliberately NOT checked:
# it is not part of the pipeline (sldl is its own Soulseek client) — it
# runs on the host purely as the user's own file-sharing presence.
# - Today's runs: playlist syncs, Bandcamp sync, manual imports, dedup
# - Weekly maintenance freshness: strip-mb-tags, strip-watermark-art,
# scrub-watermark-text, clean-sldl-index, spotify-genre
@@ -27,7 +42,10 @@
# the opposite of what a health report should do. It writes no library state,
# so there is nothing to leave half-applied.
set -u
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
# /opt/venv/bin first: `beet` lives in the app venv. Without it every beet
# probe here ("added today", "Library by format", the mp3-now count) silently
# came up empty behind its 2>/dev/null.
PATH=/opt/venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
OUT=${ALEMBIC_CONFIG_DIR:-/config}/logs/STATUS.log
TODAY=$(date +%Y%m%d)
@@ -44,12 +62,39 @@ OK=0
WARN=0
SKIP=0
mark_ok() { OK=$((OK+1)); printf " ✓ %s\n" "$*"; }
mark_warn() { WARN=$((WARN+1)); printf " ⚠ %s\n" "$*"; }
mark_skip() { SKIP=$((SKIP+1)); printf " ⊘ %s\n" "$*"; }
mark_info() { printf " %s\n" "$*"; }
# Escapes &, <, > so free text (log snippets, exception messages, filenames)
# can't be mistaken for an HTML entity by Telegram's parser.
esc() {
local s=$1
s=${s//&/&amp;}
s=${s//</&lt;}
s=${s//>/&gt;}
printf '%s' "$s"
}
section() { printf "\n▎ %s\n" "$*"; }
# One status row: a colored dot, a bold label, and (optionally) " — value".
# Plain proportional text — NO padding, NO monospace — so it reflows on mobile
# instead of wrapping mid-column. The dot is the brand's badge-state color,
# emoji being Telegram's only color channel.
_row() {
local dot=$1 label=$2 detail=${3:-}
if [[ -n "$detail" ]]; then
printf "%s <b>%s</b> — %s\n" "$dot" "$(esc "$label")" "$(esc "$detail")"
else
printf "%s <b>%s</b>\n" "$dot" "$(esc "$label")"
fi
}
mark_ok() { OK=$((OK+1)); _row "🟢" "$@"; }
mark_warn() { WARN=$((WARN+1)); _row "🟠" "$@"; }
mark_skip() { SKIP=$((SKIP+1)); _row "⚪" "$@"; }
# Free-text info line (no dot, no counter) — for sub-breakdowns.
mark_info() { printf " %s\n" "$(esc "$*")"; }
# Bold section header (reads like the web app's <h2>). No <pre> — the rows
# under it are plain reflowing text.
section() {
printf "\n<b>▎ %s</b>\n" "$(esc "$*")"
}
# syslog / `logger` is not present in the container image. Only call it if it
# exists, so these lines don't spew "logger: command not found" into the job
@@ -80,7 +125,7 @@ playlist_status() {
m3u=$(awk 'match($0, /updated with [0-9]+ tracks/) {
n=substr($0, RSTART, RLENGTH); gsub(/[^0-9]/, "", n); last=n
} END { print last }' "$log")
echo "OK|${m3u:-0} tracks in M3U"
echo "OK|${m3u:-0} tracks"
elif grep -q "=== Starting playlist run" "$log" 2>/dev/null; then
echo "WARN|started but never finished"
else
@@ -140,12 +185,11 @@ dedup_today_status() {
check_http() {
local name="$1" url="$2"
if curl -s -o /dev/null --connect-timeout 3 --max-time 6 "$url"; then
mark_ok "$(printf '%-11s reachable' "$name")"
mark_ok "$name" "reachable"
else
mark_warn "$(printf '%-11s unreachable at %s' "$name" "$url")"
mark_warn "$name" "unreachable at $url"
fi
}
check_http slskd "http://gluetun:5030/"
check_http navidrome "http://navidrome:4533/rest/ping.view"
# 2. Today's runs
@@ -186,14 +230,12 @@ dedup_today_status() {
fi
;;
esac
line=$(printf '%-13s %s' "$short" "$detail")
if [[ "$tag" == "OK" ]]; then mark_ok "$line"; else mark_warn "$line"; fi
if [[ "$tag" == "OK" ]]; then mark_ok "$short" "$detail"; else mark_warn "$short" "$detail"; fi
done
if [[ -n "$newest_dedup_today" ]]; then
any_today=1
IFS='|' read -r tag detail < <(dedup_today_status "$newest_dedup_today")
line=$(printf '%-13s %s' "dedup" "$detail")
if [[ "$tag" == "OK" ]]; then mark_ok "$line"; else mark_warn "$line"; fi
if [[ "$tag" == "OK" ]]; then mark_ok "dedup" "$detail"; else mark_warn "dedup" "$detail"; fi
fi
[[ "$any_today" -eq 0 ]] && mark_info "(nothing has run yet today)"
@@ -205,13 +247,16 @@ dedup_today_status() {
ADDED_TODAY=$(beet ls -f '$grouping' \
"added:$(date +%Y-%m-%d).." 2>/dev/null || true)
added_total=$(echo -n "$ADDED_TODAY" | grep -c '^' || true)
mark_ok "$(printf '%-13s %d new tracks in beets' 'added today' "$added_total")"
mark_ok "added today" "$added_total new tracks in beets"
if [[ "$added_total" -gt 0 ]]; then
# Per-playlist breakdown. Empty $grouping (bandcamp/manual) shown as "(none)".
# Piped through the same &/</> escaping as esc(), since this prints
# straight into the message without going through _row.
echo "$ADDED_TODAY" \
| awk '{ if ($0 == "") print "(none)"; else print }' \
| sort | uniq -c | sort -rn \
| awk '{ g=$2; for(i=3;i<=NF;i++) g=g" "$i; printf " %3d %s\n", $1, g }'
| awk '{ c=$1; g=$2; for(i=3;i<=NF;i++) g=g" "$i; printf " • %s (%d)\n", g, c }' \
| sed 's/&/\&amp;/g; s/</\&lt;/g; s/>/\&gt;/g'
fi
# 3. Maintenance freshness
@@ -284,7 +329,7 @@ PYEOF
row=$(grep "^${task}|" <<< "$maint_rows" || true)
IFS='|' read -r _ latest_status latest_age age_s log <<< "$row"
if [[ -z "$row" || "$latest_status" == "none" ]]; then
mark_skip "$(printf '%-22s never run yet' "$task")"
mark_skip "$task" "never run yet"
continue
fi
if [[ "$age_s" -lt 0 ]]; then
@@ -294,7 +339,7 @@ PYEOF
running) note="running now" ;;
*) note="$latest_status" ;;
esac
mark_warn "$(printf '%-22s never succeeded (last attempt %s: %s)' "$task" "$(human_age "$latest_age")" "$note")"
mark_warn "$task" "never succeeded (last attempt $(human_age "$latest_age"): $note)"
continue
fi
age_h=$(human_age "$age_s")
@@ -332,11 +377,12 @@ PYEOF
failed) snip="${snip:+$snip, }latest attempt failed" ;;
skipped_lock) snip="${snip:+$snip, }latest attempt skipped, lock busy" ;;
esac
line=$(printf '%-22s %-9s %s' "$task" "$age_h" "${snip:-}")
detail="$age_h"
[[ -n "$snip" ]] && detail="$age_h · $snip"
if (( age_s > MAINT_LIMIT[$task] )) || [[ "$latest_status" == "failed" ]]; then
mark_warn "$line"
mark_warn "$task" "$detail"
else
mark_ok "$line"
mark_ok "$task" "$detail"
fi
done
@@ -348,13 +394,16 @@ PYEOF
exp=$(awk -F'\t' '$6=="identity" {print $5; exit}' ${ALEMBIC_CONFIG_DIR:-/config}/pipeline/bandcamp/cookies.txt)
if [[ -n "${exp:-}" && "$exp" =~ ^[0-9]+$ && "$exp" -gt 0 ]]; then
days=$(( (exp - NOW_TS) / 86400 ))
line=$(printf '%-22s %d days left' "Bandcamp cookie" "$days")
if (( days < 14 )); then mark_warn "$line — re-export soon"; else mark_ok "$line"; fi
if (( days < 14 )); then
mark_warn "Bandcamp cookie" "$days days left — re-export soon"
else
mark_ok "Bandcamp cookie" "$days days left"
fi
else
mark_warn "Bandcamp cookie could not parse expiry"
mark_warn "Bandcamp cookie" "could not parse expiry"
fi
else
mark_warn "Bandcamp cookie missing (${ALEMBIC_CONFIG_DIR:-/config}/pipeline/bandcamp/cookies.txt)"
mark_warn "Bandcamp cookie" "missing (${ALEMBIC_CONFIG_DIR:-/config}/pipeline/bandcamp/cookies.txt)"
fi
# Qobuz Web Player token (buy-link enrich cascade #2). The token is opaque
@@ -374,12 +423,20 @@ PYEOF
-H "X-App-Id: $q_appid" -H "X-User-Auth-Token: $q_token" -w '%{http_code}' \
"https://www.qobuz.com/api.json/0.2/track/search?query=test&limit=1&app_id=$q_appid" 2>/dev/null)
case "$q_code" in
200) mark_ok "$(printf '%-22s %s' 'Qobuz token' 'valid (buy-link lookup live)')" ;;
401) mark_warn "$(printf '%-22s %s' 'Qobuz token' "EXPIRED — re-export X-User-Auth-Token to ${ALEMBIC_CONFIG_DIR:-/config}/pipeline/qobuz/token")" ;;
*) mark_warn "$(printf '%-22s %s' 'Qobuz token' "check failed (HTTP ${q_code:-none})")" ;;
200) mark_ok "Qobuz token" "valid (buy-link lookup live)" ;;
401) mark_warn "Qobuz token" "EXPIRED — re-export X-User-Auth-Token to ${ALEMBIC_CONFIG_DIR:-/config}/pipeline/qobuz/token" ;;
# A genuinely expired/bad token gets a JSON 401 from Qobuz's own API.
# 403 instead means the request never reached that code at all -- Qobuz's
# Akamai edge is rejecting the request outright (seen 2026-07-22: the
# exact same request got this "Access Denied"/edgesuite.net block from
# alembic's VPN egress IP, but a clean 401 from a non-VPN IP). Re-
# exporting the token does nothing for that -- it's the egress IP's
# reputation, not the credential. Say so, so it isn't mistaken for 401.
403) mark_warn "Qobuz token" "blocked (HTTP 403, likely Akamai/CDN, not the token) — VPN egress IP may be flagged; re-exporting the token won't fix this" ;;
*) mark_warn "Qobuz token" "check failed (HTTP ${q_code:-none})" ;;
esac
else
mark_warn "$(printf '%-22s %s' 'Qobuz token' "missing (${ALEMBIC_CONFIG_DIR:-/config}/pipeline/qobuz/token)")"
mark_warn "Qobuz token" "missing (${ALEMBIC_CONFIG_DIR:-/config}/pipeline/qobuz/token)"
fi
# VPN egress
@@ -389,17 +446,20 @@ PYEOF
# asking Docker to introspect gluetun's health, and needs no Docker socket.
egress_ip=$(curl -s --connect-timeout 5 --max-time 10 https://ifconfig.me 2>/dev/null)
if [[ -n "$egress_ip" ]]; then
mark_ok "$(printf '%-22s %s' 'VPN egress' "$egress_ip")"
mark_ok "VPN egress" "$egress_ip"
else
mark_warn "$(printf '%-22s %s' 'VPN egress' 'could not determine egress IP — VPN may be down')"
mark_warn "VPN egress" "could not determine egress IP — VPN may be down"
fi
# 5. Storage
section "Storage"
while read -r mp pct used total; do
line=$(printf '%-22s %s used (%s of %s)' "$mp" "$pct" "$used" "$total")
pct_n=${pct%\%}
if (( pct_n > 90 )); then mark_warn "$line"; else mark_ok "$line"; fi
if (( pct_n > 90 )); then
mark_warn "$mp" "$pct used ($used of $total)"
else
mark_ok "$mp" "$pct used ($used of $total)"
fi
done < <(df -h ${MUSIC_DATA_DIR:-/data/music} / 2>/dev/null | awk 'NR>1 && $1!~/tmpfs/ {print $6, $5, $3, $2}')
# 6. Navidrome
@@ -424,24 +484,40 @@ except Exception as e:
print(f\"WARN|unreachable: {e}\")
" 2>/dev/null)
state=${ND_LINE%%|*}; detail=${ND_LINE##*|}
line=$(printf '%-22s %s' "Navidrome" "$detail")
if [[ "$state" == "OK" ]]; then mark_ok "$line"; else mark_warn "$line"; fi
if [[ "$state" == "OK" ]]; then mark_ok "Navidrome" "$detail"; else mark_warn "Navidrome" "$detail"; fi
# 7. Library by format (info-only, no status pill)
# 7. Library by format (info-only, no status dot)
section "Library by format"
beet ls -f '$format' 2>/dev/null | sort | uniq -c | sort -rn \
| awk '{printf " %-6s %d tracks\n", $2, $1}'
| awk '{printf " %s — %d tracks\n", $2, $1}' \
| sed 's/&/\&amp;/g; s/</\&lt;/g; s/>/\&gt;/g'
} > "$OUT"
# ---- Build header banner and prepend ----
# Bold brand mark (⚗️ is literally the "ALEMBIC" unicode glyph — no generic
# music-note stand-in needed), a hostname chip in <code>, and an italic tally
# that reuses the same 🟢/🟠/⚪ dots as the body. When something needs
# attention, a <blockquote> callout mirrors the web app's .notice.warning
# panel; an all-clear run gets the calm .notice.success equivalent instead.
HOSTNAME_SHORT=$(hostname -s)
DATE_HUMAN=$(TZ="${TZ:-UTC}" date '+%a %Y-%m-%d %H:%M %Z')
BANNER_TITLE="🎵 Music pipeline · $HOSTNAME_SHORT · $DATE_HUMAN"
BANNER_TALLY=" $OK OK · $WARN warn · $SKIP skip"
sed -i "1c\\
${BANNER_TITLE}\\
${BANNER_TALLY}" "$OUT"
HEADER="⚗️ <b>Alembic</b> · music pipeline · <code>${HOSTNAME_SHORT}</code> · ${DATE_HUMAN}
<i>🟢 ${OK} OK 🟠 ${WARN} warn ⚪ ${SKIP} skip</i>"
if (( WARN > 0 )); then
HEADER="${HEADER}
<blockquote>🟠 <b>${WARN} issue(s)</b> need attention — see below</blockquote>"
else
HEADER="${HEADER}
<blockquote>🟢 All clear — nothing needs attention.</blockquote>"
fi
# Swap the __HEADER_PLACEHOLDER__ line for the (2-3 line) banner above. A
# straight prepend, not sed 1c — the banner's line count varies with the
# blockquote, which sed's `c` command can't take as a variable easily.
tail -n +2 "$OUT" > "${OUT}.body"
{ printf '%s\n' "$HEADER"; cat "${OUT}.body"; } > "$OUT"
rm -f "${OUT}.body"
# Always mirror the one-line summary to syslog so journalctl shows last-run
# state even if Telegram is broken. Warnings get a separate WARN tag.
@@ -450,10 +526,12 @@ if (( WARN > 0 )); then
slog "WARN: pipeline-status flagged $WARN issue(s) — see $OUT"
fi
# Send to Telegram. If it fails, log the failure to syslog so the absence of
# a message in the chat has a corresponding journal entry to grep for.
# Send to Telegram with --html: the digest is Telegram-HTML (see header
# comment), so notify-telegram.sh must send it with parse_mode=HTML. If the
# send fails, log to syslog so the absence of a message in the chat has a
# corresponding journal entry to grep for.
if [[ -x ${PIPELINE_DIR:-/app/pipeline}/lib/notify-telegram.sh ]]; then
if ! ${PIPELINE_DIR:-/app/pipeline}/lib/notify-telegram.sh < "$OUT" 2>/tmp/tg-err; then
if ! ${PIPELINE_DIR:-/app/pipeline}/lib/notify-telegram.sh --html < "$OUT" 2>/tmp/tg-err; then
slog "WARN: Telegram send failed: $(cat /tmp/tg-err 2>/dev/null | head -c 200)"
rm -f /tmp/tg-err
fi
+12 -1
View File
@@ -195,7 +195,18 @@ while IFS= read -r -d '' new_file; do
existing_id="${match%%$'\x1f'*}"
existing_container="${match#*$'\x1f'}"
existing="${MUSIC_DATA_DIR:-/data/music}/Library${existing_container#/music}"
# beets displays real ${MUSIC_DATA_DIR}/Library/... paths since the
# 2026-07-08 path rewrite -- use them as-is; only a legacy /music/...
# display path still needs the prefix swap. Blindly prepending (the old
# behavior) doubled the prefix, so every match logged [stale-db], the MP3
# never got replaced, and the leftover-import step at the bottom imported
# the staged FLAC as a NEW track -- producing exactly the flac+mp3 twins
# this script exists to prevent.
if [[ "$existing_container" == /music/* ]]; then
existing="${MUSIC_DATA_DIR:-/data/music}/Library${existing_container#/music}"
else
existing="$existing_container"
fi
if [[ ! -f "$existing" ]]; then
echo "[stale-db] beets has $existing_container but file missing" | tee -a "$LOG"
continue
+1 -1
View File
@@ -29,7 +29,7 @@ Usage:
scrub-watermark-text.py # dry run
scrub-watermark-text.py --apply # actually strip
"""
import sys, re, argparse
import os, sys, re, argparse
from pathlib import Path
from mutagen import File as MFile
from mutagen.id3 import ID3, ID3NoHeaderError
+11 -4
View File
@@ -51,10 +51,17 @@ def extract_flac_picture(flac_path):
with tempfile.NamedTemporaryFile(delete=False, suffix=".pic") as tf:
tmp = tf.name
try:
r = subprocess.run(
["metaflac", f"--export-picture-to={tmp}", flac_path],
capture_output=True, timeout=15
)
try:
r = subprocess.run(
["metaflac", f"--export-picture-to={tmp}", flac_path],
capture_output=True, timeout=60
)
except subprocess.TimeoutExpired:
# A transient IO stall on one file must not abort the whole weekly
# run (seen 2026-07-15: a healthy 190KB cover took >15s under disk
# contention and the raised TimeoutExpired failed the entire job).
print(f"[strip-art] WARN: metaflac timed out on {flac_path}, skipping file")
return None
if r.returncode != 0 or not os.path.exists(tmp):
return None
sz = os.path.getsize(tmp)
+13 -25
View File
@@ -7,7 +7,15 @@
# Run as root (cron). Logs to ${ALEMBIC_CONFIG_DIR:-/config}/logs/bandcamp-YYYYMMDD.log.
set -euo pipefail
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
# /opt/venv/bin must lead PATH -- that's where `beet` lives (the app's own
# subprocess env puts it there too, see pipeline_runner._subprocess_env()).
# This script used to run as a bare root cron job (pre-2026-07-08 cutover to
# the app scheduler) where a minimal hardened PATH made sense; omitting the
# venv here silently broke every downstream `beet import` call inside
# import-track.sh once this started running through the app instead --
# import-track.sh swallows that failure and still reports OK, so a purchase
# would sit unimported until someone noticed it missing from Navidrome.
PATH=/opt/venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
export PATH
CONFIG="${ALEMBIC_CONFIG_DIR:-/config}/pipeline/bandcamp/config.env"
@@ -87,7 +95,10 @@ shopt -u dotglob nullglob
# Hand off to the standard manual-import pipeline. No playlist tag — Bandcamp
# purchases aren't part of any Spotify playlist. The import-track.sh guard,
# albumartist fallback, beets import, and Navidrome scan all kick in normally.
# albumartist fallback, beets import, straggler/dedup safety net, and
# Navidrome scan all kick in normally (import-track.sh runs the
# replace-with-better + dedup-library safety net itself now, for every
# caller, not just this one).
log "Calling import-track.sh"
if ${PIPELINE_DIR:-/app/pipeline}/bin/import-track.sh >> "$LOG" 2>&1; then
log "import-track.sh OK"
@@ -95,28 +106,5 @@ else
log "WARN: import-track.sh exited non-zero (exit $?)"
fi
# Safety net: if anything is stranded in ${MUSIC_DATA_DIR:-/data/music}/downloads (rare with
# duplicate_action=keep set in beets config, but possible for files beets
# couldn't process), find each one's library counterpart and replace if the
# new file is higher quality, OR import as new if there's no counterpart.
log "Running replace-with-better safety pass on /downloads stragglers"
if ${PIPELINE_DIR:-/app/pipeline}/lib/replace-with-better.sh --apply >> "$LOG" 2>&1; then
log "replace-with-better OK"
else
log "WARN: replace-with-better.sh exited non-zero (exit $?)"
fi
# Now dedupe across the library. With duplicate_action=keep beets imported
# every Bandcamp file even when it conflicts with an existing track at the
# same path (it creates `.1.ext` siblings). dedup-library.sh's policy is
# "FLAC > MP3, then largest file" — Bandcamp version wins, soulseek version
# is removed from both the beets DB and disk.
log "Running dedup pass (Bandcamp FLAC wins over older Soulseek copies)"
if ${PIPELINE_DIR:-/app/pipeline}/lib/dedup-library.sh --apply >> "$LOG" 2>&1; then
log "dedup OK"
else
log "WARN: dedup-library.sh exited non-zero (exit $?)"
fi
log "=== Bandcamp sync done ==="
exit 0
+2 -1
View File
@@ -20,7 +20,8 @@
# upgrade-mp3-to-flac.sh --csv-only # just write the CSV; don't run sldl
set -euo pipefail
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
# /opt/venv/bin must lead PATH -- `beet` (used heavily below) lives there.
PATH=/opt/venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
export PATH
# These helpers are the same ones import-track.sh sources and runs under its