0.6.4: Truthful maintenance digest + scheduled jobs queue for the lock instead of skipping

The Telegram digest reported healthy weekly jobs as "never run yet": its
maintenance section still globbed for cron-era per-script log filenames
(clean-index-*.log etc.) that jobs run through pipeline_runner no longer
write. It now reads job_runs, the scheduler's own record, so it reports the
last success (age + summary snippet from that run's log), flags a newer
failed or lock-skipped attempt on the same line, and distinguishes "never
succeeded (last attempt: skipped, lock busy)" from genuinely never
scheduled. Also matched the clear-bad-genres snippet grep to the script's
current summary wording, and dropped the now-unused newest_log helper.

The DB also showed WHY two jobs had never run: on Sunday the 08:30
strip-mb-tags run took 10m19s while holding the shared pipeline lock, so
strip-watermark-art (08:35) and scrub-watermark-text (08:40) hit flock-style
instant skip and lost their only slot of the week -- the 08:40 job missed by
19 seconds. Scheduled runs now wait up to 30 minutes for the lock
(SCHEDULED_LOCK_WAIT_SECONDS) and only then record skipped_lock, so
fixed-time blocks queue instead of starving; manual "Run now" keeps the
instant skip since a person expects an immediate answer.

Tests: scheduled-run queueing, lock-wait timeout, and manual instant-skip.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
andrew
2026-07-15 10:13:02 -06:00
parent 8ecff44811
commit 7a49e2d507
4 changed files with 167 additions and 64 deletions
+32 -2
View File
@@ -25,11 +25,41 @@ def test_run_job_failure():
assert run.exit_code == 1
def test_lock_skip_when_held():
def test_manual_run_skips_immediately_when_lock_held():
async def scenario():
await pr._lock.acquire()
try:
return await pr.run_job("test:locked", ["true"], use_lock=True)
return await pr.run_job("test:locked", ["true"], use_lock=True, triggered_by="manual")
finally:
pr._lock.release()
run = _run(scenario())
assert run.status == "skipped_lock"
def test_scheduled_run_waits_for_lock_by_default(monkeypatch):
# scheduled runs must QUEUE behind a held lock (bounded), not skip --
# instant-skip starved the Sunday maintenance block (see
# SCHEDULED_LOCK_WAIT_SECONDS)
monkeypatch.setattr(pr, "SCHEDULED_LOCK_WAIT_SECONDS", 5.0)
async def scenario():
await pr._lock.acquire()
task = asyncio.ensure_future(pr.run_job("test:queued", ["true"], triggered_by="schedule"))
await asyncio.sleep(0.2) # let the job start waiting on the lock
pr._lock.release()
return await task
run = _run(scenario())
assert run.status == "success"
assert not pr._lock.locked()
def test_scheduled_run_skips_after_lock_wait_timeout():
async def scenario():
await pr._lock.acquire()
try:
return await pr.run_job("test:waited_out", ["true"], triggered_by="schedule", lock_wait=0.2)
finally:
pr._lock.release()