mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-10-05 14:11:13 +01:00
A download that hits Cloudflare hung on "Bypassing protection..." for five minutes and then died, regardless of which bypasser was configured. html_get_page() started a BypassHeartbeat thread to keep the download marked alive during a bypass, but the thread had no loop: it fired one status event and returned. Even with the loop restored it could not have worked, because update_download_status() dedupes identical (status, message) tuples and returns before refreshing _last_activity, and the heartbeat re-sent the byte-identical payload already emitted just above it. So _last_activity was frozen for the whole bypass, while both bypassers are allowed to run longer than STALL_TIMEOUT (external FlareSolverr ~394s at default settings, internal 420s per get() call). The watchdog always won. From a reporter's log: 403 at 07:04:33.390, cancelled at 07:09:33.987 - exactly 300.000s, and 41s before the bypasser would have finished and reported the real error, an HTTP 500 from FlareSolverr the user never saw. The regression is not one commit.1f093de(#536) added the heartbeat and the dedup together and refreshed activity before the dedup return, so it worked.ff094be(#832) moved the refresh below that return while tightening stall detection for #823.3a3a3ce(#845) then deleted the heartbeat's while loop to silence a B023 lint, removing the last evidence of intent. The dedup itself is correct and stays: a keep-alive that ticks on a timer proves nothing about whether an operation is progressing, so letting it refresh the stall clock would make a wedged download immortal. Split the two concerns instead. Add shelfmark/download/activity.py. A long single-shot operation declares its own upper bound once, over a sentinel status carried on the existing status_callback channel - so no new parameter has to be threaded through every handler, post-processor and output module. The orchestrator intercepts the sentinel in its per-task closure and records an absolute deadline in _activity_grace, which stall detection honours alongside STALL_TIMEOUT. The grace never extends itself and is clamped to _MAX_ACTIVITY_GRACE_SECONDS, so an operation that overruns its own declared budget is still cancelled. Each bypasser now reports max_duration_seconds() derived from its own retry and timeout settings, and http.py asks whichever is active, plus 30s of slack so the bypasser's own deadline expires first and the user sees its real failure. On that path html_get_page() also emits status_callback("error", ...) rather than silently returning an empty page. Three further fixes on the same code path: - Extract the watchdog into _find_stalled_tasks() and _cancel_stalled_task(). It was the only place holding _progress_lock across a call into book_queue, whose terminal-status hooks reach a sqlite write that gevent does not patch, blocking the hub and every download worker. It now holds the lock for dict reads only. - Bound _CDP_WORKER.run(), which waited with timeout=None while holding the module-wide LOCKED, so a single wedged in-process CDP session blocked every subsequent bypass forever on non-Docker installs. - Broaden the coordinator loop's except clause back to Exception, with escalating backoff.8d98e12(#868) narrowed it to a six-type tuple to silence BLE001, which let gevent's LoopExit and similar kill the only thread driving the download queue - undoing #832's fix for #823 and resurfacing it as #1166. GreenletExit and gevent.Timeout still propagate. Fixes #1001 Refs #1166, #823