mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-09-24 20:30:31 +01:00
3e2a7a48d5769dbe963732976bbdd28522f67c37
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3e2a7a48d5 |
fix: clear the DDoS-Guard cookie probe on AA search (#1209)
## Summary Two failure modes on the same code path, both reported this week: Anna's Archive `/search` is gated behind a DDoS-Guard cookie probe that the manual redirect follower can never satisfy. **#1202 — the cookie is dropped on every hop.** AA URLs set `allow_redirects = False`, so `html_get_page` follows redirects by hand. The 302 to `?check=1` carries a `Set-Cookie` (`__ddg*`) that has to come back on the next request. Because cookies are passed per call and `requests` keeps no jar across manual hops, it was discarded each time and the server just re-issued the same redirect until `_MAX_REDIRECTS` raised `TooManyRedirects`. The file already had the right helper — `_new_cookies()` — but only the 503 Z-Library handshake branch called it. **#1204 — the loop never reaches the bypasser.** `TooManyRedirects` isn't in `_is_retryable_error` and carries no status code, so the 403 rescue path (`status == _HTTP_STATUS_FORBIDDEN`) never fired and all attempts repeated the identical failure — ~2.5 min, surfacing as the misleading "Network restricted or mirrors are blocked". These interact, which is why #1202's fix alone isn't enough. Requests merge as `cookies={**handshake_cookies, **cookies}`, so **stale bypasser cookies override the fresh handshake ones** — once `_cf_cookies` holds an expired `__ddg*`, the probe can never clear no matter how faithfully we echo. Hence one search per restart, exactly as #1204 describes. ## Changes 1. Harvest cookies in the same-host redirect branch, the way the 503 branch already does. `_new_cookies()` returns only *new* values, so a server re-sending an identical cookie yields an empty dict and a genuine redirect loop still terminates at `_MAX_REDIRECTS`. 2. Treat a redirect loop as a detected challenge: purge the stored cookies for that host and switch to the bypasser, instead of burning the retry budget. Gated on `allow_bypasser_fallback` and `_is_cf_bypass_enabled()`, and skipped when already bypassing, so AudiobookBay (`allow_bypasser_fallback=False`) and external-bypasser setups are unaffected. The broader point in #1204 stands — the fallback would be better gated on "challenge detected" than on specific status codes, since DDoS-Guard presents at least three faces (403 js-challenge, 429, and this redirect loop). This PR fixes the two live exits without that refactor. ## Tests Two regression tests, both failing before and passing after: - `test_html_get_page_echoes_cookies_across_same_host_redirects` — the fake server only returns results if `__ddg2_` comes back on the `?check=1` hop. - `test_html_get_page_redirect_loop_purges_cookies_and_bypasses` — asserts the stored cookies are cleared, the bypasser runs, and the loop is cut short rather than repeated per attempt. `ruff check` and `ruff format` clean. `tests/download/` passes except `test_download_url_ignores_zlib_cookie_refresh_failure`, which fails identically on unmodified `main` in my environment (no `seleniumbase` — the `browser` extra isn't installed). ## Verification Applied on a live v1.3.7 install (Debian LXC, internal CDP bypasser). Before: every search timed out through 10 retries with `TooManyRedirects`, zero results. After: ``` http.py:455 - Redirect loop detected; switching to bypasser internal_bypasser.py:756 - Bypass successful using _bypass_method_cdp_gui_click internal_bypasser.py:322 - Extracted 9 protection cookies for annas-archive.pk direct_download.py:1865 - Found 24 releases via ISBN ``` ~25 s per search, results render. Note the second search still re-solves the challenge, since the freshly stored cookies go stale immediately — the design issue #1204 raises, left for the broader fix. Fixes #1202 Fixes #1204 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_012Ln3yVj3sWHG2c6T78W1we --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: CaliBrain <calibrain@l4n.xyz> |
||
|
|
0a5256ecbb |
fix(download): reconcile the two AA redirect-loop rescues (#1213)
#1210 and #1212 both added a DDoS-Guard `?check=1` rescue, and #1212 was branched before #1210 landed, so the merged result had two of them with identical guards. #1212's inline handoff returns before the raise that #1210's exception handler keys on, so the handler was shadowed and its stale-cookie purge — the substance of #1210 — never ran. Its regression test has been failing on main since the merge. Fold both into one path: - `_redirect_loop_handoff()` purges the host's stale clearance cookies, then bypasses, so the inline AA handoff and the exception handler cannot drift apart again. - The exception handler keeps its own reason to exist: non-AA hosts run with allow_redirects=True, so `requests` raises the loop itself and the manual AA follower never sees it. It now invokes the bypasser directly rather than setting a flag and continuing, which was a no-op at MAX_RETRY=1 for the same reason the 403 handoff was. - An unrescuable loop returns empty instead of raising TooManyRedirects into the retry path. That error is not retryable and carries no status, so `/dyn/md5/summary` (allow_bypasser_fallback=False) re-ran the full 6-redirect loop on all 10 attempts: 60 requests to AA and ~30s of backoff, measured. Every AA mirror shares the challenge, so there is nothing to rotate to. - `allow_bypasser_fallback` docs now describe what the flag actually gates; the old text predated #1198 and named the wrong callers. |
||
|
|
056ddd372a |
Send DDoS-Guard's ?check=1 redirect loop to the bypasser (#1210)
Fixes #1204. ## Problem #1198 sends a gated AA `/search` to the bypasser when the origin answers 403. DDoS-Guard has a second response: when the clearance cookies from an earlier solve go stale, it serves an endless `?check=1` redirect instead. `requests` follows that until `_raise_too_many_redirects`, and `TooManyRedirects` carries no status code, so `status == _HTTP_STATUS_FORBIDDEN` is false and the rescue never runs. All 10 retries re-send the same dead cookies, then the search fails as `Unable to reach download source. Network restricted or mirrors are blocked.` Direct-download search therefore works once per container start, and stays dead after the stored cookie ages out. v1.3.7 (`sha256:520715f3…`), internal bypasser, mirrors `.gl/.pk/.gd`: ``` 17:04:36 internal_bypasser.py:756 - Bypass successful using _bypass_method_cdp_gui_click ... 17:11:39 http.py:483 - Retry 1/10 for https://annas-archive.gl/search?...&check=1: TooManyRedirects: Too many redirects 17:12:12 http.py:493 - Giving up after 10 attempts 17:12:12 main.py:2870 - Release search failed for source direct_download: Unable to reach download source. Network restricted or mirrors are blocked. ``` The token is short-lived, which is what makes this reachable in normal use: ``` $ curl -sD - 'https://annas-archive.gl/search?...&check=1' HTTP/2 403 server: ddos-guard set-cookie: __ddg8_=…; Expires=Fri, 14-Aug-2026 15:39:38 GMT # issued 15:19:38, 20 min ``` ## Fix Handle the loop like the 403: drop the domain's stored cookies, then retry through the bypasser. The branch sits above the `status ==` ladder because `_get_status_code()` returns `None` for this exception. Cookies are purged only for the internal bypasser; with an external one `get_cf_cookies_for_domain()` already returns `{}`. Related but not changed here: `get_cf_cookies_for_domain()` enforces expiry for `cf_clearance` only, so `__ddg*` cookies are never evicted on age, which is why they go stale. This patch makes the rescue fire whatever the reason the cookies stopped working. ## Verification The regression test drives a real redirect loop through `html_get_page` (302 to `&check=1`, exception raised by the production path rather than faked) and asserts the cookies are purged and the bypasser runs once. - `pytest tests/download/test_http_bypasser_fallbacks.py`: 8 passed. `test_download_url_ignores_zlib_cookie_refresh_failure` fails in my checkout on a missing `seleniumbase`, unrelated to this change. - `ruff check`, `ruff format --check`: clean. - Running in production since 2026-08-14 on v1.3.7 with only this file replaced: six direct-download searches, five served, three books downloaded end to end, against one search per container start before. The rescue mid-download: ``` 19:12:14 http.py:449 - Redirect loop detected; switching to bypasser: https://annas-archive.gl/md5/cb8fba7abae800ddbae1adfb8d7699d9?&check=1 19:12:38 internal_bypasser.py:756 - Bypass successful using _bypass_method_cdp_gui_click 19:14:36 direct_download.py:1142 - Resolved download URL [aa-slow-nowait]: … 19:14:47 orchestrator.py:735 - download finished; starting post-processing ``` ## Separate issue this exposes DDoS-Guard does not accept a solved cookie from plain `requests` traffic, so after this patch the rescue runs for nearly every AA URL. `internal_bypasser.get()` serializes all solves on one module-wide lock and builds a fresh Chrome each time: 11-16 s uncontended, 43-52 s under concurrent load, measured on the host above. Correctness is cheap here, latency is not. Happy to open a separate PR for a warm browser session if that direction is welcome. Co-authored-by: Kukkerem <Kukkerem@users.noreply.github.com> |
||
|
|
03e219eb43 |
Let the bypasser solve bot challenges on Anna's Archive search (#1198)
Anna's Archive put a DDoS-Guard JS challenge in front of /search: the homepage still returns 200, but /search and /md5/<id> answer 403 on every mirror (.gl, .pk, .gd all confirmed). Search fetched both with allow_bypasser_fallback=False, which rotates mirrors on a 403 instead of invoking the bypasser, so it walked the whole mirror list, exhausted it, and surfaced "Unable to reach download source. Network restricted or mirrors are blocked." as a 503 on every query. Adding mirrors could not help — they sit behind the same gate — and neither could USE_CF_BYPASS, since search never reached that branch. Fetch search and the detail page with allow_bypasser_fallback=True so a 403 hands over to the bypasser, which already detects this challenge (DDOS_GUARD_INDICATORS matches the live page). Echoing the __ddg cookies back does not clear it; it needs real JS execution. The download-count fetch keeps allow_bypasser_fallback=False: it is decoration on the details modal and not worth holding the modal open for a browser solve. Fixes #1196 |
||
|
|
cc1a95f965 |
Fix protection bypass cancelled by stall detection at exactly 300s (#1184)
A download that hits Cloudflare hung on "Bypassing protection..." for five minutes and then died, regardless of which bypasser was configured. html_get_page() started a BypassHeartbeat thread to keep the download marked alive during a bypass, but the thread had no loop: it fired one status event and returned. Even with the loop restored it could not have worked, because update_download_status() dedupes identical (status, message) tuples and returns before refreshing _last_activity, and the heartbeat re-sent the byte-identical payload already emitted just above it. So _last_activity was frozen for the whole bypass, while both bypassers are allowed to run longer than STALL_TIMEOUT (external FlareSolverr ~394s at default settings, internal 420s per get() call). The watchdog always won. From a reporter's log: 403 at 07:04:33.390, cancelled at 07:09:33.987 - exactly 300.000s, and 41s before the bypasser would have finished and reported the real error, an HTTP 500 from FlareSolverr the user never saw. The regression is not one commit. |
||
|
|
e35b4c47a7 |
Direct source refactor (#895)
- Updated mirror selection - Removed built-in mirror options, users must provide their own configurations - Set Universal search to default, added ability to disable direct source - Updated documentation - Updated makefile |
||
|
|
d7b9f2e67f |
Backend test hardening + quality enforcement (#872)
- Reworked many tests - Enforcing lint + type checking for test suite - Fixed various issues surfaced by the new tests - CI tweaks |
||
|
|
8d98e122ec |
Linter followup (#868)
Expanded Ruff rules and completed fixes |