mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-10-01 21:15:52 +01:00
html_get_page follows Anna's Archive redirects by hand, and DDoS-Guard's gate answers /search with a 302 to the same path plus `check=1`. The follower walks that handshake by reassigning `current_url`, so every downstream handoff - the 403 branch, the 503-challenge branch, both redirect-loop rescues - passed the *probe* URL to the bypasser rather than the page we actually wanted. A solver opens that in a fresh browser holding none of the cookies the probe exists to collect, so DDoS-Guard cannot verify it automatically and serves the manual CAPTCHA page that nothing can solve. The #1292 log is exactly that: a 403 handed off on `&check=1`, FlareSolverr answering "Challenge solved!", and a 4721-byte DDOS-GUARD captcha page coming back. - `_solvable_url()` strips the probe parameter, applied at the single choke point in `_run_bypasser` so all four handoffs are covered. Scoped to the hosts whose redirects we follow manually; a URL without the parameter is returned by identity, so nothing else is re-encoded. The same reports showed three further defects, all of which stand whatever the host was reacting to: - The external bypasser logged that the solve had not cleared the protection and then returned the challenge page as a success. That skipped the one recovery left - get_bypassed_page's retry-and-rotate loop, where the next mirror is a different DDoS-Guard host - and filed the captcha page's own __ddg cookies as that host's clearance, to be replayed on every later request. It now raises ChallengeNotSolvedError before storing anything. - "Check that the bypasser is reachable and working" was the one piece of advice guaranteed to waste the reporter's time: it was reachable, it ran a full solve, and it returned a captcha. ChallengeNotSolvedError carries the marker so the search layer can name the host as the cause instead of the bypasser. - The untabled-page fingerprint logged `attempt_url`, which html_get_page has since rotated past. The #1298 bundle reported the page against annas-archive.gl when the body had come from .pk - the triage cost #1289 added the line to remove. The search now asks for the response URL and logs that. Its give-up shape is the tuple ("", url), which is truthy, so the exhaustion check reads the body rather than the response. Regression fixtures are built from the pages in the reports. The two behavioural handoff tests were checked against the unfixed code: both fail there, reproducing the reporter's log line verbatim. Refs #1292 Refs #1298