mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-09-28 22:06:05 +01:00
2c6d6a02cdfdcf610a89724eb8f52cffaeef45dc
218
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e1c3f057ab |
fix: bypass recordings, welib wrong-md5 links, footer build sha (#1364) (#1373)
Debug screen recordings never started. Every bypass logged "Capturearea
1540x1050 at position 0.0 outside the screen size 1440x1880".We ask
ffmpeg for the fingerprint screen size plus margin, the size wealso pass
SeleniumBase as xvfb_metrics. SeleniumBase builds thatdisplay with
use_xauth=True, the image ships no xauth binary, so itfalls back to a
fixed 1440x1880 Xvfb and the requested size neverexists. Drop
-video_size so x11grab records the whole screen, whateversize it turned
out to be.
welib could hand back a link for a different book. Welib
answers/md5/<md5> with a search for that md5; when it does not have the
file,the resolver took the first "Download" on the results page
(md5a2c1dc0c... resolved to auto_download/9c8cf85d...). On an
/md5/<md5>page a GET/Download link is now only taken when its href names
thatmd5; otherwise the source is reported as not having the file.
Alsoremoves _get_download_urls_from_welib and _is_source_enabled:
themd5-template branch in _get_urls_for_source always handles welib
first,so that resolver could never run.
The footer showed the build date instead of the commit. CI
stampsBUILD_VERSION as <yyyy-mm-dd>-<sha> (pr-<sha> for PR images) and
thefooter kept its first seven characters, so dev images read
"Shelfmarkmain (2026-09)". Take the trailing commit sha instead: "main
(
|
||
|
|
d978896142 | fix(auth): rename the API_KEY env var to SHELFMARK_API_KEY (#1374) | ||
|
|
3b280009ae |
feat(auth): static API_KEY (env) accepted as Bearer or X-Api-Key, cookie or key (#1366)
Supersedes #1353, per the discussion in #1352: one `API_KEY` environment variable; when set, a request carrying it is authenticated as the first admin, and cookie sessions keep working exactly as before (cookie **or** key). Nothing else changes. No table, no UI, no settings-tab switch, no per-user keys. ## What - `API_KEY` (env). Unset → the feature is off and none of the new code runs. - `Authorization: Bearer <key>` or `X-Api-Key: <key>` on any existing `/api/*` route authenticates that request as the first admin in `users.db` (`ORDER BY id`), or as a bare admin identity (`user_id="api"`, `is_admin=True`, no local user row) if the install has no admin yet. Per request only; nothing is persisted; the admin's role is read live, so deleting or demoting that user takes effect on the next request. - Both headers are checked and either may match. That is what makes the key usable behind a reverse proxy that injects its own `Authorization` header (oauth2-proxy, Authelia, forwardAuth): send the key in `X-Api-Key`. - A credential that is **not** the key is ignored and the request continues on the normal session path, so proxy-forwarded tokens are unaffected. Without a valid session such a request gets the usual `401 {"error": "Unauthorized"}`, identical to a request with no credential, so there is nothing to probe. ## How - `shelfmark/config/env.py`: `API_KEY = os.getenv("API_KEY", "").strip()`. - `shelfmark/core/api_key.py`: `extract_api_key_candidates()` (Bearer token if the scheme is Bearer, then `X-Api-Key`) and `matches_api_key()` using `hmac.compare_digest` on bytes. - `shelfmark/core/user_db.py`: `UserDB.get_first_admin()`. - `shelfmark/main.py`: `api_key_auth_middleware` (`before_request`, registered before `proxy_auth_middleware`, which early-returns for keyed requests). Only `/api/` paths; `/api/health` and `/api/auth/*` exempt; no-op when `API_KEY` is unset or the auth mode is `none`. On a match it mirrors the proxy-auth pattern: `session.clear()` then populate `user_id` / `is_admin` / `db_user_id` for this request, `permanent = False`, `modified = False`, `g.api_key_auth = True`. An `after_request` hook guarantees no `Set-Cookie` is written for a keyed request even if a handler dirties the session. - `docs/api-access.md` (new), the `API_KEY` entry in `docs/environment-variables.md`, and a README link. ## Security - Constant-time compare; the key is never logged or echoed. - Keyed requests never mint or refresh a session cookie and ignore any cookie sent with them (a non-admin cookie plus the key yields admin for that request; the browser's own session is left untouched and usable). - The mismatch path touches neither the session nor `g`, so a stray bearer on a browser request can neither log the user out nor change how their cookie is refreshed. - Store errors during the admin lookup fail closed (`500 {"error": "Authentication error"}`), never to anonymous. - Verified against Flask's `save_session` / `should_set_cookie` ordering, and under auth modes `none`, `builtin`, `proxy`. ## Tests `tests/core/test_api_key_env.py` (36): extraction and matching; first-admin lookup; middleware behaviour on a guarded route and an admin route, with and without a user_db, `X-Api-Key`, both-headers combinations, no `Set-Cookie` when a handler dirties the session, incoming non-admin cookie ignored, browser cookie still usable after a keyed request, security headers, store error → 500, mismatch → guard's 401 / cookie path / permanent cookie untouched, unset → off, exempt paths and path probes, `none` and `proxy` modes, deleted and demoted first admin, a keyed write passing the guard. Existing auth suites unchanged. All CI gates green on the fork: https://github.com/gavinmcfall/shelfmark/pull/2 (CI-only draft). Also exercised against a running instance: 47 scripted checks including 150 concurrent requests, proxy-mode switching through the key, an unset-key restart, and a log scan for the key. ## Naming `API_KEY` as discussed. If you'd rather namespace it (`SHELFMARK_API_KEY`) to avoid clashing with other tools' env vars in shared compose files, it is a one-line change; say the word. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1a5b37d9d3 |
fix(deluge): send seeding ratio limit under Deluge's own keys (#1367)
Deluge's per-torrent options are `stop_at_ratio` (bool) and `stop_ratio` (float), and `torrentmanager` checks `options['stop_at_ratio'] and get_ratio() >= options['stop_ratio']`. We were putting the indexer's float into `stop_at_ratio`, which only switched stopping on and left the daemon's global ratio (default 2.0) as the one actually enforced. A `ratio_limit` of 0 turned stopping off entirely. `stop_at_ratio_enabled` is not a Deluge option at all. Deluge has no per-torrent seeding time limit, `seed_time_limit` is a global core preference, so the value is logged as unapplied instead of sent as a key the daemon drops. qBittorrent and Transmission already honour both indexer limits, so this removes a silent difference between clients. ## Verification - `tests/prowlarr/test_deluge_client.py`: the ratio arrives as `stop_ratio` with `stop_at_ratio` set, and no key Deluge does not define is sent. Both fail on current main and pass here. - Full suite (3165), ruff, ruff format, basedpyright, vulture green. |
||
|
|
7c8e89c567 |
fix(googlebooks): page by the capped size, not the raw limit (#1370)
The Google Books search builds `maxResults` as `min(limit, 40)` because the API caps a page at 40 volumes, but advances `startIndex` by the full `limit`. The pages then stop tiling. With `limit=50`, page 1 covers items 0 to 39 and page 2 starts at 50, so items 40 to 49 are never returned and every later page drops another 10. This computes the page size once and uses it for both `maxResults` and `startIndex`. The shipped frontend asks for 40 and is unaffected. `/api/metadata/search` clamps `limit` to 100, so an API caller passing 41 to 100 was hitting it. One thing I left alone. The provider uses the base `search_paginated` heuristic, `has_more = len(books) >= options.limit`, which still reports `has_more: false` for a limit above 40 since Google can never return that many. That was already the behaviour before this change, and fixing it means either touching the shared heuristic or adding a provider override, so I kept this patch to the stride. Happy to follow up if you want it. ## Verification - `tests/metadata/test_googlebooks_parse.py`: pages 1 and 2 at `limit=50` must tile exactly, plus a guard that `limit=25` still strides by 25. The first fails on current main and passes here. - Full suite (3165), ruff, ruff format, basedpyright, vulture green. |
||
|
|
bc03ad062e |
fix(http): keep the host of a protocol-relative download link (#1368)
`get_absolute_url()` replaced both `netloc` and `scheme` whenever either one was missing. A protocol-relative href such as `//cdn.example.org/f.epub`, scraped from a page on `https://annas-archive.org/...`, parses with a netloc and an empty scheme, so it came back pointing at the page's own host. The download then 404s and the source is skipped. Each field now falls back to the base URL only when the parsed URL does not supply it. Plain relative paths resolve exactly as before, which the control test covers. This affects the Z-Library, welib and generic download link handling in `release_sources/direct_download/annas_archive.py`. ## Verification - New `tests/download/test_http_absolute_url.py`: a protocol-relative link keeps its own host, and a plain `/path` still resolves against the base. The first fails on current main and passes here. - Full suite (3165), ruff, ruff format, basedpyright, vulture green. |
||
|
|
ab3aa9a8b0 |
fix(requests): reject non-object items in the batch endpoint (#1369)
`POST /api/requests/batch` checks that `requests` is a non-empty list and then hands each element to the shared preparation helper, which calls `.get()` on it. A bare string, number or null in the list raises `AttributeError` and the caller gets a 500, while `POST /api/requests` answers 400 with a message for the same mistake. This validates the element type beside the existing list check. One bad item rejects the whole batch rather than being reported per item, which matches the endpoint's current contract: every other failure path already aborts the batch with a single error body, as `test_batch_create_requests_is_atomic` asserts. Responses for valid payloads are unchanged. ## Verification - `tests/core/test_request_routes_api.py`: a case per bad shape (int, str, null, list), plus a mixed valid and invalid batch that asserts nothing was created. All fail on current main with a 500 and pass here. - Full suite (3168), ruff, ruff format, basedpyright, vulture green. |
||
|
|
d4619be69a |
Feature: Show the AA search result stats (#1362)
When doing a direct search for a book, show the stats of the AA results. For example if we search for "The Great Gatsby", AA reports it has 240 hits and shows the first page of 50. Provide this stats info in the shelfmark webUI via ResultsSection.tsx and ReleaseModal.tsx. This table shows what should be displayed based on the total number of hits found: |Total |Display| |-------|-------| |1 |"Result 1 (1 Total)"| |6 |"Results 1-6 (6 Total)"| |144 |"Results 1-50 (144 Total)"| |500+ |"Results 1-50 (500+ Total)"| Currently, shelfmark also only shows us the first 50 search hits even if there were more available from AA. This could be added later if considered desirable. As usual, a picture is worth a 1000 words: <img width="1012" height="610" alt="direct-results-stat" src="https://github.com/user-attachments/assets/8c0c688f-444e-482a-a9a3-2dee5a5563b8" /> <img width="1013" height="741" alt="universal-results-info" src="https://github.com/user-attachments/assets/399bbb79-87f2-4432-a3d0-64937795f5f1" /> Coded with llama.cpp, opencode and 🤖 |
||
|
|
c1315a2b23 |
fix(download): check task ownership before serving queued files (#1357)
`/api/localdownload` resolves the file through the live queue and returns it before checking who owns the task; the owner check only runs on the download-history fallback, once the task has aged out of the queue. Task ids are source ids, so two users who searched the same book can end up with the same id. This applies the same rule on the queue path, reusing the `_task_owned_by_actor` helper the cancel/retry/priority routes already use, so both paths answer a non-owner with the same 404. Admin behaviour is unchanged. The 404 matches what this endpoint's history path already returns for a non-owner rather than the 403 `download_not_owned` the cancel/retry/priority routes use, happy to switch it if you prefer consistency with the siblings instead. ## Verification - `tests/core/test_activity_routes_api.py`: the owner still receives their queued file; a different user receives 404. The new case fails on current main and passes here; the existing history-fallback test is unchanged. - Full suite (3148), ruff, ruff format, basedpyright, vulture green. |
||
|
|
7934924678 |
fix(queue): don't stamp CANCELLED over a finished download (#1361)
`cancel_download` reads the task status under the queue lock, releases it, and only then writes CANCELLED through `update_status`. A download that finishes in that window has its COMPLETE overwritten. The queue and the UI show the task as cancelled while the file is already on disk, and the terminal hook fires for both statuses. The check and the write now happen in a single lock hold. Because the lock is non-reentrant and the terminal hook has to run after it is released (the stall canceller depends on that), the lock-held part of `update_status` moved into a small private helper that both paths share; `update_status` is a thin wrapper over it. A cancel arriving once the task is already terminal still returns `False`. ## Verification - `tests/core/test_queue.py::test_cancel_does_not_overwrite_a_download_that_finished_first`: a worker thread completes the download while the cancel is in flight, with the handover driven by events rather than sleeps. The task stays complete. Fails on main, passes here. - Full suite (3147), plus `tests/download/` and `tests/core/test_download_api_guardrails.py`, ruff, ruff format, basedpyright, vulture green. |
||
|
|
a6204a318e |
fix(oidc): reject backslash paths in the return_to sanitizer (#1359)
The OIDC `return_to` sanitizer rejects values starting with `//` and then relies on `urlsplit` to catch anything carrying a netloc. A value such as `/\host` has no netloc, so it is stored in the session and used as the post-login redirect target and browsers resolve the backslash as a path separator, which lands the user outside the app after a successful login. `_normalize_return_to` now also rejects values whose path contains a backslash. That matches the frontend sanitizer in `authRedirect.ts`, which parses with `URL` and already discards those forms, so the two ends agree again. The check covers the path only, so query and fragment backslashes still round-trip, and it also catches the script-root case where `/app/\host` strips to `/\host`. ## Verification - New cases in `tests/core/test_oidc_routes.py` cover the rejected forms, including under a script root, and confirm `/`, `/settings` and `/search?q=x#frag` are unaffected. They fail on current main and pass here. - Full suite (3155), ruff, ruff format, basedpyright, vulture green. |
||
|
|
545480c557 |
fix(download): default is_admin to False in the request policy guard (#1358)
`_resolve_policy_mode_for_current_user` reads `session.get("is_admin",
True)`, so a session carrying `user_id` but no `is_admin` key skips the
request policy entirely, while every other admin check in the codebase
defaults the key to `False`.
This uses the same default here. Every authenticated login path
(builtin, CWA, proxy, OIDC) writes `is_admin` into the session, and
`AUTH_METHOD=none` is already short-circuited one line earlier, so
sessions from those flows behave exactly as before.
## Verification
- `tests/core/test_request_routes_api.py::TestDownloadPolicyGuards`: a
session without `is_admin` now gets `policy_requires_request` and
nothing is queued. Fails on current main, passes here.
- Full suite (3147), ruff, ruff format, basedpyright, vulture green.
|
||
|
|
127dd82615 |
fix(users): apply user updates only after the payload validates (#1360)
`PUT /api/users/me` and `PUT /api/admin/users/<id>` write the new password hash, and then the profile fields, before the rest of the payload is checked. When the request is rejected further down as an invalid role, an admin-only setting, an invalid settings value, the route answers 400 with those writes already committed, so the caller sees an error while the password has in fact changed. Both routes now validate the whole payload before touching the database, and the password hash is folded into the same `update_user` call as the other fields so the field write is a single transaction. Error messages, status codes and the order they are reported in are unchanged. ## Verification - New tests in `tests/core/test_self_user_routes.py` and `tests/core/test_admin_users_api.py` assert that a rejected update leaves the password, profile fields and role as they were, plus a positive case that a valid payload still applies all three. They fail on current main and pass here. - Full suite (3150), ruff, ruff format, basedpyright, vulture green. |
||
|
|
acd59f7cbb |
feat(auth): provision proxy users as non-admin once an admin exists (#1356)
With `AUTH_METHOD=proxy` and no admin group configured, every user the proxy authenticates for the first time is provisioned as an admin (`is_admin = True` unless the user already exists in `users.db`). The intent to never lock an instance out makes sense, but the effect is that anyone the SSO gate lets through becomes an administrator. On an instance shared with family or a small community that is a footgun; I hit it when the first invited reader landed as an admin. This keeps the guarantee and removes the footgun: the first account is still provisioned as an admin while the instance has no admin at all, and later first-time users follow a new `PROXY_AUTH_DEFAULT_ROLE` setting (Security tab / env), default `user`. Known users keep their stored role; the `PROXY_AUTH_ADMIN_GROUP_NAME` path is unchanged and still takes precedence. I couldn't find a way with Cloudflare access to pass this along. Changes: `UserDB.has_admin()`, `_proxy_default_is_admin()` in the proxy middleware, the new `SelectField` beside the other proxy settings, the regenerated `docs/environment-variables.md` entry and a row in `docs/reverse-proxy.md`. Compatibility: the default moves from "everyone admin" to "first admin, then users". Accounts already in `users.db` are unaffected; new SSO users on an existing instance become regular users unless `PROXY_AUTH_DEFAULT_ROLE=admin` is set. If you would rather ship this purely opt-in I can flip the default to `admin`. ## Verification - `tests/core/test_auth_api.py::TestProxyProvisioningRole`: first user admin / second user not; `PROXY_AUTH_DEFAULT_ROLE=admin` restores the old behaviour; an admin from another auth source counts as "an admin exists"; a known user keeps their role whatever the default. - Full suite (3094), ruff, ruff format, basedpyright, vulture green. - Running on my own instance since 2026-09-19. |
||
|
|
44f4e13cce |
refactor: extract the per-source release search out of /api/releases (#1355)
The `/api/releases` route carries an inner `_search_source_releases` helper that builds the search plan for one source, logs the planned query type, runs the search and turns `SourceUnavailableError`/operational errors into an error message instead of raising. Anything outside the route that wants to search one source with exactly those semantics has to go through Flask today. This moves that helper into `shelfmark/core/release_search.py` as `search_source_releases()` and has the route delegate to it. Behaviour is unchanged: same plan construction (including the caller's `user_id`, so per-user default languages still apply), same logging, same error-to-message handling. It is the refactor half of #1047 by @InfiniteAvenger, split out on its own as you asked for other PRs (#1318). Their authorship is preserved on the commit; I rebased it onto current `main` and added tests. ## Verification - `tests/core/test_release_search.py`: unknown source → `"Unknown source: …"`, `SourceUnavailableError` and operational errors → `"<source>: <error>"`, success path forwards `expand_search` / `content_type` and returns the source instance, the plan receives languages / manual query / indexers / `user_id`. (These tests are type-annotated; happy to strip the annotations if you prefer the suite's bare style.) - Full suite, ruff, ruff format, basedpyright, vulture green; the existing `/api/releases` route tests are unchanged and pass. Co-authored-by: InfiniteAvenger <calebewest02@gmail.com> |
||
|
|
b7002a6eca |
feat: Add TorBox client support and settings integration (#1342)
Add **TorBox** as a torrent download client for Prowlarr releases. Users can select `TorBox` in the download client settings, configure it with the new `TORBOX_API_KEY` environment variable, and verify their credentials with the connection test button. The integration supports both magnet links and `.torrent` files. It tracks the torrent lifecycle through TorBox, downloads supported book and audiobook files from the TorBox CDN, preserves safe nested file paths, and cleans up remote and local download state. Important: Shared HTTP download logs omit full download URLs and URL-bearing exception text to avoid exposing credentials, following best practices. This applies to all clients that use the shared `download_url()` path; URLs remain available to the HTTP operations themselves. --- There is already related work in progress in #1173, which includes both torrent and direct-download support for TorBox. This PR is not intended to replace or compete with that contribution. It offers the tested torrent client functionality as a smaller, focused change that can make TorBox available to the community sooner. The direct-download integration proposed in #1173 remains valuable and could be reviewed or introduced separately. Automated tests cover configuration, connection validation, magnet and torrent-file submission, API errors, status and progress handling, file retrieval, path traversal protection, cancellation, cleanup, and sensitive URL redaction. I also validated the complete flow locally with several magnet links and `.torrent` downloads. TorBox processed the torrents and Shelfmark downloaded the resulting files as expected. AI was used to help with the implementation, with human validation. This PR and long description? Took me some good minutes at night after work, but gives me joy to open this PR to share with the community this improvement. |
||
|
|
e5dd34ae0e |
fix: unbreak main and follow up on the Blackhole handoff review (#1346)
DownloadHistoryService.record_download and updated the single production caller, but not the eleven in the test suite, leaving main red with 32 failures. Pass None, which is what the pre-#1336 behaviour recorded. For the Blackhole handoff (#1345): add_download publishes the torrent before the cancel check runs, and BlackholeClient.remove() is a no-op, so the watcher picks the file up regardless. Reporting a bare "Cancelled" hid that from the user. Name the completed handoff in the cancellation message instead, drop the _handle_cancelled_download call whose usenet branch cannot apply to a handoff-only client, and record why the orchestrator no longer verifies HandoffResult.path. Finally, make tests/direct_download a package: test_libgen_extract.py imports tests.libgen.sample_html across test directories, so without an __init__.py pytest named its modules by bare basename and a same-named module elsewhere would collide. |
||
|
|
c53545d9fe |
Add configurable word separator for naming templates (#1333)
Closes #1230 ## What Adds a "Word Separator" setting (Space / Dot / Underscore / Hyphen / Custom) that replaces internal whitespace in each naming-template placeholder's rendered value — e.g. `{Author}` renders "Arthur.Conan.Doyle" instead of "Arthur Conan Doyle" when Dot is selected. This follows option 2 from the issue rather than inventing new dotted-keyword template syntax (`{Author.}`), since it's a smaller surface: one setting applies uniformly across all four templates (books/audiobooks × rename/organize) instead of needing a parallel token for every existing one. ## How it works - Literal characters typed into the template itself (e.g. the `.` in `{Author}.-.{Title}`) are never touched — only whitespace *inside* a placeholder's resolved value is affected. - Default is "Space", which is a no-op: existing templates produce byte-identical output after this change (verified via the existing test suite, unmodified, still passing). ## Where - `shelfmark/core/naming.py` — `word_separator` param on `parse_naming_template` / `build_library_path`. - `shelfmark/download/postprocess/policy.py` — `get_word_separator()`, mirroring the existing `get_file_organization()` accessor. - `shelfmark/download/postprocess/transfer.py` — wires the resolved separator through the four existing template-rendering call sites. - `shelfmark/config/settings.py` — new `Word Separator` / `Custom Word Separator` fields next to the existing naming-template fields. - `src/frontend/.../namingTemplatePreview.ts` + `NamingTemplateField.tsx` — the settings UI has its own TS mirror of the Python renderer for the live preview; updated it in lockstep so the preview doesn't lie about what the separator will actually do. - Tests added on both sides (pytest + vitest). ## Testing - `uv run pytest tests/core/test_naming.py tests/core/test_destination_file_organization.py` — all pass, including new cases. - `uv run pytest` (full suite) — same pre-existing failures as on `main` before this change (browser/network-dependent bypass & e2e tests unrelated to this diff), everything else green. - `uv run ruff check` / `ruff format --check` / `basedpyright` — clean. - `npm run lint` / `format:check` / `typecheck` / `test:unit` (196 tests) — clean. |
||
|
|
8f608f2e64 |
fix(download): complete consumed Blackhole handoffs (#1345)
A Blackhole watcher can consume the torrent before Shelfmark checks it, leaving the task in error even though the handoff succeeded. Complete the handoff when `add_download` successfully publishes the file, and stop requiring a `HandoffResult` path to remain present. Follow-up to #1312. ## Verification - A watcher that immediately reads and removes the torrent receives the exact bytes. The task changes from ERROR before this fix to COMPLETE afterward, without running book postprocessing. - The consumed-file regression fails on current main and passes here. Resident files, write failures, cancellation, magnet rejection and normal downloads remain covered: 81 focused tests pass. - Ruff lint and formatting pass for the changed files. |
||
|
|
2bb84a17a2 |
Extract archives when zip/rar are enabled as supported formats (#1343)
The default audiobook formats include `zip` and `rar`. `scan_directory_tree` checks the supported-format list before checking for archives, so a downloaded archive lands in `book_files` and is imported as-is. The extraction branch in `collect_directory_files` is never reached. This keeps archives out of `book_files`, so they always take the archive path: extracted when extraction is allowed, imported as-is when it isn't (unchanged). Tests added in `tests/download/test_postprocess_scan_archives.py`; three of the four fail without the change. |
||
|
|
c6b70a6844 |
fix(sources): send a Referer when fetching libgen ads.php pages (#1340)
## What libgen.li's `ads.php?md5=` now returns an **empty `200`** to any request without a `Referer` — an anti-hotlinking check the mirrors added recently. Both libgen paths fetch it without one, so the page comes back blank and the download silently fails while **search keeps working** (which is exactly why it looks like rate-limiting or mirror drift rather than a bug). Same one-line cause, two call sites: the Libgen search source (`libgen/scraper.py:fetch_page`) and the AA-md5 → libgen fallback (`direct_download/annas_archive.py:_extract_libgen_download_url`). Fix: send a same-origin `Referer: <scheme>://<host>/` on the `ads.php` fetch in both. ## Worth a look in review - **The referer goes on the *resolution* fetch, not the download.** `download_url(..., referer=...)` was already correct — the blank page happens one step earlier, at the `ads.php` GET. - Reproduced against live mirrors: `ads.php` returns `Content-Length: 0` bare, the full page with a `Referer`, and resolvable files download valid bytes again. Regression tests in `tests/libgen/` and `tests/direct_download/` assert the header on both paths. Lint/format/typecheck clean. Follow-up to #1326. |
||
|
|
35b89b0d78 |
fix(sources): restore Direct Download search errors and language matches (#1339)
Fixes two regressions from the provider-driven refactor (#1337). First, the composite search caught RuntimeError, TypeError, ValueError and request errors from each provider and returned an empty list, so a failed search looked like one with no hits. It now raises the first provider failure when no provider returned releases. Second, the shared parser re-matched every row's language locally, dropping rows Anna's Archive had already matched with &lang= (free-text cells like 'English, French' or 'unknown'). parse_search_items gains a filter_languages option, which AA turns off, so AA's own language-from-path filter is again the only local one. |
||
|
|
a5cd9f0bfb |
refactor: make direct download provider-driven (#1337)
This is the refactor for the download handler |
||
|
|
af21d1da1f |
feat(sources): add Libgen as a direct catalogue search source (#1326)
## What Adds **Libgen as a search source**. Today Libgen is only a download mirror (reached by an Anna's Archive md5), so anything in Libgen but not in AA's search index is invisible — and that's where most of the CBZ/CBR comics and manga live. A Libgen search for *One Piece*, for instance, turns up ~99 volumes that AA search never shows. It's a self-contained `release_sources/libgen/` package (source + handler + settings) plus one line to register it. **No changes to `direct_download.py`** — it reuses the existing `ads.php → get.php` resolution and the mirrors already configured in `LIBGEN_MIRROR_URLS`. Plain HTTP, no bypasser needed (libgen.li isn't behind DDoS-Guard). Opt-in via a settings toggle. ## Worth a look in review - **`source_id` is `libgen:<md5>`, not the bare md5.** The download queue keys on `task_id` (= `source_id`), and `direct_download` already uses the bare md5. Since AA indexes a lot of Libgen, the same md5 shows up from both sources — a bare id would collide in the queue. The handler strips the prefix before downloading. - **Reachable like the other non-default sources** (Prowlarr, AudiobookBay, …): it appears in the per-book release search, not the free-text box (that stays wired to `direct_download`). Tests in `tests/libgen/` cover parsing (both row layouts), the source, the handler, and `get_record`. Lint/format/typecheck clean. |
||
|
|
1b17fe179a |
fix(irc): rank a surname-only result as partial, not wrong (#1332) (#1334)
"David Petrie" as "D. Petrie", then ranked the answer by the full name to recover the precision the surname gave up. The two halves disagreed. author_affinity needs two agreeing tokens before it calls a name the same person, so "Petrie" - the name on the filenames a surname search exists to reach - matched one and came back AUTHOR_MISMATCH. It therefore sorted below "Unknown" and level with "Gordon Petrie", a different author who merely shares the surname. The widened query pulled those rows in and the ranker buried them. Falling short of agreement is now separated from disagreeing with it. A name whose every token fits the one asked for is an abbreviation of it and ranks AUTHOR_PARTIAL, between agreement and "no author reported"; a name carrying a token that fits nothing still ranks AUTHOR_MISMATCH. Nothing that agreed before changes tier - "Homer"/"Homer Simpson" is still a match, since the extra token must not demote a mononym that already met its one-token requirement - so Prowlarr's #1293 ordering is unchanged except that a tracker listing a bare surname stops being read as the wrong author. Measured on the issue's own case, wanted "David Petrie": before: D Petrie, Unknown, Petrie, Gordon Petrie after: D Petrie, Petrie, Unknown, Gordon Petrie Second fix, same release: a book with no title posted the surname on its own. _build_query fell back to book.search_title or book.title, which is empty on exactly the path where the plan has no title variants, so the line reaching the channel was "@search Petrie" - not a search for anything, and the kind of bare over-broad post is_available refuses unaddressed queries to avoid. It now returns "" and the existing "No search query could be built" guard takes it. Tested with make python-lint, python-format, python-dead-code, python-typecheck and python-test. |
||
|
|
35037b35fd |
fix(irc): search by surname, and rank the answer by author (#1331) (#1332)
Fixes #1331. A search bot ANDs every term against a filename, so the given name is the term that empties the result set. Measured against irchighway's #ebooks: "Revelations David Petrie" is answered "no results", "Revelations Petrie" returns 9 matches, 6 of which parse, all filed as "D Petrie". The query now carries the title and the surname, read off the search variant so the ISBN fallback and a manual query - which set author="" on purpose - keep their current shape. Title-only, the shape #1295 settled on for Prowlarr, does not transfer: the bot caps an answer at 1000 matches, and a bare "Revelations" hits that cap with 923 parsed rows across 500 authors, so the cap itself can drop the wanted book. The surname is the token the two spellings share and it keeps the answer small. The full author then orders what comes back, reusing author_affinity from #1295, since a surname also matches a different author who shares it. It sits under server availability the way indexer priority does in #1295: a download addresses one named bot and waits 120s for it, so a match from a bot that has left the channel must not outrank a mismatch that can answer. Ranking runs on the way out rather than before the cache, because one query identity is shared by every book that produced that query. Two things found while testing: - The parser writes the literal "Unknown" when a filename has no " - " separator (parser.py:168). Ranked literally that sorts as a wrong author, so author_affinity's middle tier was unreachable here; it is now read as absent. 5 of those 923 rows are affected. - author_affinity moves to shelfmark/core/author_match.py, unchanged, so IRC does not import from the Prowlarr package. Prowlarr behaviour is untouched and its tests pass as they are. The three IRC assertions in the #1252 regression file move to the surname form. The invariant they pin - one contributor's name reaches the query, never the whole credit list - is unchanged. Tested with make python-lint, python-format, python-dead-code, python-typecheck and python-test, and end to end against irchighway with the patched source: it posts "Revelations Petrie" and returns 6 releases. |
||
|
|
99e0cfde3d |
fix: prevent Anna's Archive download countdown resets by preserving browser sessions (#1325)
## Observed bug Anna's Archive slow-download pages can return a JavaScript countdown before a download link is available. The internal browser returns that waiting-room HTML and closes its incognito session. The downloader then sleeps and fetches the URL again, which can create a new queue session and **restart the countdown instead of reaching the download link**. ## Fix - **Preserve the queue session:** keep the original browser tab open while the site's own countdown and automatic navigation finish. HTTP 200 and cached-cookie waiting-room responses enter the same flow. - **Return a consistent page:** capture HTML and readiness together in one browser evaluation so navigation cannot pair a new page's status with stale protection-page HTML. Share cache validation and page-readiness rules across their callers. - **Keep waiting cancellable and bounded:** poll cancellation while a slow browser read remains pending, rather than repeatedly cancelling and reissuing it. Apply a **300-second waiting-room limit** within the existing browser watchdog. - **Report queue timeouts accurately:** preserve the timeout across the helper-process boundary and stop the solve without restarting the browser or rotating mirrors. Waiting-room detection is limited to Anna's Archive `/slow_download/` pages containing an actual `.js-partner-countdown` element. The site controls the countdown and refresh. External-bypasser behavior and file-transfer timeouts are unchanged; the PR adds no deployment configuration or dependencies. ## Validation Validated at `c81e0a2`: | Check | Result | | --- | --- | | Full Linux unit suite | **2,978 passed** on Python 3.14 in a non-root environment with entrypoint test stubs enabled | | Focused regression coverage | **30 passed**, covering countdown completion, zero timers, navigation, both cookie-cache paths, cancellation, stuck queues, slow reads, and timeout propagation | | Navigation-race regression | Fails against the previous PR implementation and passes with the fix | | Python static checks | Ruff lint/format, BasedPyright for backend and tests, and Vulture passed | | Real Chromium fixture | Queue cookie persisted through 1.5-second DOM reads and one automatic refresh; the CDP connection survived multiple polling intervals | | Live source check | Observed **19 → 14 → 9 → 4 → download link** while retaining the browser session; the patched browser path also completed the waiting room | Full unit-suite command: ```sh pytest tests/ -n 2 --tb=short -m "not integration and not e2e" ``` The live check validates waiting-room completion and link resolution. Remote file-host availability remains a separate concern. The unit suite emitted two existing Authlib deprecation warnings. |
||
|
|
c45d342931 |
fix(prowlarr): skip indexers in Prowlarr failure back-off (#1324)
## What
Read `/api/v1/indexerstatus` once per search and skip indexers whose
`disabledTill` is still ahead. Skipped is neither attempted nor failed.
One client method, one counter on `_IndexerSearchOutcome`, ten tests.
## Why
Prowlarr's own search leaves out an indexer it has disabled after
repeated failures. Shelfmark queries each indexer through its Torznab
endpoint, which answers 429 instead:
```
Prowlarr Torznab error response: <error code="429" description="Indexer is disabled till 09/09/2026 15:00:34 due to recent failures." />
Prowlarr: 1 of 5 indexer searches failed (indexer 2 search failed: 429 Client Error: Too Many Requests ...)
Release search failed for source prowlarr: 1 of 5 indexer searches failed (...)
```
That counted as a failed search, so with one indexer in back-off and the
other four answering empty, `/api/releases?source=prowlarr` returned 503
for every book for the length of the back-off (one hour here).
`/api/v1/indexerstatus` on Prowlarr 2.5.2:
```json
[{"indexerId": 2, "disabledTill": "2026-09-09T15:00:34Z", "mostRecentFailure": "2026-09-09T14:00:34Z", "initialFailure": "2026-09-09T14:00:34Z"}]
```
## Behaviour
| indexers | before | after |
|---|---|---|
| 1 in back-off, 4 answer empty | 503 "1 of 5 indexer searches failed" |
"No releases found" |
| 1 in back-off, 1 answers with releases | releases | releases, one
Torznab call fewer |
| 1 in back-off, 1 times out, 3 answer empty | "1 of 5 failed" | "1 of 4
failed" |
| every indexer in back-off | 503 "5 of 5 failed" | "every indexer is
disabled by Prowlarr after recent failures (until ...)" |
| status endpoint unreachable | n/a | as before, nothing skipped |
Auto-expand no longer retries a pass in which nothing was asked.
## Tests
`uv run pytest tests/prowlarr`: 563 passed, 42 skipped. `ruff check` and
`ruff format` clean.
|
||
|
|
1e3fd48b8b |
fix: share rotating log file handlers (#1316)
This patch shares (for each log file) the `RotatingFileHandler` for logging across all modules, reducing the number of open file descriptors from ~78 to 1. I had originally assumed this issue was a resource leak, but it seems to just be a large fixed number of file descriptors. So this change mostly just (1) shrinks the number of open file descriptors to a reasonable level and (2) prevents two modules in the same process competing to write to a log file. |
||
|
|
c576003319 |
feat(naming): add {FirstAuthor} template token (#1322)
Closes #930. ## What New `{FirstAuthor}` naming-template token. It renders only the first author when metadata lists several ("Author1, Author2, Author3"), so multi-author books can be filed alongside the rest of that author's work instead of getting their own "Author1, Author2, ..." folder. ``` {Author} -> Terry Pratchett, Neil Gaiman {FirstAuthor} -> Terry Pratchett ``` ## How - Added to `KNOWN_TOKENS` in `shelfmark/core/naming.py`, positioned before `author` so `{FirstAuthor}` isn't parsed as literal `First` + `{Author}`. - Derived inside `parse_naming_template` from the existing `Author` value (split on `,` / `;`), so every caller — folder transfer, rename, the settings preview — picks it up with no extra wiring. An explicit `FirstAuthor` key in the metadata still wins if one is ever passed. - `{Author}` behaviour is unchanged. - Frontend `namingTemplatePreview.ts` token list + `KNOWN_TOKENS` kept in lockstep (there's a test enforcing that), with a matching `firstAuthor` helper. - Settings field descriptions + `docs/environment-variables.md` list the new token. ## Known limitation A lone author written `Last, First` is split on the comma too and renders as `Last` — the source metadata doesn't mark which form it is. Called out in the token help text and covered by a test. `{Author}` remains available for anyone who wants the raw string. ## Checks - `make python-test` — 2963 passed - `make python-lint` / `make python-format` / `make python-typecheck` / vulture — clean - `make frontend-test` — 187 passed · `frontend-lint` / `frontend-format` / `frontend-typecheck` — clean 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
265da07d7f |
feat(download): add Blackhole torrent handoff (#1312)
## Why Blackhole users need Shelfmark to hand a torrent file to their existing downloader instead of importing the downloaded book itself. ## Change - add Blackhole as a torrent client with a configurable watched directory - prefer a fetched `.torrent` file for Blackhole while preserving magnet preference for other clients - complete the queue task after the handoff without invoking book post-processing ## Verification - `uv run pytest -q tests/prowlarr/test_blackhole_client.py tests/prowlarr/test_handler.py tests/newznab/test_handler.py tests/download/test_orchestrator_lifecycle.py` - `uv run basedpyright shelfmark/download/clients/blackhole.py shelfmark/download/clients/__init__.py shelfmark/download/clients/base_handler.py shelfmark/download/clients/settings.py shelfmark/download/orchestrator.py shelfmark/release_sources/__init__.py shelfmark/release_sources/prowlarr/utils.py shelfmark/release_sources/prowlarr/handler.py shelfmark/release_sources/newznab/handler.py tests/prowlarr/test_blackhole_client.py tests/prowlarr/test_handler.py tests/newznab/test_handler.py tests/download/test_orchestrator_lifecycle.py` Fixes #1229 |
||
|
|
c18da92569 |
fix(postprocess): attach unmatched chaptered audio files to existing book group (#1176) (#1309)
### Summary Fixes #1176 When downloading an audiobook with many chaptered tracks (e.g. 250+ `.flac` or `.mp3` files), indexer XML or release metadata often caps the file list at ~100-110 entries. When the release extracts on disk, `match_plan_to_files()` matched those first ~110 files to the planned book group, while the remaining 140+ files fell into `unmatched` and triggered fallback heuristic grouping. Because heuristic grouping parsed the folder name (`Westwell - Hot & Cold (2023)`) and stripped the series/author prefix, it generated a second book titled `Hot & Cold` containing the remaining tracks, resulting in two split book folders. ### Changes - In `match_plan_to_files()` (`shelfmark/download/postprocess/packs.py`), check `unmatched` files before falling back to heuristic multi-book splitting. - If an unmatched file is chaptered audio (`.flac`, `.mp3`, `.aac`, etc.) and shares the directory with an existing book group, or if the plan was a single-book plan, append it to that group instead of creating a secondary book. - Non-chaptered standalone books (e.g. `.m4b`, `.epub`) or files in separate subfolders continue to fall back to heuristic grouping as before. - Added unit tests in `tests/download/test_packs.py` verifying: 1. Truncated track list in single folder properly appends remaining chaptered tracks without splitting. 2. Single-book plan with multi-disc audio files (`CD1`/`CD2`) groups together cleanly. 3. Multi-book packs with unmatched chaptered tracks route each track to its respective book folder. ### Testing Ran `uv run pytest tests/download/test_packs.py` and `uv run pytest tests/core/test_processing_packs.py` (all passed cleanly). Checked type annotations with `basedpyright` (0 errors) and formatting with `ruff`. Co-authored-by: amasen02 <amasen02@users.noreply.github.com> |
||
|
|
97d1bb0df4 |
fix(bypass): keep Anna's Archive's aa_ddg_check so clearance replays (#1305)
## What Add `aa_ddg_check` to the cookie-store allowlist. One name, one test. ## Why Every replay of stored clearance ends in the `?check=1` redirect loop, so each search pays a fresh browser solve. On this instance (v1.3.15, WireGuard egress, 0 VPN restarts across the traces) not one replay was accepted in three days of DEBUG logs. The `__ddg*` cookies are stored and replayed correctly. Anna's Archive also sets a cookie of its own, `aa_ddg_check`, and its `?check=1` hop only answers with the page when that cookie is present too. The allowlist keeps `cf_*` and `__ddg*` names, so this one was never stored. ## Measured, same egress IP, cookies taken from one solve | replayed | plain `requests` | `curl_cffi`, Chrome TLS fingerprint | |---|---|---| | filtered `__ddg*` only (current behaviour) | 302 → 302 → 302 … loop | 302 → 302 → 302 … loop | | filtered + `__ddg8_/9_/10_` | loop | loop | | filtered + `aa_ddg_check` | **302 → 200, real search page** | 302 → 200 | | `aa_ddg_check` alone | 302 → 403 | — | So the TLS fingerprint is not the problem, the per-check trio is not the answer, and the cookie needs the `__ddg*` clearance next to it. Cookie attributes as issued: domain `.annas-archive.gl`, path `/`, expiry 90 days. It is not bound to the query, and it is accepted with a stock Python User-Agent. ## Through the real fetch path Same process, `html_get_page`, the name allowlisted, three different queries: ``` 1st: solve expected 25.8s bypass_calls=1 title='frankenstein shelley - search - an' md5=True 2nd: other query 9.5s bypass_calls=0 title='pride and prejudice austen - searc' md5=True 3rd: third query 4.8s bypass_calls=0 title='dracula stoker - search - anna's a' md5=True ``` ## Notes - `tests/bypass/test_ddg_cookie_reuse.py` gains `test_aa_check_cookie_is_stored`; its docstring table gains the row. The bypass tests need seleniumbase to import and do not run on my macOS host, so this leans on CI. `ruff check` and `ruff format --check` pass. The logic was checked directly against `cookie_store` with the settings registry stubbed. - `__ddgmark_` carries a 24 h expiry, so the store's clearance is good for about a day before the next solve, which is what a browser would see too. - Follow-up to #1286. Same instance, same method: DEBUG trace, then a probe script inside the container. |
||
|
|
9f11e83e1f |
fix: keep polling queued Real-Debrid torrents (#1303)
Add `queued` to the existing set of non-terminal Real-Debrid torrent states so `_handle_torrent_info` returns an in-progress `DownloadStatus` and leaves the mutable download state eligible for subsequent polling. Keep the change within the existing status-classification path rather than introducing a new helper or changing the broader handling of unknown statuses. The native Real-Debrid client currently treats the documented `queued` torrent status as a terminal error because it is absent from `_STATUS_DOWNLOADING`. This occurs after a torrent has been added and its files selected, particularly for uncached torrents that wait before downloading. A torrent-info payload with `status: queued`, zero progress, and a filename returns a non-complete `DownloadState.DOWNLOADING` result rather than `DownloadState.ERROR`; After handling `queued`, the internal download state remains non-terminal so a later status poll can be processed instead of returning a cached error. Fixes #1268 Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com> |
||
|
|
9452ebc70d |
fix(bypass): stop handing solvers DDoS-Guard's ?check=1 probe URL (#1300)
html_get_page follows Anna's Archive redirects by hand, and DDoS-Guard's gate answers /search with a 302 to the same path plus `check=1`. The follower walks that handshake by reassigning `current_url`, so every downstream handoff - the 403 branch, the 503-challenge branch, both redirect-loop rescues - passed the *probe* URL to the bypasser rather than the page we actually wanted. A solver opens that in a fresh browser holding none of the cookies the probe exists to collect, so DDoS-Guard cannot verify it automatically and serves the manual CAPTCHA page that nothing can solve. The #1292 log is exactly that: a 403 handed off on `&check=1`, FlareSolverr answering "Challenge solved!", and a 4721-byte DDOS-GUARD captcha page coming back. - `_solvable_url()` strips the probe parameter, applied at the single choke point in `_run_bypasser` so all four handoffs are covered. Scoped to the hosts whose redirects we follow manually; a URL without the parameter is returned by identity, so nothing else is re-encoded. The same reports showed three further defects, all of which stand whatever the host was reacting to: - The external bypasser logged that the solve had not cleared the protection and then returned the challenge page as a success. That skipped the one recovery left - get_bypassed_page's retry-and-rotate loop, where the next mirror is a different DDoS-Guard host - and filed the captcha page's own __ddg cookies as that host's clearance, to be replayed on every later request. It now raises ChallengeNotSolvedError before storing anything. - "Check that the bypasser is reachable and working" was the one piece of advice guaranteed to waste the reporter's time: it was reachable, it ran a full solve, and it returned a captcha. ChallengeNotSolvedError carries the marker so the search layer can name the host as the cause instead of the bypasser. - The untabled-page fingerprint logged `attempt_url`, which html_get_page has since rotated past. The #1298 bundle reported the page against annas-archive.gl when the body had come from .pk - the triage cost #1289 added the line to remove. The search now asks for the response URL and logs that. Its give-up shape is the tuple ("", url), which is truthy, so the exhaustion check reads the body rather than the response. Regression fixtures are built from the pages in the reports. The two behavioural handoff tests were checked against the unfixed code: both fail there, reproducing the reporter's log line verbatim. Refs #1292 Refs #1298 |
||
|
|
cb690b45b8 |
fix(prowlarr): rank releases by author instead of querying for it (#1293) (#1295)
MyAnonamouse is the only indexer Shelfmark treats as enriched, and it
alone was sent {title} {author} while every other indexer got the title
on its own. MAM matches all search terms conjunctively, so whenever the
metadata provider spelled the author differently to the tracker -
Hardcover says Timothy Ferriss, MAM lists Tim Ferriss - the search came
back empty and the UI reported No releases found for this book, with the
release sitting on the tracker the whole time.
The enriched flag is a statement about responses: MAM returns clean
author and bookTitle attributes, which is why it earns format detection
and preferential ordering. Using that same flag to shape the request is
the actual defect, and it is why turning the flag off recovers the
search but takes format detection down with it.
So the query is title-only for every indexer now, and the author orders
the results rather than narrowing them. MAM already hands us its author
field, so agreement is judged on data we hold instead of by an AND we
cannot control. The ranking is three-way on purpose - agrees, no
metadata, disagrees - so an indexer reporting no author does not sort
below one reporting the wrong author.
A wrong verdict costs a release its position, never its visibility: a
transliteration such as Dostoevsky against Dostoyevsky sorts last
instead of vanishing. That is what makes the loose token comparison safe
to ship without a tuning knob.
Falling back to a title-only query on zero results was the alternative.
It only rescues total failure - if two of six editions happen to use the
provider's spelling, the search returns those two, no fallback fires,
and the user quietly gets a truncated list. It also spends a round trip
inside the search deadline and stacks a retry on an indexer that may
still be solving a challenge (#1249).
Manual queries skip author ranking: they are the user's own words and
should not be reordered against the metadata they were typed to
override.
|
||
|
|
3d7ea40088 |
fix(search): reach the server's deadline, query one author (#1285, #1252) (#1287)
Two independent reasons a working search reported failure to the user. 1. The client gave up before the server did (#1285) `/api/releases` bounds one release search with RELEASE_SEARCH_TIMEOUT (default 300s) and answers a spent budget with a sentence naming the real cause - the machinery added for #1276. The frontend then aborted the direct_download search at a hard-coded 180s, so it always won the race: the user saw "Request timed out. Check your network connection or proxy configuration." instead, and raising RELEASE_SEARCH_TIMEOUT changed nothing they could observe, the 180s being baked into the hashed bundle inside the image. - /api/config reports the effective (clamped) budget, and the client derives its abort from it plus a margin, so the server always answers first. - Direct-mode search shows what the server actually said. Every non-auth failure was relabelled "Unable to reach download source. Network may be restricted or mirrors blocked.", which discarded the explanation and blamed the user's network. ApiResponseError now carries `serverMessage`, set only when the server explained itself, so the status-line placeholder still falls back. Two latency fixes for the cost that made the timeout reachable at all: - Fetch each distinct AA search URL once per search. The language-filter retry re-runs every title variant, and with DIRECT_DOWNLOAD_LANGUAGE_FROM_PATH on both passes build a byte-identical URL - behind DDoS-Guard each repeat is a fresh browser solve. - Drop the solve-only bypass method. `_bypass_method_cdp_gui_click` opens with exactly that call and returns the moment it works, so the entry ahead of it could only repeat the half that had already failed, plus the backoff before the method that does work started. Reported at 0/19 successes and ~5.5s of each ~26s solve against DDoS-Guard. 2. The query carried every contributor, not one author (#1252) `_pick_search_author` returned `book.search_author` verbatim while the authors[] fallback beside it deliberately narrowed to the first name before a comma. Both fields routinely arrive holding every contributor joined with ", ": the frontend builds `book.author` as `authors.join(', ')` for display (bookTransformers.ts) and the release modal sends that display string straight back as the `author` parameter, and `browse_record_to_book_metadata` and the manual-search branch both split the joined text into `authors` while still passing the unsplit string as `search_author`, so the split was never used. A book whose metadata lists translators was therefore searched for as Blindness Jose Saramago, Giovanni Pontiero, <persian translator> which matches nothing on Anna's Archive. The bypass succeeds, the search comes back empty, and the user is told the book has no releases. Narrowed in one place, `search_plan.first_author`, so the two branches cannot drift apart again, and applied to the IRC source, which built its query with the same verbatim preference. Hardcover is unaffected: it already sets `search_author` from `_simplify_author_for_search(authors[0])`, which resolves "Last, First" itself and never yields a multi-author string. |
||
|
|
633004ecf0 |
fix(search): stop reading real Anna's Archive pages as unsolved challenges (#1294)
`_looks_like_challenge_page` substring-matched "ddos-guard"/"cloudflare" over the whole document. DDoS-Guard-fronted sites carry those strings on their own pages - Anna's Archive ships a `DDOS-GUARD` comment in the inline JS it serves on every page - so every real AA response that was not a results table was reported as an unsolved protection challenge, sending users off to fix a bypasser that had just succeeded. Measured against live pages: a served AA page (HTTP 200) is 182,685 bytes and matched the old detector; the real interstitial is 902 bytes. - `_looks_like_challenge_page` now delegates to the shared `challenge_marker()`, whose 64 KB cap is what separates a few-KB interstitial from the page behind it. `download/http.py` already used it; this module carried an unguarded private copy. - `_looks_like_aa_page` is checked ahead of the challenge branch. A genuine interstitial carries no AA markers, so nothing actually blocked leaks through. Also adds the diagnostics whose absence made #1289 guesswork: the debug bundle carries no response bodies, so "unsolved protection challenge" and FlareSolverr's "Challenge solved!" were indistinguishable after the fact. - `_log_untabled_search_page()` fingerprints the one ambiguous shape at INFO - size, size-cap verdict, AA markers, challenge marker - with a bounded 700-char head at DEBUG. Best-effort: it swallows its own errors. - The external bypasser records what it actually returned, and warns when it reports success while handing back a challenge page. Regression tests use fixtures built from the live pages rather than invented ones; the previous fixtures were two-line synthetic pages with no "ddos-guard" substring, which is why nothing caught this. Closes #1289 Closes #1292 |
||
|
|
69ff0d6a78 |
fix: trim a credit list in search_author to the first name (#1290)
Fixes #1252 for the case in the second report. `_pick_search_author` returns `search_author` untouched but trims `authors[0]` to its first comma-separated name. So the same credit list searches differently depending on which field carries it: ``` via authors[0] -> "Blindness Jose Saramago" via search_author -> "Blindness Jose Saramago, Giovanni Pontiero, Zohreh Eftekhari" ``` Anna's Archive answers the second one with nothing. That is the query in @theDoz12's log, and it explains the shape of the report: the bypass succeeds, the search runs, and the UI still says no releases. Nothing in the download path is broken, the query simply cannot match. Measured against live AA on 1.3.14, same book, same source, only the field carrying the author changed: | query | releases | | --- | --- | | `Blindness Jose Saramago, Giovanni Pontiero, Zohreh Eftekhari` | 0 | | `Blindness Jose Saramago` | 49 | | `Blindness` | 50 | With the patch the second form is produced from either field, and the same search returns 49. Three regression tests added, including one that asserts both fields yield the same query. On `tests/core/test_search_plan.py` the run goes from 5 failures to 3; the 3 that remain are the language tests, which fail identically with and without this change on my machine. Worth saying what this does not cover: the first report in that issue ends with `Found 2 releases via ISBN` and still shows nothing, so that one is a different fault further along. I could not reproduce it here. |
||
|
|
d7fe28595c |
fix(bypass): wait for the solved page before reading its source (#1286)
Follow-up to #1276 with a measurement from the instance I reported there. v1.3.13 solves the challenge again, but on my setup the solve was being thrown away immediately afterwards: ``` 19:26:08 Bypass successful using _bypass_method_cdp_gui_click 19:26:16 Bypass failed (attempt 1/10): TimeoutError: Time ran out while waiting for: {html} ``` `_get()` ends with `return await page.get_page_source()`, which is `find("html", timeout=1)` in SeleniumBase. One second is enough for a page that is already sitting on its content, but Anna's Archive answers a cleared check with a redirect to the real page, so the document is not there yet. The solve is discarded, the whole attempt restarts, and the extra requests are what earn the 429 that `note_rate_limited()` then parks the host for — 120 s, then 300 s. ## Change `_read_page_source()` waits for the document itself, with a `BYPASS_PAGE_SOURCE_TIMEOUT` setting (default 20 s, min 1, max 120) in Direct Download → Cloudflare Bypass, next to the existing bypasser timeouts. ## Measured on a live instance I patched the wait in the running container (`find("html", timeout=1)` → `timeout=20` in the installed seleniumbase, which is the same effect as this PR) and re-ran the same searches on the same host, k3s behind a Surfshark WireGuard exit, internal bypasser, v1.3.13: | | 1 s wait | 20 s wait | |---|---|---| | `Time ran out while waiting for: {html}` | one per solve | none | | 429 backoffs | 2 (120 s, then 300 s) | none | | Search for a book AA has | 199 s and 200 s, both errored | 61 s, 2 epub releases | A download after that took 5 s from LibGen, so the search was the whole cost. ## Tests Two tests in `tests/bypass/test_bypass_budgets.py`, the file already covering #1276: a page that needs longer than a second still yields its HTML, and `BYPASS_PAGE_SOURCE_TIMEOUT` overrides the default. `uv run pytest tests/ --ignore=tests/e2e`: 2848 passed, 47 skipped. Ruff check and format clean. The docs table is auto-generated, but running `scripts/generate_env_docs.py` here rewrote unrelated entries (Newznab, BOOK_LANGUAGE), so I added only the new entry by hand in the generator's format rather than commit that churn. One thing I could not judge from outside: whether 20 s is the right default for hosts other than AA. It only costs anything when a solve would otherwise be discarded, but I have measured it on one site. |
||
|
|
97e289ae13 | fix: search, Prowlarr and qBittorrent follow-ups (#1276, #1283) (#1284) | ||
|
|
c95ee72ad5 |
fix(qbittorrent): keep magnets whose metadata is still pending (#1282)
## Problem
`QBittorrentClient.add_download()` waits 20 × 0.5 s for qBittorrent to
leave `metaDL`, then raises:
```
Failed to add to qbittorrent: Torrent metadata resolution was not confirmed within the visibility grace period
(response=TorrentsAddedMetadata({'added_torrent_ids': [], 'failure_count': 0, 'pending_count': 1, 'success_count': 0}))
```
The wait exists to learn qBittorrent's primary torrent ID, which for
hybrid torrents switches from the v1 hash to the truncated v2 hash once
metadata resolves. A magnet on a thin public swarm routinely needs
longer than 10 s to find a peer that will serve metadata, and the
download is then abandoned even though the add itself succeeded. The
torrent stays in qBittorrent (`base_handler` logs "leaving in
qbittorrent") and often completes minutes later with nobody watching it.
Seen on v1.3.12 with public indexers through Prowlarr: every magnet-only
release failed this way, while `.torrent` releases from a private
indexer were fine. qBittorrent showed the same torrents at `metaDL 0%
seeds=0/0`, and they resolved on their own well after shelfmark had
given up.
## Change
Return the info hash we already have instead of raising when the grace
period expires. Reads then resolve either identity:
- `get_status()` and `get_download_path()` use `_resolve_torrent()`
instead of `_get_torrent_info()`, so a v1 hash still matches after
qBittorrent re-keys the torrent to v2. `_torrent_matches_download_id`
already compares `hash`, `infohash_v1` and `infohash_v2`.
- `remove()` and `set_category()` address the torrent by its current
primary hash through a new `_current_hash()` helper, which falls back to
the ID it was given when the torrent cannot be resolved.
- The two magic numbers become `_METADATA_WAIT_POLLS` and
`_METADATA_WAIT_INTERVAL_SECONDS`.
The happy path does not change. When metadata resolves inside the grace
period the resolved primary hash comes back as before, and
`_resolve_torrent()` tries the exact-hash lookup first, so it costs no
extra request.
## Tests
`test_add_fails_when_metadata_never_resolves` asserted the old
behaviour, so it becomes
`test_add_keeps_torrent_when_metadata_never_resolves` and asserts the
info hash is returned.
`test_get_status_resolves_hash_after_metadata_switch` is new: it reads
status by the v1 hash after qBittorrent reports the torrent under its v2
hash.
`uv run pytest tests/ --ignore=tests/e2e` gives the same 55 failures
with and without this change (they are all in `tests/bypass/` and need
Chrome, which my machine has no headless setup for), and
`tests/prowlarr/` is green at 524 passed. Ruff check and format are
clean. I have not run this branch against a live qBittorrent, so a
second pair of eyes on the `remove()` path would help.
|
||
|
|
b25acdb2ad |
fix(packs): don't disrupt normal downloads when inspecting for packs (#1274)
Follow-ups to the multi-book pack feature (#1270), which inspects every release before download. Two behaviours leaked into the ordinary single-book flow and are corrected here: - A flat folder of chaptered audio (`01 - Chapter.mp3`, `02 - ...`) was detected as a pack, because each track name parses to a series position, so clicking download popped the review panel for one normal audiobook. Flat folders are now split one-book-per-file only with real evidence of distinct books: two or more series positions, more than one title, and no chaptered audio (only the single-file m4b/m4a containers and ebook formats qualify). Subfolder packs and flat m4b/m4a packs are unchanged. - Every release that couldn't be inspected (usenet, magnet-only, sources without a list_files hook, ABB single-file) showed an info toast on download. That is now a console.warn, so a normal download is silent again. Adds regression tests for the chaptered-mp3 cases. |
||
|
|
f441b85da2 |
feat(packs): inspect multi-book releases and file each book separately (#1270)
## Multi-book packs: inspect a release before download and file each book separately Closes #576 ### Problem One queued release is always treated as one book. When a torrent is actually a whole series (`Series/Book 1 - Title/…`, or a flat folder of `Series 1.0 - Title.m4b` files), post-processing walks the whole tree, flattens every file into one list and renames them `Title - 01…10` under the searched book's `{Author}/{Title}`. Audiobookshelf then sees a single 10-file "book" and the user has to re-file everything by hand. ### What this does Most releases expose their file list *before* anything is downloaded, so the split is decided up front and approved by the user, then the download is fire-and-forget: 1. **Inspect** – clicking a release's download button now calls `POST /api/releases/inspect` first. A new optional `DownloadHandler.list_files(release_data)` hook returns the release's files without downloading: - **AudiobookBay** reads the torrent file table off the detail page it already fetches (the page is now cached for 120 s, so inspect + download cost ABB one request). - **Prowlarr** parses `info.files` from the `.torrent` it already fetches (the existing 120 s torrent-fetch cache is reused). Magnet-only and usenet releases report "can't inspect". - Other sources default to `None`. 2. **Review** – if the plan contains more than one book, the Find Releases modal swaps the list for a review panel: one row per book with editable title / series position / year, expandable file lists, non-book sidecars (`.txt`, covers) shown as ignored, a "Treat as a single book" switch, and **Download N books**. Single-book releases queue immediately, exactly as before. 3. **File** – the approved plan travels with the task (`DownloadTask.book_plan`, retry-safe) and post-processing files each book through the existing transfer code, one book at a time (`dataclasses.replace(task, title=…, series_position=…, year=…)`), so organize/rename templates, part numbering (now scoped per book), hardlinks, torrent copy-preserve and usenet handling are unchanged. Status reads `Complete (N books, M files)`. 4. **Fallback** – when a release can't be inspected the user gets a toast, and a small "Multi-book pack" toggle in the modal header forces a heuristic split (subfolder = book, or one book per file when the file names carry series positions). Planning lives in `shelfmark/download/postprocess/packs.py` and is shared by the inspect endpoint and post-processing, so what the user approved is what gets filed. The name parser strips `Book 3 -`, `03 -`, `1.0 -`, `3.`, `[03]`, `#3`, a leading series name, labels like "An Expanse Novella -", repeated titles (`Gods of Risk 2.5 - Gods of Risk`) and a trailing `(Year)`; author and series name come from the book that was searched, and the searched book's own series position is never applied to its siblings. ### Files - `shelfmark/download/postprocess/packs.py` (new) – `PackFile/PackBook/PackPlan`, `plan_pack`, `parse_pack_book_name`, `group_files_into_books`, `match_plan_to_files` - `shelfmark/core/release_inspect_routes.py` (new) – `POST /api/releases/inspect` - `shelfmark/release_sources/__init__.py` – `DownloadHandler.list_files` hook - `shelfmark/release_sources/audiobookbay/{scraper,handler}.py` – detail-page cache, `extract_file_list`, `list_files` - `shelfmark/release_sources/prowlarr/handler.py`, `download/clients/torrent_utils.py` – `extract_file_list_from_torrent`, `list_files` - `shelfmark/core/models.py`, `download/orchestrator.py` – `multi_book` / `book_plan` fields, queue + retry serialization - `shelfmark/download/postprocess/transfer.py`, `pipeline.py`, `outputs/folder.py` – per-book transfer branch and status message - `src/frontend`: `components/PackReviewPanel.tsx` (new), `ReleaseModal.tsx`, `App.tsx`, `services/api.ts`, `types/index.ts`, `utils/releasePayload.ts` (payload builder moved out of `App.tsx`), `utils/packReview.ts` - `docs/dev/release-sources-plugin-guide.md` – documents the `list_files` hook ### Out of scope (follow-ups) - Listing files from an NZB (Shelfmark already fetches the bytes; `<file subject>` names are noisy) - Inspecting magnet links via qBittorrent's files API after a paused add - BookLore / email outputs (they ignore `book_plan`; noted in code) - The combined ebook + audiobook flow ### Testing **Automated** (`make checks`, `make python-test`, `make frontend-test` all green; the only failures on my machine are the pre-existing `tests/config/test_entrypoint_permissions.py` cases, which need bash ≥ 4 and fail identically on `main` under macOS bash 3.2): - `tests/download/test_packs.py` – name parsing (markers, series name, novella labels, repeated titles, bare numeric titles like `1984`), nested / flat / mixed / deeper-nested packs, single wrapping folder not treated as a pack, plan-to-disk matching with basename fallback - `tests/core/test_processing_packs.py` – full `post_process_download` runs on a real temp filesystem: approved plan files each book under its own `{Author}/{Title}`, heuristic split of a nested pack, searched book's series position does not leak, multi-file book inside a pack keeps `- 01/- 02` per book, hardlinked torrent pack leaves the seeding tree intact, no pack fields ⇒ behaviour unchanged, single group degrades to the searched title, status message - `tests/core/test_release_inspect_routes.py` – plan response, not-inspectable, handler errors never 500, unknown source / missing `source_id` ⇒ 400, login required - `tests/audiobookbay/test_file_list.py` – file-table scraping from real ABB markup (multi-file and single-file pages), handler host validation, one page fetch shared by magnet + file list - `tests/prowlarr/test_torrent_file_list.py` – multi-file / single-file `.torrent` parsing, handler behaviour for torrent URL vs magnet vs usenet vs cache miss - `tests/download/test_orchestrator_pack_fields.py` – queue-time parsing and retry round-trip - Frontend: `releasePayload.test.ts`, `packReview.test.ts` (vitest) **Manual, on a real deployment** (arm64 image built from this branch, run as a side container next to production with the same qBittorrent / Audiobookshelf setup, `FILE_ORGANIZATION_AUDIOBOOK=organize`, hardlinks on): - AudiobookBay "The Expanse Complete 2.0" (7.87 GB, 36 files): clicking download opened the review panel in ~1 s showing **18 books · 18 files · 18 files ignored** (the `.txt` sidecars), with series positions 0.1–9.5 and years parsed from the file names; novella labels stripped ("The Churn", "The Butcher of Anderson Station"). Editing a title in the panel works. Confirming queued one task; the magnet resolved from the cached page in ~30 ms; after the download the task reported `Complete (18 books, 18 files)`, 18 hardlinks landed as `audiobooks/James S. A. Corey/<Title>/<Title>.m4b`, the torrent kept seeding, and Audiobookshelf scanned each folder as its own book (title, author, embedded chapters). - A second pack ("Expanse [01 - 9.5]", `Title N - Title` naming) was inspected to verify the repeated-title rule and the Back button, without downloading. - Single-book releases still queue immediately with no extra UI. |
||
|
|
02b7e9d958 |
feat(newznab): support multiple named indexers (#1271)
## Summary - add a named Newznab indexer table with per-indexer URL and API key settings - search every configured indexer and retain the originating indexer name on each result - namespace cached release IDs across connections and isolate individual indexer failures - preserve the legacy single-indexer settings as a fallback - support masked API-key cells and trusted SABnzbd prefetching for named indexers ## Validation - 121 Newznab and SABnzbd backend tests passed on Python 3.14 - Ruff passed for all changed Python files - frontend TypeScript and strict lint checks passed - all 134 frontend unit tests passed - frontend formatting check passed ## Compatibility Existing `NEWZNAB_URL` and `NEWZNAB_API_KEY` configurations continue to work whenever `NEWZNAB_INDEXERS` is empty. Co-authored-by: Ryan <zab1996@users.noreply.github.com> |
||
|
|
ff06a1a581 |
fix(search): follow-ups to per-user book languages (#1267)
Review follow-ups to #1255, all in the code that PR touched. Drop the dead user_id from the Prowlarr retry path. ProwlarrSource.search never reads plan.languages, and _refresh_release builds a synthetic book with no titles_by_language, so the title variants came out identical with and without it. It also should not language-filter: it re-finds one exact release by its guid. Pin the tab move in tests. BOOK_LANGUAGE moved from the General tab to Search Mode with no migration, which only works because both tabs persist into the same settings.json. Nothing asserted that, so splitting the files later would silently reset every install to ["en"]. Covers the stored value, a fresh install, and ENV precedence. Stop the UI inventing a default language. An empty BOOK_LANGUAGE is a deliberate "no default filter" that the backend preserves, but the two frontend call sites replaced it with the first supported language, so the filter said English where the server filtered nothing. resolveDefaultLanguageCodes now falls back only when the value is absent. Keep the normalized value for every validated search key. validate_user_settings gated the write-back on a hand-maintained subset of the keys the search validator recognises, so METADATA_PROVIDER_COMBINED, SHOW_COMBINED_SELECTOR and FORCE_COMBINED_SEARCH were validated and then stored raw -- a padded provider name was accepted and persisted with its padding. Reuse the validator's own key set instead. Skip blank language entries rather than rejecting them, so "" and "en," mean the same as [] and ["en"] instead of erroring on an unnamed language. Extract resolveListOverride for the list-override detection that was copy-pasted between the two user-settings sections, and mention languages in the Search Preferences section description. |
||
|
|
463ef49ac3 |
feat(search): let each user pick their own default book languages (#1255)
## Why `BOOK_LANGUAGE` is a per-reader property, not a per-instance one. On a shared install one household member searches in German while another wants English and German — today whoever changes the setting changes it for everyone, and the only escape is re-picking languages in the filter on every single search. The per-user override machinery already carries `SEARCH_MODE`, the metadata providers and the default release sources, so the language default mostly had to opt into it. ## What changed **The field.** `BOOK_LANGUAGE` becomes `user_overridable` and moves from the **General** tab to **Search Mode**, next to the other user-overridable search defaults (per [review](https://github.com/calibrain/shelfmark/pull/1255#issuecomment-5391189094) — the first version had the Search section span two tabs, this one doesn't). Admins set it per user in the user editor, users set it in **My Account → Search Preferences**, and the Search Mode tab carries the usual "N users override this" summary. **No migration for the move.** `general` and `search_mode` both persist into `settings.json`, and a field's value is resolved through `load_config_file(tab)` for the tab it's declared on — so an install that already stores `BOOK_LANGUAGE` keeps its value. Checked against a `settings.json` written while the field still lived on General: the stored value resolves unchanged, a fresh install still gets `["en"]`, and `BOOK_LANGUAGE` in the environment still overrides both. **The two places the default is read.** - `/api/config` seeds the frontend's language filter, so it now resolves `BOOK_LANGUAGE` for the session user. - `build_release_search_plan` falls back to the default whenever a request carries no language filter — which is exactly what the filter's "Default" option sends. It takes an optional `user_id`, passed by `/api/releases` from the session and by the Prowlarr retry path from `task.user_id`, so a retry re-searches in the languages of whoever queued the download. **Validation.** Overrides go through `normalize_language()`, so `"German"`, `"ger"` and `"de"` all store as `de`, and an unknown language is rejected with a message naming it instead of being silently searched for. An empty list stays an empty list (a deliberate "no default filter"), `null` clears the override as everywhere else, and ENV still wins: with `BOOK_LANGUAGE` set in the environment the field reports `fromEnv` and overrides are ignored. **Scope.** Only the language default becomes overridable. The two format lists left behind under "Default Search Filters" stay admin-only — they describe what the library and its post-processing accept, not what a reader wants to read. There's a test pinning that. ## Verification - 2681 unit tests pass (2670 before, 11 added) - `ruff check`, `ruff format`, `basedpyright` over backend and tests, and `vulture` all clean; frontend lint, format, typecheck and 126 unit tests clean - `docs/environment-variables.md` regenerated via `scripts/generate_env_docs.py` (the `BOOK_LANGUAGE` row follows the field into the Search Mode section) - Manually against a two-user instance with builtin auth (first round, before the tab move): with user A on German and user B on English+German, `/api/config` returns each reader their own `default_language` and an unfiltered `/api/releases` plans the matching languages; an admin can set and read the same override for another user; clearing it falls back to the global value; a stray `"klingon"` is rejected; and `BOOK_LANGUAGE` in the environment overrides both users with the field marked `fromEnv` - After the tab move I re-ran the suites above plus the stored-value/fresh-install/ENV check described under "No migration for the move"; the behaviour it exercises is what the move could have broken Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: CaliBrain <calibrain@l4n.xyz> |
||
|
|
a5595cf9f1 | Change test for fake extension that wont work (#1266) | ||
|
|
9bcf595111 |
feat(prowlarr): warn when an indexer declares a format Shelfmark can't process (#1265)
## Problem Companion to #1264, but general rather than mp4-specific. MyAnonamouse titles carry a structured `[LANG / FORMATS]` bracket that `_extract_mam_formats` parses. When every token in it is something Shelfmark doesn't know — e.g. `The Martian by Andy Weir [ENG / MP4]` — the release is rendered with **no format chip at all**, just the generic headphones/book icon with an "Audiobook" tooltip. To a user that looks like an ordinary result. It downloads fine and then fails post-processing with *"No book files found in download"*. The backend already *had* the signal (a format token it couldn't map); it just threw it away. ## Change **Backend** (`shelfmark/release_sources/prowlarr/source.py`) - `_split_mam_formats(raw_title) -> (recognized, unrecognized)` replaces the body of `_extract_mam_formats`, which is kept as a thin wrapper returning `recognized` so nothing else changes. - Releases gain `extra["unrecognized_formats"]` (list, or `None` when empty / when format detection is off). **Frontend** - `getUnrecognizedReleaseFormats(release)` in `utils/releaseFormats.ts` (normalised + deduped, same shape as `getReleaseFormats`). - `ReleaseCell` `format_content_type`: when there is **no** recognised format but the indexer named one, render an amber `MP4 Unsupported` badge (compact view: amber `MP4`) with tooltip *"Unsupported format (MP4) - Shelfmark cannot process this release"*. When a recognised format exists the existing badge is untouched, even if extra unknown tokens were present. Only the chip changes — the download button still works, so a user can still grab and hand-process the files if they want to. Happy to disable the button instead if you'd prefer. ## Tests - `tests/prowlarr/test_source.py`: `TestSplitMamFormats` (recognised / unrecognised / mixed / no bracket / wrapper compat) and `TestUnrecognizedFormatOnRelease` (lands in `extra`, empty when recognised, absent without format detection). - `src/frontend/src/tests/releaseFormats.test.ts`: 3 cases for the new helper. - `ruff check` clean; `pytest tests/prowlarr -m "not integration"` 511 passed; `tsc --noEmit`, `oxlint --deny warnings`, `vitest` all clean. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
65e2e3be20 |
feat(audiobook): recognise .mp4 as an audiobook format (#1264)
## Problem Some trackers — MyAnonamouse in particular — distribute AAC audiobooks as per-chapter `.mp4` files. That's the same ISO-BMFF container as `.m4a`/`.m4b`, just with the generic extension (`ftyp isom`, audio-only). Today those releases: 1. show up in Prowlarr search results with **no format chip** — only the generic "Audiobook" icon, because no format could be inferred; 2. download successfully; then 3. fail post-processing with **"No book files found in download"**, because `.mp4` isn't in `AUDIOBOOK_FORMATS` (`shelfmark/core/utils.py`). Real example: MAM #627978, *The Martian* (Andy Weir, 2020 edition) — 142 files `0001 … 0142 Andy Weir (2020) The Martian.mp4` + `cover.jpg`, 305 MB. Every file is a valid AAC-in-MP4 chapter. Adding `mp4` to `SUPPORTED_AUDIOBOOK_FORMATS` in `settings.json` doesn't help since the hard-coded tuple is what post-processing scans against. ## Change - Add `"mp4"` to `AUDIOBOOK_FORMATS` (single source of truth — settings UI, Prowlarr parsing, IRC parser, archive extraction and post-download scan all derive from it), with a comment explaining why. - Add `".mp4"` to the two hand-maintained debrid `_BOOK_EXTENSIONS` lists (AllDebrid / Real-Debrid) so file selection matches. - Slot `mp4` into the IRC `AUDIOBOOK_FORMAT_PRIORITY` table right after `m4a` (same container family). - Update the documented default in `docs/environment-variables.md`. - New regression test `test_audiobook_multifile_mp4_chapters_are_book_files` modelled on the existing multi-file usenet test. ### Note for existing installs The legacy-default migration only widens configs that still hold the old `m4b,mp3` list, so users on the current widened default won't pick up `mp4` automatically — they'll need to tick it in Settings → Audiobook formats. New installs get it by default. Happy to extend the migration if you'd rather it be automatic. ## Testing - `ruff check` / `ruff format --check`: clean - `pytest tests/core tests/config tests/irc tests/prowlarr tests/download -m "not integration and not e2e"`: 2296 passed, new test + `test_audiobook_format_consistency.py` all green. The 10 failures in `test_entrypoint_permissions.py` / `test_orchestrator_stall.py` reproduce identically on untouched `main` on macOS and are unrelated. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |