mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-10-04 10:11:13 +01:00
main
239
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9321096032 |
fix(clients): correct Debrid-Link file listing and rate-limit handling (#1425)
Follow-up to #1380. - Fetch the file list with ?ids=, the documented parameter. There is no ?id=, so a many-file torrent was not expanded and the call could fail with badArguments at the last step of every download. - Back off status checks for 60s after floodDetected, doubling to a 240s cap, instead of an hour. The orchestrator cancels a download after five minutes without a change, so the hour-long pause cancelled every active download. Only a flood on the status check starts it, and a successful check resets it. - Drop the client-side one-hour stall timer. The orchestrator stall timer always fires first. - Keep progress at 50% when file retrieval starts instead of dropping to 0% until the first file arrives. - Report queued, paused and verifying torrents as QUEUED, PAUSED and CHECKING, so a queued torrent gets the handler queue grace. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
cad18b8019 |
feat(clients): add Debrid-Link as a torrent client (#1380)
calibrain said in #1315 that he'd review and merge a Debrid-Link client if someone took a stab at it, so here it is. It's the fourth debrid service alongside AllDebrid, Real-Debrid and TorBox and it uses the same DownloadClient interface as the other three. No frontend changes, the settings UI picks it up from the backend field definitions. **I can't test this against the live API.** I don't use debrid and I don't have a Debrid-Link key. Everything here is built from their v2 API docs and cross-checked against two independent reference implementations, so the request shapes should be right, but nobody has run a real magnet through it yet. @Tatsu941 you opened the issue and mentioned they give developers free premium. Would you be able to try it? The things most worth confirming are the field names on a seedbox torrent, that `premiumLeft` is the right premium check, and that an API error comes back the way the client expects. Happy to fix whatever turns up. Two things about their API shaped how this is written. Every v2 endpoint wraps its payload in a `{success, value}` envelope and a failure can arrive with an HTTP 200, so `_request_value` decides the outcome from the body instead of the status code. A completed seedbox torrent already carries a direct `downloadUrl` on every file. The other three clients each need a separate unrestrict or link-request call per file and this one doesn't, which is why it came out shorter than them despite doing the same job. Endpoints used: | Endpoint | Used for | |---|---| | `GET account/infos` | test the key, confirm premium time remains | | `POST seedbox/add` | magnet as JSON, .torrent as multipart | | `GET seedbox/list` | poll progress | | `DELETE seedbox/{id}/remove` | clean up | Ruff and basedpyright are clean and the 36 new tests pass. The two failures in `tests/e2e/test_proxy_auth_middleware.py` are already on main and fail the same way on a pristine checkout. They assume the first proxy-authenticated user becomes an admin, which only holds while the user table has no admin in it, so they depend on how many users earlier tests created. One thing for review: the download loop uses `download_url` like the other three clients, which holds each file in memory. I kept the existing pattern rather than do something different in a new-provider PR, but say so if you'd rather it streamed here. Closes #1315 |
||
|
|
f773ac6029 |
fix(sabnzbd): strip Prowlarr API key on cross-host NZB redirects (#1424)
_fetch_nzb_content sent X-Api-Key then followed redirects with the header still attached, leaking the Prowlarr key when Prowlarr 302s to the indexer download link. Follow redirects manually with allow_redirects=False and re-evaluate headers per hop, mirroring the torrent path. |
||
|
|
755f28a6a4 |
feat(api): add a read-only API key and a /api/stats endpoint (#1419)
This is the read-only API key from #1410, where you said to go ahead. I run Shelfmark behind Homepage and wanted more on the dashboard tile than up or down, without putting an admin API key in the dashboard's config. This adds a second key, `SHELFMARK_API_KEY_READONLY`, that can read one new endpoint. `GET /api/stats` returns counts only: books added over the last 7 and 30 days by format, the queue, requests by outcome, and download failures over the last 7 days. No titles and no user names. The read-only key, the admin key or an admin session can read it. The read-only key gets a 403 on every other path and on any write, including a POST to `/api/stats` itself, and it never gets a session or a cookie. If a request carries both keys, the admin one wins, so a proxy's own `Authorization` header can't downgrade a correct `X-Api-Key`. Unset means off, like the existing key. The endpoint exists either way, but without a key only an admin session can read it. It also adds a `db_path` property on `UserDB`, so the stats module can open the database read-only. Tests cover the key scope, the 403s for the read-only key, both keys at once, who can reach the endpoint, and the counters. The full suite passes, and I broke each rule on purpose to check a test catches it. A version of this has been running on my own install behind Homepage. If you'd rather have the stats in a different shape or under a different path, I'm happy to change it. |
||
|
|
32abaff21e |
feat(naming): {Narrator} placeholder and MyAnonamouse series fallback (#1407)
## Human-Written explanation:
Following up on previous PR to use MAM ID to get nararrator and series
to show up in search, this PR will allow the series and nararrator
fields to be added to the path when saving an audiobook.
I have tested my ghcr.io image on my instance and it seemed to work
properly.
All the text below is written by Claude.
# feat(naming): {Narrator} placeholder and MyAnonamouse series fallback
Follow-up to #1390 and #1399. Related to #605 and #934 (narrator
support).
## Problem
Several narrations of the same audiobook currently land in the same
folder, because nothing in the path template tells them apart. Keeping
more than one version means renaming folders by hand after every
download.
Since #1390, MyAnonamouse results carry the release's narrator and
series in `extra`, but the naming templates can't use them: there is no
`{Narrator}` placeholder, and `{Series}` / `{SeriesPosition}` only come
from the metadata provider.
## Changes
- **`{Narrator}` placeholder** (`core/naming.py`,
`download/postprocess/transfer.py`, `core/models.py`): `DownloadTask`
gains `narrator`, read from `narrator` or `extra.narrator` when a
release is queued, and kept in the restart-safe retry payload. It's
available in audiobook templates like any other placeholder, including
prefix/suffix blocks such as `{ - Narrator}`.
- **Audiobookshelf folder style**: Audiobookshelf reads the narrator
from folder names like `Title {Narrator}`. The template `{Title}
{{Narrator}}` already matches as `{` + `{Narrator}`; the parser now also
drops the closing brace when the narrator is empty, so the folder is
`Title` rather than `Title }`.
- **No stray spaces around `/`**: an empty placeholder at the start or
end of a folder name no longer leaves a space there (e.g. `Title
/Title`). This applies to all templates.
- **Series fallback** (`release_sources/prowlarr/mam.py`, `source.py`,
`download/orchestrator.py`): MAM enrichment also stores the first
series' name and number as `extra.series_name` /
`extra.series_position`, which queueing already falls back to when the
metadata provider has no series. The provider's series still wins. A
release-level series number is only used when it names the same series
as the provider, so a provider series is never paired with the number of
a different MAM series (an omnibus, for example).
- **Settings and preview** (`config/settings.py`,
`namingTemplatePreview.ts`): the audiobook template descriptions list
`{Narrator}` and explain `{{Narrator}}`. The settings page preview
mirrors the parser changes and offers `{Narrator}` under "Insert
variable" for audiobook templates only.
- **Docs**: regenerated `environment-variables.md`. This also picked up
the Library Check settings, which weren't in the generated docs yet;
happy to drop that part to keep the diff focused.
## Example
Audiobook Path Template `{Author}/{Title} {{Narrator}}/{Title}`:
| Narrator | Result |
|----------|--------|
| Samuel Roukin | `Christopher Ruocchio/Empire of Silence {Samuel
Roukin}/Empire of Silence.m4b` |
| none | `Christopher Ruocchio/Empire of Silence/Empire of Silence.m4b`
|
## Testing
- `tests/core/test_narrator_template_variable.py` (new, 16 tests):
`{{Narrator}}` with and without a value, the `{ {Narrator}}` prefix
form, prefix/suffix blocks, sanitizing, `build_library_path`, task
metadata, retry payload round trip, MAM series name/number parsing
(including a `1-3` range having no number), and queueing (narrator from
`extra`, series fallback, no cross-series number, blank narrator).
- `namingTemplatePreview.test.ts`: the same `{{Narrator}}` cases for the
preview.
- Updated `test_generate_env_docs.py` for the new template description.
- Full `pytest -m "not integration and not e2e"` compared with an
upstream `main` worktree on the same machine: no new failures (the
remaining ones are Windows-only on both).
- `ruff check`, `ruff format --check`, `basedpyright` (0 errors) and
`vulture` on touched files; frontend `tsc --noEmit`, `oxlint`, `oxfmt
--check`, `vitest` (218 passed).
- Checked the settings page locally: the audiobook preview renders `...
{Simon Vance}.mp3` and `{Narrator}` is listed for audiobook templates
only.
No behavior change for existing templates, apart from spaces next to `/`
being trimmed.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
119d3374e3 | fix(downloads): reconcile downloads a restart interrupted (#1423) | ||
|
|
95d9f1da57 |
fix(sabnzbd): prefetch NZBs served from the indexer's own domain (#1416)
Indexers often serve NZB downloads from a different host than their API (e.g. file.indexer.example for indexer.example/api). The exact-origin check added in #967 rejected those, so SABnzbd silently fell back to addurl. Trust any host in the configured indexer's registrable domain, using the Public Suffix List so shared suffixes like co.uk or duckdns.org never widen trust. Scheme and port must still match, IP literals and single-label hosts still require an exact match, and the Prowlarr API key header is unchanged. This closes this issue: https://github.com/calibrain/shelfmark/issues/1411 Disclaimer: Implemented by Claude and Co-Authored by me. Tested and verified by me alone (Why is it always this way around and not the other) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
8392b43dd3 |
fix(downloads): reconcile downloads a restart interrupted (#1409)
I found this while grabbing new books and working on my fork. I redeployed the container while downloads were running, and it left a mess in my queue. It's an edge case, but a few rows were stuck until I cleared them by hand. The download queue lives in memory, so stopping the process with a download in flight leaves its history row "active" and its request "queued", with nothing working on either. The activity API shows that row as "Interrupted" when it's read, but nothing is stored and nothing looks at it again. Upstream only reopens an errored request when no manual retry is offered, and an interrupted row always offers one, so the request just sits there. This closes those rows out once at startup, before the coordinator starts, as an error with the message "Interrupted". The retry payload stays, so the manual retry still works. A request whose download was queued in the last 24 hours goes back to pending through `reopen_failed_request`. Older ones get closed out but their requests are left alone, because I didn't think a startup should reopen a backlog or fill an admin's approval queue with old requests. A request that already arrived is left alone, and running it twice changes nothing. It assumes a single worker process, which `entrypoint.sh` already sets. On my install the first start after deploying it closed 53 rows that had been orphaned for nine days. Tests are in `tests/core/test_startup_reconcile.py`, and the full suite passes locally. If the 24 hour window feels off, or you'd rather handle this another way, I'm happy to change it or drop the PR. Co-authored-by: CaliBrain <calibrain@l4n.xyz> |
||
|
|
7888217380 |
feat(download): add Remove & Delete Files torrent completion action (#1401)
Closes #1027 ## Summary - add a **Remove & Delete Files** (`remove_and_delete`) option to the Torrent Completion Action setting - after a successful import, remove the torrent and ask the client to delete its downloaded data (`delete_files=True`), matching what the usenet "move" flow already does - clarify the setting description so it says **Remove** keeps the downloaded files (raised in the issue comments) - regenerate the `PROWLARR_TORRENT_ACTION` entry in `docs/environment-variables.md` ## Behavior The deletion runs from `post_process_cleanup`, the same place as the existing Remove and Change Category actions, so it only happens after output transfer and post-processing have succeeded. At that point every file has already been copied or hardlinked into the library, and transfer size checks have passed. A failed import never removes or deletes anything. The torrent client deletes its own data, so Shelfmark does not delete paths itself and remote path mappings are not involved. Keep, Remove, and Change Category behave as before, and Keep is still the default. If a request matched a completed torrent that was already in the client, this option deletes that torrent's data too. That is the same scope the existing Remove action already applies to. ## Validation - `make python-checks` (ruff check, ruff format, basedpyright on backend and tests, vulture): passed - `make python-test`: 3363 passed, 5 skipped, 9 failed. All 9 are in `tests/config/test_entrypoint_permissions.py` and happen because macOS `/bin/bash` 3.2 does not support `${1,,}` in `entrypoint.sh`. They do not touch this change. - new tests in `tests/prowlarr/test_handler.py`: - Remove passes `delete_files=False` and Remove & Delete Files passes `delete_files=True` - a failed import with Remove & Delete Files does not call the client - the delete case fails without the handler change - manual smoke run: `PROWLARR_TORRENT_ACTION=remove_and_delete` set through real config loading, with a stub client that deletes its folder. The torrent folder was deleted and the hardlinked library file stayed intact. - pre-commit hooks (prek) passed --------- Co-authored-by: Evan Kazakin <evan@Evans-MacBook-Pro.local> |
||
|
|
d99e9d7c4f |
fix(download): remove a cancelled torrent that never started (#1421)
Cancelling a torrent download leaves the torrent in the client. `_safe_remove_download` documents the rule: "torrents: never remove or delete client data (avoid breaking seeding)". That makes sense for a torrent that has data. A torrent that is still at 0% has nothing to seed and nothing to resume, and leaving it behind holds a download slot in qBittorrent's queue or sits there as a dead entry. I hit this on a qBittorrent that is shared with Sonarr, Radarr and Lidarr and has a download limit. Shelfmark cancelled downloads that were queued (#1420) and left them in the client. Three torrents that never fetched metadata held every active slot, and each torrent added after them waited behind them. Removing those by hand freed the slots. This removes the torrent and its files when a download is cancelled, Shelfmark added the torrent, and the client reports 0% progress. A torrent with any data is left alone as before. So is one that was already in the client when the download started, since Shelfmark didn't add it. If the removal fails it is logged and the cancel carries on. Tests cover removal at 0% for a queued and a downloading torrent, leaving one that has data, leaving one the user already had, and a failing removal. The full suite passes, and I broke each rule on purpose to check a test catches it. This changes a rule you wrote down, so I've kept it apart from the queue fix. If you'd rather have it as a setting, or only remove a torrent that is still queued or fetching metadata, I'm happy to rework it. |
||
|
|
88f73e3bc3 |
fix(download): do not cancel a queued torrent as stalled (#1420)
qBittorrent only runs a few downloads at once (`max_active_downloads` defaults to 3) and holds the rest as `queuedDL`. Shelfmark's stall timer cancels a download after five minutes without a changed status or progress. A queued torrent reports "Queued" on every poll, so it looks idle, and Shelfmark cancels a download that was only waiting for a slot. I run Shelfmark against the same qBittorrent that Sonarr, Radarr and Lidarr use, which I assume is a common setup. Their torrents and Shelfmark's share one limit, so anything that fills the slots makes Shelfmark's torrent wait. In my case three torrents that never fetched metadata held every slot. I'd expect a busy Sonarr queue to hold them the same way, but I haven't reproduced that. Once the stall timer had cancelled the waiting downloads they stayed in the client, still queued behind the same blockers. On my instance 10 AudioBookBay downloads were cancelled while qBittorrent still had them queued. This changes the torrent poll loop. When the client reports a queued torrent, it asks for an activity grace (the mechanism `html_get_page` already uses for protection bypasses) and releases it as soon as the torrent leaves the queue. The grace is renewed every ten minutes and stops at two hours, so a queue that never moves still ends. A torrent that is downloading and moving is timed as before until the new window below applies. The second commit sets the stall message before the cancel runs. The terminal hook copies the message into the history row, and it was set afterwards, so a download the timer ended was recorded with the last poll's "Queued". That left no way to tell it apart from a person cancelling. The last two commits are about slow torrents. The timer only counts a change in progress, so a torrent that needs more than five minutes to fetch metadata or find its first peer is cancelled before it has any progress to count, and one that moves a few megabytes at a time can be cancelled between bursts. I hit this with a torrent that had a single peer. A torrent now gets one 15 minute window the first time it is seen outside the queue, and each step forward in progress renews it. One that makes no progress for 15 minutes is still cancelled, so a dead torrent now takes about 15 minutes to end instead of five. That trade is the part I'm least sure about, so the number is easy to change. Tests cover the grace being requested once, released when the torrent starts, renewed, and capped, the order of the message and the cancel, and the start and movement windows. The full suite passes, and I broke each rule on purpose to check a test catches it. The two hour ceiling is a guess. If you'd rather have a setting for it, or a different number, I'm happy to change it. |
||
|
|
bac0b93566 |
fix: keep https:// for qBittorrent and hide release sources from non-admins (#1422)
qBittorrent over HTTPS (#1417) qbittorrent-api ignores the scheme in `host` and probes HTTP and HTTPS itself. When the HTTPS probe fails, for instance on a certificate Shelfmark doesn't trust, it falls back to plain HTTP against the TLS port, and a reverse proxy answers "400 The plain HTTP request was sent to HTTPS port". The download client and the Test Connection button now pass FORCE_SCHEME_FROM_HOST for https:// URLs, so the configured scheme is used as-is and a certificate problem is reported as one. http:// and bare host:port URLs keep the probe: normalization adds http:// to a bare host, and the probe follows a proxy's redirect to HTTPS. Release sources leaked to non-admins (#1418) A download queued from an approved request carries the release an admin picked, and the requester could read it from their own activity feed: the request's release_data (source_id, indexer, info URL, torrent attributes), the download's full server path, and the download id itself, which for Prowlarr is "<indexer id>:<guid>" and for a private tracker is a URL into it. That id keys every status payload, the /api/localdownload query and the cover proxy URL, so hiding release_data alone would not have closed the leak. For non-admin viewers: - request rows keep only the release fields the activity cards display - download_path is reduced to the file name, which the browser download reveals anyway - request-linked downloads are addressed by an opaque id (a keyed HMAC of the task id) in the snapshot, history, /api/status, queue order, active downloads, dismissed keys, cover URLs and the websocket status and progress events. Routes that take a download id translate it back, searching only the caller's own downloads. Downloads a user queued directly keep their real id: the user picked that release from search results that already showed it, and the release list matches its buttons to the queue by that id. Admin views are unchanged. The activity sidebar now links a fulfilled request to its download by request_id instead of release_data.source_id. Closes #1417 Closes #1418 |
||
|
|
41a4df01c4 |
fix: hand DiamWall 513 challenges to the bypasser (#1400)
## Summary Recognize DiamWall interstitials in both the HTTP retry path and the internal browser helper. HTTP 513 challenge responses go directly to the configured bypasser, while ordinary 513 responses keep the existing error behavior. Fixes #1386 ## Testing - `uv run --frozen --extra browser pytest -q -n 0 tests/download/test_http_challenge_513.py tests/bypass/test_internal_bypasser.py`: 27 passed - Ruff lint and format checks passed |
||
|
|
893f7bdb92 |
fix(mam): keep the session ID on MAM and rerun Prowlarr's exact search (#1399)
Follow-up to #1390. - The origin came from each result's infoUrl, matched by a regex that did not check the host, and the mam_id cookie had no domain. Any Prowlarr indexer returning a URL like https://evil.example?myanonamouse.net/t/1 sent the session ID to evil.example. Requests now always go to https://www.myanonamouse.net (Prowlarr's only MAM URL), with the cookie as a header and redirects off. Only results from Prowlarr's MyAnonamouse indexer are looked up, and their URLs must be on myanonamouse.net. - The lookup searched every category and read one page, so for common titles most of Prowlarr's results were missed (a "Dune" audiobook search: 56 audiobooks on the first 100 of 328 matches). It now reruns Prowlarr's exact search: the same query clean-up, the MAM main categories behind the Torznab categories searched (13/15/16 for audiobooks, 14 for e-books, all once expanded), and the MAM indexer's own search type, search-in options and languages. Further pages are read while IDs are missing, page 1 of every title first, at most 4 requests per search. - Failed requests back off for 1, 2, 4 ... up to 30 minutes. The 10th consecutive failure stops enrichment until Test MAM Session passes, the session ID changes, or Shelfmark restarts. - The detail cache prunes expired entries instead of growing for as long as Shelfmark runs. |
||
|
|
a05e002305 |
feat: narrator, series and bitrate columns for MyAnonamouse results (#1390)
# feat: narrator, series and bitrate columns for MyAnonamouse results Addresses #605 and #934 (narrator in the release list). ## Problem Both requests were closed because the narrator isn't in Torznab results, which is correct. Prowlarr's MyAnonamouse indexer reads `author_info` but drops the `narrator_info` and `series_info` MAM returns next to it, and neither Prowlarr's `ReleaseInfo` nor the Torznab output has a field for them. Shelfmark's size tooltip already looks for a `narrator` Torznab attribute, but nothing ever sends one. When a book has several narrations, choosing one means going back and forth between Shelfmark and the tracker. ## Approach Optional, opt-in enrichment using the user's own MAM session (`mam_id`), the same approach AudioBookRequest's MAM indexer uses: 1. After the Prowlarr search, releases whose info URL is `…myanonamouse.net/t/<id>` are collected. 2. Shelfmark sends the same query text to MAM's JSON search (`/tor/js/loadSearchJSONbasic.php`, normally one request, at most 3), and matches torrents back to Prowlarr's results by torrent ID. The MAM origin is taken from the result URL, so a custom Prowlarr MAM base URL is respected, and the cookie is only ever sent to `*.myanonamouse.net`. 3. Adds `extra.narrator`, `extra.series` (`The Sun Eater #1`) and `extra.bitrate` / `extra.bitrate_value`. MAM has no bitrate field, so it's parsed from the uploader's free-text tags (`64 kbps`) and some releases won't have one. Lookups are cached per torrent for an hour, respect the existing Prowlarr search deadline, go through the configured proxy (`get_proxies`), and never fail the search: a 403, timeout or bad JSON is logged and the list renders without the extra details. ## Changes - **`release_sources/prowlarr/mam.py`** (new): small MAM client (`search`, and `get_username` for the test button), `narrator_info` / `series_info` / tags parsing, cached best-effort `lookup_torrent_details()`. A 403 includes MAM's reply and a note about the IP/ASN lock. - **`release_sources/prowlarr/source.py`**: enrichment after the result loop. Series, Narrator and Bitrate columns only when a MAM ID is configured, since otherwise they would be empty for every row. A Torznab `bitrate` attribute from other indexers is also mapped to `extra.bitrate`. - **`release_sources/prowlarr/settings.py`**: "MyAnonamouse Enrichment" section with `PROWLARR_MAM_ID` (password field, env-overridable like every setting) and a **Test MAM Session** button. - **`release_sources/__init__.py`**: `ColumnSchema` gains optional `setting_key` and `content_types`. The new `apply_column_visibility()` drops gated columns and their grid tracks. Neither field is serialized. - **`main.py`**: `/api/releases` applies `apply_column_visibility()` with the request's content type and the user's effective settings. - **`config/settings.py` / `users_settings.py`**: Search Mode › "Release List Columns" with `SHOW_SERIES_COLUMN`, `SHOW_NARRATOR_COLUMN` and `SHOW_BITRATE_COLUMN`, all default on and user-overridable. Narrator and bitrate are audiobook-only, series shows for both. AudiobookBay's existing bitrate column now follows the bitrate toggle. - **Frontend**: text cells truncate with a hover title; the mobile info line wraps and skips empty text/number cells so blank optional columns don't leave orphan `·` separators; the size tooltip no longer lists Bitrate twice. - **Docs**: new `docs/myanonamouse-enrichment.md` (linked from the index), and a regenerated `environment-variables.md`. The regeneration also picked up a few pre-existing drifts from `main` (the Libgen section, `AA_DEFAULT_SORT` default, a duplicate `BOOK_LANGUAGE` row). I can drop those if you'd rather keep this diff focused. ## ⚠️ MAM sessions are IP/ASN-locked MyAnonamouse locks each session to one IP or ASN. Reusing the session Prowlarr (or a seedbox script) uses will often **403**. **A separate MAM session for Shelfmark will likely be needed** when Shelfmark reaches MAM from a different IP (another host, a VPN container, or a proxy in Shelfmark's Network settings), or when the existing session is ASN-locked to another network. The setting's description, the error message and the new doc all say so. ## Testing - `tests/prowlarr/test_mam_enrichment.py` (new, 19 tests): parsing (narrator dedupe, multiple series, missing numbers, malformed JSON, tag bitrate), lookup (stops once all IDs are found, cache, 403 and connection errors return empty, expired deadline skips the request), only MAM releases enriched, Torznab bitrate mapping, column config with and without a MAM ID for audiobook and ebook, toggles, grid-track removal, and gates not serialized. - `tests/core/test_admin_users_api.py`: the curated search-preference key list now includes the three toggles. - Full `pytest -m "not integration and not e2e"` compared with an upstream `main` worktree on the same machine: no new failures. The remaining ~115 failures on both are Windows-only (tor/entrypoint shell tests, path separators). - `ruff check` / `ruff format --check` / `basedpyright` (0 errors) / `vulture` on touched files; frontend `tsc --noEmit`, `oxlint`, `oxfmt --check`, `vitest` (201 passed). - Manually verified with a real MAM account on a Docker build of this branch: the test button, then narrator, series and bitrate on MyAnonamouse audiobook results. No behavior change unless `PROWLARR_MAM_ID` is set, apart from the bitrate toggle on AudiobookBay (default on, same as today). 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: CaliBrain <calibrain@l4n.xyz> |
||
|
|
dce91e6972 |
fix(audiobookbay): reuse the resolved magnet when retrying (#1388) (#1398)
Retrying an AudiobookBay download re-scraped the detail page for the magnet link before checking the torrent client, so a torrent that had already finished in the client could not be imported while AudiobookBay was down. The handler now keeps the magnet it resolved in the task's retry context, which is persisted with the download history, and a retry hands that magnet straight to the client's existing-download check. The context is handler-owned, so a client-supplied download URL cannot seed it. Fixes #1388 |
||
|
|
ece6b8f341 |
fix(auth): fail closed when auth prerequisites are missing (#1387) (#1397)
Auth mode resolution fell back to "none" (anonymous full admin) whenever the configured method's prerequisites were missing: no local password admin for builtin/OIDC, no Calibre-Web database, a blank proxy header, an unrecognized AUTH_METHOD (including "OIDC" in uppercase), or any error while reading the config. Deleting or demoting the last local admin was allowed on purpose because of that fallback, which exposed OIDC instances publicly. - Only an explicit AUTH_METHOD=none disables authentication. A configured method stays active when its prerequisites are missing, so sign-in fails instead of opening up. - An unrecognized or unreadable AUTH_METHOD resolves to "unavailable", which still requires a session and accepts no login. Values are normalized, so AUTH_METHOD=OIDC works. - Restore the guard against deleting or demoting the last local password admin while builtin/OIDC is active (unless DISABLE_LOCAL_AUTH is set). - Require a local admin before enabling Local auth, as OIDC already did. - Log a recovery hint at startup when builtin/OIDC runs without a local admin, and document recovery via AUTH_METHOD=none. - Drop the "will fall back to No Authentication" UI toasts and hints. Fixes https://github.com/calibrain/shelfmark/issues/1387 |
||
|
|
37a77e9562 |
feat(library): mark search results already in a Calibre library (#1377)
Per discussion #1372, where you said you were fine with this specific implementation: check whether metadata.db exists, read it if so, and show a check mark saying the book is already there. Searching for a book you already own gives no hint that you own it, so the easiest way to end up with a second copy is to not remember you have the first. This reads a Calibre `metadata.db`, read only, and marks matching results with an **In library** badge in the card, list and compact views and in the details dialog. Off by default. It sits in Settings, General beside the existing Library URL, with a test button that reports how many books it indexed. No HTTP call, no token, nothing written back. Matching runs most to least confident: a shared external id, then an ISBN compared in both ISBN-10 and ISBN-13 form, then fuzzy title tokens plus the author surname. The check fails open, so an unreadable database degrades the badge and never blocks a search, and entries are cached for ten minutes with an early refresh when the file changes, so a large library costs one read rather than one per search. `text_match.py` is new and shared by the index and the provider, so title, author and ISBN matching stays consistent in one place. ## On the provider interface `library_index` talks only to a `LibraryProvider` protocol and knows nothing about Calibre. That is deliberate but it is not speculative generality, it is what let me send you the Calibre half on its own: I run an Audiobookshelf provider on the same interface in my fork, which is where the audiobook side of the badge comes from. I have left that out because it is a new service integration rather than something already in the codebase, which is the line your non-goals draw. Happy to send it separately if you ever want it, and equally happy for the answer to be no. Adding a library is a module with the `LibraryProvider` shape plus one line in `all_providers()`. ## Verification - `tests/core/test_library_index.py`: id, ISBN and fuzzy matching, per-content-type provider selection, fail-open on provider errors, stale-cache reuse, TTL and fingerprint refresh, per-provider cache isolation, and the test-connection path including unsaved form values. - `tests/core/test_text_match.py`: ISBN variants and token matching. - `src/frontend/src/tests/libraryBadge.test.ts` and the added cases in `bookTransformers.test.ts`. - Python suite (3269) and frontend suite (206) green, plus ruff, ruff format, basedpyright, vulture, tsc, oxlint, oxfmt and the production build. |
||
|
|
fe99d4bb5b |
feat(search): add a configurable default content type (#1371)
Closes #1018. Worth correcting the issue first: the search tab is not hardcoded. `useContentTypePreferences` has persisted the user's choice to localStorage since #564, so a browser that has picked a tab already keeps it. What is missing is the other half the issue asks for, a default for a browser that has stored nothing, and a per-user override. `DEFAULT_CONTENT_TYPE` is a user overridable select next to `BOOK_LANGUAGE`, so it follows the same path: global value in Settings, per-user value in Search Preferences, resolved with `user_id` in `/api/config`. The frontend uses it only when this browser has no stored choice, which is captured before the existing effect writes one, so nothing changes for anyone who has already picked a tab. The resolution is a pure function in `utils/contentTypePreference.ts` rather than logic inside the hook, since vitest here has no jsdom and the existing tests cover resolvers like `resolveDefaultLanguageCodes` the same way. One small move in `App.tsx`: the `config` state was declared below the hook that now reads it, so it moved above it. ## Verification - `tests/core/test_config_api.py`: the payload carries `default_content_type` and reads it with the user's id. - `tests/core/test_admin_users_api.py`: the key appears in the per-user search preferences list. - `src/frontend/src/tests/contentTypePreference.test.ts`: a stored tab wins, combined mode survives, the server default applies when nothing is stored, and an unrecognised value falls back to ebook. - Python suite (3163) and frontend suite (201) green, plus ruff, ruff format, basedpyright, vulture, tsc, oxlint, oxfmt and the production build. |
||
|
|
ce1092db7f |
test(auth): stop proxy provisioning tests depending on run order (#1381)
I hit this while building the Debrid-Link client in #1380. The new tests there changed the test count, that reshuffled the xdist workers, and these two went red. They were related to this branch, but were tests related to another PR I did which didn't have proper tests.. `test_sets_session_from_header` and `test_reads_remote_user_wsgi_fallback` both assert that a proxy-authenticated user comes back with `is_admin is True`. Since #1356 that only holds for the bootstrap account, because `_proxy_default_is_admin` returns True only while `has_admin()` is False. Both tests assume they're provisioning the first account, and only one of them can be. Run with `-n 0` so nothing is sharded: | | | |---|---| | either test alone | passes | | both, file order | the second fails | | both, reversed | the second fails | | both, across 2 workers | both pass | Reversing the order moving which one breaks is what makes it an ordering problem rather than a real one. It stays green on CI because the suite runs with `-n auto` and `tests/conftest.py` calls `mkdtemp` at module level, which runs once per worker process. Each worker gets its own `CONFIG_DIR` and its own `users.db`, the two tests land on different workers, and each one is genuinely first in its own database. That passed but it's luck rather than design, and any change to the test count can put them back together. So this gives every test in the file an empty user table and stops the question of who ran first from mattering. While I was fixed that I added coverage for the rule itself, which I had failed to test for: - the bootstrap account is an admin and the next one isn't - `PROXY_AUTH_DEFAULT_ROLE=admin` promotes later accounts - a user already in the database keeps its stored role instead of picking up the default Tests only, no source changes. The full suite passes serially now, where it had those two failures before, and it's still green under `-n auto`. |
||
|
|
b690832659 |
fix(download): stream a completed book instead of buffering it in RAM - lowering memory needs significantly (#1378)
Credit where it is due: this defect was found and measured by **@DrNgo** in [DrNgo/shelfmark-fork@ae2185c](https://github.com/DrNgo/shelfmark-fork/commit/ae2185c8a83d537e55f1dc03f745ab844b5cdcdb). He recorded a 493 MB audiobook peaking near 963 MB and OOM-killing a 1 GiB container, with the proxy access log showing a single `GET /api/localdownload` returning 502 at the exact second of the kill. The analysis is his; I am sending it because it is still open here. Serving a completed book from the live queue calls `get_book_data`, which reads the whole file into bytes so the route can wrap it in a `BytesIO` for `send_file`. That is two copies of the book, with one resident for the length of the client transfer, to hand over a file that is already sitting on disk. `get_book_path` returns the path the task already holds and `send_file` streams it. The history fallback a few lines up in the same route has always worked this way, so this makes the two paths consistent rather than introducing anything new. `get_book_data` stays for callers that genuinely want the bytes, now documented as the expensive option. ## The fetch side has the same problem, and this PR does not fix it I should note: `download_url` in `shelfmark/download/http.py` still builds the inbound file in a `BytesIO` and returns it, so a large download is fully resident while it fetches. @DrNgo's commit fixes that too, with a `tempfile.SpooledTemporaryFile(max_size=...)` so small payloads stay in memory exactly as they do now and large ones spill to disk. I left it out deliberately. It changes the return type of `download_url` from `BytesIO` to a file object, which touches several callers, and it lands in a file that has been reworked around bypass handling, waiting rooms and resume since his branch point. That deserves its own PR rather than riding along with a two-function change. I may send it later; if @DrNgo sends it first, his should win, and if you would rather have both together say so and I will hold this one. ## Verification - `tests/download/test_orchestrator_retry.py`: `get_book_path` returns the path without opening the file (the test fails the run if it does), and reports a missing file rather than handing back a dead path. - The existing `/api/localdownload` tests still pass unchanged, including the history fallback and the ownership checks. - Full suite (3222), ruff, ruff format, basedpyright, vulture green. |
||
|
|
e1c3f057ab |
fix: bypass recordings, welib wrong-md5 links, footer build sha (#1364) (#1373)
Debug screen recordings never started. Every bypass logged "Capturearea
1540x1050 at position 0.0 outside the screen size 1440x1880".We ask
ffmpeg for the fingerprint screen size plus margin, the size wealso pass
SeleniumBase as xvfb_metrics. SeleniumBase builds thatdisplay with
use_xauth=True, the image ships no xauth binary, so itfalls back to a
fixed 1440x1880 Xvfb and the requested size neverexists. Drop
-video_size so x11grab records the whole screen, whateversize it turned
out to be.
welib could hand back a link for a different book. Welib
answers/md5/<md5> with a search for that md5; when it does not have the
file,the resolver took the first "Download" on the results page
(md5a2c1dc0c... resolved to auto_download/9c8cf85d...). On an
/md5/<md5>page a GET/Download link is now only taken when its href names
thatmd5; otherwise the source is reported as not having the file.
Alsoremoves _get_download_urls_from_welib and _is_source_enabled:
themd5-template branch in _get_urls_for_source always handles welib
first,so that resolver could never run.
The footer showed the build date instead of the commit. CI
stampsBUILD_VERSION as <yyyy-mm-dd>-<sha> (pr-<sha> for PR images) and
thefooter kept its first seven characters, so dev images read
"Shelfmarkmain (2026-09)". Take the trailing commit sha instead: "main
(
|
||
|
|
d978896142 | fix(auth): rename the API_KEY env var to SHELFMARK_API_KEY (#1374) | ||
|
|
3b280009ae |
feat(auth): static API_KEY (env) accepted as Bearer or X-Api-Key, cookie or key (#1366)
Supersedes #1353, per the discussion in #1352: one `API_KEY` environment variable; when set, a request carrying it is authenticated as the first admin, and cookie sessions keep working exactly as before (cookie **or** key). Nothing else changes. No table, no UI, no settings-tab switch, no per-user keys. ## What - `API_KEY` (env). Unset → the feature is off and none of the new code runs. - `Authorization: Bearer <key>` or `X-Api-Key: <key>` on any existing `/api/*` route authenticates that request as the first admin in `users.db` (`ORDER BY id`), or as a bare admin identity (`user_id="api"`, `is_admin=True`, no local user row) if the install has no admin yet. Per request only; nothing is persisted; the admin's role is read live, so deleting or demoting that user takes effect on the next request. - Both headers are checked and either may match. That is what makes the key usable behind a reverse proxy that injects its own `Authorization` header (oauth2-proxy, Authelia, forwardAuth): send the key in `X-Api-Key`. - A credential that is **not** the key is ignored and the request continues on the normal session path, so proxy-forwarded tokens are unaffected. Without a valid session such a request gets the usual `401 {"error": "Unauthorized"}`, identical to a request with no credential, so there is nothing to probe. ## How - `shelfmark/config/env.py`: `API_KEY = os.getenv("API_KEY", "").strip()`. - `shelfmark/core/api_key.py`: `extract_api_key_candidates()` (Bearer token if the scheme is Bearer, then `X-Api-Key`) and `matches_api_key()` using `hmac.compare_digest` on bytes. - `shelfmark/core/user_db.py`: `UserDB.get_first_admin()`. - `shelfmark/main.py`: `api_key_auth_middleware` (`before_request`, registered before `proxy_auth_middleware`, which early-returns for keyed requests). Only `/api/` paths; `/api/health` and `/api/auth/*` exempt; no-op when `API_KEY` is unset or the auth mode is `none`. On a match it mirrors the proxy-auth pattern: `session.clear()` then populate `user_id` / `is_admin` / `db_user_id` for this request, `permanent = False`, `modified = False`, `g.api_key_auth = True`. An `after_request` hook guarantees no `Set-Cookie` is written for a keyed request even if a handler dirties the session. - `docs/api-access.md` (new), the `API_KEY` entry in `docs/environment-variables.md`, and a README link. ## Security - Constant-time compare; the key is never logged or echoed. - Keyed requests never mint or refresh a session cookie and ignore any cookie sent with them (a non-admin cookie plus the key yields admin for that request; the browser's own session is left untouched and usable). - The mismatch path touches neither the session nor `g`, so a stray bearer on a browser request can neither log the user out nor change how their cookie is refreshed. - Store errors during the admin lookup fail closed (`500 {"error": "Authentication error"}`), never to anonymous. - Verified against Flask's `save_session` / `should_set_cookie` ordering, and under auth modes `none`, `builtin`, `proxy`. ## Tests `tests/core/test_api_key_env.py` (36): extraction and matching; first-admin lookup; middleware behaviour on a guarded route and an admin route, with and without a user_db, `X-Api-Key`, both-headers combinations, no `Set-Cookie` when a handler dirties the session, incoming non-admin cookie ignored, browser cookie still usable after a keyed request, security headers, store error → 500, mismatch → guard's 401 / cookie path / permanent cookie untouched, unset → off, exempt paths and path probes, `none` and `proxy` modes, deleted and demoted first admin, a keyed write passing the guard. Existing auth suites unchanged. All CI gates green on the fork: https://github.com/gavinmcfall/shelfmark/pull/2 (CI-only draft). Also exercised against a running instance: 47 scripted checks including 150 concurrent requests, proxy-mode switching through the key, an unset-key restart, and a log scan for the key. ## Naming `API_KEY` as discussed. If you'd rather namespace it (`SHELFMARK_API_KEY`) to avoid clashing with other tools' env vars in shared compose files, it is a one-line change; say the word. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1a5b37d9d3 |
fix(deluge): send seeding ratio limit under Deluge's own keys (#1367)
Deluge's per-torrent options are `stop_at_ratio` (bool) and `stop_ratio` (float), and `torrentmanager` checks `options['stop_at_ratio'] and get_ratio() >= options['stop_ratio']`. We were putting the indexer's float into `stop_at_ratio`, which only switched stopping on and left the daemon's global ratio (default 2.0) as the one actually enforced. A `ratio_limit` of 0 turned stopping off entirely. `stop_at_ratio_enabled` is not a Deluge option at all. Deluge has no per-torrent seeding time limit, `seed_time_limit` is a global core preference, so the value is logged as unapplied instead of sent as a key the daemon drops. qBittorrent and Transmission already honour both indexer limits, so this removes a silent difference between clients. ## Verification - `tests/prowlarr/test_deluge_client.py`: the ratio arrives as `stop_ratio` with `stop_at_ratio` set, and no key Deluge does not define is sent. Both fail on current main and pass here. - Full suite (3165), ruff, ruff format, basedpyright, vulture green. |
||
|
|
7c8e89c567 |
fix(googlebooks): page by the capped size, not the raw limit (#1370)
The Google Books search builds `maxResults` as `min(limit, 40)` because the API caps a page at 40 volumes, but advances `startIndex` by the full `limit`. The pages then stop tiling. With `limit=50`, page 1 covers items 0 to 39 and page 2 starts at 50, so items 40 to 49 are never returned and every later page drops another 10. This computes the page size once and uses it for both `maxResults` and `startIndex`. The shipped frontend asks for 40 and is unaffected. `/api/metadata/search` clamps `limit` to 100, so an API caller passing 41 to 100 was hitting it. One thing I left alone. The provider uses the base `search_paginated` heuristic, `has_more = len(books) >= options.limit`, which still reports `has_more: false` for a limit above 40 since Google can never return that many. That was already the behaviour before this change, and fixing it means either touching the shared heuristic or adding a provider override, so I kept this patch to the stride. Happy to follow up if you want it. ## Verification - `tests/metadata/test_googlebooks_parse.py`: pages 1 and 2 at `limit=50` must tile exactly, plus a guard that `limit=25` still strides by 25. The first fails on current main and passes here. - Full suite (3165), ruff, ruff format, basedpyright, vulture green. |
||
|
|
bc03ad062e |
fix(http): keep the host of a protocol-relative download link (#1368)
`get_absolute_url()` replaced both `netloc` and `scheme` whenever either one was missing. A protocol-relative href such as `//cdn.example.org/f.epub`, scraped from a page on `https://annas-archive.org/...`, parses with a netloc and an empty scheme, so it came back pointing at the page's own host. The download then 404s and the source is skipped. Each field now falls back to the base URL only when the parsed URL does not supply it. Plain relative paths resolve exactly as before, which the control test covers. This affects the Z-Library, welib and generic download link handling in `release_sources/direct_download/annas_archive.py`. ## Verification - New `tests/download/test_http_absolute_url.py`: a protocol-relative link keeps its own host, and a plain `/path` still resolves against the base. The first fails on current main and passes here. - Full suite (3165), ruff, ruff format, basedpyright, vulture green. |
||
|
|
ab3aa9a8b0 |
fix(requests): reject non-object items in the batch endpoint (#1369)
`POST /api/requests/batch` checks that `requests` is a non-empty list and then hands each element to the shared preparation helper, which calls `.get()` on it. A bare string, number or null in the list raises `AttributeError` and the caller gets a 500, while `POST /api/requests` answers 400 with a message for the same mistake. This validates the element type beside the existing list check. One bad item rejects the whole batch rather than being reported per item, which matches the endpoint's current contract: every other failure path already aborts the batch with a single error body, as `test_batch_create_requests_is_atomic` asserts. Responses for valid payloads are unchanged. ## Verification - `tests/core/test_request_routes_api.py`: a case per bad shape (int, str, null, list), plus a mixed valid and invalid batch that asserts nothing was created. All fail on current main with a 500 and pass here. - Full suite (3168), ruff, ruff format, basedpyright, vulture green. |
||
|
|
d4619be69a |
Feature: Show the AA search result stats (#1362)
When doing a direct search for a book, show the stats of the AA results. For example if we search for "The Great Gatsby", AA reports it has 240 hits and shows the first page of 50. Provide this stats info in the shelfmark webUI via ResultsSection.tsx and ReleaseModal.tsx. This table shows what should be displayed based on the total number of hits found: |Total |Display| |-------|-------| |1 |"Result 1 (1 Total)"| |6 |"Results 1-6 (6 Total)"| |144 |"Results 1-50 (144 Total)"| |500+ |"Results 1-50 (500+ Total)"| Currently, shelfmark also only shows us the first 50 search hits even if there were more available from AA. This could be added later if considered desirable. As usual, a picture is worth a 1000 words: <img width="1012" height="610" alt="direct-results-stat" src="https://github.com/user-attachments/assets/8c0c688f-444e-482a-a9a3-2dee5a5563b8" /> <img width="1013" height="741" alt="universal-results-info" src="https://github.com/user-attachments/assets/399bbb79-87f2-4432-a3d0-64937795f5f1" /> Coded with llama.cpp, opencode and 🤖 |
||
|
|
c1315a2b23 |
fix(download): check task ownership before serving queued files (#1357)
`/api/localdownload` resolves the file through the live queue and returns it before checking who owns the task; the owner check only runs on the download-history fallback, once the task has aged out of the queue. Task ids are source ids, so two users who searched the same book can end up with the same id. This applies the same rule on the queue path, reusing the `_task_owned_by_actor` helper the cancel/retry/priority routes already use, so both paths answer a non-owner with the same 404. Admin behaviour is unchanged. The 404 matches what this endpoint's history path already returns for a non-owner rather than the 403 `download_not_owned` the cancel/retry/priority routes use, happy to switch it if you prefer consistency with the siblings instead. ## Verification - `tests/core/test_activity_routes_api.py`: the owner still receives their queued file; a different user receives 404. The new case fails on current main and passes here; the existing history-fallback test is unchanged. - Full suite (3148), ruff, ruff format, basedpyright, vulture green. |
||
|
|
7934924678 |
fix(queue): don't stamp CANCELLED over a finished download (#1361)
`cancel_download` reads the task status under the queue lock, releases it, and only then writes CANCELLED through `update_status`. A download that finishes in that window has its COMPLETE overwritten. The queue and the UI show the task as cancelled while the file is already on disk, and the terminal hook fires for both statuses. The check and the write now happen in a single lock hold. Because the lock is non-reentrant and the terminal hook has to run after it is released (the stall canceller depends on that), the lock-held part of `update_status` moved into a small private helper that both paths share; `update_status` is a thin wrapper over it. A cancel arriving once the task is already terminal still returns `False`. ## Verification - `tests/core/test_queue.py::test_cancel_does_not_overwrite_a_download_that_finished_first`: a worker thread completes the download while the cancel is in flight, with the handover driven by events rather than sleeps. The task stays complete. Fails on main, passes here. - Full suite (3147), plus `tests/download/` and `tests/core/test_download_api_guardrails.py`, ruff, ruff format, basedpyright, vulture green. |
||
|
|
a6204a318e |
fix(oidc): reject backslash paths in the return_to sanitizer (#1359)
The OIDC `return_to` sanitizer rejects values starting with `//` and then relies on `urlsplit` to catch anything carrying a netloc. A value such as `/\host` has no netloc, so it is stored in the session and used as the post-login redirect target and browsers resolve the backslash as a path separator, which lands the user outside the app after a successful login. `_normalize_return_to` now also rejects values whose path contains a backslash. That matches the frontend sanitizer in `authRedirect.ts`, which parses with `URL` and already discards those forms, so the two ends agree again. The check covers the path only, so query and fragment backslashes still round-trip, and it also catches the script-root case where `/app/\host` strips to `/\host`. ## Verification - New cases in `tests/core/test_oidc_routes.py` cover the rejected forms, including under a script root, and confirm `/`, `/settings` and `/search?q=x#frag` are unaffected. They fail on current main and pass here. - Full suite (3155), ruff, ruff format, basedpyright, vulture green. |
||
|
|
545480c557 |
fix(download): default is_admin to False in the request policy guard (#1358)
`_resolve_policy_mode_for_current_user` reads `session.get("is_admin",
True)`, so a session carrying `user_id` but no `is_admin` key skips the
request policy entirely, while every other admin check in the codebase
defaults the key to `False`.
This uses the same default here. Every authenticated login path
(builtin, CWA, proxy, OIDC) writes `is_admin` into the session, and
`AUTH_METHOD=none` is already short-circuited one line earlier, so
sessions from those flows behave exactly as before.
## Verification
- `tests/core/test_request_routes_api.py::TestDownloadPolicyGuards`: a
session without `is_admin` now gets `policy_requires_request` and
nothing is queued. Fails on current main, passes here.
- Full suite (3147), ruff, ruff format, basedpyright, vulture green.
|
||
|
|
127dd82615 |
fix(users): apply user updates only after the payload validates (#1360)
`PUT /api/users/me` and `PUT /api/admin/users/<id>` write the new password hash, and then the profile fields, before the rest of the payload is checked. When the request is rejected further down as an invalid role, an admin-only setting, an invalid settings value, the route answers 400 with those writes already committed, so the caller sees an error while the password has in fact changed. Both routes now validate the whole payload before touching the database, and the password hash is folded into the same `update_user` call as the other fields so the field write is a single transaction. Error messages, status codes and the order they are reported in are unchanged. ## Verification - New tests in `tests/core/test_self_user_routes.py` and `tests/core/test_admin_users_api.py` assert that a rejected update leaves the password, profile fields and role as they were, plus a positive case that a valid payload still applies all three. They fail on current main and pass here. - Full suite (3150), ruff, ruff format, basedpyright, vulture green. |
||
|
|
acd59f7cbb |
feat(auth): provision proxy users as non-admin once an admin exists (#1356)
With `AUTH_METHOD=proxy` and no admin group configured, every user the proxy authenticates for the first time is provisioned as an admin (`is_admin = True` unless the user already exists in `users.db`). The intent to never lock an instance out makes sense, but the effect is that anyone the SSO gate lets through becomes an administrator. On an instance shared with family or a small community that is a footgun; I hit it when the first invited reader landed as an admin. This keeps the guarantee and removes the footgun: the first account is still provisioned as an admin while the instance has no admin at all, and later first-time users follow a new `PROXY_AUTH_DEFAULT_ROLE` setting (Security tab / env), default `user`. Known users keep their stored role; the `PROXY_AUTH_ADMIN_GROUP_NAME` path is unchanged and still takes precedence. I couldn't find a way with Cloudflare access to pass this along. Changes: `UserDB.has_admin()`, `_proxy_default_is_admin()` in the proxy middleware, the new `SelectField` beside the other proxy settings, the regenerated `docs/environment-variables.md` entry and a row in `docs/reverse-proxy.md`. Compatibility: the default moves from "everyone admin" to "first admin, then users". Accounts already in `users.db` are unaffected; new SSO users on an existing instance become regular users unless `PROXY_AUTH_DEFAULT_ROLE=admin` is set. If you would rather ship this purely opt-in I can flip the default to `admin`. ## Verification - `tests/core/test_auth_api.py::TestProxyProvisioningRole`: first user admin / second user not; `PROXY_AUTH_DEFAULT_ROLE=admin` restores the old behaviour; an admin from another auth source counts as "an admin exists"; a known user keeps their role whatever the default. - Full suite (3094), ruff, ruff format, basedpyright, vulture green. - Running on my own instance since 2026-09-19. |
||
|
|
44f4e13cce |
refactor: extract the per-source release search out of /api/releases (#1355)
The `/api/releases` route carries an inner `_search_source_releases` helper that builds the search plan for one source, logs the planned query type, runs the search and turns `SourceUnavailableError`/operational errors into an error message instead of raising. Anything outside the route that wants to search one source with exactly those semantics has to go through Flask today. This moves that helper into `shelfmark/core/release_search.py` as `search_source_releases()` and has the route delegate to it. Behaviour is unchanged: same plan construction (including the caller's `user_id`, so per-user default languages still apply), same logging, same error-to-message handling. It is the refactor half of #1047 by @InfiniteAvenger, split out on its own as you asked for other PRs (#1318). Their authorship is preserved on the commit; I rebased it onto current `main` and added tests. ## Verification - `tests/core/test_release_search.py`: unknown source → `"Unknown source: …"`, `SourceUnavailableError` and operational errors → `"<source>: <error>"`, success path forwards `expand_search` / `content_type` and returns the source instance, the plan receives languages / manual query / indexers / `user_id`. (These tests are type-annotated; happy to strip the annotations if you prefer the suite's bare style.) - Full suite, ruff, ruff format, basedpyright, vulture green; the existing `/api/releases` route tests are unchanged and pass. Co-authored-by: InfiniteAvenger <calebewest02@gmail.com> |
||
|
|
b7002a6eca |
feat: Add TorBox client support and settings integration (#1342)
Add **TorBox** as a torrent download client for Prowlarr releases. Users can select `TorBox` in the download client settings, configure it with the new `TORBOX_API_KEY` environment variable, and verify their credentials with the connection test button. The integration supports both magnet links and `.torrent` files. It tracks the torrent lifecycle through TorBox, downloads supported book and audiobook files from the TorBox CDN, preserves safe nested file paths, and cleans up remote and local download state. Important: Shared HTTP download logs omit full download URLs and URL-bearing exception text to avoid exposing credentials, following best practices. This applies to all clients that use the shared `download_url()` path; URLs remain available to the HTTP operations themselves. --- There is already related work in progress in #1173, which includes both torrent and direct-download support for TorBox. This PR is not intended to replace or compete with that contribution. It offers the tested torrent client functionality as a smaller, focused change that can make TorBox available to the community sooner. The direct-download integration proposed in #1173 remains valuable and could be reviewed or introduced separately. Automated tests cover configuration, connection validation, magnet and torrent-file submission, API errors, status and progress handling, file retrieval, path traversal protection, cancellation, cleanup, and sensitive URL redaction. I also validated the complete flow locally with several magnet links and `.torrent` downloads. TorBox processed the torrents and Shelfmark downloaded the resulting files as expected. AI was used to help with the implementation, with human validation. This PR and long description? Took me some good minutes at night after work, but gives me joy to open this PR to share with the community this improvement. |
||
|
|
e5dd34ae0e |
fix: unbreak main and follow up on the Blackhole handoff review (#1346)
DownloadHistoryService.record_download and updated the single production caller, but not the eleven in the test suite, leaving main red with 32 failures. Pass None, which is what the pre-#1336 behaviour recorded. For the Blackhole handoff (#1345): add_download publishes the torrent before the cancel check runs, and BlackholeClient.remove() is a no-op, so the watcher picks the file up regardless. Reporting a bare "Cancelled" hid that from the user. Name the completed handoff in the cancellation message instead, drop the _handle_cancelled_download call whose usenet branch cannot apply to a handoff-only client, and record why the orchestrator no longer verifies HandoffResult.path. Finally, make tests/direct_download a package: test_libgen_extract.py imports tests.libgen.sample_html across test directories, so without an __init__.py pytest named its modules by bare basename and a same-named module elsewhere would collide. |
||
|
|
c53545d9fe |
Add configurable word separator for naming templates (#1333)
Closes #1230 ## What Adds a "Word Separator" setting (Space / Dot / Underscore / Hyphen / Custom) that replaces internal whitespace in each naming-template placeholder's rendered value — e.g. `{Author}` renders "Arthur.Conan.Doyle" instead of "Arthur Conan Doyle" when Dot is selected. This follows option 2 from the issue rather than inventing new dotted-keyword template syntax (`{Author.}`), since it's a smaller surface: one setting applies uniformly across all four templates (books/audiobooks × rename/organize) instead of needing a parallel token for every existing one. ## How it works - Literal characters typed into the template itself (e.g. the `.` in `{Author}.-.{Title}`) are never touched — only whitespace *inside* a placeholder's resolved value is affected. - Default is "Space", which is a no-op: existing templates produce byte-identical output after this change (verified via the existing test suite, unmodified, still passing). ## Where - `shelfmark/core/naming.py` — `word_separator` param on `parse_naming_template` / `build_library_path`. - `shelfmark/download/postprocess/policy.py` — `get_word_separator()`, mirroring the existing `get_file_organization()` accessor. - `shelfmark/download/postprocess/transfer.py` — wires the resolved separator through the four existing template-rendering call sites. - `shelfmark/config/settings.py` — new `Word Separator` / `Custom Word Separator` fields next to the existing naming-template fields. - `src/frontend/.../namingTemplatePreview.ts` + `NamingTemplateField.tsx` — the settings UI has its own TS mirror of the Python renderer for the live preview; updated it in lockstep so the preview doesn't lie about what the separator will actually do. - Tests added on both sides (pytest + vitest). ## Testing - `uv run pytest tests/core/test_naming.py tests/core/test_destination_file_organization.py` — all pass, including new cases. - `uv run pytest` (full suite) — same pre-existing failures as on `main` before this change (browser/network-dependent bypass & e2e tests unrelated to this diff), everything else green. - `uv run ruff check` / `ruff format --check` / `basedpyright` — clean. - `npm run lint` / `format:check` / `typecheck` / `test:unit` (196 tests) — clean. |
||
|
|
8f608f2e64 |
fix(download): complete consumed Blackhole handoffs (#1345)
A Blackhole watcher can consume the torrent before Shelfmark checks it, leaving the task in error even though the handoff succeeded. Complete the handoff when `add_download` successfully publishes the file, and stop requiring a `HandoffResult` path to remain present. Follow-up to #1312. ## Verification - A watcher that immediately reads and removes the torrent receives the exact bytes. The task changes from ERROR before this fix to COMPLETE afterward, without running book postprocessing. - The consumed-file regression fails on current main and passes here. Resident files, write failures, cancellation, magnet rejection and normal downloads remain covered: 81 focused tests pass. - Ruff lint and formatting pass for the changed files. |
||
|
|
2bb84a17a2 |
Extract archives when zip/rar are enabled as supported formats (#1343)
The default audiobook formats include `zip` and `rar`. `scan_directory_tree` checks the supported-format list before checking for archives, so a downloaded archive lands in `book_files` and is imported as-is. The extraction branch in `collect_directory_files` is never reached. This keeps archives out of `book_files`, so they always take the archive path: extracted when extraction is allowed, imported as-is when it isn't (unchanged). Tests added in `tests/download/test_postprocess_scan_archives.py`; three of the four fail without the change. |
||
|
|
c6b70a6844 |
fix(sources): send a Referer when fetching libgen ads.php pages (#1340)
## What libgen.li's `ads.php?md5=` now returns an **empty `200`** to any request without a `Referer` — an anti-hotlinking check the mirrors added recently. Both libgen paths fetch it without one, so the page comes back blank and the download silently fails while **search keeps working** (which is exactly why it looks like rate-limiting or mirror drift rather than a bug). Same one-line cause, two call sites: the Libgen search source (`libgen/scraper.py:fetch_page`) and the AA-md5 → libgen fallback (`direct_download/annas_archive.py:_extract_libgen_download_url`). Fix: send a same-origin `Referer: <scheme>://<host>/` on the `ads.php` fetch in both. ## Worth a look in review - **The referer goes on the *resolution* fetch, not the download.** `download_url(..., referer=...)` was already correct — the blank page happens one step earlier, at the `ads.php` GET. - Reproduced against live mirrors: `ads.php` returns `Content-Length: 0` bare, the full page with a `Referer`, and resolvable files download valid bytes again. Regression tests in `tests/libgen/` and `tests/direct_download/` assert the header on both paths. Lint/format/typecheck clean. Follow-up to #1326. |
||
|
|
35b89b0d78 |
fix(sources): restore Direct Download search errors and language matches (#1339)
Fixes two regressions from the provider-driven refactor (#1337). First, the composite search caught RuntimeError, TypeError, ValueError and request errors from each provider and returned an empty list, so a failed search looked like one with no hits. It now raises the first provider failure when no provider returned releases. Second, the shared parser re-matched every row's language locally, dropping rows Anna's Archive had already matched with &lang= (free-text cells like 'English, French' or 'unknown'). parse_search_items gains a filter_languages option, which AA turns off, so AA's own language-from-path filter is again the only local one. |
||
|
|
a5cd9f0bfb |
refactor: make direct download provider-driven (#1337)
This is the refactor for the download handler |
||
|
|
af21d1da1f |
feat(sources): add Libgen as a direct catalogue search source (#1326)
## What Adds **Libgen as a search source**. Today Libgen is only a download mirror (reached by an Anna's Archive md5), so anything in Libgen but not in AA's search index is invisible — and that's where most of the CBZ/CBR comics and manga live. A Libgen search for *One Piece*, for instance, turns up ~99 volumes that AA search never shows. It's a self-contained `release_sources/libgen/` package (source + handler + settings) plus one line to register it. **No changes to `direct_download.py`** — it reuses the existing `ads.php → get.php` resolution and the mirrors already configured in `LIBGEN_MIRROR_URLS`. Plain HTTP, no bypasser needed (libgen.li isn't behind DDoS-Guard). Opt-in via a settings toggle. ## Worth a look in review - **`source_id` is `libgen:<md5>`, not the bare md5.** The download queue keys on `task_id` (= `source_id`), and `direct_download` already uses the bare md5. Since AA indexes a lot of Libgen, the same md5 shows up from both sources — a bare id would collide in the queue. The handler strips the prefix before downloading. - **Reachable like the other non-default sources** (Prowlarr, AudiobookBay, …): it appears in the per-book release search, not the free-text box (that stays wired to `direct_download`). Tests in `tests/libgen/` cover parsing (both row layouts), the source, the handler, and `get_record`. Lint/format/typecheck clean. |
||
|
|
1b17fe179a |
fix(irc): rank a surname-only result as partial, not wrong (#1332) (#1334)
"David Petrie" as "D. Petrie", then ranked the answer by the full name to recover the precision the surname gave up. The two halves disagreed. author_affinity needs two agreeing tokens before it calls a name the same person, so "Petrie" - the name on the filenames a surname search exists to reach - matched one and came back AUTHOR_MISMATCH. It therefore sorted below "Unknown" and level with "Gordon Petrie", a different author who merely shares the surname. The widened query pulled those rows in and the ranker buried them. Falling short of agreement is now separated from disagreeing with it. A name whose every token fits the one asked for is an abbreviation of it and ranks AUTHOR_PARTIAL, between agreement and "no author reported"; a name carrying a token that fits nothing still ranks AUTHOR_MISMATCH. Nothing that agreed before changes tier - "Homer"/"Homer Simpson" is still a match, since the extra token must not demote a mononym that already met its one-token requirement - so Prowlarr's #1293 ordering is unchanged except that a tracker listing a bare surname stops being read as the wrong author. Measured on the issue's own case, wanted "David Petrie": before: D Petrie, Unknown, Petrie, Gordon Petrie after: D Petrie, Petrie, Unknown, Gordon Petrie Second fix, same release: a book with no title posted the surname on its own. _build_query fell back to book.search_title or book.title, which is empty on exactly the path where the plan has no title variants, so the line reaching the channel was "@search Petrie" - not a search for anything, and the kind of bare over-broad post is_available refuses unaddressed queries to avoid. It now returns "" and the existing "No search query could be built" guard takes it. Tested with make python-lint, python-format, python-dead-code, python-typecheck and python-test. |
||
|
|
35037b35fd |
fix(irc): search by surname, and rank the answer by author (#1331) (#1332)
Fixes #1331. A search bot ANDs every term against a filename, so the given name is the term that empties the result set. Measured against irchighway's #ebooks: "Revelations David Petrie" is answered "no results", "Revelations Petrie" returns 9 matches, 6 of which parse, all filed as "D Petrie". The query now carries the title and the surname, read off the search variant so the ISBN fallback and a manual query - which set author="" on purpose - keep their current shape. Title-only, the shape #1295 settled on for Prowlarr, does not transfer: the bot caps an answer at 1000 matches, and a bare "Revelations" hits that cap with 923 parsed rows across 500 authors, so the cap itself can drop the wanted book. The surname is the token the two spellings share and it keeps the answer small. The full author then orders what comes back, reusing author_affinity from #1295, since a surname also matches a different author who shares it. It sits under server availability the way indexer priority does in #1295: a download addresses one named bot and waits 120s for it, so a match from a bot that has left the channel must not outrank a mismatch that can answer. Ranking runs on the way out rather than before the cache, because one query identity is shared by every book that produced that query. Two things found while testing: - The parser writes the literal "Unknown" when a filename has no " - " separator (parser.py:168). Ranked literally that sorts as a wrong author, so author_affinity's middle tier was unreachable here; it is now read as absent. 5 of those 923 rows are affected. - author_affinity moves to shelfmark/core/author_match.py, unchanged, so IRC does not import from the Prowlarr package. Prowlarr behaviour is untouched and its tests pass as they are. The three IRC assertions in the #1252 regression file move to the surname form. The invariant they pin - one contributor's name reaches the query, never the whole credit list - is unchanged. Tested with make python-lint, python-format, python-dead-code, python-typecheck and python-test, and end to end against irchighway with the patched source: it posts "Revelations Petrie" and returns 6 releases. |
||
|
|
99e0cfde3d |
fix: prevent Anna's Archive download countdown resets by preserving browser sessions (#1325)
## Observed bug Anna's Archive slow-download pages can return a JavaScript countdown before a download link is available. The internal browser returns that waiting-room HTML and closes its incognito session. The downloader then sleeps and fetches the URL again, which can create a new queue session and **restart the countdown instead of reaching the download link**. ## Fix - **Preserve the queue session:** keep the original browser tab open while the site's own countdown and automatic navigation finish. HTTP 200 and cached-cookie waiting-room responses enter the same flow. - **Return a consistent page:** capture HTML and readiness together in one browser evaluation so navigation cannot pair a new page's status with stale protection-page HTML. Share cache validation and page-readiness rules across their callers. - **Keep waiting cancellable and bounded:** poll cancellation while a slow browser read remains pending, rather than repeatedly cancelling and reissuing it. Apply a **300-second waiting-room limit** within the existing browser watchdog. - **Report queue timeouts accurately:** preserve the timeout across the helper-process boundary and stop the solve without restarting the browser or rotating mirrors. Waiting-room detection is limited to Anna's Archive `/slow_download/` pages containing an actual `.js-partner-countdown` element. The site controls the countdown and refresh. External-bypasser behavior and file-transfer timeouts are unchanged; the PR adds no deployment configuration or dependencies. ## Validation Validated at `c81e0a2`: | Check | Result | | --- | --- | | Full Linux unit suite | **2,978 passed** on Python 3.14 in a non-root environment with entrypoint test stubs enabled | | Focused regression coverage | **30 passed**, covering countdown completion, zero timers, navigation, both cookie-cache paths, cancellation, stuck queues, slow reads, and timeout propagation | | Navigation-race regression | Fails against the previous PR implementation and passes with the fix | | Python static checks | Ruff lint/format, BasedPyright for backend and tests, and Vulture passed | | Real Chromium fixture | Queue cookie persisted through 1.5-second DOM reads and one automatic refresh; the CDP connection survived multiple polling intervals | | Live source check | Observed **19 → 14 → 9 → 4 → download link** while retaining the browser session; the patched browser path also completed the waiting room | Full unit-suite command: ```sh pytest tests/ -n 2 --tb=short -m "not integration and not e2e" ``` The live check validates waiting-room completion and link resolution. Remote file-host availability remains a separate concern. The unit suite emitted two existing Authlib deprecation warnings. |
||
|
|
c45d342931 |
fix(prowlarr): skip indexers in Prowlarr failure back-off (#1324)
## What
Read `/api/v1/indexerstatus` once per search and skip indexers whose
`disabledTill` is still ahead. Skipped is neither attempted nor failed.
One client method, one counter on `_IndexerSearchOutcome`, ten tests.
## Why
Prowlarr's own search leaves out an indexer it has disabled after
repeated failures. Shelfmark queries each indexer through its Torznab
endpoint, which answers 429 instead:
```
Prowlarr Torznab error response: <error code="429" description="Indexer is disabled till 09/09/2026 15:00:34 due to recent failures." />
Prowlarr: 1 of 5 indexer searches failed (indexer 2 search failed: 429 Client Error: Too Many Requests ...)
Release search failed for source prowlarr: 1 of 5 indexer searches failed (...)
```
That counted as a failed search, so with one indexer in back-off and the
other four answering empty, `/api/releases?source=prowlarr` returned 503
for every book for the length of the back-off (one hour here).
`/api/v1/indexerstatus` on Prowlarr 2.5.2:
```json
[{"indexerId": 2, "disabledTill": "2026-09-09T15:00:34Z", "mostRecentFailure": "2026-09-09T14:00:34Z", "initialFailure": "2026-09-09T14:00:34Z"}]
```
## Behaviour
| indexers | before | after |
|---|---|---|
| 1 in back-off, 4 answer empty | 503 "1 of 5 indexer searches failed" |
"No releases found" |
| 1 in back-off, 1 answers with releases | releases | releases, one
Torznab call fewer |
| 1 in back-off, 1 times out, 3 answer empty | "1 of 5 failed" | "1 of 4
failed" |
| every indexer in back-off | 503 "5 of 5 failed" | "every indexer is
disabled by Prowlarr after recent failures (until ...)" |
| status endpoint unreachable | n/a | as before, nothing skipped |
Auto-expand no longer retries a pass in which nothing was asked.
## Tests
`uv run pytest tests/prowlarr`: 563 passed, 42 skipped. `ruff check` and
`ruff format` clean.
|
||
|
|
1e3fd48b8b |
fix: share rotating log file handlers (#1316)
This patch shares (for each log file) the `RotatingFileHandler` for logging across all modules, reducing the number of open file descriptors from ~78 to 1. I had originally assumed this issue was a resource leak, but it seems to just be a large fixed number of file descriptors. So this change mostly just (1) shrinks the number of open file descriptors to a reasonable level and (2) prevents two modules in the same process competing to write to a log file. |