Files
shelfmark/shelfmark/release_sources/prowlarr/handler.py
T
Lance Marks f441b85da2 feat(packs): inspect multi-book releases and file each book separately (#1270)
## Multi-book packs: inspect a release before download and file each
book separately

Closes #576

### Problem

One queued release is always treated as one book. When a torrent is
actually a whole series
(`Series/Book 1 - Title/…`, or a flat folder of `Series 1.0 - Title.m4b`
files), post-processing
walks the whole tree, flattens every file into one list and renames them
`Title - 01…10` under the
searched book's `{Author}/{Title}`. Audiobookshelf then sees a single
10-file "book" and the user
has to re-file everything by hand.

### What this does

Most releases expose their file list *before* anything is downloaded, so
the split is decided up
front and approved by the user, then the download is fire-and-forget:

1. **Inspect** – clicking a release's download button now calls `POST
/api/releases/inspect`
first. A new optional `DownloadHandler.list_files(release_data)` hook
returns the release's
   files without downloading:
- **AudiobookBay** reads the torrent file table off the detail page it
already fetches (the
page is now cached for 120 s, so inspect + download cost ABB one
request).
- **Prowlarr** parses `info.files` from the `.torrent` it already
fetches (the existing 120 s
torrent-fetch cache is reused). Magnet-only and usenet releases report
"can't inspect".
   - Other sources default to `None`.
2. **Review** – if the plan contains more than one book, the Find
Releases modal swaps the list
for a review panel: one row per book with editable title / series
position / year, expandable
file lists, non-book sidecars (`.txt`, covers) shown as ignored, a
"Treat as a single book"
switch, and **Download N books**. Single-book releases queue
immediately, exactly as before.
3. **File** – the approved plan travels with the task
(`DownloadTask.book_plan`, retry-safe) and
post-processing files each book through the existing transfer code, one
book at a time
(`dataclasses.replace(task, title=…, series_position=…, year=…)`), so
organize/rename
templates, part numbering (now scoped per book), hardlinks, torrent
copy-preserve and usenet
   handling are unchanged. Status reads `Complete (N books, M files)`.
4. **Fallback** – when a release can't be inspected the user gets a
toast, and a small
"Multi-book pack" toggle in the modal header forces a heuristic split
(subfolder = book, or
   one book per file when the file names carry series positions).

Planning lives in `shelfmark/download/postprocess/packs.py` and is
shared by the inspect endpoint
and post-processing, so what the user approved is what gets filed. The
name parser strips
`Book 3 -`, `03 -`, `1.0 -`, `3.`, `[03]`, `#3`, a leading series name,
labels like
"An Expanse Novella -", repeated titles (`Gods of Risk 2.5 - Gods of
Risk`) and a trailing
`(Year)`; author and series name come from the book that was searched,
and the searched book's
own series position is never applied to its siblings.

### Files

- `shelfmark/download/postprocess/packs.py` (new) –
`PackFile/PackBook/PackPlan`, `plan_pack`,
`parse_pack_book_name`, `group_files_into_books`, `match_plan_to_files`
- `shelfmark/core/release_inspect_routes.py` (new) – `POST
/api/releases/inspect`
- `shelfmark/release_sources/__init__.py` – `DownloadHandler.list_files`
hook
- `shelfmark/release_sources/audiobookbay/{scraper,handler}.py` –
detail-page cache,
  `extract_file_list`, `list_files`
- `shelfmark/release_sources/prowlarr/handler.py`,
`download/clients/torrent_utils.py` –
  `extract_file_list_from_torrent`, `list_files`
- `shelfmark/core/models.py`, `download/orchestrator.py` – `multi_book`
/ `book_plan` fields,
  queue + retry serialization
- `shelfmark/download/postprocess/transfer.py`, `pipeline.py`,
`outputs/folder.py` – per-book
  transfer branch and status message
- `src/frontend`: `components/PackReviewPanel.tsx` (new),
`ReleaseModal.tsx`, `App.tsx`,
`services/api.ts`, `types/index.ts`, `utils/releasePayload.ts` (payload
builder moved out of
  `App.tsx`), `utils/packReview.ts`
- `docs/dev/release-sources-plugin-guide.md` – documents the
`list_files` hook

### Out of scope (follow-ups)

- Listing files from an NZB (Shelfmark already fetches the bytes; `<file
subject>` names are noisy)
- Inspecting magnet links via qBittorrent's files API after a paused add
- BookLore / email outputs (they ignore `book_plan`; noted in code)
- The combined ebook + audiobook flow

### Testing

**Automated** (`make checks`, `make python-test`, `make frontend-test`
all green; the only
failures on my machine are the pre-existing
`tests/config/test_entrypoint_permissions.py` cases,
which need bash ≥ 4 and fail identically on `main` under macOS bash
3.2):

- `tests/download/test_packs.py` – name parsing (markers, series name,
novella labels, repeated
titles, bare numeric titles like `1984`), nested / flat / mixed /
deeper-nested packs, single
wrapping folder not treated as a pack, plan-to-disk matching with
basename fallback
- `tests/core/test_processing_packs.py` – full `post_process_download`
runs on a real temp
filesystem: approved plan files each book under its own
`{Author}/{Title}`, heuristic split
of a nested pack, searched book's series position does not leak,
multi-file book inside a pack
keeps `- 01/- 02` per book, hardlinked torrent pack leaves the seeding
tree intact, no pack
fields ⇒ behaviour unchanged, single group degrades to the searched
title, status message
- `tests/core/test_release_inspect_routes.py` – plan response,
not-inspectable, handler errors
  never 500, unknown source / missing `source_id` ⇒ 400, login required
- `tests/audiobookbay/test_file_list.py` – file-table scraping from real
ABB markup (multi-file
and single-file pages), handler host validation, one page fetch shared
by magnet + file list
- `tests/prowlarr/test_torrent_file_list.py` – multi-file / single-file
`.torrent` parsing,
  handler behaviour for torrent URL vs magnet vs usenet vs cache miss
- `tests/download/test_orchestrator_pack_fields.py` – queue-time parsing
and retry round-trip
- Frontend: `releasePayload.test.ts`, `packReview.test.ts` (vitest)

**Manual, on a real deployment** (arm64 image built from this branch,
run as a side container
next to production with the same qBittorrent / Audiobookshelf setup,
`FILE_ORGANIZATION_AUDIOBOOK=organize`,
hardlinks on):

- AudiobookBay "The Expanse Complete 2.0" (7.87 GB, 36 files): clicking
download opened the review
panel in ~1 s showing **18 books · 18 files · 18 files ignored** (the
`.txt` sidecars), with
series positions 0.1–9.5 and years parsed from the file names; novella
labels stripped
("The Churn", "The Butcher of Anderson Station"). Editing a title in the
panel works.
Confirming queued one task; the magnet resolved from the cached page in
~30 ms; after the
download the task reported `Complete (18 books, 18 files)`, 18 hardlinks
landed as
`audiobooks/James S. A. Corey/<Title>/<Title>.m4b`, the torrent kept
seeding, and
Audiobookshelf scanned each folder as its own book (title, author,
embedded chapters).
- A second pack ("Expanse [01 - 9.5]", `Title N - Title` naming) was
inspected to verify the
  repeated-title rule and the Back button, without downloading.
- Single-book releases still queue immediately with no extra UI.
2026-08-27 00:40:08 -04:00

402 lines
16 KiB
Python

"""Prowlarr download handler - resolves releases and delegates lifecycle to shared clients."""
from typing import TYPE_CHECKING, Any
from urllib.parse import urlparse
import requests
from shelfmark.core.config import config
from shelfmark.core.logger import setup_logger
from shelfmark.core.request_helpers import normalize_optional_text
from shelfmark.core.search_plan import build_release_search_plan
from shelfmark.core.utils import normalize_http_url
from shelfmark.download.clients import (
DownloadClient,
get_client,
list_configured_clients,
)
from shelfmark.download.clients.base_handler import (
COMPLETED_PATH_MAX_ATTEMPTS as _DEFAULT_COMPLETED_PATH_MAX_ATTEMPTS,
)
from shelfmark.download.clients.base_handler import (
COMPLETED_PATH_RETRY_INTERVAL as _DEFAULT_COMPLETED_PATH_RETRY_INTERVAL,
)
from shelfmark.download.clients.base_handler import (
POLL_INTERVAL as _DEFAULT_POLL_INTERVAL,
)
from shelfmark.download.clients.base_handler import (
DownloadRequest,
ExternalClientHandler,
)
from shelfmark.download.clients.torrent_utils import (
extract_file_list_from_torrent,
extract_torrent_info,
)
from shelfmark.metadata_providers import BookMetadata
from shelfmark.release_sources import register_handler
from shelfmark.release_sources.prowlarr.api import IndexerSeedSettings, ProwlarrClient
from shelfmark.release_sources.prowlarr.cache import cache_release, get_release, remove_release
from shelfmark.release_sources.prowlarr.source import ProwlarrSource
from shelfmark.release_sources.prowlarr.utils import (
build_source_id,
coerce_int_like,
get_preferred_download_url,
get_protocol,
sanitize_download_url,
)
if TYPE_CHECKING:
from collections.abc import Callable
from shelfmark.core.models import DownloadTask
from shelfmark.download.postprocess.packs import PackFile
logger = setup_logger(__name__)
# Errors that ProwlarrClient can raise when fetching indexer settings.
_SEED_SETTINGS_FALLBACK_ERRORS = (
requests.exceptions.RequestException,
OSError,
RuntimeError,
TypeError,
ValueError,
)
__all__ = [
"ProwlarrHandler",
"POLL_INTERVAL",
"COMPLETED_PATH_RETRY_INTERVAL",
"COMPLETED_PATH_MAX_ATTEMPTS",
"config",
]
# Backwards-compat constants for tests patching this module.
POLL_INTERVAL = _DEFAULT_POLL_INTERVAL
COMPLETED_PATH_RETRY_INTERVAL = _DEFAULT_COMPLETED_PATH_RETRY_INTERVAL
COMPLETED_PATH_MAX_ATTEMPTS = _DEFAULT_COMPLETED_PATH_MAX_ATTEMPTS
EXPIRED_LINK_REFRESH_ERROR = (
"The indexer download link expired and the release could not be refreshed. "
"Search again for a fresh result."
)
HASH_DETECTION_ERROR = "Could not determine torrent hash from URL"
def _coerce_positive_minutes(raw_minutes: object) -> int | None:
minutes = coerce_int_like(raw_minutes)
if minutes is None:
return None
return minutes if minutes > 0 else None
@register_handler("prowlarr")
class ProwlarrHandler(ExternalClientHandler):
"""Handler for Prowlarr downloads via configured torrent or usenet client."""
@staticmethod
def _build_prowlarr_client() -> ProwlarrClient | None:
"""Build a ProwlarrClient from config, or None if not configured."""
raw_url = config.get("PROWLARR_URL", "")
raw_api_key = config.get("PROWLARR_API_KEY", "")
url = normalize_optional_text(raw_url) if isinstance(raw_url, str) else None
api_key = normalize_optional_text(raw_api_key) if isinstance(raw_api_key, str) else None
if not url or not api_key:
return None
normalized_url = normalize_http_url(url)
if not normalized_url:
return None
return ProwlarrClient(normalized_url, api_key)
def _fetch_seed_settings_fallback(self, raw_indexer_id: object) -> IndexerSeedSettings | None:
"""Fetch share limits for one indexer directly from Prowlarr.
Used when the cached release is missing its search-time seed-limit
enrichment so that transient failures during search cannot cause a
torrent to be added without its configured share limits.
"""
indexer_id = coerce_int_like(raw_indexer_id)
if indexer_id is None:
return None
client = self._build_prowlarr_client()
if client is None:
return None
try:
settings = client.get_indexer_seed_settings(restrict_to=[indexer_id])
except _SEED_SETTINGS_FALLBACK_ERRORS:
logger.warning(
"Grab-time seed settings fallback failed for indexerId=%s",
indexer_id,
exc_info=True,
)
return None
return settings.get(indexer_id)
def list_files(self, release_data: dict[str, Any]) -> list[PackFile] | None:
"""List a cached torrent release's files from its .torrent, without downloading.
Magnet-only and usenet releases cannot be listed ahead of time.
"""
source_id = str(release_data.get("source_id") or "")
prowlarr_result = get_release(source_id) if source_id else None
if not prowlarr_result or get_protocol(prowlarr_result) != "torrent":
return None
download_url = sanitize_download_url(str(prowlarr_result.get("downloadUrl") or "").strip())
if not download_url or download_url.startswith("magnet:"):
return None
expected_hash = str(prowlarr_result.get("infoHash") or "").strip() or None
info = extract_torrent_info(download_url, expected_hash=expected_hash)
if not info.torrent_data:
return None
return extract_file_list_from_torrent(info.torrent_data)
def _get_client(self, protocol: str) -> DownloadClient | None:
"""Compatibility shim so module-level patching still works in tests."""
return get_client(protocol)
def _list_configured_clients(self) -> list[str]:
"""Compatibility shim so module-level patching still works in tests."""
return list_configured_clients()
def _poll_interval(self) -> float:
return POLL_INTERVAL
def _completed_path_retry_interval(self) -> float:
return COMPLETED_PATH_RETRY_INTERVAL
def _completed_path_max_attempts(self) -> int:
return COMPLETED_PATH_MAX_ATTEMPTS
def build_retry_resolution_fields(self, release_data: dict[str, Any]) -> dict[str, Any]:
source_id = normalize_optional_text(release_data.get("source_id"))
extra = release_data.get("extra")
if not isinstance(extra, dict):
extra = {}
retry_source_context: dict[str, Any] = {}
indexer_id = release_data.get("indexer_id") or extra.get("indexer_id")
if indexer_id is not None:
retry_source_context["indexer_id"] = indexer_id
indexer = normalize_optional_text(release_data.get("indexer") or extra.get("indexer"))
if indexer is not None and indexer.lower() != "unknown":
retry_source_context["indexer"] = indexer
info_url = normalize_optional_text(release_data.get("info_url") or extra.get("info_url"))
if info_url is not None:
retry_source_context["info_url"] = info_url
if source_id is not None:
retry_source_context["source_id"] = source_id
return {
"retry_download_url": None,
"retry_download_protocol": None,
"retry_source_context": retry_source_context,
}
@classmethod
def _restore_download_request_from_task(cls, task: DownloadTask) -> DownloadRequest | None:
"""Rebuild a DownloadRequest when the in-memory Prowlarr cache is gone."""
retry_download_url = normalize_optional_text(getattr(task, "retry_download_url", None))
retry_download_protocol = normalize_optional_text(
getattr(task, "retry_download_protocol", None)
)
if retry_download_url is None or retry_download_protocol is None:
return None
protocol = retry_download_protocol.lower()
if protocol not in {"torrent", "usenet"}:
return None
ratio_limit = getattr(task, "retry_ratio_limit", None)
if not isinstance(ratio_limit, (int, float)) or isinstance(ratio_limit, bool):
ratio_limit = None
seeding_time_limit = getattr(task, "retry_seeding_time_limit_minutes", None)
if not isinstance(seeding_time_limit, int) or isinstance(seeding_time_limit, bool):
seeding_time_limit = None
return DownloadRequest(
url=retry_download_url,
protocol=protocol,
release_name=(
normalize_optional_text(getattr(task, "retry_release_name", None))
or task.title
or "Unknown"
),
expected_hash=normalize_optional_text(getattr(task, "retry_expected_hash", None)),
seeding_time_limit=seeding_time_limit,
ratio_limit=float(ratio_limit) if ratio_limit is not None else None,
)
def _resolve_download(
self,
task: DownloadTask,
status_callback: Callable[[str, str | None], None],
) -> DownloadRequest | None:
"""Resolve Prowlarr cache entry into download request parameters."""
# Look up the cached release
prowlarr_result = get_release(task.task_id)
if not prowlarr_result:
logger.info("Prowlarr release cache miss, refreshing: %s", task.task_id)
prowlarr_result = self._refresh_release(task)
if prowlarr_result is None:
logger.warning("Prowlarr release refresh failed: %s", task.task_id)
status_callback("error", EXPIRED_LINK_REFRESH_ERROR)
return None
# Extract download URL
download_url = get_preferred_download_url(prowlarr_result)
if not download_url:
status_callback("error", "No download URL available")
return None
# Determine protocol
protocol = get_protocol(prowlarr_result)
if protocol == "unknown":
status_callback("error", "Could not determine download protocol")
return None
release_name = prowlarr_result.get("title") or task.title or "Unknown"
expected_hash = str(prowlarr_result.get("infoHash") or "").strip() or None
seeding_time_limit = None
ratio_limit = None
if config.get("PROWLARR_USE_SEED_PREFERENCES", False):
raw_configured_seed_time = prowlarr_result.get("configuredSeedTimeMinutes")
raw_configured_ratio = prowlarr_result.get("configuredRatioLimit")
seeding_time_limit = _coerce_positive_minutes(raw_configured_seed_time)
ratio_limit = float(raw_configured_ratio) if raw_configured_ratio is not None else None
# Fallback: search-time enrichment can be missing when the indexer
# settings fetch transiently failed during the search (#795).
# Re-resolve the limits from Prowlarr at grab time so torrents are
# never sent to the client without their configured share limits.
if seeding_time_limit is None and ratio_limit is None and protocol == "torrent":
fallback = self._fetch_seed_settings_fallback(prowlarr_result.get("indexerId"))
if fallback:
seeding_time_limit = _coerce_positive_minutes(
fallback.get("seeding_time_limit_minutes")
)
raw_ratio = fallback.get("ratio_limit")
ratio_limit = float(raw_ratio) if raw_ratio is not None else None
if seeding_time_limit is None and ratio_limit is None and protocol == "torrent":
logger.warning(
"Prowlarr seed preferences are enabled but no share limits "
"could be resolved for release '%s' (indexerId=%s); the "
"torrent will use the client's global limits",
release_name,
prowlarr_result.get("indexerId"),
)
return DownloadRequest(
url=download_url,
protocol=protocol,
release_name=release_name,
expected_hash=expected_hash,
seeding_time_limit=seeding_time_limit,
ratio_limit=ratio_limit,
)
def _refresh_release(self, task: DownloadTask) -> dict[str, Any] | None:
"""Re-query Prowlarr and cache the exact original release if it still exists."""
title = normalize_optional_text(task.title)
if title is None:
return None
context = getattr(task, "retry_source_context", None)
if not isinstance(context, dict):
context = {}
indexer = normalize_optional_text(context.get("indexer"))
book = BookMetadata(
provider="shelfmark",
provider_id=task.task_id,
title=title,
authors=[task.author] if task.author else [],
search_title=title,
search_author=task.author,
)
# No language default here on purpose: this re-finds one exact release by its
# guid, and Prowlarr does not filter on plan.languages anyway.
plan = build_release_search_plan(
book,
indexers=[indexer] if indexer is not None else None,
)
source = ProwlarrSource()
results = source.search(book, plan, content_type=task.content_type or "ebook")
for release in results:
raw_release = get_release(release.source_id)
if raw_release is None:
continue
if not self._raw_release_matches_task(raw_release, task.task_id):
continue
cache_release(task.task_id, raw_release)
logger.info("Refreshed Prowlarr release: %s", task.task_id)
return raw_release
return None
@staticmethod
def _raw_release_matches_task(raw_release: dict[str, Any], task_id: str) -> bool:
wanted = normalize_optional_text(task_id)
if wanted is None:
return False
bare = [
identity
for identity in (
normalize_optional_text(raw_release.get("guid")),
normalize_optional_text(raw_release.get("infoUrl")),
)
if identity is not None
]
identities = [*bare, build_source_id(raw_release)]
indexer_id = coerce_int_like(raw_release.get("indexerId"))
if indexer_id is not None:
identities.extend(f"{indexer_id}:{identity}" for identity in bare)
return wanted in identities
def _refresh_download_request_after_add_failure(
self,
*,
task: DownloadTask,
request: DownloadRequest,
error: Exception,
status_callback: Callable[[str, str | None], None],
) -> DownloadRequest | None:
"""Refresh once when a cached Prowlarr torrent proxy URL has expired."""
if request.protocol != "torrent":
return None
if HASH_DETECTION_ERROR not in str(error):
return None
parsed = urlparse(request.url)
if parsed.scheme.lower() not in {"http", "https"}:
return None
logger.info("Refreshing stale Prowlarr torrent URL for %s", task.task_id)
remove_release(task.task_id)
refreshed_request = self._resolve_download(task, status_callback)
if refreshed_request is None:
raise RuntimeError(EXPIRED_LINK_REFRESH_ERROR) from error
return refreshed_request
def _on_download_complete(self, task: DownloadTask) -> None:
"""Remove completed release from the Prowlarr cache."""
remove_release(task.task_id)
def cancel(self, task_id: str) -> bool:
"""Cancel download and clean up cache. Primary cancellation is via cancel_flag."""
logger.debug("Cancel requested for Prowlarr task: %s", task_id)
remove_release(task_id)
return super().cancel(task_id)