mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-10-05 18:41:04 +01:00
## Multi-book packs: inspect a release before download and file each book separately Closes #576 ### Problem One queued release is always treated as one book. When a torrent is actually a whole series (`Series/Book 1 - Title/…`, or a flat folder of `Series 1.0 - Title.m4b` files), post-processing walks the whole tree, flattens every file into one list and renames them `Title - 01…10` under the searched book's `{Author}/{Title}`. Audiobookshelf then sees a single 10-file "book" and the user has to re-file everything by hand. ### What this does Most releases expose their file list *before* anything is downloaded, so the split is decided up front and approved by the user, then the download is fire-and-forget: 1. **Inspect** – clicking a release's download button now calls `POST /api/releases/inspect` first. A new optional `DownloadHandler.list_files(release_data)` hook returns the release's files without downloading: - **AudiobookBay** reads the torrent file table off the detail page it already fetches (the page is now cached for 120 s, so inspect + download cost ABB one request). - **Prowlarr** parses `info.files` from the `.torrent` it already fetches (the existing 120 s torrent-fetch cache is reused). Magnet-only and usenet releases report "can't inspect". - Other sources default to `None`. 2. **Review** – if the plan contains more than one book, the Find Releases modal swaps the list for a review panel: one row per book with editable title / series position / year, expandable file lists, non-book sidecars (`.txt`, covers) shown as ignored, a "Treat as a single book" switch, and **Download N books**. Single-book releases queue immediately, exactly as before. 3. **File** – the approved plan travels with the task (`DownloadTask.book_plan`, retry-safe) and post-processing files each book through the existing transfer code, one book at a time (`dataclasses.replace(task, title=…, series_position=…, year=…)`), so organize/rename templates, part numbering (now scoped per book), hardlinks, torrent copy-preserve and usenet handling are unchanged. Status reads `Complete (N books, M files)`. 4. **Fallback** – when a release can't be inspected the user gets a toast, and a small "Multi-book pack" toggle in the modal header forces a heuristic split (subfolder = book, or one book per file when the file names carry series positions). Planning lives in `shelfmark/download/postprocess/packs.py` and is shared by the inspect endpoint and post-processing, so what the user approved is what gets filed. The name parser strips `Book 3 -`, `03 -`, `1.0 -`, `3.`, `[03]`, `#3`, a leading series name, labels like "An Expanse Novella -", repeated titles (`Gods of Risk 2.5 - Gods of Risk`) and a trailing `(Year)`; author and series name come from the book that was searched, and the searched book's own series position is never applied to its siblings. ### Files - `shelfmark/download/postprocess/packs.py` (new) – `PackFile/PackBook/PackPlan`, `plan_pack`, `parse_pack_book_name`, `group_files_into_books`, `match_plan_to_files` - `shelfmark/core/release_inspect_routes.py` (new) – `POST /api/releases/inspect` - `shelfmark/release_sources/__init__.py` – `DownloadHandler.list_files` hook - `shelfmark/release_sources/audiobookbay/{scraper,handler}.py` – detail-page cache, `extract_file_list`, `list_files` - `shelfmark/release_sources/prowlarr/handler.py`, `download/clients/torrent_utils.py` – `extract_file_list_from_torrent`, `list_files` - `shelfmark/core/models.py`, `download/orchestrator.py` – `multi_book` / `book_plan` fields, queue + retry serialization - `shelfmark/download/postprocess/transfer.py`, `pipeline.py`, `outputs/folder.py` – per-book transfer branch and status message - `src/frontend`: `components/PackReviewPanel.tsx` (new), `ReleaseModal.tsx`, `App.tsx`, `services/api.ts`, `types/index.ts`, `utils/releasePayload.ts` (payload builder moved out of `App.tsx`), `utils/packReview.ts` - `docs/dev/release-sources-plugin-guide.md` – documents the `list_files` hook ### Out of scope (follow-ups) - Listing files from an NZB (Shelfmark already fetches the bytes; `<file subject>` names are noisy) - Inspecting magnet links via qBittorrent's files API after a paused add - BookLore / email outputs (they ignore `book_plan`; noted in code) - The combined ebook + audiobook flow ### Testing **Automated** (`make checks`, `make python-test`, `make frontend-test` all green; the only failures on my machine are the pre-existing `tests/config/test_entrypoint_permissions.py` cases, which need bash ≥ 4 and fail identically on `main` under macOS bash 3.2): - `tests/download/test_packs.py` – name parsing (markers, series name, novella labels, repeated titles, bare numeric titles like `1984`), nested / flat / mixed / deeper-nested packs, single wrapping folder not treated as a pack, plan-to-disk matching with basename fallback - `tests/core/test_processing_packs.py` – full `post_process_download` runs on a real temp filesystem: approved plan files each book under its own `{Author}/{Title}`, heuristic split of a nested pack, searched book's series position does not leak, multi-file book inside a pack keeps `- 01/- 02` per book, hardlinked torrent pack leaves the seeding tree intact, no pack fields ⇒ behaviour unchanged, single group degrades to the searched title, status message - `tests/core/test_release_inspect_routes.py` – plan response, not-inspectable, handler errors never 500, unknown source / missing `source_id` ⇒ 400, login required - `tests/audiobookbay/test_file_list.py` – file-table scraping from real ABB markup (multi-file and single-file pages), handler host validation, one page fetch shared by magnet + file list - `tests/prowlarr/test_torrent_file_list.py` – multi-file / single-file `.torrent` parsing, handler behaviour for torrent URL vs magnet vs usenet vs cache miss - `tests/download/test_orchestrator_pack_fields.py` – queue-time parsing and retry round-trip - Frontend: `releasePayload.test.ts`, `packReview.test.ts` (vitest) **Manual, on a real deployment** (arm64 image built from this branch, run as a side container next to production with the same qBittorrent / Audiobookshelf setup, `FILE_ORGANIZATION_AUDIOBOOK=organize`, hardlinks on): - AudiobookBay "The Expanse Complete 2.0" (7.87 GB, 36 files): clicking download opened the review panel in ~1 s showing **18 books · 18 files · 18 files ignored** (the `.txt` sidecars), with series positions 0.1–9.5 and years parsed from the file names; novella labels stripped ("The Churn", "The Butcher of Anderson Station"). Editing a title in the panel works. Confirming queued one task; the magnet resolved from the cached page in ~30 ms; after the download the task reported `Complete (18 books, 18 files)`, 18 hardlinks landed as `audiobooks/James S. A. Corey/<Title>/<Title>.m4b`, the torrent kept seeding, and Audiobookshelf scanned each folder as its own book (title, author, embedded chapters). - A second pack ("Expanse [01 - 9.5]", `Title N - Title` naming) was inspected to verify the repeated-title rule and the Back button, without downloading. - Single-book releases still queue immediately with no extra UI.
402 lines
16 KiB
Python
402 lines
16 KiB
Python
"""Prowlarr download handler - resolves releases and delegates lifecycle to shared clients."""
|
|
|
|
from typing import TYPE_CHECKING, Any
|
|
from urllib.parse import urlparse
|
|
|
|
import requests
|
|
|
|
from shelfmark.core.config import config
|
|
from shelfmark.core.logger import setup_logger
|
|
from shelfmark.core.request_helpers import normalize_optional_text
|
|
from shelfmark.core.search_plan import build_release_search_plan
|
|
from shelfmark.core.utils import normalize_http_url
|
|
from shelfmark.download.clients import (
|
|
DownloadClient,
|
|
get_client,
|
|
list_configured_clients,
|
|
)
|
|
from shelfmark.download.clients.base_handler import (
|
|
COMPLETED_PATH_MAX_ATTEMPTS as _DEFAULT_COMPLETED_PATH_MAX_ATTEMPTS,
|
|
)
|
|
from shelfmark.download.clients.base_handler import (
|
|
COMPLETED_PATH_RETRY_INTERVAL as _DEFAULT_COMPLETED_PATH_RETRY_INTERVAL,
|
|
)
|
|
from shelfmark.download.clients.base_handler import (
|
|
POLL_INTERVAL as _DEFAULT_POLL_INTERVAL,
|
|
)
|
|
from shelfmark.download.clients.base_handler import (
|
|
DownloadRequest,
|
|
ExternalClientHandler,
|
|
)
|
|
from shelfmark.download.clients.torrent_utils import (
|
|
extract_file_list_from_torrent,
|
|
extract_torrent_info,
|
|
)
|
|
from shelfmark.metadata_providers import BookMetadata
|
|
from shelfmark.release_sources import register_handler
|
|
from shelfmark.release_sources.prowlarr.api import IndexerSeedSettings, ProwlarrClient
|
|
from shelfmark.release_sources.prowlarr.cache import cache_release, get_release, remove_release
|
|
from shelfmark.release_sources.prowlarr.source import ProwlarrSource
|
|
from shelfmark.release_sources.prowlarr.utils import (
|
|
build_source_id,
|
|
coerce_int_like,
|
|
get_preferred_download_url,
|
|
get_protocol,
|
|
sanitize_download_url,
|
|
)
|
|
|
|
if TYPE_CHECKING:
|
|
from collections.abc import Callable
|
|
|
|
from shelfmark.core.models import DownloadTask
|
|
from shelfmark.download.postprocess.packs import PackFile
|
|
|
|
logger = setup_logger(__name__)
|
|
|
|
# Errors that ProwlarrClient can raise when fetching indexer settings.
|
|
_SEED_SETTINGS_FALLBACK_ERRORS = (
|
|
requests.exceptions.RequestException,
|
|
OSError,
|
|
RuntimeError,
|
|
TypeError,
|
|
ValueError,
|
|
)
|
|
|
|
__all__ = [
|
|
"ProwlarrHandler",
|
|
"POLL_INTERVAL",
|
|
"COMPLETED_PATH_RETRY_INTERVAL",
|
|
"COMPLETED_PATH_MAX_ATTEMPTS",
|
|
"config",
|
|
]
|
|
|
|
# Backwards-compat constants for tests patching this module.
|
|
POLL_INTERVAL = _DEFAULT_POLL_INTERVAL
|
|
COMPLETED_PATH_RETRY_INTERVAL = _DEFAULT_COMPLETED_PATH_RETRY_INTERVAL
|
|
COMPLETED_PATH_MAX_ATTEMPTS = _DEFAULT_COMPLETED_PATH_MAX_ATTEMPTS
|
|
EXPIRED_LINK_REFRESH_ERROR = (
|
|
"The indexer download link expired and the release could not be refreshed. "
|
|
"Search again for a fresh result."
|
|
)
|
|
HASH_DETECTION_ERROR = "Could not determine torrent hash from URL"
|
|
|
|
|
|
def _coerce_positive_minutes(raw_minutes: object) -> int | None:
|
|
minutes = coerce_int_like(raw_minutes)
|
|
if minutes is None:
|
|
return None
|
|
return minutes if minutes > 0 else None
|
|
|
|
|
|
@register_handler("prowlarr")
|
|
class ProwlarrHandler(ExternalClientHandler):
|
|
"""Handler for Prowlarr downloads via configured torrent or usenet client."""
|
|
|
|
@staticmethod
|
|
def _build_prowlarr_client() -> ProwlarrClient | None:
|
|
"""Build a ProwlarrClient from config, or None if not configured."""
|
|
raw_url = config.get("PROWLARR_URL", "")
|
|
raw_api_key = config.get("PROWLARR_API_KEY", "")
|
|
url = normalize_optional_text(raw_url) if isinstance(raw_url, str) else None
|
|
api_key = normalize_optional_text(raw_api_key) if isinstance(raw_api_key, str) else None
|
|
if not url or not api_key:
|
|
return None
|
|
normalized_url = normalize_http_url(url)
|
|
if not normalized_url:
|
|
return None
|
|
return ProwlarrClient(normalized_url, api_key)
|
|
|
|
def _fetch_seed_settings_fallback(self, raw_indexer_id: object) -> IndexerSeedSettings | None:
|
|
"""Fetch share limits for one indexer directly from Prowlarr.
|
|
|
|
Used when the cached release is missing its search-time seed-limit
|
|
enrichment so that transient failures during search cannot cause a
|
|
torrent to be added without its configured share limits.
|
|
"""
|
|
indexer_id = coerce_int_like(raw_indexer_id)
|
|
if indexer_id is None:
|
|
return None
|
|
|
|
client = self._build_prowlarr_client()
|
|
if client is None:
|
|
return None
|
|
|
|
try:
|
|
settings = client.get_indexer_seed_settings(restrict_to=[indexer_id])
|
|
except _SEED_SETTINGS_FALLBACK_ERRORS:
|
|
logger.warning(
|
|
"Grab-time seed settings fallback failed for indexerId=%s",
|
|
indexer_id,
|
|
exc_info=True,
|
|
)
|
|
return None
|
|
|
|
return settings.get(indexer_id)
|
|
|
|
def list_files(self, release_data: dict[str, Any]) -> list[PackFile] | None:
|
|
"""List a cached torrent release's files from its .torrent, without downloading.
|
|
|
|
Magnet-only and usenet releases cannot be listed ahead of time.
|
|
"""
|
|
source_id = str(release_data.get("source_id") or "")
|
|
prowlarr_result = get_release(source_id) if source_id else None
|
|
if not prowlarr_result or get_protocol(prowlarr_result) != "torrent":
|
|
return None
|
|
download_url = sanitize_download_url(str(prowlarr_result.get("downloadUrl") or "").strip())
|
|
if not download_url or download_url.startswith("magnet:"):
|
|
return None
|
|
expected_hash = str(prowlarr_result.get("infoHash") or "").strip() or None
|
|
info = extract_torrent_info(download_url, expected_hash=expected_hash)
|
|
if not info.torrent_data:
|
|
return None
|
|
return extract_file_list_from_torrent(info.torrent_data)
|
|
|
|
def _get_client(self, protocol: str) -> DownloadClient | None:
|
|
"""Compatibility shim so module-level patching still works in tests."""
|
|
return get_client(protocol)
|
|
|
|
def _list_configured_clients(self) -> list[str]:
|
|
"""Compatibility shim so module-level patching still works in tests."""
|
|
return list_configured_clients()
|
|
|
|
def _poll_interval(self) -> float:
|
|
return POLL_INTERVAL
|
|
|
|
def _completed_path_retry_interval(self) -> float:
|
|
return COMPLETED_PATH_RETRY_INTERVAL
|
|
|
|
def _completed_path_max_attempts(self) -> int:
|
|
return COMPLETED_PATH_MAX_ATTEMPTS
|
|
|
|
def build_retry_resolution_fields(self, release_data: dict[str, Any]) -> dict[str, Any]:
|
|
source_id = normalize_optional_text(release_data.get("source_id"))
|
|
extra = release_data.get("extra")
|
|
if not isinstance(extra, dict):
|
|
extra = {}
|
|
|
|
retry_source_context: dict[str, Any] = {}
|
|
indexer_id = release_data.get("indexer_id") or extra.get("indexer_id")
|
|
if indexer_id is not None:
|
|
retry_source_context["indexer_id"] = indexer_id
|
|
|
|
indexer = normalize_optional_text(release_data.get("indexer") or extra.get("indexer"))
|
|
if indexer is not None and indexer.lower() != "unknown":
|
|
retry_source_context["indexer"] = indexer
|
|
|
|
info_url = normalize_optional_text(release_data.get("info_url") or extra.get("info_url"))
|
|
if info_url is not None:
|
|
retry_source_context["info_url"] = info_url
|
|
|
|
if source_id is not None:
|
|
retry_source_context["source_id"] = source_id
|
|
|
|
return {
|
|
"retry_download_url": None,
|
|
"retry_download_protocol": None,
|
|
"retry_source_context": retry_source_context,
|
|
}
|
|
|
|
@classmethod
|
|
def _restore_download_request_from_task(cls, task: DownloadTask) -> DownloadRequest | None:
|
|
"""Rebuild a DownloadRequest when the in-memory Prowlarr cache is gone."""
|
|
retry_download_url = normalize_optional_text(getattr(task, "retry_download_url", None))
|
|
retry_download_protocol = normalize_optional_text(
|
|
getattr(task, "retry_download_protocol", None)
|
|
)
|
|
if retry_download_url is None or retry_download_protocol is None:
|
|
return None
|
|
|
|
protocol = retry_download_protocol.lower()
|
|
if protocol not in {"torrent", "usenet"}:
|
|
return None
|
|
|
|
ratio_limit = getattr(task, "retry_ratio_limit", None)
|
|
if not isinstance(ratio_limit, (int, float)) or isinstance(ratio_limit, bool):
|
|
ratio_limit = None
|
|
|
|
seeding_time_limit = getattr(task, "retry_seeding_time_limit_minutes", None)
|
|
if not isinstance(seeding_time_limit, int) or isinstance(seeding_time_limit, bool):
|
|
seeding_time_limit = None
|
|
|
|
return DownloadRequest(
|
|
url=retry_download_url,
|
|
protocol=protocol,
|
|
release_name=(
|
|
normalize_optional_text(getattr(task, "retry_release_name", None))
|
|
or task.title
|
|
or "Unknown"
|
|
),
|
|
expected_hash=normalize_optional_text(getattr(task, "retry_expected_hash", None)),
|
|
seeding_time_limit=seeding_time_limit,
|
|
ratio_limit=float(ratio_limit) if ratio_limit is not None else None,
|
|
)
|
|
|
|
def _resolve_download(
|
|
self,
|
|
task: DownloadTask,
|
|
status_callback: Callable[[str, str | None], None],
|
|
) -> DownloadRequest | None:
|
|
"""Resolve Prowlarr cache entry into download request parameters."""
|
|
# Look up the cached release
|
|
prowlarr_result = get_release(task.task_id)
|
|
if not prowlarr_result:
|
|
logger.info("Prowlarr release cache miss, refreshing: %s", task.task_id)
|
|
prowlarr_result = self._refresh_release(task)
|
|
if prowlarr_result is None:
|
|
logger.warning("Prowlarr release refresh failed: %s", task.task_id)
|
|
status_callback("error", EXPIRED_LINK_REFRESH_ERROR)
|
|
return None
|
|
|
|
# Extract download URL
|
|
download_url = get_preferred_download_url(prowlarr_result)
|
|
if not download_url:
|
|
status_callback("error", "No download URL available")
|
|
return None
|
|
|
|
# Determine protocol
|
|
protocol = get_protocol(prowlarr_result)
|
|
if protocol == "unknown":
|
|
status_callback("error", "Could not determine download protocol")
|
|
return None
|
|
|
|
release_name = prowlarr_result.get("title") or task.title or "Unknown"
|
|
expected_hash = str(prowlarr_result.get("infoHash") or "").strip() or None
|
|
|
|
seeding_time_limit = None
|
|
ratio_limit = None
|
|
if config.get("PROWLARR_USE_SEED_PREFERENCES", False):
|
|
raw_configured_seed_time = prowlarr_result.get("configuredSeedTimeMinutes")
|
|
raw_configured_ratio = prowlarr_result.get("configuredRatioLimit")
|
|
|
|
seeding_time_limit = _coerce_positive_minutes(raw_configured_seed_time)
|
|
ratio_limit = float(raw_configured_ratio) if raw_configured_ratio is not None else None
|
|
|
|
# Fallback: search-time enrichment can be missing when the indexer
|
|
# settings fetch transiently failed during the search (#795).
|
|
# Re-resolve the limits from Prowlarr at grab time so torrents are
|
|
# never sent to the client without their configured share limits.
|
|
if seeding_time_limit is None and ratio_limit is None and protocol == "torrent":
|
|
fallback = self._fetch_seed_settings_fallback(prowlarr_result.get("indexerId"))
|
|
if fallback:
|
|
seeding_time_limit = _coerce_positive_minutes(
|
|
fallback.get("seeding_time_limit_minutes")
|
|
)
|
|
raw_ratio = fallback.get("ratio_limit")
|
|
ratio_limit = float(raw_ratio) if raw_ratio is not None else None
|
|
|
|
if seeding_time_limit is None and ratio_limit is None and protocol == "torrent":
|
|
logger.warning(
|
|
"Prowlarr seed preferences are enabled but no share limits "
|
|
"could be resolved for release '%s' (indexerId=%s); the "
|
|
"torrent will use the client's global limits",
|
|
release_name,
|
|
prowlarr_result.get("indexerId"),
|
|
)
|
|
|
|
return DownloadRequest(
|
|
url=download_url,
|
|
protocol=protocol,
|
|
release_name=release_name,
|
|
expected_hash=expected_hash,
|
|
seeding_time_limit=seeding_time_limit,
|
|
ratio_limit=ratio_limit,
|
|
)
|
|
|
|
def _refresh_release(self, task: DownloadTask) -> dict[str, Any] | None:
|
|
"""Re-query Prowlarr and cache the exact original release if it still exists."""
|
|
title = normalize_optional_text(task.title)
|
|
if title is None:
|
|
return None
|
|
|
|
context = getattr(task, "retry_source_context", None)
|
|
if not isinstance(context, dict):
|
|
context = {}
|
|
|
|
indexer = normalize_optional_text(context.get("indexer"))
|
|
book = BookMetadata(
|
|
provider="shelfmark",
|
|
provider_id=task.task_id,
|
|
title=title,
|
|
authors=[task.author] if task.author else [],
|
|
search_title=title,
|
|
search_author=task.author,
|
|
)
|
|
# No language default here on purpose: this re-finds one exact release by its
|
|
# guid, and Prowlarr does not filter on plan.languages anyway.
|
|
plan = build_release_search_plan(
|
|
book,
|
|
indexers=[indexer] if indexer is not None else None,
|
|
)
|
|
|
|
source = ProwlarrSource()
|
|
results = source.search(book, plan, content_type=task.content_type or "ebook")
|
|
for release in results:
|
|
raw_release = get_release(release.source_id)
|
|
if raw_release is None:
|
|
continue
|
|
if not self._raw_release_matches_task(raw_release, task.task_id):
|
|
continue
|
|
|
|
cache_release(task.task_id, raw_release)
|
|
logger.info("Refreshed Prowlarr release: %s", task.task_id)
|
|
return raw_release
|
|
|
|
return None
|
|
|
|
@staticmethod
|
|
def _raw_release_matches_task(raw_release: dict[str, Any], task_id: str) -> bool:
|
|
wanted = normalize_optional_text(task_id)
|
|
if wanted is None:
|
|
return False
|
|
|
|
bare = [
|
|
identity
|
|
for identity in (
|
|
normalize_optional_text(raw_release.get("guid")),
|
|
normalize_optional_text(raw_release.get("infoUrl")),
|
|
)
|
|
if identity is not None
|
|
]
|
|
|
|
identities = [*bare, build_source_id(raw_release)]
|
|
indexer_id = coerce_int_like(raw_release.get("indexerId"))
|
|
if indexer_id is not None:
|
|
identities.extend(f"{indexer_id}:{identity}" for identity in bare)
|
|
|
|
return wanted in identities
|
|
|
|
def _refresh_download_request_after_add_failure(
|
|
self,
|
|
*,
|
|
task: DownloadTask,
|
|
request: DownloadRequest,
|
|
error: Exception,
|
|
status_callback: Callable[[str, str | None], None],
|
|
) -> DownloadRequest | None:
|
|
"""Refresh once when a cached Prowlarr torrent proxy URL has expired."""
|
|
if request.protocol != "torrent":
|
|
return None
|
|
if HASH_DETECTION_ERROR not in str(error):
|
|
return None
|
|
|
|
parsed = urlparse(request.url)
|
|
if parsed.scheme.lower() not in {"http", "https"}:
|
|
return None
|
|
|
|
logger.info("Refreshing stale Prowlarr torrent URL for %s", task.task_id)
|
|
remove_release(task.task_id)
|
|
refreshed_request = self._resolve_download(task, status_callback)
|
|
if refreshed_request is None:
|
|
raise RuntimeError(EXPIRED_LINK_REFRESH_ERROR) from error
|
|
return refreshed_request
|
|
|
|
def _on_download_complete(self, task: DownloadTask) -> None:
|
|
"""Remove completed release from the Prowlarr cache."""
|
|
remove_release(task.task_id)
|
|
|
|
def cancel(self, task_id: str) -> bool:
|
|
"""Cancel download and clean up cache. Primary cancellation is via cancel_flag."""
|
|
logger.debug("Cancel requested for Prowlarr task: %s", task_id)
|
|
remove_release(task_id)
|
|
return super().cancel(task_id)
|