mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-09-29 10:14:54 +01:00
## Backport bug fixes from `NemesisHubris/litfinder` Forwards a curated set of bug fixes from [NemesisHubris/litfinder](https://github.com/NemesisHubris/litfinder) — a community fork of this project — that address open issues here. All commits preserve original authorship via `git cherry-pick`; this PR is a backport rather than original work. Each fix has been reviewed locally, lint/format-cleaned to match this repo's existing ruff config, and verified with the test suite. Rebrand strings, license switches, and features have been deliberately excluded. ### Upstream issues addressed - **#999** — Mirror URLs with query params no longer break search requests (strip query string/fragment in `normalize_http_url`) - **#956** — Apprise notifications now respect the configured proxy (proxy env vars injected before dispatch) - **#1025** — rTorrent: separate `RTORRENT_AUDIOBOOK_LABEL` setting, falls back to book label if unset - **#1010** — Stop button in Activity no longer makes the panel disappear (snapshot refresh on cancel) - **#1021** — Anna's Archive slow-download countdown now caps retries instead of looping forever - **#1040** — Empty destination directory cleaned up when write probe fails - **PR #1031** — Language detection from Anna's Archive distant path when listing metadata is missing ### Additional fixes (no open issue but clear bugs) - **fix: Python 2 `except` syntax across 27 files** — `except X, Y:` is a SyntaxError in Python 3 and prevents affected modules from importing at runtime. Mechanical sweep to `except (X, Y):`. - **fix(abb): info hash validation with magnet fallback** — adds SHA-1/SHA-256 hex validation on extracted info hashes; falls back to scanning the full page for a magnet link (e.g. posted in comments) when the table value is malformed. Also extends the exact-phrase fallback to manual queries and defaults the ABB listing language to `en` when missing, preventing valid results from being hidden by the language filter. Includes a small test-fixture fix (`test(abb): use valid hex info hashes in scraper test fixtures`) since the existing fixtures used non-hex placeholders that the new validation correctly rejects. - **fix: Anna's Archive title parser** — handles nested edition spans and filters `lgli` catalog descriptor entries (e.g. "Book/Online Audio") that were polluting search results. ### Deliberately not included - LitFinder rebranding (UI strings, Apprise app ID, logo). The `fix: three upstream bugs` commit (#999/#956/#1025) was cherry-picked with Apprise app-id, description, and logo-URL strings reverted from "LitFinder" back to "Shelfmark"; noted in the commit body. - Features from the LitFinder fork (multi-variant title search, multi-book flat-folder grouping, fuzzy text matching, "Leave in Place" output handler, admin display name, custom-source plugin system). These are larger behavior changes that each warrant their own focused review — happy to send any of them separately if of interest. - LitFinder-specific test environment and CI infrastructure. ### Verification - Backend: **1879 passed**, 96 skipped (1 preexisting failure on `seleniumbase`-dependent test in local venv; runs fine in the standard Docker image with the `browser` extra) - Lint, format, dead-code: all clean against this repo's existing ruff/vulture config - One follow-up cleanup commit (`style: ruff lint and format fixes for ported commits`) brings the cherry-picked code into compliance with this repo's ruff settings — no behavior changes there ### Etiquette / credit Per-commit authorship preserved by cherry-pick. The only edits to the original commits are: - `fix: three upstream bugs` — Apprise rebrand strings reverted to "Shelfmark" (noted in commit body, original author retained as `Co-Authored-By` via cherry-pick) - One follow-up `style:` commit for ruff config alignment Big thanks to [@NemesisHubris](https://github.com/NemesisHubris) for the original work in LitFinder; this PR exists to make sure these fixes reach Shelfmark's wider user base. Happy to revise scope, split into smaller PRs, or split off the Py2 cleanup separately if that's preferable. --------- Co-authored-by: NemesisHubris <155838970+NemesisHubris@users.noreply.github.com> Co-authored-by: CaliBrain <calibrain@l4n.xyz>
334 lines
10 KiB
Python
334 lines
10 KiB
Python
"""Shared utility functions for the Shelfmark."""
|
|
|
|
import base64
|
|
import importlib
|
|
import os
|
|
import re
|
|
import sqlite3
|
|
from pathlib import Path
|
|
from threading import Lock
|
|
from typing import TYPE_CHECKING
|
|
from urllib.parse import urlparse
|
|
|
|
from shelfmark.core.request_helpers import normalize_optional_text
|
|
|
|
if TYPE_CHECKING:
|
|
from types import ModuleType
|
|
|
|
|
|
def normalize_http_url(
|
|
url: str | None,
|
|
*,
|
|
default_scheme: str = "http",
|
|
strip_trailing_slash: bool = True,
|
|
allow_special: tuple[str, ...] = (),
|
|
) -> str:
|
|
"""Normalize a configured HTTP URL for requests and links."""
|
|
if not isinstance(url, str):
|
|
return ""
|
|
|
|
normalized = url.strip()
|
|
if not normalized:
|
|
return ""
|
|
|
|
if (normalized.startswith('"') and normalized.endswith('"')) or (
|
|
normalized.startswith("'") and normalized.endswith("'")
|
|
):
|
|
normalized = normalized[1:-1].strip()
|
|
if not normalized:
|
|
return ""
|
|
|
|
if allow_special:
|
|
special_map = {value.lower(): value for value in allow_special if isinstance(value, str)}
|
|
special_match = special_map.get(normalized.lower())
|
|
if special_match is not None:
|
|
return special_match
|
|
|
|
if normalized.startswith(("/", "./", "../")):
|
|
return normalized
|
|
|
|
if "://" not in normalized:
|
|
scheme = default_scheme.strip().rstrip(":/")
|
|
if scheme:
|
|
normalized = f"{scheme}://{normalized}"
|
|
|
|
# Strip query string and fragment — mirrors are used as base URLs for
|
|
# constructing search requests; params/fragments on the configured URL
|
|
# produce malformed URLs when paths are appended (issue #999).
|
|
parsed = urlparse(normalized)
|
|
if parsed.query or parsed.fragment:
|
|
normalized = parsed._replace(query="", fragment="").geturl()
|
|
|
|
if strip_trailing_slash:
|
|
normalized = normalized.rstrip("/")
|
|
|
|
return normalized
|
|
|
|
|
|
_xmlrpc_patch_lock = Lock()
|
|
_xmlrpc_patch_applied = False
|
|
_XMLRPC_PATCH_ERRORS = (ImportError, AttributeError, OSError, RuntimeError)
|
|
|
|
|
|
def get_hardened_xmlrpc_client() -> ModuleType:
|
|
"""Return ``xmlrpc.client`` after best-effort defusedxml monkey patching."""
|
|
global _xmlrpc_patch_applied
|
|
if not _xmlrpc_patch_applied:
|
|
with _xmlrpc_patch_lock:
|
|
if not _xmlrpc_patch_applied:
|
|
try:
|
|
from defusedxml.xmlrpc import monkey_patch
|
|
|
|
monkey_patch()
|
|
_xmlrpc_patch_applied = True
|
|
except _XMLRPC_PATCH_ERRORS:
|
|
# Keep runtime behavior unchanged if defusedxml is unavailable.
|
|
_xmlrpc_patch_applied = False
|
|
|
|
return importlib.import_module("xmlrpc.client")
|
|
|
|
|
|
def normalize_base_path(value: str | None) -> str:
|
|
"""Normalize a URL base path for reverse proxy subpath deployments."""
|
|
if not isinstance(value, str):
|
|
return ""
|
|
|
|
path = value.strip()
|
|
if not path:
|
|
return ""
|
|
|
|
if "://" in path:
|
|
parsed = urlparse(path)
|
|
path = parsed.path or ""
|
|
|
|
if not path or path == "/":
|
|
return ""
|
|
|
|
if not path.startswith("/"):
|
|
path = "/" + path
|
|
|
|
return path.rstrip("/")
|
|
|
|
|
|
def is_audiobook(content_type: str | None) -> bool:
|
|
"""Check if content type indicates an audiobook."""
|
|
return bool(content_type and "audiobook" in content_type.lower())
|
|
|
|
|
|
CONTENT_TYPES = [
|
|
"book (fiction)",
|
|
"book (non-fiction)",
|
|
"book (unknown)",
|
|
"magazine",
|
|
"comic book",
|
|
"audiobook",
|
|
"standards document",
|
|
"musical score",
|
|
"other",
|
|
]
|
|
|
|
# Maps AA content types to their config keys for content-type routing
|
|
# Used when AA_CONTENT_TYPE_ROUTING is enabled
|
|
_AA_CONTENT_TYPE_TO_CONFIG_KEY = {
|
|
"book (fiction)": "AA_CONTENT_TYPE_DIR_FICTION",
|
|
"book (non-fiction)": "AA_CONTENT_TYPE_DIR_NON_FICTION",
|
|
"book (unknown)": "AA_CONTENT_TYPE_DIR_UNKNOWN",
|
|
"magazine": "AA_CONTENT_TYPE_DIR_MAGAZINE",
|
|
"comic book": "AA_CONTENT_TYPE_DIR_COMIC",
|
|
"audiobook": "AA_CONTENT_TYPE_DIR_AUDIOBOOK",
|
|
"standards document": "AA_CONTENT_TYPE_DIR_STANDARDS",
|
|
"musical score": "AA_CONTENT_TYPE_DIR_MUSICAL_SCORE",
|
|
"other": "AA_CONTENT_TYPE_DIR_OTHER",
|
|
}
|
|
|
|
# Legacy mapping - kept for backwards compatibility during migration
|
|
_LEGACY_CONTENT_TYPE_TO_CONFIG_KEY = {
|
|
"book (fiction)": "INGEST_DIR_BOOK_FICTION",
|
|
"book (non-fiction)": "INGEST_DIR_BOOK_NON_FICTION",
|
|
"book (unknown)": "INGEST_DIR_BOOK_UNKNOWN",
|
|
"magazine": "INGEST_DIR_MAGAZINE",
|
|
"comic book": "INGEST_DIR_COMIC_BOOK",
|
|
"audiobook": "INGEST_DIR_AUDIOBOOK",
|
|
"standards document": "INGEST_DIR_STANDARDS_DOCUMENT",
|
|
"musical score": "INGEST_DIR_MUSICAL_SCORE",
|
|
"other": "INGEST_DIR_OTHER",
|
|
}
|
|
|
|
_USER_PLACEHOLDER_PATTERN = re.compile(r"\{user\}", re.IGNORECASE)
|
|
_INVALID_USER_PATH_CHARS = re.compile(r'[\\/:*?"<>|]')
|
|
|
|
|
|
def _sanitize_user_for_path(username: str) -> str:
|
|
"""Sanitize username for path usage in destination placeholders."""
|
|
sanitized = _INVALID_USER_PATH_CHARS.sub("_", username.strip())
|
|
return sanitized.strip(" .")
|
|
|
|
|
|
def _resolve_destination_username(
|
|
user_id: int | None = None,
|
|
username: str | None = None,
|
|
) -> str:
|
|
explicit = str(username or "").strip()
|
|
if explicit:
|
|
return explicit
|
|
|
|
if user_id is None:
|
|
return ""
|
|
|
|
try:
|
|
from shelfmark.core.user_db import UserDB
|
|
|
|
user_db = UserDB(str(Path(os.environ.get("CONFIG_DIR", "/config")) / "users.db"))
|
|
user_db.initialize()
|
|
user = user_db.get_user(user_id=user_id)
|
|
if not user:
|
|
return ""
|
|
return str(user.get("username") or "").strip()
|
|
except ImportError, OSError, sqlite3.Error:
|
|
return ""
|
|
|
|
|
|
def _expand_user_destination_placeholder(
|
|
path_value: str,
|
|
user_id: int | None = None,
|
|
username: str | None = None,
|
|
) -> str:
|
|
"""Expand `{User}` placeholders in destination paths."""
|
|
if not isinstance(path_value, str):
|
|
return path_value
|
|
|
|
if not _USER_PLACEHOLDER_PATTERN.search(path_value):
|
|
return path_value
|
|
|
|
resolved_username = _sanitize_user_for_path(
|
|
_resolve_destination_username(user_id=user_id, username=username)
|
|
)
|
|
return _USER_PLACEHOLDER_PATTERN.sub(resolved_username, path_value)
|
|
|
|
|
|
def get_destination(
|
|
*,
|
|
is_audiobook: bool = False,
|
|
user_id: int | None = None,
|
|
username: str | None = None,
|
|
) -> Path:
|
|
"""Get base destination directory. Audiobooks fall back to main destination."""
|
|
from shelfmark.core.config import config
|
|
|
|
if is_audiobook:
|
|
# Audiobook destination with fallback to main destination
|
|
audiobook_dest = config.get("DESTINATION_AUDIOBOOK", "", user_id=user_id)
|
|
if audiobook_dest:
|
|
return Path(
|
|
_expand_user_destination_placeholder(
|
|
str(audiobook_dest),
|
|
user_id=user_id,
|
|
username=username,
|
|
)
|
|
)
|
|
|
|
# Main destination (also fallback for audiobooks)
|
|
# Check new setting first, then legacy INGEST_DIR
|
|
destination = config.get("DESTINATION", "", user_id=user_id) or config.get(
|
|
"INGEST_DIR", "/books"
|
|
)
|
|
return Path(
|
|
_expand_user_destination_placeholder(
|
|
str(destination),
|
|
user_id=user_id,
|
|
username=username,
|
|
)
|
|
)
|
|
|
|
|
|
def get_aa_content_type_dir(content_type: str | None = None) -> Path | None:
|
|
"""Get override directory for AA content-type routing if configured."""
|
|
from shelfmark.core.config import config
|
|
|
|
# Check if content-type routing is enabled (new or legacy setting)
|
|
if not config.get("AA_CONTENT_TYPE_ROUTING", False) and not config.get(
|
|
"USE_CONTENT_TYPE_DIRECTORIES", False
|
|
):
|
|
return None
|
|
|
|
if not content_type:
|
|
return None
|
|
|
|
content_type_lower = content_type.lower().strip()
|
|
|
|
# Try new AA-specific config keys first, then legacy keys
|
|
for mapping in (_AA_CONTENT_TYPE_TO_CONFIG_KEY, _LEGACY_CONTENT_TYPE_TO_CONFIG_KEY):
|
|
config_key = mapping.get(content_type_lower)
|
|
if config_key:
|
|
custom_dir = _coerce_config_path(config.get(config_key, ""))
|
|
if custom_dir is not None:
|
|
return custom_dir
|
|
|
|
return None
|
|
|
|
|
|
def get_ingest_dir(content_type: str | None = None) -> Path:
|
|
"""Return the legacy ingest directory for a content type."""
|
|
from shelfmark.core.config import config
|
|
|
|
# Check new DESTINATION setting first, then legacy INGEST_DIR
|
|
default_ingest_dir = _coerce_config_path(config.get("DESTINATION", "")) or _coerce_config_path(
|
|
config.get("INGEST_DIR", "/books")
|
|
)
|
|
if default_ingest_dir is None:
|
|
default_ingest_dir = Path("/books")
|
|
|
|
if not content_type:
|
|
return default_ingest_dir
|
|
|
|
# Check for content-type override
|
|
override_dir = get_aa_content_type_dir(content_type)
|
|
if override_dir:
|
|
return override_dir
|
|
|
|
return default_ingest_dir
|
|
|
|
|
|
def transform_cover_url(cover_url: str | None, cache_id: str) -> str | None:
|
|
"""Transform external cover URL to local proxy URL when caching is enabled."""
|
|
if not cover_url:
|
|
return cover_url
|
|
|
|
# Skip if already a local URL (starts with /)
|
|
if cover_url.startswith("/"):
|
|
return cover_url
|
|
|
|
# Check if cover caching is enabled
|
|
from shelfmark.config.env import is_covers_cache_enabled
|
|
|
|
if not is_covers_cache_enabled():
|
|
return cover_url
|
|
|
|
from shelfmark.core.config import config as app_config
|
|
|
|
# Encode the original URL and create a proxy URL
|
|
encoded_url = base64.urlsafe_b64encode(cover_url.encode()).decode()
|
|
base_path = normalize_base_path(normalize_optional_text(app_config.get("URL_BASE", "")))
|
|
if base_path:
|
|
return f"{base_path}/api/covers/{cache_id}?url={encoded_url}"
|
|
return f"/api/covers/{cache_id}?url={encoded_url}"
|
|
|
|
|
|
def _coerce_config_path(value: object) -> Path | None:
|
|
if isinstance(value, os.PathLike):
|
|
path_value = os.fspath(value)
|
|
if isinstance(path_value, str):
|
|
normalized = path_value.strip()
|
|
if normalized:
|
|
return Path(normalized)
|
|
return None
|
|
|
|
if not isinstance(value, str):
|
|
return None
|
|
|
|
normalized = value.strip()
|
|
if not normalized:
|
|
return None
|
|
|
|
return Path(normalized)
|