Files
shelfmark/shelfmark/download/postprocess/custom_script.py
T
816a735cde Add a {Language} naming template variable, and consolidate language resolution (#1142)
Fixes #1138
Fixes #1141

## Problem

Two language editions of one book resolve to the same canonical title,
so they render to the same path and the second gets a `_1` collision
suffix. Audiobookshelf treats a folder as exactly one library item, so
the pair becomes a single book with both files as tracks and a summed
runtime.

Shelfmark already parses and displays the language. It just never
reached the template engine.

## `{Language}` template variable

A template like `{Author}/{Title}{ (Language)}/{Author} - {Title}` now
yields:

```
/library/J K Rowling/Harry Potter (sv)/J K Rowling - Harry Potter.m4b
/library/J K Rowling/Harry Potter/J K Rowling - Harry Potter.m4b
```

The untagged edition's path is byte-identical to today, so no existing
layout shifts.

Three details worth flagging:

**The value is casefolded.** On a case-insensitive filesystem `(SV)` and
`(sv)` would collapse back into one folder, reintroducing the exact
collision being fixed.

**Values meaning "we don't know" render nothing** rather than producing
`Project Hail Mary (unknown)` folders. Anna's Archive reports that
string literally (`direct_download.py`, `language = detected or
"unknown"`).

**The frontend wasn't sending the release language at all**, so the
token would have stayed empty for exactly the audiobook sources in the
report. Prowlarr and AudiobookBay do not put language in `extra` the way
`direct_download` does, hence the payload plumbing. It reads
`release.language`, never `book.language` — the latter is the provider's
canonical edition and would mislabel a translation, with a regression
test for that specifically.

Not gated to audiobooks: Calibre-Web-Automated stages ingested files by
basename and discards folder structure, so the rename (filename)
template is the only lever those users have. Verified that form works:
`J K Rowling - Harry Potter (sv).epub`.

## Language consolidation (#1141)

Three release sources each carried their own alias map, all resolving to
the same ISO 639-1 codes, alongside a bundled database that only one of
them used. Adding a language meant editing three places.

Aliases now live in `data/book-languages.json` beside the code and name
they belong to, and `shelfmark/core/languages.py` resolves any of them —
two-letter code, ISO 639-2 three-letter in either the bibliographic or
terminological form, or English name. Prowlarr and AudiobookBay drop
their tables. Direct Download keeps its own path-parsing heuristics,
including the ambiguous short codes that collide with English words
(`de`, `en`, `no`, `in`), and takes only the alias data.

This also closes a coverage gap. MyAnonamouse offers 62 languages;
Prowlarr mapped 37, and an unmapped code is *dropped* rather than passed
through, so the other 25 carried no language at all — leaving
`{Language}` empty and the collision unfixed for Latin, Farsi, Tamil,
Urdu and the rest. Seven languages MAM offers had no database entry at
all: Bosnian, Burmese, Estonian, Icelandic, Manx, Scottish Gaelic,
Sanskrit.

Also fixes the Traditional Chinese code, which used a U+2011
non-breaking hyphen. Nothing compares against the ASCII spelling today
so it was latent, but it would silently defeat the first thing that did.

## Validation

Verified end to end against a live Prowlarr and MyAnonamouse, not just
unit tests. A real search returning both an English and a Swedish
edition, through the actual `queue_release` → `DownloadTask` → naming
path:

```
STEP 1  real MAM search        -> 37 releases, languages: ['en', 'sv']
STEP 3  queue_release          -> task.language='sv'
STEP 4  build_metadata_dict    -> metadata['Language']='sv'
STEP 5  build_library_path     -> /library/J K Rowling/Harry Potter (sv)/...
two language editions resolve to DIFFERENT folders: True
```

The refactor is pinned by a snapshot of both per-source maps taken
*before* they were deleted. All 131 aliases are asserted to still
resolve to the same code, one parametrised test each, so a regression
names the specific alias.

Also verified: the filename-only template, the retry round-trip
(`serialize_task_for_retry` → `_restore_task_from_retry_payload`, plus a
legacy payload with no `language` key), and placeholder handling.

Added a `KNOWN_TOKENS` ordering invariant test — `find_placeholder()`
does a substring `.find()` in list order and nothing protected that
contract, so a future token in the wrong position could silently shadow
an existing one. And a lockstep guard on the frontend, since
`KNOWN_TOKENS` is hand-duplicated in TypeScript.

**One caveat worth stating.** Three MAM codes are confirmed by
observation (`ENG`→`en`, `SWE`→`sv`, `MAL`→`ml`, the last from a real
`[MAL / EPUB]` Tagore release). The remaining ~59 are derived from ISO
639-2 rather than observed, because MAM's catalogue is overwhelmingly
English — enabling 27 extra languages still yielded only one non-English
hit across 258 results. Mitigated rather than closed: both 639-2
variants are present for every language where they differ, and a wrong
alias is an unused entry while a missing one loses the language. Happy
to correct any code a maintainer knows differs.

## Test results

2056 Python tests pass (up from 1906). Frontend typecheck, lint, format
and 126 unit tests pass.

Pre-existing failures on my machine, unchanged by this branch and
unrelated: `tests/bypass/` needs `seleniumbase`, and
`tests/config/test_entrypoint_permissions.py` uses bash-4 syntax that
macOS bash 3.2 rejects.

---------

Co-authored-by: delize <4028612+delize@users.noreply.github.com>
Co-authored-by: CaliBrain <calibrain@l4n.xyz>
2026-07-28 15:19:59 -04:00

326 lines
10 KiB
Python

"""Custom script execution helpers for post-processing hooks."""
from __future__ import annotations
import json
import os
import subprocess
from dataclasses import dataclass, field
from pathlib import Path
from typing import TYPE_CHECKING, Any
import shelfmark.core.config as core_config
from shelfmark.core.logger import setup_logger
from shelfmark.download.fs import run_blocking_io
from .steps import log_plan_steps, record_step
if TYPE_CHECKING:
from collections.abc import Callable
from shelfmark.core.models import DownloadTask
from .types import PlanStep
logger = setup_logger(__name__)
DEFAULT_CUSTOM_SCRIPT_TIMEOUT_SECONDS = 300 # 5 minutes
def resolve_custom_script_target(target_path: Path, destination: Path, path_mode: str) -> Path:
"""Resolve the path that should be passed as the custom script argument.
In absolute mode, we pass the full target path.
In relative mode, we pass a path relative to the destination folder. If the
target is not within the destination, fall back to just the filename to
avoid leaking unrelated absolute paths.
"""
mode = (path_mode or "absolute").strip().lower()
if mode != "relative":
return target_path
try:
return target_path.relative_to(destination)
except ValueError:
if target_path.is_absolute():
return Path(target_path.name)
return target_path
@dataclass(frozen=True)
class CustomScriptExecution:
"""Resolved command inputs for a single custom script run."""
script_path: str
target_arg: Path
target_abs: Path
destination: Path
mode: str
phase: str
payload_json: str | None = None
@dataclass(frozen=True)
class CustomScriptTransferSummary:
"""Transfer metadata exposed to custom post-process scripts."""
op_counts: dict[str, int]
use_hardlink: bool
is_torrent: bool
preserve_source: bool
@dataclass(frozen=True)
class CustomScriptContext:
"""Runtime context exposed to custom post-process scripts."""
task: DownloadTask
phase: str
output_mode: str
destination: Path | None = None
final_paths: list[Path] = field(default_factory=list)
target_path: Path | None = None
organization_mode: str | None = None
transfer: CustomScriptTransferSummary | None = None
output_details: dict[str, Any] = field(default_factory=dict)
def prepare_custom_script_execution(
script_path: str,
*,
target_path: Path,
destination: Path,
path_mode: str,
phase: str,
payload: dict[str, Any] | None = None,
) -> CustomScriptExecution:
"""Resolve script arguments and payload for a custom hook invocation."""
mode = (path_mode or "absolute").strip().lower()
if mode != "relative":
mode = "absolute"
target_arg = resolve_custom_script_target(target_path, destination, mode)
return CustomScriptExecution(
script_path=str(script_path),
target_arg=target_arg,
target_abs=target_path,
destination=destination,
mode=mode,
phase=phase,
payload_json=json.dumps(payload, indent=2, sort_keys=True) + "\n" if payload else None,
)
def run_custom_script(
execution: CustomScriptExecution,
*,
task_id: str,
status_callback: Callable[[str, str | None], None],
timeout_seconds: int = DEFAULT_CUSTOM_SCRIPT_TIMEOUT_SECONDS,
) -> bool:
"""Run a prepared custom script and report success."""
cwd: str | None = None
if execution.mode == "relative":
# Make relative paths unambiguous by running the script from the destination folder.
cwd = str(execution.destination)
logger.info(
"Task %s: running custom script %s on %s (%s)",
task_id,
execution.script_path,
execution.target_arg,
execution.phase,
)
try:
run_kwargs: dict[str, Any] = {
"check": True,
"timeout": timeout_seconds,
"capture_output": True,
"text": True,
"cwd": cwd,
}
# If we are not sending a JSON payload, close stdin so scripts that try
# to read it won't block indefinitely. When we do send a payload, let
# subprocess.run manage stdin implicitly via `input=` to avoid passing
# both arguments at once.
if execution.payload_json is None:
run_kwargs["stdin"] = subprocess.DEVNULL
else:
run_kwargs["input"] = execution.payload_json
result = run_blocking_io(
subprocess.run,
[execution.script_path, str(execution.target_arg)],
**run_kwargs,
)
if result.stdout:
logger.debug("Task %s: custom script stdout: %s", task_id, result.stdout.strip())
except FileNotFoundError:
logger.exception("Task %s: custom script not found: %s", task_id, execution.script_path)
status_callback("error", f"Custom script not found: {execution.script_path}")
return False
except PermissionError:
logger.exception(
"Task %s: custom script not executable: %s", task_id, execution.script_path
)
status_callback("error", f"Custom script not executable: {execution.script_path}")
return False
except subprocess.TimeoutExpired:
logger.exception(
"Task %s: custom script timed out after %ss: %s",
task_id,
timeout_seconds,
execution.script_path,
)
status_callback("error", "Custom script timed out")
return False
except subprocess.CalledProcessError as exc:
stderr = exc.stderr.strip() if exc.stderr else "No error output"
logger.exception(
"Task %s: custom script failed (exit code %s): %s",
task_id,
exc.returncode,
stderr,
)
status_callback("error", f"Custom script failed: {stderr[:100]}")
return False
else:
return True
def _choose_custom_script_target(
*,
explicit_target: Path | None,
destination: Path | None,
final_paths: list[Path],
) -> Path | None:
if explicit_target is not None:
return explicit_target
if len(final_paths) == 1:
return final_paths[0]
if len(final_paths) > 1:
try:
return Path(os.path.commonpath([str(p.parent) for p in final_paths]))
except ValueError:
return destination or final_paths[0].parent
return destination
def _build_custom_script_payload(
context: CustomScriptContext, *, target_path: Path
) -> dict[str, Any]:
payload: dict[str, Any] = {
"version": 1,
"phase": context.phase,
"task": {
"task_id": context.task.task_id,
"source": context.task.source,
"search_mode": context.task.search_mode.value if context.task.search_mode else None,
"title": context.task.title,
"author": context.task.author,
"year": context.task.year,
"format": context.task.format,
"content_type": context.task.content_type,
"series_name": context.task.series_name,
"series_position": context.task.series_position,
"subtitle": context.task.subtitle,
"language": context.task.language,
"original_download_path": context.task.original_download_path,
},
"output": {
"mode": context.output_mode,
"organization_mode": context.organization_mode,
},
"paths": {
"destination": str(context.destination) if context.destination else None,
"target": str(target_path),
"final_paths": [str(p) for p in context.final_paths],
},
}
if context.output_details:
payload["output"]["details"] = context.output_details
if context.transfer:
payload["transfer"] = {
"op_counts": context.transfer.op_counts,
"use_hardlink": context.transfer.use_hardlink,
"is_torrent": context.transfer.is_torrent,
"preserve_source": context.transfer.preserve_source,
}
return payload
def maybe_run_custom_script(
context: CustomScriptContext,
*,
status_callback: Callable[[str, str | None], None],
steps: list[PlanStep] | None = None,
) -> bool:
"""Run the custom script hook (if configured).
The output handler provides a `CustomScriptContext` describing what it did.
This function is responsible for choosing the script target, building the
optional JSON payload, and executing the script.
"""
script_path = getattr(core_config.config, "CUSTOM_SCRIPT", None)
if not isinstance(script_path, str) or not script_path.strip():
return True
target_path = _choose_custom_script_target(
explicit_target=context.target_path,
destination=context.destination,
final_paths=context.final_paths,
)
if not target_path:
logger.warning(
"Task %s: custom script configured but no target could be determined; skipping",
context.task.task_id,
)
return True
configured_path_mode = core_config.config.get("CUSTOM_SCRIPT_PATH_MODE", "absolute")
path_mode = configured_path_mode if isinstance(configured_path_mode, str) else "absolute"
payload: dict[str, Any] | None = None
if core_config.config.get("CUSTOM_SCRIPT_JSON_PAYLOAD", False):
payload = _build_custom_script_payload(context, target_path=target_path)
# If no destination is available for this output, fall back to the target's
# parent directory so the script can still run consistently.
execution_destination = context.destination or target_path.parent
execution = prepare_custom_script_execution(
script_path,
target_path=target_path,
destination=execution_destination,
path_mode=path_mode,
phase=context.phase,
payload=payload,
)
if steps is not None:
payload_bytes = len(execution.payload_json.encode("utf-8")) if execution.payload_json else 0
record_step(
steps,
"custom_script",
script=str(execution.script_path),
target=str(execution.target_arg),
target_abs=str(execution.target_abs),
mode=str(execution.mode),
phase=str(execution.phase),
payload_stdin=bool(execution.payload_json),
payload_bytes=payload_bytes,
)
log_plan_steps(context.task.task_id, steps)
return run_custom_script(
execution, task_id=context.task.task_id, status_callback=status_callback
)