mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-10-05 18:51:05 +01:00
Fixes #1138 Fixes #1141 ## Problem Two language editions of one book resolve to the same canonical title, so they render to the same path and the second gets a `_1` collision suffix. Audiobookshelf treats a folder as exactly one library item, so the pair becomes a single book with both files as tracks and a summed runtime. Shelfmark already parses and displays the language. It just never reached the template engine. ## `{Language}` template variable A template like `{Author}/{Title}{ (Language)}/{Author} - {Title}` now yields: ``` /library/J K Rowling/Harry Potter (sv)/J K Rowling - Harry Potter.m4b /library/J K Rowling/Harry Potter/J K Rowling - Harry Potter.m4b ``` The untagged edition's path is byte-identical to today, so no existing layout shifts. Three details worth flagging: **The value is casefolded.** On a case-insensitive filesystem `(SV)` and `(sv)` would collapse back into one folder, reintroducing the exact collision being fixed. **Values meaning "we don't know" render nothing** rather than producing `Project Hail Mary (unknown)` folders. Anna's Archive reports that string literally (`direct_download.py`, `language = detected or "unknown"`). **The frontend wasn't sending the release language at all**, so the token would have stayed empty for exactly the audiobook sources in the report. Prowlarr and AudiobookBay do not put language in `extra` the way `direct_download` does, hence the payload plumbing. It reads `release.language`, never `book.language` — the latter is the provider's canonical edition and would mislabel a translation, with a regression test for that specifically. Not gated to audiobooks: Calibre-Web-Automated stages ingested files by basename and discards folder structure, so the rename (filename) template is the only lever those users have. Verified that form works: `J K Rowling - Harry Potter (sv).epub`. ## Language consolidation (#1141) Three release sources each carried their own alias map, all resolving to the same ISO 639-1 codes, alongside a bundled database that only one of them used. Adding a language meant editing three places. Aliases now live in `data/book-languages.json` beside the code and name they belong to, and `shelfmark/core/languages.py` resolves any of them — two-letter code, ISO 639-2 three-letter in either the bibliographic or terminological form, or English name. Prowlarr and AudiobookBay drop their tables. Direct Download keeps its own path-parsing heuristics, including the ambiguous short codes that collide with English words (`de`, `en`, `no`, `in`), and takes only the alias data. This also closes a coverage gap. MyAnonamouse offers 62 languages; Prowlarr mapped 37, and an unmapped code is *dropped* rather than passed through, so the other 25 carried no language at all — leaving `{Language}` empty and the collision unfixed for Latin, Farsi, Tamil, Urdu and the rest. Seven languages MAM offers had no database entry at all: Bosnian, Burmese, Estonian, Icelandic, Manx, Scottish Gaelic, Sanskrit. Also fixes the Traditional Chinese code, which used a U+2011 non-breaking hyphen. Nothing compares against the ASCII spelling today so it was latent, but it would silently defeat the first thing that did. ## Validation Verified end to end against a live Prowlarr and MyAnonamouse, not just unit tests. A real search returning both an English and a Swedish edition, through the actual `queue_release` → `DownloadTask` → naming path: ``` STEP 1 real MAM search -> 37 releases, languages: ['en', 'sv'] STEP 3 queue_release -> task.language='sv' STEP 4 build_metadata_dict -> metadata['Language']='sv' STEP 5 build_library_path -> /library/J K Rowling/Harry Potter (sv)/... two language editions resolve to DIFFERENT folders: True ``` The refactor is pinned by a snapshot of both per-source maps taken *before* they were deleted. All 131 aliases are asserted to still resolve to the same code, one parametrised test each, so a regression names the specific alias. Also verified: the filename-only template, the retry round-trip (`serialize_task_for_retry` → `_restore_task_from_retry_payload`, plus a legacy payload with no `language` key), and placeholder handling. Added a `KNOWN_TOKENS` ordering invariant test — `find_placeholder()` does a substring `.find()` in list order and nothing protected that contract, so a future token in the wrong position could silently shadow an existing one. And a lockstep guard on the frontend, since `KNOWN_TOKENS` is hand-duplicated in TypeScript. **One caveat worth stating.** Three MAM codes are confirmed by observation (`ENG`→`en`, `SWE`→`sv`, `MAL`→`ml`, the last from a real `[MAL / EPUB]` Tagore release). The remaining ~59 are derived from ISO 639-2 rather than observed, because MAM's catalogue is overwhelmingly English — enabling 27 extra languages still yielded only one non-English hit across 258 results. Mitigated rather than closed: both 639-2 variants are present for every language where they differ, and a wrong alias is an unused entry while a missing one loses the language. Happy to correct any code a maintainer knows differs. ## Test results 2056 Python tests pass (up from 1906). Frontend typecheck, lint, format and 126 unit tests pass. Pre-existing failures on my machine, unchanged by this branch and unrelated: `tests/bypass/` needs `seleniumbase`, and `tests/config/test_entrypoint_permissions.py` uses bash-4 syntax that macOS bash 3.2 rejects. --------- Co-authored-by: delize <4028612+delize@users.noreply.github.com> Co-authored-by: CaliBrain <calibrain@l4n.xyz>
139 lines
4.9 KiB
Python
139 lines
4.9 KiB
Python
"""Canonical language resolution shared by every release source.
|
||
|
||
Release sources report a language in whatever shape their upstream uses: a
|
||
two-letter code, an ISO 639-2 three-letter code in either the bibliographic or
|
||
terminological form, or an English name. They all need the same ISO 639-1 code
|
||
out the other side, so the aliases live in one place (``data/book-languages.json``)
|
||
and adding a language means editing one file.
|
||
"""
|
||
|
||
import json
|
||
import threading
|
||
import unicodedata
|
||
from pathlib import Path
|
||
|
||
from shelfmark.core.logger import setup_logger
|
||
|
||
logger = setup_logger(__name__)
|
||
|
||
LANGUAGE_DATA_PATH = Path(__file__).resolve().parents[1].parent / "data" / "book-languages.json"
|
||
|
||
# Values a source uses to mean "we could not tell".
|
||
LANGUAGE_PLACEHOLDERS = frozenset({"", "-", "--", "unknown", "unk", "n/a", "na", "none", "null"})
|
||
|
||
_ALIAS_TO_CODE: dict[str, str] | None = None
|
||
_CODE_TO_NAME: dict[str, str] | None = None
|
||
_LOCK = threading.Lock()
|
||
|
||
|
||
# Separators that stand in for the hyphen in a subtag. The dashes turn up in
|
||
# codes copied from web pages -- "zh‑Hant" used U+2011, which renders close
|
||
# enough to both a hyphen and an underscore to go unnoticed -- and the
|
||
# underscore is the spelling Direct Download accepted before this module existed.
|
||
_SUBTAG_SEPARATORS = dict.fromkeys(map(ord, "‐‑‒–—―−﹘﹣-_"), "-")
|
||
|
||
|
||
def _fold(value: str) -> str:
|
||
"""Casefold, strip accents, and normalize subtag separators, so 'Español'
|
||
and 'espanol', or 'zh-Hant', 'zh‑Hant' and 'zh_Hant', all match."""
|
||
decomposed = unicodedata.normalize("NFKD", value).translate(_SUBTAG_SEPARATORS)
|
||
stripped = "".join(ch for ch in decomposed if not unicodedata.combining(ch))
|
||
return " ".join(stripped.split()).casefold()
|
||
|
||
|
||
def _load() -> tuple[dict[str, str], dict[str, str]]:
|
||
global _ALIAS_TO_CODE, _CODE_TO_NAME
|
||
|
||
if _ALIAS_TO_CODE is not None and _CODE_TO_NAME is not None:
|
||
return _ALIAS_TO_CODE, _CODE_TO_NAME
|
||
|
||
with _LOCK:
|
||
if _ALIAS_TO_CODE is not None and _CODE_TO_NAME is not None:
|
||
return _ALIAS_TO_CODE, _CODE_TO_NAME
|
||
|
||
alias_to_code: dict[str, str] = {}
|
||
code_to_name: dict[str, str] = {}
|
||
|
||
try:
|
||
raw = json.loads(LANGUAGE_DATA_PATH.read_text(encoding="utf-8"))
|
||
except OSError, ValueError:
|
||
logger.exception("Failed to load language data from %s", LANGUAGE_DATA_PATH)
|
||
raw = []
|
||
|
||
if not isinstance(raw, list):
|
||
logger.warning("Language data at %s is not a list", LANGUAGE_DATA_PATH)
|
||
raw = []
|
||
|
||
for item in raw:
|
||
if not isinstance(item, dict):
|
||
continue
|
||
code = str(item.get("code") or "").strip()
|
||
name = str(item.get("language") or "").strip()
|
||
if not code:
|
||
continue
|
||
|
||
code_to_name.setdefault(code, name or code)
|
||
|
||
for candidate in (code, name, *(item.get("aliases") or [])):
|
||
folded = _fold(str(candidate))
|
||
if folded and folded not in LANGUAGE_PLACEHOLDERS:
|
||
alias_to_code.setdefault(folded, code)
|
||
|
||
_ALIAS_TO_CODE = alias_to_code
|
||
_CODE_TO_NAME = code_to_name
|
||
return alias_to_code, code_to_name
|
||
|
||
|
||
def normalize_language(value: object) -> str | None:
|
||
"""Resolve any known spelling of a language to its ISO 639-1 code.
|
||
|
||
Accepts a two-letter code, an ISO 639-2 three-letter code in either the
|
||
bibliographic or terminological form, or an English name. Returns None for
|
||
anything unrecognised or for the placeholders a source uses to say it does
|
||
not know, so callers can treat "no language" uniformly.
|
||
"""
|
||
if value is None:
|
||
return None
|
||
|
||
folded = _fold(str(value))
|
||
if not folded or folded in LANGUAGE_PLACEHOLDERS:
|
||
return None
|
||
|
||
alias_to_code, _ = _load()
|
||
return alias_to_code.get(folded)
|
||
|
||
|
||
def language_name(code: str | None) -> str | None:
|
||
"""Return the English name for a language code, or None if unknown."""
|
||
if not code:
|
||
return None
|
||
|
||
_, code_to_name = _load()
|
||
return code_to_name.get(str(code).strip())
|
||
|
||
|
||
def language_alias_map() -> dict[str, str]:
|
||
"""Every known alias mapped to its code, for callers doing their own matching.
|
||
|
||
Direct Download scans free-text paths and needs the whole alias set up front
|
||
to look for, rather than resolving one candidate at a time.
|
||
"""
|
||
alias_to_code, _ = _load()
|
||
return dict(alias_to_code)
|
||
|
||
|
||
def supported_book_languages() -> list[dict[str, str]]:
|
||
"""The selectable languages, as ``{"language": ..., "code": ...}``.
|
||
|
||
Aliases are an implementation detail of resolution, so they are left out of
|
||
what the settings dropdown and the API hand to clients.
|
||
"""
|
||
_, code_to_name = _load()
|
||
return [{"language": name, "code": code} for code, name in code_to_name.items()]
|
||
|
||
|
||
def known_language_codes() -> frozenset[str]:
|
||
"""Every ISO 639-1 code the bundled language data defines."""
|
||
_, code_to_name = _load()
|
||
return frozenset(code_to_name)
|