mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-09-24 20:30:31 +01:00
Fixes #1138 Fixes #1141 ## Problem Two language editions of one book resolve to the same canonical title, so they render to the same path and the second gets a `_1` collision suffix. Audiobookshelf treats a folder as exactly one library item, so the pair becomes a single book with both files as tracks and a summed runtime. Shelfmark already parses and displays the language. It just never reached the template engine. ## `{Language}` template variable A template like `{Author}/{Title}{ (Language)}/{Author} - {Title}` now yields: ``` /library/J K Rowling/Harry Potter (sv)/J K Rowling - Harry Potter.m4b /library/J K Rowling/Harry Potter/J K Rowling - Harry Potter.m4b ``` The untagged edition's path is byte-identical to today, so no existing layout shifts. Three details worth flagging: **The value is casefolded.** On a case-insensitive filesystem `(SV)` and `(sv)` would collapse back into one folder, reintroducing the exact collision being fixed. **Values meaning "we don't know" render nothing** rather than producing `Project Hail Mary (unknown)` folders. Anna's Archive reports that string literally (`direct_download.py`, `language = detected or "unknown"`). **The frontend wasn't sending the release language at all**, so the token would have stayed empty for exactly the audiobook sources in the report. Prowlarr and AudiobookBay do not put language in `extra` the way `direct_download` does, hence the payload plumbing. It reads `release.language`, never `book.language` — the latter is the provider's canonical edition and would mislabel a translation, with a regression test for that specifically. Not gated to audiobooks: Calibre-Web-Automated stages ingested files by basename and discards folder structure, so the rename (filename) template is the only lever those users have. Verified that form works: `J K Rowling - Harry Potter (sv).epub`. ## Language consolidation (#1141) Three release sources each carried their own alias map, all resolving to the same ISO 639-1 codes, alongside a bundled database that only one of them used. Adding a language meant editing three places. Aliases now live in `data/book-languages.json` beside the code and name they belong to, and `shelfmark/core/languages.py` resolves any of them — two-letter code, ISO 639-2 three-letter in either the bibliographic or terminological form, or English name. Prowlarr and AudiobookBay drop their tables. Direct Download keeps its own path-parsing heuristics, including the ambiguous short codes that collide with English words (`de`, `en`, `no`, `in`), and takes only the alias data. This also closes a coverage gap. MyAnonamouse offers 62 languages; Prowlarr mapped 37, and an unmapped code is *dropped* rather than passed through, so the other 25 carried no language at all — leaving `{Language}` empty and the collision unfixed for Latin, Farsi, Tamil, Urdu and the rest. Seven languages MAM offers had no database entry at all: Bosnian, Burmese, Estonian, Icelandic, Manx, Scottish Gaelic, Sanskrit. Also fixes the Traditional Chinese code, which used a U+2011 non-breaking hyphen. Nothing compares against the ASCII spelling today so it was latent, but it would silently defeat the first thing that did. ## Validation Verified end to end against a live Prowlarr and MyAnonamouse, not just unit tests. A real search returning both an English and a Swedish edition, through the actual `queue_release` → `DownloadTask` → naming path: ``` STEP 1 real MAM search -> 37 releases, languages: ['en', 'sv'] STEP 3 queue_release -> task.language='sv' STEP 4 build_metadata_dict -> metadata['Language']='sv' STEP 5 build_library_path -> /library/J K Rowling/Harry Potter (sv)/... two language editions resolve to DIFFERENT folders: True ``` The refactor is pinned by a snapshot of both per-source maps taken *before* they were deleted. All 131 aliases are asserted to still resolve to the same code, one parametrised test each, so a regression names the specific alias. Also verified: the filename-only template, the retry round-trip (`serialize_task_for_retry` → `_restore_task_from_retry_payload`, plus a legacy payload with no `language` key), and placeholder handling. Added a `KNOWN_TOKENS` ordering invariant test — `find_placeholder()` does a substring `.find()` in list order and nothing protected that contract, so a future token in the wrong position could silently shadow an existing one. And a lockstep guard on the frontend, since `KNOWN_TOKENS` is hand-duplicated in TypeScript. **One caveat worth stating.** Three MAM codes are confirmed by observation (`ENG`→`en`, `SWE`→`sv`, `MAL`→`ml`, the last from a real `[MAL / EPUB]` Tagore release). The remaining ~59 are derived from ISO 639-2 rather than observed, because MAM's catalogue is overwhelmingly English — enabling 27 extra languages still yielded only one non-English hit across 258 results. Mitigated rather than closed: both 639-2 variants are present for every language where they differ, and a wrong alias is an unused entry while a missing one loses the language. Happy to correct any code a maintainer knows differs. ## Test results 2056 Python tests pass (up from 1906). Frontend typecheck, lint, format and 126 unit tests pass. Pre-existing failures on my machine, unchanged by this branch and unrelated: `tests/bypass/` needs `seleniumbase`, and `tests/config/test_entrypoint_permissions.py` uses bash-4 syntax that macOS bash 3.2 rejects. --------- Co-authored-by: delize <4028612+delize@users.noreply.github.com> Co-authored-by: CaliBrain <calibrain@l4n.xyz>
197 lines
7.7 KiB
Python
197 lines
7.7 KiB
Python
"""Tests for the {Language} template variable.
|
|
|
|
Different-language editions of one book resolve to the same title, so without a
|
|
language token they render to the same path and land in one folder. Audiobookshelf
|
|
treats a folder as exactly one library item, so the two editions become a single
|
|
book with both files as tracks (calibrain/shelfmark#1138).
|
|
"""
|
|
|
|
import pytest
|
|
|
|
from shelfmark.core.models import DownloadTask
|
|
from shelfmark.core.naming import (
|
|
KNOWN_TOKENS,
|
|
normalize_language_code,
|
|
parse_naming_template,
|
|
)
|
|
from shelfmark.download.orchestrator import (
|
|
_restore_task_from_retry_payload,
|
|
serialize_task_for_retry,
|
|
)
|
|
from shelfmark.download.postprocess.transfer import build_metadata_dict
|
|
|
|
|
|
class TestLanguageInKnownTokens:
|
|
def test_language_in_known_tokens(self):
|
|
assert "language" in KNOWN_TOKENS
|
|
|
|
def test_language_token_parsed(self):
|
|
assert parse_naming_template("{Language}", {"Language": "sv"}) == "sv"
|
|
|
|
def test_language_token_case_insensitive(self):
|
|
assert parse_naming_template("{language}", {"Language": "sv"}) == "sv"
|
|
|
|
|
|
class TestKnownTokensOrdering:
|
|
"""find_placeholder() does a substring find over KNOWN_TOKENS in list order.
|
|
|
|
Nothing else guards this contract, so a future token added in the wrong
|
|
position would silently shadow an existing one.
|
|
"""
|
|
|
|
def test_tokens_are_ordered_longest_first(self):
|
|
lengths = [len(token) for token in KNOWN_TOKENS]
|
|
assert lengths == sorted(lengths, reverse=True)
|
|
|
|
def test_no_token_is_shadowed_by_an_earlier_substring(self):
|
|
for shorter_index, shorter in enumerate(KNOWN_TOKENS):
|
|
for longer_index, longer in enumerate(KNOWN_TOKENS):
|
|
if shorter is longer or shorter not in longer:
|
|
continue
|
|
assert shorter_index > longer_index, (
|
|
f"{shorter!r} precedes {longer!r} and would shadow it"
|
|
)
|
|
|
|
|
|
class TestLanguageTemplateSubstitution:
|
|
"""The acceptance cases from the issue."""
|
|
|
|
TEMPLATE = "{Author}/{Title}{ (Language)}/{Author} - {Title}"
|
|
BASE = {"Author": "Andy Weir", "Title": "Project Hail Mary"}
|
|
|
|
def test_translated_edition_gets_its_own_folder(self):
|
|
result = parse_naming_template(
|
|
self.TEMPLATE, {**self.BASE, "Language": "sv"}, allow_path_separators=True
|
|
)
|
|
assert result == "Andy Weir/Project Hail Mary (sv)/Andy Weir - Project Hail Mary"
|
|
|
|
def test_untagged_edition_is_unchanged(self):
|
|
result = parse_naming_template(
|
|
self.TEMPLATE, {**self.BASE, "Language": None}, allow_path_separators=True
|
|
)
|
|
assert result == "Andy Weir/Project Hail Mary/Andy Weir - Project Hail Mary"
|
|
|
|
def test_the_two_editions_do_not_collide(self):
|
|
english = parse_naming_template(
|
|
self.TEMPLATE, {**self.BASE, "Language": None}, allow_path_separators=True
|
|
)
|
|
swedish = parse_naming_template(
|
|
self.TEMPLATE, {**self.BASE, "Language": "sv"}, allow_path_separators=True
|
|
)
|
|
assert english != swedish
|
|
|
|
def test_language_as_a_leading_folder(self):
|
|
result = parse_naming_template(
|
|
"{Language/}{Author}/{Title}",
|
|
{**self.BASE, "Language": "sv"},
|
|
allow_path_separators=True,
|
|
)
|
|
assert result == "sv/Andy Weir/Project Hail Mary"
|
|
|
|
def test_language_in_a_filename_template(self):
|
|
result = parse_naming_template(
|
|
"{Author} - {Title}{ (Language)}", {**self.BASE, "Language": "sv"}
|
|
)
|
|
assert result == "Andy Weir - Project Hail Mary (sv)"
|
|
|
|
def test_language_is_sanitized(self):
|
|
result = parse_naming_template("{Title}{ (Language)}", {"Title": "Book", "Language": "s/v"})
|
|
assert "/" not in result
|
|
|
|
|
|
class TestNormalizeLanguageCode:
|
|
def test_lowercases(self):
|
|
assert normalize_language_code("EN") == "en"
|
|
assert normalize_language_code("Sv") == "sv"
|
|
|
|
def test_strips_whitespace(self):
|
|
assert normalize_language_code(" sv ") == "sv"
|
|
|
|
def test_placeholders_render_empty(self):
|
|
for placeholder in ("unknown", "unk", "n/a", "na", "-", "--", "none", "null", ""):
|
|
assert normalize_language_code(placeholder) == "", placeholder
|
|
|
|
def test_placeholders_are_matched_case_insensitively(self):
|
|
assert normalize_language_code("Unknown") == ""
|
|
|
|
def test_none_renders_empty(self):
|
|
assert normalize_language_code(None) == ""
|
|
|
|
|
|
class TestBuildMetadataWithLanguage:
|
|
def test_language_reaches_the_template_metadata(self):
|
|
task = DownloadTask(task_id="t", source="prowlarr", title="Book", language="sv")
|
|
assert build_metadata_dict(task)["Language"] == "sv"
|
|
|
|
def test_language_is_normalized_on_the_way_out(self):
|
|
task = DownloadTask(task_id="t", source="prowlarr", title="Book", language="SV")
|
|
assert build_metadata_dict(task)["Language"] == "sv"
|
|
|
|
def test_placeholder_language_does_not_reach_the_path(self):
|
|
# Anna's Archive reports the literal string "unknown" when it cannot tell.
|
|
task = DownloadTask(task_id="t", source="direct", title="Book", language="unknown")
|
|
assert build_metadata_dict(task)["Language"] == ""
|
|
|
|
def test_missing_language_renders_empty(self):
|
|
task = DownloadTask(task_id="t", source="prowlarr", title="Book")
|
|
assert build_metadata_dict(task)["Language"] == ""
|
|
|
|
|
|
class TestLanguageSurvivesRetry:
|
|
"""DownloadTask is not rebuilt from dataclasses.fields(), so each of the
|
|
three orchestrator sites has to carry the field explicitly."""
|
|
|
|
def test_roundtrip_preserves_language(self):
|
|
task = DownloadTask(task_id="t", source="prowlarr", title="Book", language="sv")
|
|
restored = _restore_task_from_retry_payload(serialize_task_for_retry(task))
|
|
assert restored is not None
|
|
assert restored.language == "sv"
|
|
|
|
def test_legacy_payload_without_language_restores_cleanly(self):
|
|
task = DownloadTask(task_id="t", source="prowlarr", title="Book", language="sv")
|
|
payload = serialize_task_for_retry(task)
|
|
del payload["language"]
|
|
|
|
restored = _restore_task_from_retry_payload(payload)
|
|
|
|
assert restored is not None
|
|
assert restored.language is None
|
|
|
|
|
|
class TestEverySpellingCollapsesToOneFolder:
|
|
"""Sources report the same language differently; if the token rendered each
|
|
spelling verbatim they would land in separate folders, which is the exact
|
|
collision this token exists to prevent (reported on PR #1142)."""
|
|
|
|
TEMPLATE = "{Author}/{Title}{ (Language)}/{Title}"
|
|
|
|
def _folder(self, language):
|
|
task = DownloadTask(
|
|
task_id="t", source="prowlarr", title="Dune", author="Frank Herbert", language=language
|
|
)
|
|
return parse_naming_template(
|
|
self.TEMPLATE, build_metadata_dict(task), allow_path_separators=True
|
|
)
|
|
|
|
@pytest.mark.parametrize(
|
|
"spellings",
|
|
[
|
|
("en", "eng", "English", "english", "ENG", " Eng "),
|
|
("sv", "swe", "Swedish"),
|
|
("de", "ger", "deu", "German"),
|
|
("ml", "mal", "Malayalam"),
|
|
("fa", "per", "fas", "Farsi", "Persian"),
|
|
],
|
|
)
|
|
def test_all_spellings_of_a_language_share_one_folder(self, spellings):
|
|
rendered = {self._folder(spelling) for spelling in spellings}
|
|
assert len(rendered) == 1, f"{spellings} produced {sorted(rendered)}"
|
|
|
|
def test_a_language_we_cannot_resolve_is_kept_rather_than_dropped(self):
|
|
# It still separates editions, and cannot collide with a resolved code
|
|
# precisely because nothing resolves to it.
|
|
assert "klingon" in self._folder("Klingon")
|
|
|
|
def test_different_languages_still_get_different_folders(self):
|
|
assert self._folder("English") != self._folder("Swedish")
|