Files
shelfmark/tests/core/test_language_template_variable.py
T
816a735cde Add a {Language} naming template variable, and consolidate language resolution (#1142)
Fixes #1138
Fixes #1141

## Problem

Two language editions of one book resolve to the same canonical title,
so they render to the same path and the second gets a `_1` collision
suffix. Audiobookshelf treats a folder as exactly one library item, so
the pair becomes a single book with both files as tracks and a summed
runtime.

Shelfmark already parses and displays the language. It just never
reached the template engine.

## `{Language}` template variable

A template like `{Author}/{Title}{ (Language)}/{Author} - {Title}` now
yields:

```
/library/J K Rowling/Harry Potter (sv)/J K Rowling - Harry Potter.m4b
/library/J K Rowling/Harry Potter/J K Rowling - Harry Potter.m4b
```

The untagged edition's path is byte-identical to today, so no existing
layout shifts.

Three details worth flagging:

**The value is casefolded.** On a case-insensitive filesystem `(SV)` and
`(sv)` would collapse back into one folder, reintroducing the exact
collision being fixed.

**Values meaning "we don't know" render nothing** rather than producing
`Project Hail Mary (unknown)` folders. Anna's Archive reports that
string literally (`direct_download.py`, `language = detected or
"unknown"`).

**The frontend wasn't sending the release language at all**, so the
token would have stayed empty for exactly the audiobook sources in the
report. Prowlarr and AudiobookBay do not put language in `extra` the way
`direct_download` does, hence the payload plumbing. It reads
`release.language`, never `book.language` — the latter is the provider's
canonical edition and would mislabel a translation, with a regression
test for that specifically.

Not gated to audiobooks: Calibre-Web-Automated stages ingested files by
basename and discards folder structure, so the rename (filename)
template is the only lever those users have. Verified that form works:
`J K Rowling - Harry Potter (sv).epub`.

## Language consolidation (#1141)

Three release sources each carried their own alias map, all resolving to
the same ISO 639-1 codes, alongside a bundled database that only one of
them used. Adding a language meant editing three places.

Aliases now live in `data/book-languages.json` beside the code and name
they belong to, and `shelfmark/core/languages.py` resolves any of them —
two-letter code, ISO 639-2 three-letter in either the bibliographic or
terminological form, or English name. Prowlarr and AudiobookBay drop
their tables. Direct Download keeps its own path-parsing heuristics,
including the ambiguous short codes that collide with English words
(`de`, `en`, `no`, `in`), and takes only the alias data.

This also closes a coverage gap. MyAnonamouse offers 62 languages;
Prowlarr mapped 37, and an unmapped code is *dropped* rather than passed
through, so the other 25 carried no language at all — leaving
`{Language}` empty and the collision unfixed for Latin, Farsi, Tamil,
Urdu and the rest. Seven languages MAM offers had no database entry at
all: Bosnian, Burmese, Estonian, Icelandic, Manx, Scottish Gaelic,
Sanskrit.

Also fixes the Traditional Chinese code, which used a U+2011
non-breaking hyphen. Nothing compares against the ASCII spelling today
so it was latent, but it would silently defeat the first thing that did.

## Validation

Verified end to end against a live Prowlarr and MyAnonamouse, not just
unit tests. A real search returning both an English and a Swedish
edition, through the actual `queue_release` → `DownloadTask` → naming
path:

```
STEP 1  real MAM search        -> 37 releases, languages: ['en', 'sv']
STEP 3  queue_release          -> task.language='sv'
STEP 4  build_metadata_dict    -> metadata['Language']='sv'
STEP 5  build_library_path     -> /library/J K Rowling/Harry Potter (sv)/...
two language editions resolve to DIFFERENT folders: True
```

The refactor is pinned by a snapshot of both per-source maps taken
*before* they were deleted. All 131 aliases are asserted to still
resolve to the same code, one parametrised test each, so a regression
names the specific alias.

Also verified: the filename-only template, the retry round-trip
(`serialize_task_for_retry` → `_restore_task_from_retry_payload`, plus a
legacy payload with no `language` key), and placeholder handling.

Added a `KNOWN_TOKENS` ordering invariant test — `find_placeholder()`
does a substring `.find()` in list order and nothing protected that
contract, so a future token in the wrong position could silently shadow
an existing one. And a lockstep guard on the frontend, since
`KNOWN_TOKENS` is hand-duplicated in TypeScript.

**One caveat worth stating.** Three MAM codes are confirmed by
observation (`ENG`→`en`, `SWE`→`sv`, `MAL`→`ml`, the last from a real
`[MAL / EPUB]` Tagore release). The remaining ~59 are derived from ISO
639-2 rather than observed, because MAM's catalogue is overwhelmingly
English — enabling 27 extra languages still yielded only one non-English
hit across 258 results. Mitigated rather than closed: both 639-2
variants are present for every language where they differ, and a wrong
alias is an unused entry while a missing one loses the language. Happy
to correct any code a maintainer knows differs.

## Test results

2056 Python tests pass (up from 1906). Frontend typecheck, lint, format
and 126 unit tests pass.

Pre-existing failures on my machine, unchanged by this branch and
unrelated: `tests/bypass/` needs `seleniumbase`, and
`tests/config/test_entrypoint_permissions.py` uses bash-4 syntax that
macOS bash 3.2 rejects.

---------

Co-authored-by: delize <4028612+delize@users.noreply.github.com>
Co-authored-by: CaliBrain <calibrain@l4n.xyz>
2026-07-28 15:19:59 -04:00

197 lines
7.7 KiB
Python

"""Tests for the {Language} template variable.
Different-language editions of one book resolve to the same title, so without a
language token they render to the same path and land in one folder. Audiobookshelf
treats a folder as exactly one library item, so the two editions become a single
book with both files as tracks (calibrain/shelfmark#1138).
"""
import pytest
from shelfmark.core.models import DownloadTask
from shelfmark.core.naming import (
KNOWN_TOKENS,
normalize_language_code,
parse_naming_template,
)
from shelfmark.download.orchestrator import (
_restore_task_from_retry_payload,
serialize_task_for_retry,
)
from shelfmark.download.postprocess.transfer import build_metadata_dict
class TestLanguageInKnownTokens:
def test_language_in_known_tokens(self):
assert "language" in KNOWN_TOKENS
def test_language_token_parsed(self):
assert parse_naming_template("{Language}", {"Language": "sv"}) == "sv"
def test_language_token_case_insensitive(self):
assert parse_naming_template("{language}", {"Language": "sv"}) == "sv"
class TestKnownTokensOrdering:
"""find_placeholder() does a substring find over KNOWN_TOKENS in list order.
Nothing else guards this contract, so a future token added in the wrong
position would silently shadow an existing one.
"""
def test_tokens_are_ordered_longest_first(self):
lengths = [len(token) for token in KNOWN_TOKENS]
assert lengths == sorted(lengths, reverse=True)
def test_no_token_is_shadowed_by_an_earlier_substring(self):
for shorter_index, shorter in enumerate(KNOWN_TOKENS):
for longer_index, longer in enumerate(KNOWN_TOKENS):
if shorter is longer or shorter not in longer:
continue
assert shorter_index > longer_index, (
f"{shorter!r} precedes {longer!r} and would shadow it"
)
class TestLanguageTemplateSubstitution:
"""The acceptance cases from the issue."""
TEMPLATE = "{Author}/{Title}{ (Language)}/{Author} - {Title}"
BASE = {"Author": "Andy Weir", "Title": "Project Hail Mary"}
def test_translated_edition_gets_its_own_folder(self):
result = parse_naming_template(
self.TEMPLATE, {**self.BASE, "Language": "sv"}, allow_path_separators=True
)
assert result == "Andy Weir/Project Hail Mary (sv)/Andy Weir - Project Hail Mary"
def test_untagged_edition_is_unchanged(self):
result = parse_naming_template(
self.TEMPLATE, {**self.BASE, "Language": None}, allow_path_separators=True
)
assert result == "Andy Weir/Project Hail Mary/Andy Weir - Project Hail Mary"
def test_the_two_editions_do_not_collide(self):
english = parse_naming_template(
self.TEMPLATE, {**self.BASE, "Language": None}, allow_path_separators=True
)
swedish = parse_naming_template(
self.TEMPLATE, {**self.BASE, "Language": "sv"}, allow_path_separators=True
)
assert english != swedish
def test_language_as_a_leading_folder(self):
result = parse_naming_template(
"{Language/}{Author}/{Title}",
{**self.BASE, "Language": "sv"},
allow_path_separators=True,
)
assert result == "sv/Andy Weir/Project Hail Mary"
def test_language_in_a_filename_template(self):
result = parse_naming_template(
"{Author} - {Title}{ (Language)}", {**self.BASE, "Language": "sv"}
)
assert result == "Andy Weir - Project Hail Mary (sv)"
def test_language_is_sanitized(self):
result = parse_naming_template("{Title}{ (Language)}", {"Title": "Book", "Language": "s/v"})
assert "/" not in result
class TestNormalizeLanguageCode:
def test_lowercases(self):
assert normalize_language_code("EN") == "en"
assert normalize_language_code("Sv") == "sv"
def test_strips_whitespace(self):
assert normalize_language_code(" sv ") == "sv"
def test_placeholders_render_empty(self):
for placeholder in ("unknown", "unk", "n/a", "na", "-", "--", "none", "null", ""):
assert normalize_language_code(placeholder) == "", placeholder
def test_placeholders_are_matched_case_insensitively(self):
assert normalize_language_code("Unknown") == ""
def test_none_renders_empty(self):
assert normalize_language_code(None) == ""
class TestBuildMetadataWithLanguage:
def test_language_reaches_the_template_metadata(self):
task = DownloadTask(task_id="t", source="prowlarr", title="Book", language="sv")
assert build_metadata_dict(task)["Language"] == "sv"
def test_language_is_normalized_on_the_way_out(self):
task = DownloadTask(task_id="t", source="prowlarr", title="Book", language="SV")
assert build_metadata_dict(task)["Language"] == "sv"
def test_placeholder_language_does_not_reach_the_path(self):
# Anna's Archive reports the literal string "unknown" when it cannot tell.
task = DownloadTask(task_id="t", source="direct", title="Book", language="unknown")
assert build_metadata_dict(task)["Language"] == ""
def test_missing_language_renders_empty(self):
task = DownloadTask(task_id="t", source="prowlarr", title="Book")
assert build_metadata_dict(task)["Language"] == ""
class TestLanguageSurvivesRetry:
"""DownloadTask is not rebuilt from dataclasses.fields(), so each of the
three orchestrator sites has to carry the field explicitly."""
def test_roundtrip_preserves_language(self):
task = DownloadTask(task_id="t", source="prowlarr", title="Book", language="sv")
restored = _restore_task_from_retry_payload(serialize_task_for_retry(task))
assert restored is not None
assert restored.language == "sv"
def test_legacy_payload_without_language_restores_cleanly(self):
task = DownloadTask(task_id="t", source="prowlarr", title="Book", language="sv")
payload = serialize_task_for_retry(task)
del payload["language"]
restored = _restore_task_from_retry_payload(payload)
assert restored is not None
assert restored.language is None
class TestEverySpellingCollapsesToOneFolder:
"""Sources report the same language differently; if the token rendered each
spelling verbatim they would land in separate folders, which is the exact
collision this token exists to prevent (reported on PR #1142)."""
TEMPLATE = "{Author}/{Title}{ (Language)}/{Title}"
def _folder(self, language):
task = DownloadTask(
task_id="t", source="prowlarr", title="Dune", author="Frank Herbert", language=language
)
return parse_naming_template(
self.TEMPLATE, build_metadata_dict(task), allow_path_separators=True
)
@pytest.mark.parametrize(
"spellings",
[
("en", "eng", "English", "english", "ENG", " Eng "),
("sv", "swe", "Swedish"),
("de", "ger", "deu", "German"),
("ml", "mal", "Malayalam"),
("fa", "per", "fas", "Farsi", "Persian"),
],
)
def test_all_spellings_of_a_language_share_one_folder(self, spellings):
rendered = {self._folder(spelling) for spelling in spellings}
assert len(rendered) == 1, f"{spellings} produced {sorted(rendered)}"
def test_a_language_we_cannot_resolve_is_kept_rather_than_dropped(self):
# It still separates editions, and cannot collide with a resolved code
# precisely because nothing resolves to it.
assert "klingon" in self._folder("Klingon")
def test_different_languages_still_get_different_folders(self):
assert self._folder("English") != self._folder("Swedish")