mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-09-24 13:40:21 +01:00
Fixes #1138 Fixes #1141 ## Problem Two language editions of one book resolve to the same canonical title, so they render to the same path and the second gets a `_1` collision suffix. Audiobookshelf treats a folder as exactly one library item, so the pair becomes a single book with both files as tracks and a summed runtime. Shelfmark already parses and displays the language. It just never reached the template engine. ## `{Language}` template variable A template like `{Author}/{Title}{ (Language)}/{Author} - {Title}` now yields: ``` /library/J K Rowling/Harry Potter (sv)/J K Rowling - Harry Potter.m4b /library/J K Rowling/Harry Potter/J K Rowling - Harry Potter.m4b ``` The untagged edition's path is byte-identical to today, so no existing layout shifts. Three details worth flagging: **The value is casefolded.** On a case-insensitive filesystem `(SV)` and `(sv)` would collapse back into one folder, reintroducing the exact collision being fixed. **Values meaning "we don't know" render nothing** rather than producing `Project Hail Mary (unknown)` folders. Anna's Archive reports that string literally (`direct_download.py`, `language = detected or "unknown"`). **The frontend wasn't sending the release language at all**, so the token would have stayed empty for exactly the audiobook sources in the report. Prowlarr and AudiobookBay do not put language in `extra` the way `direct_download` does, hence the payload plumbing. It reads `release.language`, never `book.language` — the latter is the provider's canonical edition and would mislabel a translation, with a regression test for that specifically. Not gated to audiobooks: Calibre-Web-Automated stages ingested files by basename and discards folder structure, so the rename (filename) template is the only lever those users have. Verified that form works: `J K Rowling - Harry Potter (sv).epub`. ## Language consolidation (#1141) Three release sources each carried their own alias map, all resolving to the same ISO 639-1 codes, alongside a bundled database that only one of them used. Adding a language meant editing three places. Aliases now live in `data/book-languages.json` beside the code and name they belong to, and `shelfmark/core/languages.py` resolves any of them — two-letter code, ISO 639-2 three-letter in either the bibliographic or terminological form, or English name. Prowlarr and AudiobookBay drop their tables. Direct Download keeps its own path-parsing heuristics, including the ambiguous short codes that collide with English words (`de`, `en`, `no`, `in`), and takes only the alias data. This also closes a coverage gap. MyAnonamouse offers 62 languages; Prowlarr mapped 37, and an unmapped code is *dropped* rather than passed through, so the other 25 carried no language at all — leaving `{Language}` empty and the collision unfixed for Latin, Farsi, Tamil, Urdu and the rest. Seven languages MAM offers had no database entry at all: Bosnian, Burmese, Estonian, Icelandic, Manx, Scottish Gaelic, Sanskrit. Also fixes the Traditional Chinese code, which used a U+2011 non-breaking hyphen. Nothing compares against the ASCII spelling today so it was latent, but it would silently defeat the first thing that did. ## Validation Verified end to end against a live Prowlarr and MyAnonamouse, not just unit tests. A real search returning both an English and a Swedish edition, through the actual `queue_release` → `DownloadTask` → naming path: ``` STEP 1 real MAM search -> 37 releases, languages: ['en', 'sv'] STEP 3 queue_release -> task.language='sv' STEP 4 build_metadata_dict -> metadata['Language']='sv' STEP 5 build_library_path -> /library/J K Rowling/Harry Potter (sv)/... two language editions resolve to DIFFERENT folders: True ``` The refactor is pinned by a snapshot of both per-source maps taken *before* they were deleted. All 131 aliases are asserted to still resolve to the same code, one parametrised test each, so a regression names the specific alias. Also verified: the filename-only template, the retry round-trip (`serialize_task_for_retry` → `_restore_task_from_retry_payload`, plus a legacy payload with no `language` key), and placeholder handling. Added a `KNOWN_TOKENS` ordering invariant test — `find_placeholder()` does a substring `.find()` in list order and nothing protected that contract, so a future token in the wrong position could silently shadow an existing one. And a lockstep guard on the frontend, since `KNOWN_TOKENS` is hand-duplicated in TypeScript. **One caveat worth stating.** Three MAM codes are confirmed by observation (`ENG`→`en`, `SWE`→`sv`, `MAL`→`ml`, the last from a real `[MAL / EPUB]` Tagore release). The remaining ~59 are derived from ISO 639-2 rather than observed, because MAM's catalogue is overwhelmingly English — enabling 27 extra languages still yielded only one non-English hit across 258 results. Mitigated rather than closed: both 639-2 variants are present for every language where they differ, and a wrong alias is an unused entry while a missing one loses the language. Happy to correct any code a maintainer knows differs. ## Test results 2056 Python tests pass (up from 1906). Frontend typecheck, lint, format and 126 unit tests pass. Pre-existing failures on my machine, unchanged by this branch and unrelated: `tests/bypass/` needs `seleniumbase`, and `tests/config/test_entrypoint_permissions.py` uses bash-4 syntax that macOS bash 3.2 rejects. --------- Co-authored-by: delize <4028612+delize@users.noreply.github.com> Co-authored-by: CaliBrain <calibrain@l4n.xyz>
60 lines
1.9 KiB
Python
60 lines
1.9 KiB
Python
from scripts.generate_env_docs import generate_env_docs
|
|
|
|
|
|
def test_generated_env_docs_use_canonical_mirror_env_vars() -> None:
|
|
docs = generate_env_docs()
|
|
|
|
for canonical_var in (
|
|
"AA_BASE_URL",
|
|
"AA_MIRROR_URLS",
|
|
"LIBGEN_MIRROR_URLS",
|
|
"ZLIB_MIRROR_URLS",
|
|
"WELIB_MIRROR_URLS",
|
|
):
|
|
assert f"`{canonical_var}`" in docs
|
|
|
|
for legacy_var in (
|
|
"AA_ADDITIONAL_URLS",
|
|
"LIBGEN_ADDITIONAL_URLS",
|
|
"ZLIB_PRIMARY_URL",
|
|
"ZLIB_ADDITIONAL_URLS",
|
|
"WELIB_PRIMARY_URL",
|
|
"WELIB_ADDITIONAL_URLS",
|
|
):
|
|
assert f"`{legacy_var}`" not in docs
|
|
|
|
assert "https://annas-archive.gl" not in docs
|
|
|
|
|
|
def test_generated_env_docs_describe_mirror_lists_as_comma_separated_strings() -> None:
|
|
docs = generate_env_docs()
|
|
|
|
assert (
|
|
"| `AA_MIRROR_URLS` | List the Anna's Archive mirror URLs you want Shelfmark to use. "
|
|
"Type a URL and press Enter to add it. Order matters when Auto is selected. | "
|
|
"string (comma-separated) | _empty list_ |"
|
|
) in docs
|
|
assert (
|
|
"| `LIBGEN_MIRROR_URLS` | Mirrors are tried in the order you add them until one works. | "
|
|
"string (comma-separated) | _empty list_ |"
|
|
) in docs
|
|
|
|
|
|
def test_generated_env_docs_include_custom_component_value_fields() -> None:
|
|
docs = generate_env_docs()
|
|
|
|
for env_var in (
|
|
"TEMPLATE_RENAME",
|
|
"TEMPLATE_ORGANIZE",
|
|
"TEMPLATE_AUDIOBOOK_RENAME",
|
|
"TEMPLATE_AUDIOBOOK_ORGANIZE",
|
|
):
|
|
assert f"`{env_var}`" in docs
|
|
|
|
assert (
|
|
"| `TEMPLATE_AUDIOBOOK_ORGANIZE` | Use / to create folders. Variables: "
|
|
"{Author}, {Title}, {Year}, {Language}, {User}, {OriginalName} "
|
|
"(source filename without extension), {Series}, {SeriesPosition}, {Subtitle}, "
|
|
"{PrimaryTitle}, {PartNumber}. Use arbitrary prefix/suffix:"
|
|
) in docs
|