mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-09-30 13:04:58 +01:00
Fixes #1138 Fixes #1141 ## Problem Two language editions of one book resolve to the same canonical title, so they render to the same path and the second gets a `_1` collision suffix. Audiobookshelf treats a folder as exactly one library item, so the pair becomes a single book with both files as tracks and a summed runtime. Shelfmark already parses and displays the language. It just never reached the template engine. ## `{Language}` template variable A template like `{Author}/{Title}{ (Language)}/{Author} - {Title}` now yields: ``` /library/J K Rowling/Harry Potter (sv)/J K Rowling - Harry Potter.m4b /library/J K Rowling/Harry Potter/J K Rowling - Harry Potter.m4b ``` The untagged edition's path is byte-identical to today, so no existing layout shifts. Three details worth flagging: **The value is casefolded.** On a case-insensitive filesystem `(SV)` and `(sv)` would collapse back into one folder, reintroducing the exact collision being fixed. **Values meaning "we don't know" render nothing** rather than producing `Project Hail Mary (unknown)` folders. Anna's Archive reports that string literally (`direct_download.py`, `language = detected or "unknown"`). **The frontend wasn't sending the release language at all**, so the token would have stayed empty for exactly the audiobook sources in the report. Prowlarr and AudiobookBay do not put language in `extra` the way `direct_download` does, hence the payload plumbing. It reads `release.language`, never `book.language` — the latter is the provider's canonical edition and would mislabel a translation, with a regression test for that specifically. Not gated to audiobooks: Calibre-Web-Automated stages ingested files by basename and discards folder structure, so the rename (filename) template is the only lever those users have. Verified that form works: `J K Rowling - Harry Potter (sv).epub`. ## Language consolidation (#1141) Three release sources each carried their own alias map, all resolving to the same ISO 639-1 codes, alongside a bundled database that only one of them used. Adding a language meant editing three places. Aliases now live in `data/book-languages.json` beside the code and name they belong to, and `shelfmark/core/languages.py` resolves any of them — two-letter code, ISO 639-2 three-letter in either the bibliographic or terminological form, or English name. Prowlarr and AudiobookBay drop their tables. Direct Download keeps its own path-parsing heuristics, including the ambiguous short codes that collide with English words (`de`, `en`, `no`, `in`), and takes only the alias data. This also closes a coverage gap. MyAnonamouse offers 62 languages; Prowlarr mapped 37, and an unmapped code is *dropped* rather than passed through, so the other 25 carried no language at all — leaving `{Language}` empty and the collision unfixed for Latin, Farsi, Tamil, Urdu and the rest. Seven languages MAM offers had no database entry at all: Bosnian, Burmese, Estonian, Icelandic, Manx, Scottish Gaelic, Sanskrit. Also fixes the Traditional Chinese code, which used a U+2011 non-breaking hyphen. Nothing compares against the ASCII spelling today so it was latent, but it would silently defeat the first thing that did. ## Validation Verified end to end against a live Prowlarr and MyAnonamouse, not just unit tests. A real search returning both an English and a Swedish edition, through the actual `queue_release` → `DownloadTask` → naming path: ``` STEP 1 real MAM search -> 37 releases, languages: ['en', 'sv'] STEP 3 queue_release -> task.language='sv' STEP 4 build_metadata_dict -> metadata['Language']='sv' STEP 5 build_library_path -> /library/J K Rowling/Harry Potter (sv)/... two language editions resolve to DIFFERENT folders: True ``` The refactor is pinned by a snapshot of both per-source maps taken *before* they were deleted. All 131 aliases are asserted to still resolve to the same code, one parametrised test each, so a regression names the specific alias. Also verified: the filename-only template, the retry round-trip (`serialize_task_for_retry` → `_restore_task_from_retry_payload`, plus a legacy payload with no `language` key), and placeholder handling. Added a `KNOWN_TOKENS` ordering invariant test — `find_placeholder()` does a substring `.find()` in list order and nothing protected that contract, so a future token in the wrong position could silently shadow an existing one. And a lockstep guard on the frontend, since `KNOWN_TOKENS` is hand-duplicated in TypeScript. **One caveat worth stating.** Three MAM codes are confirmed by observation (`ENG`→`en`, `SWE`→`sv`, `MAL`→`ml`, the last from a real `[MAL / EPUB]` Tagore release). The remaining ~59 are derived from ISO 639-2 rather than observed, because MAM's catalogue is overwhelmingly English — enabling 27 extra languages still yielded only one non-English hit across 258 results. Mitigated rather than closed: both 639-2 variants are present for every language where they differ, and a wrong alias is an unused entry while a missing one loses the language. Happy to correct any code a maintainer knows differs. ## Test results 2056 Python tests pass (up from 1906). Frontend typecheck, lint, format and 126 unit tests pass. Pre-existing failures on my machine, unchanged by this branch and unrelated: `tests/bypass/` needs `seleniumbase`, and `tests/config/test_entrypoint_permissions.py` uses bash-4 syntax that macOS bash 3.2 rejects. --------- Co-authored-by: delize <4028612+delize@users.noreply.github.com> Co-authored-by: CaliBrain <calibrain@l4n.xyz>
281 lines
5.7 KiB
JSON
281 lines
5.7 KiB
JSON
{
|
||
"_comment": "Frozen snapshot of the per-source language handling as it was before consolidation into shelfmark.core.languages. prowlarr_three_letter and audiobookbay_names were explicit maps; direct_download_derived is what its loader built from the data file, including the underscore spelling of a hyphenated code. Every alias here must still resolve to the same code. The one deliberate change: Traditional Chinese was canonically 'zh‑Hant' with a U+2011 non-breaking hyphen and is now the ASCII 'zh-Hant', so expectations for it name the new code while the old spelling remains a resolvable alias.",
|
||
"audiobookbay_names": {
|
||
"afrikaans": "af",
|
||
"arabic": "ar",
|
||
"bangla": "bn",
|
||
"bengali": "bn",
|
||
"bosnian": "bs",
|
||
"bulgarian": "bg",
|
||
"burmese": "my",
|
||
"catalan": "ca",
|
||
"chinese": "zh",
|
||
"croatian": "hr",
|
||
"czech": "cs",
|
||
"danish": "da",
|
||
"dutch": "nl",
|
||
"english": "en",
|
||
"estonian": "et",
|
||
"farsi": "fa",
|
||
"filipino": "fil",
|
||
"finnish": "fi",
|
||
"french": "fr",
|
||
"german": "de",
|
||
"greek": "el",
|
||
"gujarati": "gu",
|
||
"hebrew": "he",
|
||
"hindi": "hi",
|
||
"hungarian": "hu",
|
||
"icelandic": "is",
|
||
"indonesian": "id",
|
||
"irish": "ga",
|
||
"italian": "it",
|
||
"japanese": "ja",
|
||
"javanese": "jv",
|
||
"kannada": "kn",
|
||
"korean": "ko",
|
||
"latin": "la",
|
||
"latvian": "lv",
|
||
"lithuanian": "lt",
|
||
"malay": "ms",
|
||
"malayalam": "ml",
|
||
"manx": "gv",
|
||
"marathi": "mr",
|
||
"norwegian": "no",
|
||
"persian": "fa",
|
||
"polish": "pl",
|
||
"portuguese": "pt",
|
||
"punjabi": "pa",
|
||
"romanian": "ro",
|
||
"russian": "ru",
|
||
"sanskrit": "sa",
|
||
"scottish gaelic": "gd",
|
||
"serbian": "sr",
|
||
"slovenian": "sl",
|
||
"spanish": "es",
|
||
"swedish": "sv",
|
||
"tagalog": "fil",
|
||
"tamil": "ta",
|
||
"telugu": "te",
|
||
"thai": "th",
|
||
"turkish": "tr",
|
||
"ukrainian": "uk",
|
||
"urdu": "ur",
|
||
"vietnamese": "vi"
|
||
},
|
||
"direct_download_derived": {
|
||
"af": "af",
|
||
"afrikaans": "af",
|
||
"albanian": "sq",
|
||
"ar": "ar",
|
||
"arabic": "ar",
|
||
"armenian": "hy",
|
||
"az": "az",
|
||
"azerbaijani": "az",
|
||
"ba": "ba",
|
||
"bangla": "bn",
|
||
"bashkir": "ba",
|
||
"be": "be",
|
||
"belarusian": "be",
|
||
"bg": "bg",
|
||
"bn": "bn",
|
||
"bo": "bo",
|
||
"bulgarian": "bg",
|
||
"ca": "ca",
|
||
"catalan": "ca",
|
||
"chinese": "zh",
|
||
"croatian": "hr",
|
||
"cs": "cs",
|
||
"czech": "cs",
|
||
"da": "da",
|
||
"danish": "da",
|
||
"de": "de",
|
||
"dutch": "nl",
|
||
"el": "el",
|
||
"en": "en",
|
||
"english": "en",
|
||
"eo": "eo",
|
||
"es": "es",
|
||
"esperanto": "eo",
|
||
"fa": "fa",
|
||
"fi": "fi",
|
||
"fil": "fil",
|
||
"filipino": "fil",
|
||
"finnish": "fi",
|
||
"fr": "fr",
|
||
"french": "fr",
|
||
"ga": "ga",
|
||
"galician": "gl",
|
||
"georgian": "ka",
|
||
"german": "de",
|
||
"gl": "gl",
|
||
"greek": "el",
|
||
"gu": "gu",
|
||
"gujarati": "gu",
|
||
"he": "he",
|
||
"hebrew": "he",
|
||
"hi": "hi",
|
||
"hindi": "hi",
|
||
"hr": "hr",
|
||
"hu": "hu",
|
||
"hungarian": "hu",
|
||
"hy": "hy",
|
||
"id": "id",
|
||
"indonesian": "id",
|
||
"irish": "ga",
|
||
"it": "it",
|
||
"italian": "it",
|
||
"ja": "ja",
|
||
"japanese": "ja",
|
||
"javanese": "jv",
|
||
"jv": "jv",
|
||
"ka": "ka",
|
||
"kannada": "kn",
|
||
"kazakh": "kk",
|
||
"kinyarwanda": "rw",
|
||
"kk": "kk",
|
||
"kn": "kn",
|
||
"ko": "ko",
|
||
"korean": "ko",
|
||
"ky": "ky",
|
||
"kyrgyz": "ky",
|
||
"la": "la",
|
||
"latin": "la",
|
||
"latvian": "lv",
|
||
"lithuanian": "lt",
|
||
"lt": "lt",
|
||
"lv": "lv",
|
||
"malay": "ms",
|
||
"malayalam": "ml",
|
||
"marathi": "mr",
|
||
"ml": "ml",
|
||
"mn": "mn",
|
||
"mongolian": "mn",
|
||
"mr": "mr",
|
||
"ms": "ms",
|
||
"nl": "nl",
|
||
"no": "no",
|
||
"norwegian": "no",
|
||
"pa": "pa",
|
||
"persian": "fa",
|
||
"pl": "pl",
|
||
"polish": "pl",
|
||
"portuguese": "pt",
|
||
"pt": "pt",
|
||
"punjabi": "pa",
|
||
"qu": "qu",
|
||
"quechua": "qu",
|
||
"ro": "ro",
|
||
"romanian": "ro",
|
||
"ru": "ru",
|
||
"russian": "ru",
|
||
"rw": "rw",
|
||
"serbian": "sr",
|
||
"shan": "shn",
|
||
"shn": "shn",
|
||
"sk": "sk",
|
||
"sl": "sl",
|
||
"slovak": "sk",
|
||
"slovenian": "sl",
|
||
"spanish": "es",
|
||
"sq": "sq",
|
||
"sr": "sr",
|
||
"sv": "sv",
|
||
"sw": "sw",
|
||
"swahili": "sw",
|
||
"swedish": "sv",
|
||
"ta": "ta",
|
||
"tamil": "ta",
|
||
"te": "te",
|
||
"telugu": "te",
|
||
"th": "th",
|
||
"thai": "th",
|
||
"tibetan": "bo",
|
||
"tr": "tr",
|
||
"traditional chinese": "zh-Hant",
|
||
"turkish": "tr",
|
||
"ug": "ug",
|
||
"uk": "uk",
|
||
"ukrainian": "uk",
|
||
"ur": "ur",
|
||
"urdu": "ur",
|
||
"uyghur": "ug",
|
||
"vi": "vi",
|
||
"vietnamese": "vi",
|
||
"zh": "zh",
|
||
"zh‑hant": "zh-Hant"
|
||
},
|
||
"prowlarr_three_letter": {
|
||
"afr": "af",
|
||
"ara": "ar",
|
||
"ben": "bn",
|
||
"bos": "bs",
|
||
"bul": "bg",
|
||
"bur": "my",
|
||
"cat": "ca",
|
||
"ces": "cs",
|
||
"chi": "zh",
|
||
"cze": "cs",
|
||
"dan": "da",
|
||
"deu": "de",
|
||
"dut": "nl",
|
||
"ell": "el",
|
||
"eng": "en",
|
||
"est": "et",
|
||
"fas": "fa",
|
||
"fin": "fi",
|
||
"fra": "fr",
|
||
"fre": "fr",
|
||
"ger": "de",
|
||
"gla": "gd",
|
||
"gle": "ga",
|
||
"glv": "gv",
|
||
"gre": "el",
|
||
"guj": "gu",
|
||
"heb": "he",
|
||
"hin": "hi",
|
||
"hrv": "hr",
|
||
"hun": "hu",
|
||
"ice": "is",
|
||
"ind": "id",
|
||
"isl": "is",
|
||
"ita": "it",
|
||
"jap": "ja",
|
||
"jav": "jv",
|
||
"jpn": "ja",
|
||
"kan": "kn",
|
||
"kor": "ko",
|
||
"lat": "la",
|
||
"lav": "lv",
|
||
"lit": "lt",
|
||
"mal": "ml",
|
||
"mar": "mr",
|
||
"may": "ms",
|
||
"msa": "ms",
|
||
"mya": "my",
|
||
"nld": "nl",
|
||
"nor": "no",
|
||
"pan": "pa",
|
||
"per": "fa",
|
||
"pol": "pl",
|
||
"por": "pt",
|
||
"rom": "ro",
|
||
"ron": "ro",
|
||
"rus": "ru",
|
||
"san": "sa",
|
||
"slv": "sl",
|
||
"spa": "es",
|
||
"srp": "sr",
|
||
"swe": "sv",
|
||
"tam": "ta",
|
||
"tel": "te",
|
||
"tgl": "fil",
|
||
"tha": "th",
|
||
"tur": "tr",
|
||
"ukr": "uk",
|
||
"urd": "ur",
|
||
"vie": "vi",
|
||
"zho": "zh"
|
||
}
|
||
}
|