mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-09-28 22:06:05 +01:00
First of all, I don't know if you even want to merge a scraper-based metadata provider. I made this just for my use-case. If you'd rather not, I completely understand it. An alternative would be adopting [Audiobookshelf's Metadata Provider API](https://audiobookshelf.org/docs/documentation/community/community-providers) which I contributed to it for exactly the reason to not have scrapers. ## What Adds [Moly.hu](https://moly.hu) — the Hungarian community book catalog — as a metadata provider, following the existing provider plugin architecture (`@register_provider` + settings tab with enable checkbox and Test Connection button, disabled by default). ## Why None of the current providers cover Hungarian editions well: Hardcover and Open Library rarely index them, and Google Books coverage is spotty. Moly.hu is the de-facto catalog for Hungarian books (local editions *and* Hungarian translations of foreign works). With this provider, Universal mode works end-to-end for Hungarian titles: moly search → localized title/author feed the release search → indexers that carry Hungarian content can actually match. Related pain points: #595 (books missing from metadata providers), #1035 (interest in niche sources). ## How - HTML scraping with BeautifulSoup (already a dependency), no API key needed - Scraping approach (search URL, page structure, language-tag mapping) adapted from the long-lived Calibre `Moly_hu` plugin (GPL v3, credited in the module docstring), with fallback selector chains inherited from it - Sliding-window rate limit (30 req/min) to stay polite to a small community site - Standard `@cacheable` decorators; fetch failures return `None` so they are not cached (same behavior as the Google Books provider) - Search results carry cover thumbnails, rating and series info as display fields; `get_book` parses title (zero-width chars stripped, nested series link excluded), authors, ISBN-13/10, publisher, publish year, description (spoiler-warning prefix stripped), tags/genres, cover, and language (from moly's language tags, defaulting to `hu`) - ISBN search resolves through moly's site search ## Testing - `tests/metadata/test_moly_parse.py`: offline tests with fixture HTML mirroring live moly.hu markup — search parsing/dedup, pagination guard, failure-not-cached behavior, book-page parsing, ISBN resolution, ISBN validation helper - `uv run pytest tests/metadata` green (45 passed), `ruff check` / `ruff format` clean - Verified live against moly.hu (search, get_book, ISBN lookup) and running in Docker alongside Hardcover --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>