mirror of
https://github.com/calibrain/shelfmark.git
synced 2026-10-03 12:15:47 +01:00
Per discussion #1372, where you said you were fine with this specific implementation: check whether metadata.db exists, read it if so, and show a check mark saying the book is already there. Searching for a book you already own gives no hint that you own it, so the easiest way to end up with a second copy is to not remember you have the first. This reads a Calibre `metadata.db`, read only, and marks matching results with an **In library** badge in the card, list and compact views and in the details dialog. Off by default. It sits in Settings, General beside the existing Library URL, with a test button that reports how many books it indexed. No HTTP call, no token, nothing written back. Matching runs most to least confident: a shared external id, then an ISBN compared in both ISBN-10 and ISBN-13 form, then fuzzy title tokens plus the author surname. The check fails open, so an unreadable database degrades the badge and never blocks a search, and entries are cached for ten minutes with an early refresh when the file changes, so a large library costs one read rather than one per search. `text_match.py` is new and shared by the index and the provider, so title, author and ISBN matching stays consistent in one place. ## On the provider interface `library_index` talks only to a `LibraryProvider` protocol and knows nothing about Calibre. That is deliberate but it is not speculative generality, it is what let me send you the Calibre half on its own: I run an Audiobookshelf provider on the same interface in my fork, which is where the audiobook side of the badge comes from. I have left that out because it is a new service integration rather than something already in the codebase, which is the line your non-goals draw. Happy to send it separately if you ever want it, and equally happy for the answer to be no. Adding a library is a module with the `LibraryProvider` shape plus one line in `all_providers()`. ## Verification - `tests/core/test_library_index.py`: id, ISBN and fuzzy matching, per-content-type provider selection, fail-open on provider errors, stale-cache reuse, TTL and fingerprint refresh, per-provider cache isolation, and the test-connection path including unsaved form values. - `tests/core/test_text_match.py`: ISBN variants and token matching. - `src/frontend/src/tests/libraryBadge.test.ts` and the added cases in `bookTransformers.test.ts`. - Python suite (3269) and frontend suite (206) green, plus ruff, ruff format, basedpyright, vulture, tsc, oxlint, oxfmt and the production build.
78 lines
2.3 KiB
Python
78 lines
2.3 KiB
Python
"""Tests for the shared token/ISBN/surname matching helpers."""
|
|
|
|
import pytest
|
|
|
|
from shelfmark.core import text_match
|
|
|
|
_TWENTY_WORDS = " ".join(f"word{i:02d}" for i in range(20))
|
|
|
|
|
|
def test_tokens_lowercases_and_splits_on_non_alphanumerics():
|
|
assert text_match.tokens("Dungeon Crawler Carl: Book 1!") == [
|
|
"dungeon",
|
|
"crawler",
|
|
"carl",
|
|
"book",
|
|
"1",
|
|
]
|
|
|
|
|
|
@pytest.mark.parametrize("value", [None, "", " --- "])
|
|
def test_tokens_empty_input(value):
|
|
assert text_match.tokens(value) == []
|
|
|
|
|
|
def test_significant_tokens_drops_stopwords_and_single_characters():
|
|
assert text_match.significant_tokens("The Way of Kings, Part A 2") == ["way", "kings", "part"]
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
("author", "expected"),
|
|
[
|
|
("Sanderson, Brandon", "sanderson"),
|
|
("Brandon Sanderson", "sanderson"),
|
|
("Dinniman, Matt J.", "dinniman"),
|
|
("", None),
|
|
(None, None),
|
|
("A", None),
|
|
],
|
|
)
|
|
def test_author_surname(author, expected):
|
|
assert text_match.author_surname(author) == expected
|
|
|
|
|
|
def test_title_tokens_match_default_threshold_edges():
|
|
words = _TWENTY_WORDS.split()
|
|
|
|
assert text_match.title_tokens_match(_TWENTY_WORDS, set(words[:17])) is True # 17/20 == 0.85
|
|
assert text_match.title_tokens_match(_TWENTY_WORDS, set(words[:16])) is False
|
|
|
|
|
|
def test_title_tokens_match_explicit_threshold_is_inclusive():
|
|
assert text_match.title_tokens_match("alpha beta", {"alpha"}, threshold=0.5) is True
|
|
assert text_match.title_tokens_match("alpha beta", {"alpha"}, threshold=0.51) is False
|
|
|
|
|
|
def test_title_tokens_match_ignores_stopwords_in_title():
|
|
assert text_match.title_tokens_match("The Name of the Wind", {"name", "wind"}) is True
|
|
|
|
|
|
@pytest.mark.parametrize("title", [None, "", "the of a"])
|
|
def test_title_tokens_match_without_significant_tokens_is_false(title):
|
|
assert text_match.title_tokens_match(title, {"the", "of"}) is False
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
("value", "expected"),
|
|
[
|
|
("978-0-59-382024-7", "9780593820247"),
|
|
("0-306-40615-x", "030640615X"),
|
|
(" 0306406152 ", "0306406152"),
|
|
(9780593820247, "9780593820247"),
|
|
(None, ""),
|
|
("", ""),
|
|
],
|
|
)
|
|
def test_normalize_isbn(value, expected):
|
|
assert text_match.normalize_isbn(value) == expected
|