Files
shelfmark/tests/core/test_text_match.py
T
splitsec2 37a77e9562 feat(library): mark search results already in a Calibre library (#1377)
Per discussion #1372, where you said you were fine with this specific
implementation: check whether metadata.db exists, read it if so, and
show a check mark saying the book is already there.

Searching for a book you already own gives no hint that you own it, so
the easiest way to end up with a second copy is to not remember you have
the first. This reads a Calibre `metadata.db`, read only, and marks
matching results with an **In library** badge in the card, list and
compact views and in the details dialog.

Off by default. It sits in Settings, General beside the existing Library
URL, with a test button that reports how many books it indexed. No HTTP
call, no token, nothing written back.

Matching runs most to least confident: a shared external id, then an
ISBN compared in both ISBN-10 and ISBN-13 form, then fuzzy title tokens
plus the author surname. The check fails open, so an unreadable database
degrades the badge and never blocks a search, and entries are cached for
ten minutes with an early refresh when the file changes, so a large
library costs one read rather than one per search.

`text_match.py` is new and shared by the index and the provider, so
title, author and ISBN matching stays consistent in one place.

## On the provider interface

`library_index` talks only to a `LibraryProvider` protocol and knows
nothing about Calibre. That is deliberate but it is not speculative
generality, it is what let me send you the Calibre half on its own: I
run an Audiobookshelf provider on the same interface in my fork, which
is where the audiobook side of the badge comes from. I have left that
out because it is a new service integration rather than something
already in the codebase, which is the line your non-goals draw. Happy to
send it separately if you ever want it, and equally happy for the answer
to be no.

Adding a library is a module with the `LibraryProvider` shape plus one
line in `all_providers()`.

## Verification

- `tests/core/test_library_index.py`: id, ISBN and fuzzy matching,
per-content-type provider selection, fail-open on provider errors,
stale-cache reuse, TTL and fingerprint refresh, per-provider cache
isolation, and the test-connection path including unsaved form values.
- `tests/core/test_text_match.py`: ISBN variants and token matching.
- `src/frontend/src/tests/libraryBadge.test.ts` and the added cases in
`bookTransformers.test.ts`.
- Python suite (3269) and frontend suite (206) green, plus ruff, ruff
format, basedpyright, vulture, tsc, oxlint, oxfmt and the production
build.
2026-09-25 18:18:10 -04:00

78 lines
2.3 KiB
Python

"""Tests for the shared token/ISBN/surname matching helpers."""
import pytest
from shelfmark.core import text_match
_TWENTY_WORDS = " ".join(f"word{i:02d}" for i in range(20))
def test_tokens_lowercases_and_splits_on_non_alphanumerics():
assert text_match.tokens("Dungeon Crawler Carl: Book 1!") == [
"dungeon",
"crawler",
"carl",
"book",
"1",
]
@pytest.mark.parametrize("value", [None, "", " --- "])
def test_tokens_empty_input(value):
assert text_match.tokens(value) == []
def test_significant_tokens_drops_stopwords_and_single_characters():
assert text_match.significant_tokens("The Way of Kings, Part A 2") == ["way", "kings", "part"]
@pytest.mark.parametrize(
("author", "expected"),
[
("Sanderson, Brandon", "sanderson"),
("Brandon Sanderson", "sanderson"),
("Dinniman, Matt J.", "dinniman"),
("", None),
(None, None),
("A", None),
],
)
def test_author_surname(author, expected):
assert text_match.author_surname(author) == expected
def test_title_tokens_match_default_threshold_edges():
words = _TWENTY_WORDS.split()
assert text_match.title_tokens_match(_TWENTY_WORDS, set(words[:17])) is True # 17/20 == 0.85
assert text_match.title_tokens_match(_TWENTY_WORDS, set(words[:16])) is False
def test_title_tokens_match_explicit_threshold_is_inclusive():
assert text_match.title_tokens_match("alpha beta", {"alpha"}, threshold=0.5) is True
assert text_match.title_tokens_match("alpha beta", {"alpha"}, threshold=0.51) is False
def test_title_tokens_match_ignores_stopwords_in_title():
assert text_match.title_tokens_match("The Name of the Wind", {"name", "wind"}) is True
@pytest.mark.parametrize("title", [None, "", "the of a"])
def test_title_tokens_match_without_significant_tokens_is_false(title):
assert text_match.title_tokens_match(title, {"the", "of"}) is False
@pytest.mark.parametrize(
("value", "expected"),
[
("978-0-59-382024-7", "9780593820247"),
("0-306-40615-x", "030640615X"),
(" 0306406152 ", "0306406152"),
(9780593820247, "9780593820247"),
(None, ""),
("", ""),
],
)
def test_normalize_isbn(value, expected):
assert text_match.normalize_isbn(value) == expected