fix(irc): search by surname, and rank the answer by author (#1331) (#1332)

Fixes #1331.

A search bot ANDs every term against a filename, so the given name is
the term
that empties the result set. Measured against irchighway's #ebooks:
"Revelations
David Petrie" is answered "no results", "Revelations Petrie" returns 9
matches,
6 of which parse, all filed as "D Petrie".

The query now carries the title and the surname, read off the search
variant so
the ISBN fallback and a manual query - which set author="" on purpose -
keep
their current shape.

Title-only, the shape #1295 settled on for Prowlarr, does not transfer:
the bot
caps an answer at 1000 matches, and a bare "Revelations" hits that cap
with 923
parsed rows across 500 authors, so the cap itself can drop the wanted
book. The
surname is the token the two spellings share and it keeps the answer
small.

The full author then orders what comes back, reusing author_affinity
from #1295,
since a surname also matches a different author who shares it. It sits
under
server availability the way indexer priority does in #1295: a download
addresses
one named bot and waits 120s for it, so a match from a bot that has left
the
channel must not outrank a mismatch that can answer. Ranking runs on the
way out
rather than before the cache, because one query identity is shared by
every book
that produced that query.

Two things found while testing:

- The parser writes the literal "Unknown" when a filename has no " - "
separator
  (parser.py:168). Ranked literally that sorts as a wrong author, so
author_affinity's middle tier was unreachable here; it is now read as
absent.
  5 of those 923 rows are affected.
- author_affinity moves to shelfmark/core/author_match.py, unchanged, so
IRC
does not import from the Prowlarr package. Prowlarr behaviour is
untouched and
  its tests pass as they are.

The three IRC assertions in the #1252 regression file move to the
surname form.
The invariant they pin - one contributor's name reaches the query, never
the
whole credit list - is unchanged.

Tested with make python-lint, python-format, python-dead-code,
python-typecheck
and python-test, and end to end against irchighway with the patched
source: it
posts "Revelations Petrie" and returns 6 releases.
This commit is contained in:
Zoltán Szabó
2026-09-11 22:14:37 -04:00
committed by GitHub
parent 8c902d7f7a
commit 35037b35fd
8 changed files with 391 additions and 93 deletions
+18 -6
View File
@@ -7,6 +7,18 @@ from shelfmark.release_sources.irc.parser import SearchResult
from shelfmark.release_sources.irc.source import IRCReleaseSource
def _plan(title, author="", manual_query=None):
"""A plan stub for tests that only need one to reach search().
Tests that care how a plan is built use build_release_search_plan instead.
"""
return SimpleNamespace(
title_variants=[SimpleNamespace(title=title, author=author)],
author=author,
manual_query=manual_query,
)
def test_convert_to_releases_marks_audiobook_results_and_sorts_audio_before_archives():
source = IRCReleaseSource()
source._online_servers = set()
@@ -73,7 +85,7 @@ def test_search_uses_cached_results_without_opening_a_connection(monkeypatch):
monkeypatch.setattr(irc_source, "_emit_status", lambda *_args, **_kwargs: None)
book = BookMetadata(provider="hardcover", provider_id="123", title="Cached Book")
plan = SimpleNamespace(primary_query="Cached Book")
plan = _plan("Cached Book")
releases = source.search(book, plan)
@@ -144,7 +156,7 @@ def test_search_no_dcc_offer_releases_connection_and_caches_empty_result(monkeyp
)
book = BookMetadata(provider="hardcover", provider_id="abc", title="Missing Result")
plan = SimpleNamespace(primary_query="Missing Result")
plan = _plan("Missing Result")
releases = source.search(book, plan, content_type="audiobook")
@@ -225,7 +237,7 @@ def test_audiobook_search_routes_to_configured_audiobook_channel_and_bot(monkeyp
)
book = BookMetadata(provider="hardcover", provider_id="ab", title="Audio Book")
plan = SimpleNamespace(primary_query="Audio Book")
plan = _plan("Audio Book")
source.search(book, plan, content_type="audiobook")
@@ -292,7 +304,7 @@ def test_audiobook_search_reuses_main_bot_when_only_channel_configured(monkeypat
)
book = BookMetadata(provider="hardcover", provider_id="ab2", title="Audio Book")
plan = SimpleNamespace(primary_query="Audio Book")
plan = _plan("Audio Book")
source.search(book, plan, content_type="audiobook")
@@ -332,7 +344,7 @@ def test_search_without_search_bot_never_posts_to_channel(monkeypatch):
)
book = BookMetadata(provider="hardcover", provider_id="nobot", title="No Bot")
plan = SimpleNamespace(primary_query="No Bot")
plan = _plan("No Bot")
assert source.search(book, plan) == []
@@ -400,7 +412,7 @@ def test_search_send_budget_blocks_repost_and_returns_cache(monkeypatch):
irc_source._record_message_sent(send_key)
book = BookMetadata(provider="hardcover", provider_id="cd", title="Budget Book")
plan = SimpleNamespace(primary_query="Budget Book")
plan = _plan("Budget Book")
# expand_search=True bypasses the top-level cache, forcing the budget path.
releases = source.search(book, plan, expand_search=True)