Backport bug fixes from NemesisHubris/litfinder (addresses #999, #956, #1010, #1021, #1025, #1040) (#1066)

## Backport bug fixes from `NemesisHubris/litfinder`

Forwards a curated set of bug fixes from
[NemesisHubris/litfinder](https://github.com/NemesisHubris/litfinder) —
a community fork of this project — that address open issues here. All
commits preserve original authorship via `git cherry-pick`; this PR is a
backport rather than original work. Each fix has been reviewed locally,
lint/format-cleaned to match this repo's existing ruff config, and
verified with the test suite. Rebrand strings, license switches, and
features have been deliberately excluded.

### Upstream issues addressed

- **#999** — Mirror URLs with query params no longer break search
requests (strip query string/fragment in `normalize_http_url`)
- **#956** — Apprise notifications now respect the configured proxy
(proxy env vars injected before dispatch)
- **#1025** — rTorrent: separate `RTORRENT_AUDIOBOOK_LABEL` setting,
falls back to book label if unset
- **#1010** — Stop button in Activity no longer makes the panel
disappear (snapshot refresh on cancel)
- **#1021** — Anna's Archive slow-download countdown now caps retries
instead of looping forever
- **#1040** — Empty destination directory cleaned up when write probe
fails
- **PR #1031** — Language detection from Anna's Archive distant path
when listing metadata is missing

### Additional fixes (no open issue but clear bugs)

- **fix: Python 2 `except` syntax across 27 files** — `except X, Y:` is
a SyntaxError in Python 3 and prevents affected modules from importing
at runtime. Mechanical sweep to `except (X, Y):`.
- **fix(abb): info hash validation with magnet fallback** — adds
SHA-1/SHA-256 hex validation on extracted info hashes; falls back to
scanning the full page for a magnet link (e.g. posted in comments) when
the table value is malformed. Also extends the exact-phrase fallback to
manual queries and defaults the ABB listing language to `en` when
missing, preventing valid results from being hidden by the language
filter. Includes a small test-fixture fix (`test(abb): use valid hex
info hashes in scraper test fixtures`) since the existing fixtures used
non-hex placeholders that the new validation correctly rejects.
- **fix: Anna's Archive title parser** — handles nested edition spans
and filters `lgli` catalog descriptor entries (e.g. "Book/Online Audio")
that were polluting search results.

### Deliberately not included

- LitFinder rebranding (UI strings, Apprise app ID, logo). The `fix:
three upstream bugs` commit (#999/#956/#1025) was cherry-picked with
Apprise app-id, description, and logo-URL strings reverted from
"LitFinder" back to "Shelfmark"; noted in the commit body.
- Features from the LitFinder fork (multi-variant title search,
multi-book flat-folder grouping, fuzzy text matching, "Leave in Place"
output handler, admin display name, custom-source plugin system). These
are larger behavior changes that each warrant their own focused review —
happy to send any of them separately if of interest.
- LitFinder-specific test environment and CI infrastructure.

### Verification

- Backend: **1879 passed**, 96 skipped (1 preexisting failure on
`seleniumbase`-dependent test in local venv; runs fine in the standard
Docker image with the `browser` extra)
- Lint, format, dead-code: all clean against this repo's existing
ruff/vulture config
- One follow-up cleanup commit (`style: ruff lint and format fixes for
ported commits`) brings the cherry-picked code into compliance with this
repo's ruff settings — no behavior changes there

### Etiquette / credit

Per-commit authorship preserved by cherry-pick. The only edits to the
original commits are:
- `fix: three upstream bugs` — Apprise rebrand strings reverted to
"Shelfmark" (noted in commit body, original author retained as
`Co-Authored-By` via cherry-pick)
- One follow-up `style:` commit for ruff config alignment

Big thanks to [@NemesisHubris](https://github.com/NemesisHubris) for the
original work in LitFinder; this PR exists to make sure these fixes
reach Shelfmark's wider user base. Happy to revise scope, split into
smaller PRs, or split off the Py2 cleanup separately if that's
preferable.

---------

Co-authored-by: NemesisHubris <155838970+NemesisHubris@users.noreply.github.com>
Co-authored-by: CaliBrain <calibrain@l4n.xyz>
This commit is contained in:
spindrift
2026-06-15 00:28:43 -04:00
committed by GitHub
co-authored by NemesisHubris CaliBrain
parent 6231678c28
commit aba1a68dda
18 changed files with 821 additions and 31 deletions
+5 -5
View File
@@ -58,7 +58,7 @@ SAMPLE_DETAIL_HTML = """
<table>
<tr>
<td>Info Hash</td>
<td>ABC123DEF456GHI789JKL012MNO345PQR678STU</td>
<td>ABC123DEF456789012345678901234567890ABCD</td>
</tr>
<tr>
<td>Tracker 1</td>
@@ -83,7 +83,7 @@ DETAIL_HTML_NO_TRACKERS = """
<table>
<tr>
<td>Info Hash</td>
<td>ABC123DEF456GHI789JKL012MNO345PQR678STU</td>
<td>ABC123DEF456789012345678901234567890ABCD</td>
</tr>
</table>
</body>
@@ -360,7 +360,7 @@ class TestExtractMagnetLink:
assert magnet_link is not None
assert magnet_link.startswith("magnet:?xt=urn:btih:")
assert "ABC123DEF456GHI789JKL012MNO345PQR678STU" in magnet_link
assert "ABC123DEF456789012345678901234567890ABCD" in magnet_link
assert "udp%3A//tracker.openbittorrent.com%3A80" in magnet_link
assert "http%3A//tracker.example.com%3A8080" in magnet_link
assert mock_html_get.call_count == 2
@@ -395,7 +395,7 @@ class TestExtractMagnetLink:
assert magnet_link is not None
assert magnet_link.startswith("magnet:?xt=urn:btih:")
assert "ABC123DEF456GHI789JKL012MNO345PQR678STU" in magnet_link
assert "ABC123DEF456789012345678901234567890ABCD" in magnet_link
# Should contain default trackers
assert "udp%3A//tracker.openbittorrent.com%3A80" in magnet_link
@@ -441,7 +441,7 @@ class TestExtractMagnetLink:
<table>
<tr>
<td>Info Hash</td>
<td>ABC 123 DEF 456</td>
<td>ABC 123 DEF 456 789 012 345 678 901 234 567 890 ABC D</td>
</tr>
</table>
</body>
+14
View File
@@ -368,6 +368,20 @@ def test_download_source_settings_include_direct_download_toggle():
assert "Add your own mirror URLs" in toggle_field.description
def test_download_source_settings_include_distant_path_language_toggle():
from shelfmark.config.settings import download_source_settings
fields = download_source_settings()
toggle_field = next(
field
for field in fields
if getattr(field, "key", None) == "DIRECT_DOWNLOAD_LANGUAGE_FROM_PATH"
)
assert toggle_field.default is False
assert "distant path" in toggle_field.description.lower()
def test_fast_source_options_lock_entries_without_mirror_or_donator_requirements(monkeypatch):
from shelfmark.config.settings import _get_fast_source_options
+69
View File
@@ -597,3 +597,72 @@ def test_resolve_user_routes_expands_multiselect_event_rows(monkeypatch):
{"event": "request_fulfilled", "url": "ntfys://ntfy.sh/user-main"},
{"event": "all", "url": "ntfys://ntfy.sh/user-all"},
]
class TestAppriseProxyEnv:
"""Regression tests for issue #956 — proxy settings ignored for notifications."""
def _patch_config(self, monkeypatch, values):
from shelfmark.core import config as config_module
def _fake_get(key, default="", **_kwargs):
return values.get(key, default)
monkeypatch.setattr(config_module.config, "get", _fake_get)
def test_http_proxy_mode_injects_proxy_env(self, monkeypatch):
self._patch_config(
monkeypatch,
{
"PROXY_MODE": "http",
"HTTP_PROXY": "http://proxy.example.com:8080",
"HTTPS_PROXY": "",
"NO_PROXY": "",
},
)
monkeypatch.delenv("HTTP_PROXY", raising=False)
monkeypatch.delenv("HTTPS_PROXY", raising=False)
result = notifications_module._apprise_proxy_env()
assert result["HTTP_PROXY"] == "http://proxy.example.com:8080"
assert result["HTTPS_PROXY"] == "http://proxy.example.com:8080"
def test_socks5_proxy_mode_injects_socks_env(self, monkeypatch):
self._patch_config(
monkeypatch,
{
"PROXY_MODE": "socks5",
"SOCKS5_PROXY": "socks5://proxy.example.com:1080",
"NO_PROXY": "",
},
)
monkeypatch.delenv("HTTP_PROXY", raising=False)
monkeypatch.delenv("HTTPS_PROXY", raising=False)
result = notifications_module._apprise_proxy_env()
assert result["HTTP_PROXY"] == "socks5://proxy.example.com:1080"
assert result["HTTPS_PROXY"] == "socks5://proxy.example.com:1080"
def test_no_proxy_mode_returns_empty_dict(self, monkeypatch):
self._patch_config(monkeypatch, {"PROXY_MODE": ""})
result = notifications_module._apprise_proxy_env()
assert result == {}
def test_does_not_override_already_set_env_vars(self, monkeypatch):
self._patch_config(
monkeypatch,
{
"PROXY_MODE": "http",
"HTTP_PROXY": "http://new-proxy.example.com:8080",
"NO_PROXY": "",
},
)
monkeypatch.setenv("HTTP_PROXY", "http://existing-proxy.example.com:3128")
result = notifications_module._apprise_proxy_env()
assert "HTTP_PROXY" not in result
+25
View File
@@ -5,6 +5,31 @@ import types
import xmlrpc.client as stdlib_xmlrpc_client
from shelfmark.core import utils
from shelfmark.core.utils import normalize_http_url
class TestNormalizeHttpUrlQueryStripping:
"""Regression tests for issue #999 — mirror URLs with query params/fragments."""
def test_strips_query_string_from_configured_url(self) -> None:
result = normalize_http_url("http://mirror.example.com/search?token=abc123")
assert result == "http://mirror.example.com/search"
def test_strips_fragment_from_configured_url(self) -> None:
result = normalize_http_url("http://mirror.example.com/search#section")
assert result == "http://mirror.example.com/search"
def test_strips_both_query_and_fragment(self) -> None:
result = normalize_http_url("https://mirror.example.com/path?key=val&x=1#top")
assert result == "https://mirror.example.com/path"
def test_plain_url_unchanged(self) -> None:
result = normalize_http_url("http://mirror.example.com/search")
assert result == "http://mirror.example.com/search"
def test_trailing_slash_still_stripped_after_query_removal(self) -> None:
result = normalize_http_url("http://mirror.example.com/?token=x")
assert result == "http://mirror.example.com"
def test_get_hardened_xmlrpc_client_tolerates_patch_runtime_error(monkeypatch) -> None:
@@ -182,3 +182,170 @@ class TestDirectDownloadSearchQueries:
("mistborn custom query", ["en"], ["epub"]),
("mistborn custom query", None, ["epub"]),
]
# --- Distant-path language detection tests ---
def _patch_path_language(monkeypatch, enabled: bool = True):
import shelfmark.release_sources.direct_download as dd
original_get = dd.config.get
def _fake_get(key: str, default=None, user_id=None):
del user_id
if key == "DIRECT_DOWNLOAD_LANGUAGE_FROM_PATH":
return enabled
return original_get(key, default)
monkeypatch.setattr(dd.config, "get", _fake_get)
return dd
def _row_from_html(html: str):
from bs4 import BeautifulSoup
return BeautifulSoup(html, "html.parser").find("tr")
def _make_row(distant_path: str, language: str = "", record_id: str = "rec-1") -> str:
return rf"""
<tr>
<td><a href="/md5/{record_id}"><img src="cover.jpg"></a></td>
<td><span>A Book Title</span></td>
<td><span>Author Name</span></td>
<td><span>Publisher</span></td>
<td><span>2024</span></td>
<td><span>-</span></td>
<td><span>-</span></td>
<td><span>{language}</span></td>
<td><span>fiction</span></td>
<td><span>epub</span></td>
<td><span>1 mb</span></td>
<td><span>{distant_path}</span></td>
</tr>
"""
def test_detects_bracketed_language_from_distant_path(monkeypatch):
dd = _patch_path_language(monkeypatch)
row = _row_from_html(_make_row(r"lgli/N:\comics1\emule\2021.08.01\[BD FR] Scrameustache.cbz"))
record = dd._parse_search_result_row(row)
assert record is not None
assert record.language == "fr"
assert record.download_path is not None
def test_detects_mixed_case_bracketed_language(monkeypatch):
dd = _patch_path_language(monkeypatch)
row = _row_from_html(_make_row(r"lgli/V:\comics\_0DAY3\[Fr]\BDs [Fr]\!Pdf\S\Book.pdf"))
record = dd._parse_search_result_row(row)
assert record is not None
assert record.language == "fr"
def test_overrides_unknown_language_with_path_detection(monkeypatch):
dd = _patch_path_language(monkeypatch)
row = _row_from_html(_make_row(r"lgli/V:\comics\_0DAY3\[Fr]\Book.pdf", language="unknown"))
record = dd._parse_search_result_row(row)
assert record is not None
assert record.language == "fr"
def test_sets_unknown_when_path_has_no_language(monkeypatch):
dd = _patch_path_language(monkeypatch)
row = _row_from_html(_make_row(r"lgli/N:\comics1\emule\NoLanguageHere.epub"))
record = dd._parse_search_result_row(row)
assert record is not None
assert record.language == "unknown"
def test_avoids_en_false_positive_when_french_present(monkeypatch):
dd = _patch_path_language(monkeypatch)
row = _row_from_html(
_make_row(r"lgli/V:\comics\_0DAY2\Stripboeken Frans - BD en Français\[BD Fr] Book.cbr")
)
record = dd._parse_search_result_row(row)
assert record is not None
assert record.language == "fr"
def test_keeps_row_with_missing_language_when_toggle_disabled(monkeypatch):
dd = _patch_path_language(monkeypatch, enabled=False)
row = _row_from_html(_make_row(r"lgli/N:\comics1\[BD FR] Scrameustache.cbz"))
record = dd._parse_search_result_row(row)
assert record is not None
assert record.language is None
def test_keeps_sparse_lgli_row(monkeypatch):
"""lgli rows missing author/publisher/year must not be dropped."""
dd = _patch_path_language(monkeypatch)
html = r"""
<tr>
<td><a href="/md5/sparse-1"><img src="cover.jpg"></a></td>
<td><span>Gos - 1978 - Le scrameustache T06.cbz</span></td>
<td></td><td></td><td></td><td></td><td></td><td></td>
<td><span>Comic book</span></td>
<td><span>cbz</span></td>
<td><span>17.4MB</span></td>
<td><span>lgli/N:\comics1\ftp\[BD.FR] French Comics\Book.cbz</span></td>
</tr>
"""
record = dd._parse_search_result_row(_row_from_html(html))
assert record is not None
assert record.id == "sparse-1"
assert record.language == "fr"
assert record.author is None
def test_search_books_filters_locally_when_path_language_enabled(monkeypatch):
dd = _patch_path_language(monkeypatch)
monkeypatch.setattr(dd.network, "get_aa_base_url", lambda: "https://mirror.example")
monkeypatch.setattr(dd.network, "AAMirrorSelector", lambda: object())
captured_url: dict[str, str] = {}
def _fake_html_get_page(url: str, selector, allow_bypasser_fallback=False):
del selector, allow_bypasser_fallback
captured_url["url"] = url
return r"""
<table>
<tr>
<td><a href="/md5/rec-fr"><img src="c.jpg"></a></td>
<td><span>Livre FR</span></td><td><span>Auteur</span></td>
<td><span>Editeur</span></td><td><span>2025</span></td>
<td><span>-</span></td><td><span>-</span></td><td></td>
<td><span>fiction</span></td><td><span>pdf</span></td>
<td><span>2 mb</span></td>
<td><span>lgli/V:\comics\_0DAY3\[Fr]\Book FR.pdf</span></td>
</tr>
<tr>
<td><a href="/md5/rec-en"><img src="c.jpg"></a></td>
<td><span>Book EN</span></td><td><span>Author</span></td>
<td><span>Publisher</span></td><td><span>2025</span></td>
<td><span>-</span></td><td><span>-</span></td><td></td>
<td><span>fiction</span></td><td><span>pdf</span></td>
<td><span>2 mb</span></td>
<td><span>lgli/V:\comics\_0DAY3\[En]\Book EN.pdf</span></td>
</tr>
</table>
"""
monkeypatch.setattr(dd.downloader, "html_get_page", _fake_html_get_page)
records = dd.search_books("demo", SearchFilters(lang=["fr"], format=["pdf"]))
assert "&lang=" not in captured_url["url"]
assert len(records) == 1
assert records[0].id == "rec-fr"
assert records[0].language == "fr"
def test_book_matches_requested_languages_logic():
import shelfmark.release_sources.direct_download as dd
assert dd._book_matches_requested_languages(None, {"fr"}) is True
assert dd._book_matches_requested_languages(None, set()) is True
assert dd._book_matches_requested_languages("en", {"fr"}) is False
assert dd._book_matches_requested_languages("fr", {"fr"}) is True
+99
View File
@@ -349,6 +349,105 @@ class TestRTorrentClientAddDownload:
assert "RPC Error" in str(excinfo.value)
class TestRTorrentClientAudiobookLabel:
"""Regression tests for issue #1025 — rTorrent audiobook label selection."""
def _make_client(self, monkeypatch, config_values):
monkeypatch.setattr(
"shelfmark.download.clients.rtorrent.config.get",
make_config_getter(config_values),
)
mock_rpc = MagicMock()
mock_xmlrpc = create_mock_xmlrpc_module()
mock_xmlrpc.ServerProxy.return_value = mock_rpc
mock_torrent_info = MagicMock()
mock_torrent_info.torrent_data = None
mock_torrent_info.magnet_url = "magnet:?xt=urn:btih:abc123"
mock_torrent_info.info_hash = "abc123"
mock_torrent_info.is_magnet = True
return mock_rpc, mock_xmlrpc, mock_torrent_info
def test_uses_audiobook_label_when_content_type_is_audiobook(self, monkeypatch):
config_values = {
"RTORRENT_URL": "http://localhost:8080/RPC2",
"RTORRENT_LABEL": "books",
"RTORRENT_AUDIOBOOK_LABEL": "audiobooks",
"RTORRENT_DOWNLOAD_DIR": "/downloads",
}
mock_rpc, mock_xmlrpc, mock_torrent_info = self._make_client(monkeypatch, config_values)
with patch.dict("sys.modules", {"xmlrpc.client": mock_xmlrpc}):
with patch(
"shelfmark.download.clients.torrent_utils.extract_torrent_info",
return_value=mock_torrent_info,
):
if "shelfmark.download.clients.rtorrent" in sys.modules:
del sys.modules["shelfmark.download.clients.rtorrent"]
from shelfmark.download.clients.rtorrent import RTorrentClient
client = RTorrentClient()
client.add_download(
"magnet:?xt=urn:btih:abc123", "Test Audiobook", content_type="audiobook"
)
args = mock_rpc.load.start.call_args[0]
assert "d.custom1.set=audiobooks" in args[2]
assert "d.custom1.set=books" not in args[2]
def test_falls_back_to_book_label_when_audiobook_label_not_set(self, monkeypatch):
config_values = {
"RTORRENT_URL": "http://localhost:8080/RPC2",
"RTORRENT_LABEL": "books",
"RTORRENT_AUDIOBOOK_LABEL": "",
"RTORRENT_DOWNLOAD_DIR": "/downloads",
}
mock_rpc, mock_xmlrpc, mock_torrent_info = self._make_client(monkeypatch, config_values)
with patch.dict("sys.modules", {"xmlrpc.client": mock_xmlrpc}):
with patch(
"shelfmark.download.clients.torrent_utils.extract_torrent_info",
return_value=mock_torrent_info,
):
if "shelfmark.download.clients.rtorrent" in sys.modules:
del sys.modules["shelfmark.download.clients.rtorrent"]
from shelfmark.download.clients.rtorrent import RTorrentClient
client = RTorrentClient()
client.add_download(
"magnet:?xt=urn:btih:abc123", "Test Audiobook", content_type="audiobook"
)
args = mock_rpc.load.start.call_args[0]
assert "d.custom1.set=books" in args[2]
def test_uses_book_label_for_non_audiobook_content(self, monkeypatch):
config_values = {
"RTORRENT_URL": "http://localhost:8080/RPC2",
"RTORRENT_LABEL": "books",
"RTORRENT_AUDIOBOOK_LABEL": "audiobooks",
"RTORRENT_DOWNLOAD_DIR": "/downloads",
}
mock_rpc, mock_xmlrpc, mock_torrent_info = self._make_client(monkeypatch, config_values)
with patch.dict("sys.modules", {"xmlrpc.client": mock_xmlrpc}):
with patch(
"shelfmark.download.clients.torrent_utils.extract_torrent_info",
return_value=mock_torrent_info,
):
if "shelfmark.download.clients.rtorrent" in sys.modules:
del sys.modules["shelfmark.download.clients.rtorrent"]
from shelfmark.download.clients.rtorrent import RTorrentClient
client = RTorrentClient()
client.add_download("magnet:?xt=urn:btih:abc123", "Test Book")
args = mock_rpc.load.start.call_args[0]
assert "d.custom1.set=books" in args[2]
assert "d.custom1.set=audiobooks" not in args[2]
class TestRTorrentClientGetStatus:
"""Tests for RTorrentClient.get_status()."""
@@ -383,6 +383,42 @@ class TestTransmissionClientGetStatus:
assert status.complete is True
assert "/downloads/Test Torrent" in status.file_path
def test_get_status_stopped_treated_as_complete(self, monkeypatch):
"""Regression: torrents stopped after seeding ratio/idle limit must show complete."""
config_values = {
"TRANSMISSION_URL": "http://localhost:9091",
"TRANSMISSION_USERNAME": "admin",
"TRANSMISSION_PASSWORD": "password",
"TRANSMISSION_CATEGORY": "test",
}
monkeypatch.setattr(
"shelfmark.download.clients.transmission.config.get",
make_config_getter(config_values),
)
mock_torrent = MockTorrent(
percent_done=1.0,
status="stopped",
download_dir="/downloads",
)
mock_client_instance = MagicMock()
mock_client_instance.get_torrent.return_value = mock_torrent
mock_transmission_rpc = create_mock_transmission_rpc_module()
mock_transmission_rpc.Client.return_value = mock_client_instance
with patch.dict("sys.modules", {"transmission_rpc": mock_transmission_rpc}):
if "shelfmark.download.clients.transmission" in sys.modules:
del sys.modules["shelfmark.download.clients.transmission"]
from shelfmark.download.clients.transmission import TransmissionClient
client = TransmissionClient()
status = client.get_status("abc123")
assert status.complete is True
assert status.progress == 100.0
def test_get_status_not_found(self, monkeypatch):
"""Test status for non-existent torrent."""
config_values = {