feat(api): read rss results from indexers #1

Merged
thibault merged 1 commit from feat/indexer-rss into 0.2 2026-10-08 18:30:24 +02:00
Owner

What

PR D of the indexer engine: the rss results type.

  • src/indexers/rss.rs: reads a feed with quick-xml's plain Reader (no serde, no namespace resolution, so names stay as written: ex:seeders). Each <item> (RSS 2.0 / 1.0) or <entry> (Atom) becomes a small JSON value: an element is its trimmed text (CDATA, predefined entities and character references decoded), or, when it has attributes or children, an object of its children by qualified name, @attr keys and its own text under #text. Repeated elements: the first one wins. The root must be rss, feed or RDF; malformed XML, unclosed elements or another document (an HTML page) is an error.
  • extract::rss runs the existing field reading on each item, so sizes, counts, dates (RFC 2822 pubDate), infohash from the field or the magnet, magnet built from infohash + trackers, and the "left out" count all work unchanged. lookup learns the @attr segment, and a path leading to an element with attributes reads its #text.
  • The engine fetches rss with the same limits, cache and filters; the "rss is not supported yet" skip is gone, so /api/indexers/check reports rss indexers like json ones.
  • .torrent links stay plain links: nothing downloads them, an item needs an infohash or a magnet holding one.

Tests

  • rss.rs: namespaced elements, attributes, CDATA, entities, repeated elements, Atom entries, malformed / unclosed / HTML / empty answers.
  • Fake indexer: /rss (namespaced infohash, CDATA title, 1.4 GiB, RFC 2822 date, .torrent enclosure via enclosure.@url, an item with only a base32 magnet in <link>, an item with neither), /rss-broken (malformed), /rss-named (the fake catalog's releases as a feed). An rss indexer pointed at an HTML page now fails instead of being skipped.
  • API: an anime search through an rss indexer keeps the batch and the season pack and drops the single episode, as from JSON.

124 tests pass; fmt and clippy clean.

Docs

docs/INDEXERS.md: status line, a complete fictional RSS feed and definition, an "RSS results" section (paths, attributes, decoding, first wins, UTF-8, Atom), new troubleshooting errors. README mentions RSS; ROADMAP phase 7 gets an rss step ✅.

New dependency

quick-xml = "0.41": already in Cargo.lock through librqbit (librqbit-upnp), same version, so nothing new is downloaded or built. A small, fast, widely used XML pull parser; the plain reader keeps the code short and gives prefixed names as written, which definitions use.

🤖 Generated with Claude Code

## What PR D of the indexer engine: the `rss` results type. - `src/indexers/rss.rs`: reads a feed with quick-xml's plain `Reader` (no serde, no namespace resolution, so names stay as written: `ex:seeders`). Each `<item>` (RSS 2.0 / 1.0) or `<entry>` (Atom) becomes a small JSON value: an element is its trimmed text (CDATA, predefined entities and character references decoded), or, when it has attributes or children, an object of its children by qualified name, `@attr` keys and its own text under `#text`. Repeated elements: the first one wins. The root must be `rss`, `feed` or `RDF`; malformed XML, unclosed elements or another document (an HTML page) is an error. - `extract::rss` runs the existing field reading on each item, so sizes, counts, dates (RFC 2822 `pubDate`), infohash from the field or the magnet, magnet built from infohash + trackers, and the "left out" count all work unchanged. `lookup` learns the `@attr` segment, and a path leading to an element with attributes reads its `#text`. - The engine fetches rss with the same limits, cache and filters; the "rss is not supported yet" skip is gone, so `/api/indexers/check` reports rss indexers like json ones. - `.torrent` links stay plain links: nothing downloads them, an item needs an infohash or a magnet holding one. ## Tests - rss.rs: namespaced elements, attributes, CDATA, entities, repeated elements, Atom entries, malformed / unclosed / HTML / empty answers. - Fake indexer: `/rss` (namespaced infohash, CDATA title, `1.4 GiB`, RFC 2822 date, `.torrent` enclosure via `enclosure.@url`, an item with only a base32 magnet in `<link>`, an item with neither), `/rss-broken` (malformed), `/rss-named` (the fake catalog's releases as a feed). An rss indexer pointed at an HTML page now fails instead of being skipped. - API: an anime search through an rss indexer keeps the batch and the season pack and drops the single episode, as from JSON. 124 tests pass; fmt and clippy clean. ## Docs docs/INDEXERS.md: status line, a complete fictional RSS feed and definition, an "RSS results" section (paths, attributes, decoding, first wins, UTF-8, Atom), new troubleshooting errors. README mentions RSS; ROADMAP phase 7 gets an `rss` step ✅. ## New dependency `quick-xml = "0.41"`: already in Cargo.lock through librqbit (librqbit-upnp), same version, so nothing new is downloaded or built. A small, fast, widely used XML pull parser; the plain reader keeps the code short and gives prefixed names as written, which definitions use. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
feat(api): read rss results from indexers
Some checks failed
CI / test (pull_request) Has been cancelled
CI / web (pull_request) Has been cancelled
0bd815ff04
Indexers whose results type is rss are now searched and checked like json
ones. Each <item> (or Atom <entry>) of the feed becomes a small JSON value,
read with the same paths, format inference, magnet building and filters.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01947SGTYxD1CJLsk2PcULRA
ci: trigger the checks again
Some checks failed
CI / test (pull_request) Has been cancelled
CI / web (pull_request) Has been cancelled
5d7533f920
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
thibault force-pushed feat/indexer-rss from 5d7533f920
Some checks failed
CI / test (pull_request) Has been cancelled
CI / web (pull_request) Has been cancelled
to 0bd815ff04
Some checks failed
CI / test (pull_request) Has been cancelled
CI / web (pull_request) Has been cancelled
2026-10-08 18:19:25 +02:00
Compare
thibault force-pushed feat/indexer-rss from 0bd815ff04
Some checks failed
CI / test (pull_request) Has been cancelled
CI / web (pull_request) Has been cancelled
to 46b732d124
All checks were successful
CI / web (pull_request) Successful in 7s
CI / test (pull_request) Successful in 25s
2026-10-08 18:19:48 +02:00
Compare
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
thibault/plankton!1
No description provided.