Skip to content

Repository files navigation

rss-parser

Typed, pydantic-powered RSS/Atom parsing for Python.

PyPI version Python versions Downloads Wheel status License

CI Docs PyPi publish

rss-parser turns RSS/Atom XML into typed pydantic models — autocomplete, validation, and clear errors instead of digging through nested dicts.

Documentation

At a glance

Ecosystem Python 3.10+ — installed with pip install rss-parser, imported as rss_parser. This is not the npm package of the same name; there is no JavaScript/Node.js distribution.
Feed formats RSS 2.0, RSS 0.91/0.92, Atom 1.0, RSS 1.0 (RDF), plus typed Apple Podcasts (itunes:*) extensions
Runtime dependencies pydantic v2 (>=2.7), xmltodict, typing-extensions
Repository github.com/dhvcc/rss-parser
Documentation dhvcc.github.io/rss-parser
Package pypi.org/project/rss-parser
License GPL-3.0
Issues github.com/dhvcc/rss-parser/issues
Security Report vulnerabilities privately — see SECURITY.md. Do not open a public issue.

Installation

pip install rss-parser

For AI coding agents

skills/rss-parser/SKILL.md is an Agent Skill: the API model, recipes and the pitfalls that trip agents up, in one file. Install it into Claude Code, Cursor, Copilot, Codex and friends with:

npx skills add dhvcc/rss-parser

The library also ships context7.json, so Context7 serves these docs with Python-specific rules.

Command line

Installing the package installs an rss-parser command — handy from a shell or an agent, with validate as the verb that has no substitute (nothing else knows the three feed schemas):

rss-parser validate feed.xml            # exit 0 ok, 1 rejected; errors on stderr
rss-parser validate --json feed.xml     # {"valid": true, "feed_type": "rss", "items": 36}
rss-parser items feed.xml | jq -r '.content.title.content'   # NDJSON, one item per line
rss-parser jsonfeed feed.xml            # JSON Feed 1.1 document (lossy - see the docs)
curl -sSL "$url" | rss-parser validate -                     # it never fetches for you

Full reference, exit codes and caveats: Command line interface.

Parsing from a URL

rss-parser does not fetch anything — it parses text you already have, so there is no parseURL/parseString (that is the npm package). Bring your own HTTP client:

import requests
from rss_parser import parse

feed = parse(requests.get(url, timeout=10).text)

parse() accepts str or bytes — pass response.content and the feed's own encoding declaration is honored, which matters for feeds that are not UTF-8. Polling, conditional GET, deduplication by guid/id and normalizing across RSS/Atom/RDF are covered in Fetching feeds from a URL.

Quickstart

from rss_parser import parse
from requests import get  # noqa

rss_url = "https://rss.art19.com/apology-line"
response = get(rss_url)

feed = parse(response.content)  # detects RSS 2.0 / 0.9x, Atom 1.0 or RSS 1.0 (RDF)

print("Language", feed.channel.language)
print("RSS", feed.version)

for item in feed.channel.items:
    print(item.title)
    print(str(item.description)[:50])

# Language en
# RSS 2.0
# Wondery Presents - Flipping The Bird: Elon vs Twitter
# <p>When Elon Musk posted a video of himself arrivi
# Introducing: The Apology Line
# <p>If you could call a number and say you’re sorry

parse() picks the right parser from the XML root element and raises UnknownFeedTypeError if the document is not a feed. If you already know the feed type, use the explicit parsers: RSSParser, AtomParser, RDFParser, PodcastParser.

Podcasts

itunes:* tags are supported out of the box, fully typed:

from rss_parser import PodcastParser

podcast = PodcastParser.parse(feed_xml)
channel = podcast.channel.content

channel.itunes_author                    # 'Wondery'
channel.itunes_owner.content.email       # 'iwonder@wondery.com'
channel.itunes_image.attributes["href"]  # artwork url

episode = channel.items[0].content
episode.itunes_duration                  # '00:05:01'
episode.itunes_episode_type              # 'trailer'

Custom fields: one subclass away

The models are generic, so extending the schema doesn't require re-declaring the whole tree:

from typing import Optional
from pydantic import Field

from rss_parser import RSSParser
from rss_parser.models.rss import RSS, Channel, Item
from rss_parser.models.types import Tag


class MyItem(Item):
    dc_creator: Optional[Tag[str]] = Field(alias="dc:creator", default=None)


rss = RSSParser.parse(data, schema=RSS[Channel[MyItem]])

rss.channel.items[0].content.dc_creator

And even without a custom schema, unknown tags are never dropped — they're kept in model_extra:

rss = RSSParser.parse(podcast_xml)
rss.channel.content.model_extra["itunes:author"]  # 'Wondery'

See Customizing the schema for mixins, repeatable tags, and the field types cheat sheet.

Migrating from 3.x

4.0 removes the legacy pydantic v1 models, fixes several RSS 2.0 spec violations, and makes the models generic and lossless. See the migration guide for the full list.

Contributing

Pull requests are welcome. For major changes, please open an issue first to discuss what you would like to change.

Install dependencies with uv sync (install uv).

Using pre-commit is highly recommended. To install hooks, run:

uv run pre-commit install -t=pre-commit -t=pre-push

See Contributing for tests, snapshots, and docs.

License

GPLv3

About

Typed pythonic RSS/Atom parser: RSS 2.0/0.9x, Atom 1.0, RSS 1.0 (RDF) and podcasts into pydantic v2 models, plus a CLI

Topics

Resources

Contributing

Security policy

Stars

Watchers

Forks

Releases

Used by

Contributors

Languages