Skip to content

feat: Generate RFC and RFC subseries BibXML - #11695

Open
kesara wants to merge 35 commits into
mainfrom
feat/bibgen
Open

kesara wants to merge 35 commits into
mainfrom
feat/bibgen

Conversation

@kesara

@kesara kesara commented Sep 2, 2026

Copy link
Copy Markdown
Member

Fixes #10830
Fixes #10914

kesara and others added 28 commits May 18, 2026 19:46
* feat: Add method to save BibXML for a RFC

* chore: Add k8s settings

* style: ruff ruff

* chore: Add celery task to recreate BibXML for all RFCs

* fix: Unhandle exceptions

* fix: Fail proof URL construction

* refactor: Use itertools.batched

* style: ruff ruff

* test: Fix sync tests

* fix: Improve urljoin
* refactor: Generalize create_rfc_bibxml method

* feat: Ability to create BibXML for BCPs

* feat: Ability to create BibXML for STDs

* feat: Ability to create BibXML for FYIs

* feat: Add task to recreate rfcsubseries BibXML
* ci: numeric settings -> int for searchindex cfg (#10904)

* fix: remove no longer needed notification to the rfc-editor (#10911)

* fix: meeting_stats -> 404 for nonexistent meeting (#10896)

* fix: meeting_stats->404 for nonexistent number

* test: add test case

* chore: fix lint

* test: fix new test

* fix: set searchindex hiddenDefault flag (#10916)

* fix: set searchindex hiddenDefault flag

* chore: adjust docstring comments

* ci: pin xym until we can adapt to new required arguments (#10939)

* ci: update base image target version to 20260527T1529

* fix: adjust searchindex abstract sanitiziation (#10950)

* fix: adjust searchindex abstract sanitiziation

* test: leading/trailing whitespace to be stripped

* fix: strip each line of the abstract

* chore: add items to tracked yarn cache (#10964)

* chore: configurable I-D submit timeout for k8s (#10971)

* feat: update rfc json (#10951)

* feat: update rfc json

* fix: resilience around fetch of publication status

* fix: guards against malformed RFC document objects

* chore: fix log msg typos

* test: mock more async calls

* chore: nginx -> 1.30 (#10978)

* chore: "nginxinc" GH org is now "nginx" (#10981)

* chore: use redis for django caches (#10975)

* chore: drop memcache, add redis (WIP) (#10940)

* chore(dev): use redis for dev sessions

Adds a redis container with a persistent volume and uses this for sessions instead of the database. Results in the same general behavior as before - logins should be persistent until the docker compose stack is torn down.

* chore(deps): install redis/hiredis python pkgs

* chore: remove memcached (mostly)

Still keeps some config references and a custom cache backend

* chore: switch to django-redis for sentinel support

* chore: production redis cache config

Uses sentinel. Untested.

* chore: telemetry: add redis, drop memcache

* chore: remove now-unused cache.py

* chore: adjust imports in gunicorn.conf.py

* chore: adjust redis sentinel settings

* feat: size limit on redis values

* style: ruff ruff

* chore(dev): SizeLimitingRedisClient for dev

Does not make a functional difference in most cases, but
might as well exercise the class.

* test: use db-backed sessions for tests

* chore: typing nit

* ci: remove DEPLOY_STRATEGY

"strategery"

* ci: remove strategy from dt manifests altogether

* ci: update base image target version to 20260604T1454

* ci: remove guard against missing CACHES in prod (#10986)

* chore: revert switch to redis (#10991)

* ci: update base image target version to 20260605T2314

* feat: add group type to search index (#11001)

* feat: self-serve queries for inputs to reporting and survey purposes (#10976)

* feat: utilities to count I-D submitters and authors by year

* chore: fix comment typo

* fix: return querysets and sets rather than lists. Improve docstrings.

* refactor: don't flatten sets to lists

* fix: filter to posted submissions and produce report

* chore: black

* chore: boilerplate

* fix: re-re-re fix the ^M problem in the issue template

* feat: self-serve queries for inputs to reporting and survey purposes

* Update ietf/utils/tests_reports.py

Co-authored-by: Jennifer Richards <jeni@borkbork.org>

---------

Co-authored-by: Jennifer Richards <jeni@borkbork.org>

* ci: remove strategy from dt manifests (#11004)

* fix: tweak logging settings to avoid spurious keyerrors (#11009)

* fix: improved source and keywords for rfc json (#11008)

* test: make check for April 1 resilient across line-breaks (#11006)

* test: make check for April 1 resilient across line-breaks

* test: be resilient against faker creating a 63 character long title

* chore(dev): fix nginx config (#11014)

keepalive 0 is not supported by the version of nginx
in the dev container

---------

Co-authored-by: Jennifer Richards <jennifer@staff.ietf.org>
Co-authored-by: Robert Sparks <rjsparks@nostrum.com>
Co-authored-by: jennifer-richards <19472766+jennifer-richards@users.noreply.github.com>
Co-authored-by: Jennifer Richards <jeni@borkbork.org>
* fix(bibxml): Add subseries info

* fix: Harden subseries filtering

* refactor: Refactor get_rfc_bibxml method
* fix(bibxml): Derive orgs

* style: Ruff ruff

Good boy!

* fix: Use a lookup table

* chore: Remove Mitra from ORG_LOOKUP
* fix(bibxml): Remove XML declaration

* fix(bibxml): Include date after author elements
ci: merge main to feat/bibgen
ci: Merge main to feat/bibgen
…11386)

* fix(bibxml): Avoid generating BibXML for historic subseries entries

* fix(bibxml): Include save_bibxml call inside try block
ci: Merge main to feat/bibgen
* fix(bibxml): Replace double spaces in abstract with new line

* fix(bibxml): Match bib.ietf.org abstract whitespace

Abstracts as stored separate paragraphs with a blank line and sentences
with two spaces. bib.ietf.org emits one <t> per paragraph and collapses
the whitespace within a paragraph to single spaces; do the same.

Replacing the double spaces with a newline, as 384d384 did, matched
neither: measured against the 9112 RFCs whose abstract bib.ietf.org also
publishes, 1299 abstracts matched before this change and 8946 after.

* fix(bibxml): Omit the abstract element when there is none

707 RFCs have no abstract. An <abstract> holding an empty <t> is not
valid BibXML -- it requires at least one non-empty <t> -- and is not what
bib.ietf.org publishes for those RFCs, which is no element at all.

---------

Co-authored-by: Robert Sparks <rjsparks@nostrum.com>
ci: Merge main to feat/bibgen
* fix(bibxml): Update orgs

* fix(bibxml): Update IANA entry
chore: Merge changes from main to feat/bibgen
* fix(bibxml): Improve author name processing

* fix(bibxml): More improvements
* chore: Restore sync/tasks.py from main

* chore: Re-commit feat/bibgen changes
chore: Main to feat/bibgen

@jennifer-richards jennifer-richards left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some quick comments

Comment thread ietf/sync/bibxml.py Outdated
Comment thread ietf/sync/tasks.py
Comment thread ietf/sync/bibxml.py
chore: Merge main to feat/bibgen
kesara and others added 6 commits October 2, 2026 11:11
ci: merge main to feat/bibgen
* feat: Backward compatability for 4 digit only RFCs

* test: cover 4-digit zero-padded RFC bibxml for RFC < 1000

Add tests for the four_digits anchor padding and the dual (padded and
unpadded) output of recreate_rfc_bibxml() for RFCs below 1000, and fix the
existing recreate tests to tolerate multiple RFC fixtures.

Co-Authored-By: Claude <noreply@anthropic.com>
Model: claude-opus-4-8 (thinking: medium)
ci: Merge main to feat/bibgen
@kesara
kesara marked this pull request as ready for review October 7, 2026 11:16
Comment thread ietf/settings.py
Comment on lines +829 to 833
"bibxml_bucket": {
"BACKEND": "django.core.files.storage.InMemoryStorage",
"OPTIONS": {"location": "bibxml_bucket"},
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry to notice this late in the process, but datatracker will own these bibxml files. I think should be using a blobdb storage + replication. Instead of setting up a separate bucket, we should add one to the ARTIFACT_STORAGE_NAMES and go through store_str() and the other storage_utils methods instead of accessing storages directly.

I think we probably also want to combine bibxml-ids and the new bucket here, separating the I-D from RFC files in paths rather than using two separate buckets. Will have to give that some thought because it's a refactor + adjustment to deployed data. I don't think we have anything reading from the buckets yet so it should be fairly low stress to do the switch.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, though bibxml-ids corresponds to a module we serve via rsync. That might lead to complication if we combine the rfc and id data.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add a Celery task to generate BibXML files for RFC subseries (BCPs, STDs, FYIs) Add a Celery task to generate BibXML files for RFCs into blobstore

2 participants