Skip to content

Keep the loader cache the size of the spec, and off every MCP call - #101

Merged
AndersRobstad merged 2 commits into
mainfrom
fiber-loader-manifest-bloat
Sep 29, 2026
Merged

AndersRobstad merged 2 commits into
mainfrom
fiber-loader-manifest-bloat

Conversation

@AndersRobstad

Copy link
Copy Markdown
Collaborator

Problem

Refreshing a collection whose OpenAPI spec has schemas that refer to each other in a ring (Object → Function → Operand → Object) wrote loaders/<section>.json at ~2.5 GB. After that every MCP tool call hung, list_sections included, while initialize still answered.

openapi::resolve_schema expanded every $ref in place and stopped only where a reference met itself on the current path. The number of such paths through a densely connected ring grows factorially, and every endpoint stored its own expanded copy of both its request and its response schema. The MCP server already cached the parsed file per section, but the first read loaded the whole 2.5 GB file to count endpoints and never finished.

Fix

Schemas stored once, expanded on request.

  • Schemas move to loaders/<id>.schemas.json. Each is stored as the document wrote it, with every definition it reaches kept once, keyed by $ref.
  • openapi::bundle makes one endpoint's schema self-contained when get_endpoint or the editor asks for it. A definition used once is written in place. Anything used more than once is written a single time under $defs, so the output can't exceed the schema plus the definitions it reaches.
  • Ajv in the editor already reads $defs, and the old output left cyclic refs dangling so it couldn't validate them. That now works; there's a new test in json-schema.spec.ts.

Calls that don't need schemas never read them.

  • <id>.json holds only loadedAt and endpoints, which is all listing, searching and send_request's access decision need.
  • The MCP server caches schemas separately by file stamp, and only get_endpoint loads them.
  • The endpoint cache now loads under its lock, so concurrent calls wait for one read instead of each starting their own.

Older caches. The old format serialized endpoints before schemas, so read_cache streams and stops once it has both fields. A 2.5 GB file is read only as far as its ~250 KB endpoint list. Its schemas aren't served: get_endpoint returns null until the next refresh writes both files.

Request bodies. Body skeletons get a 1,000-node budget, and a body that doesn't fit is rebuilt one level shallower. The depth limit alone let the test spec produce a 440 KB body per endpoint.

Verification

  • New fixture src-tauri/tests/fixtures/cyclic-union.json (19.5 KB): a union whose branches contain the union again, over 7 schemas that all reference each other, shared by 20 endpoints.
  • Rust tests cover the minimal A → oneOf[B, C] cycle, the cache staying smaller than the spec, reading an old-format file, and that listing, searching and access decisions never load schemas.
  • cargo fmt --check, clippy -D warnings with and without default features, and cargo test with and without default features all pass. json-schema.spec.ts passes 19/19.

Through the real MCP server over stdio:

0.16.2 this branch
Saved data dir with the two 2.5 GB caches, list_sections no answer in 45 s 0.2 s
Test spec, cache written by refresh_endpoints 748 MB 6.5 KB schemas + 657 KB endpoint list
Test spec, list_sections / get_endpoint after restart 19.4 s / 31 s (22 MB response) 0.14 s / 0.18 s

A spec whose schemas refer to each other in a ring wrote a 2.5 GB loader
cache: every endpoint's schemas were stored with each $ref expanded along
every path through the ring. Every MCP call read that file first, so the
server answered `initialize` and then hung.

Schemas now live in <id>.schemas.json as the document wrote them, each
reachable definition once, and are bundled per endpoint on request with
shared definitions under $defs. <id>.json holds only the endpoint list,
which is all listing, searching and access decisions read. A cache from
before the split is read only as far as its endpoint list. Request-body
skeletons get a node budget, since the depth limit alone still let them
branch into hundreds of kilobytes.
@AndersRobstad
AndersRobstad merged commit a7ec8ac into main Sep 29, 2026
4 checks passed
@AndersRobstad
AndersRobstad deleted the fiber-loader-manifest-bloat branch September 29, 2026 12:37
@github-actions github-actions Bot mentioned this pull request Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant