Repository navigation
24h MCP session cleanup can be outpaced by high-churn ChatGPT reconnects #256
Description
Activity
Local Linux reproduction on the current desktop branch confirms the same failure mode.
- devspace-tunnel.service uptime: ~16h
- MCP sessions created: 963
- MCP sessions closed: 0
- DevSpace Node RSS/PSS: ~1.52 GB / ~1.51 GB
- service cgroup memory peak: ~18.3 GiB
- current policy on this branch: 24h idle timeout, 5m sweep, no capacity bound
I am implementing the bounded-stateful mitigation here: capacity 256, active-response/SSE protection, LRU idle eviction, reservation during initialize to prevent concurrent oversubscription, 503 + Retry-After when all capacity is protected, plus focused regression coverage. I will post verification/commit details here when complete.
Implementation checkpoint:
b572a23(fix(server): bound MCP session retention).What changed:
- hard stateful MCP capacity: 256 retained sessions
- pending initialize requests reserve capacity immediately, preventing concurrent oversubscription
- least-recently-used idle sessions are closed/evicted at capacity
- active HTTP responses / live SSE sessions are protected from eviction
- when every slot is protected, initialize returns HTTP 503 with
Retry-After: 5instead of growing memory - eviction close failures remain capacity-counted and are logged instead of silently freeing a slot
- existing 24h idle cleanup remains as age cleanup, but is no longer the memory bound
- structured logs now include session count/capacity and capacity eviction/exhaustion events
Regression coverage includes 1,024-session churn against a 256-session cap, active-session protection, concurrent initialize reservations, close failure, idle cleanup, and shutdown cleanup.
Verification:
- focused MCP session tests: pass
npm run typecheck: pass- full
npm test: pass npm run build: passgit diff --check: pass- generated
dist/server.jscontains the capacity guard used by the installeddist/cli.js servepath
The live service still needs a restart to discard its currently retained pre-fix sessions and load this build; I am doing that after this checkpoint.
OOM postmortem found a second independent failure mode that the session cap does not cover:
- 2026-08-29 15:21:51 CEST: global OOM killed
python3indevspace-tunnel.service, anon RSS ~49.5 GB - 2026-08-29 17:20:42 CEST: global OOM killed another
python3in the same service cgroup, anon RSS ~40.4 GB - both map to DevSpace workspace
ws_83c8a94a1b(/home/Maczuga/Dokumenty/New World Bot) and journal entries identify the invoking tool as DevSpacebash/python3 - current unit has
MemoryHigh=infinity,MemoryMax=infinity,MemorySwapMax=infinity,OOMPolicy=stop
So the follow-up in this same OOM hardening wave is an aggregate systemd cgroup guard for the desktop-managed background service. This is deliberately service-level (not just a per-command limit), so multiple concurrent tool children cannot collectively drive the machine into global OOM. The limit will be derived from physical RAM with a conservative ceiling and an explicit desktop env override/disable path.
- 2026-08-29 15:21:51 CEST: global OOM killed
Follow-up OOM hardening is complete in
a113077(fix(desktop): cap background service memory).This addresses the second failure mode from the August 29 OOMs: runaway tool children inside
devspace-tunnel.service, not just retained MCP transport state.What changed:
- desktop now installs a systemd drop-in for an existing
devspace-tunnel.service - aggregate
MemoryMaxdefaults to 20% of physical RAM, clamped to 4-16 GiB MemoryHighis 75% of the hard limitMemorySwapMaxis separately bounded (up to 2 GiB)- persistent desktop override uses
OOMPolicy=continue, so a cgroup OOM does not automatically take down the whole unit DEVSPACE_DESKTOP_SERVICE_MEMORY_MAX_MBallows an explicit hard-limit overrideDEVSPACE_DESKTOP_SERVICE_MEMORY_GUARD=0disables desktop management- README + focused regression coverage added
Verification:
npm run desktop:check: 56/56 passnpm run typecheck: pass- full
npm test: pass npm run build: passgit diff --check: pass
Live Fedora verification on the affected machine:
- runtime
MemoryHigh=12884901888(12 GiB) - runtime
MemoryMax=17179869184(16 GiB) - runtime
MemorySwapMax=2147483648(2 GiB) - current service memory ~3.0 GiB while concurrent DevSpace work is active
memory.events: high=0, max=0, oom=0, oom_kill=0 after applying the guard
For immediate safety I applied those three memory limits through a runtime-only systemd control drop-in, so the active service is protected without restarting the desktop/chat. I deliberately removed the persistent
systemctl set-propertylayer becauseuser.controlhas higher precedence and would defeat the desktop override/disable knobs. The live unit therefore keeps its pre-existingOOMPolicy=stopuntil the updated desktop app next applies its normal persistent80-devspace-desktop-memory.conf; the committed override setsOOMPolicy=continue.- desktop now installs a systemd drop-in for an existing
@Maczuga you can raise PR against your current work, leme have a quick look while you are iterating on this one
Done
#201 is now merged and removes the retained MCP transport-session path at the root. Modern MCP is request-scoped, and older 2025-era clients use the SDK’s stateless compatibility path.
That removes the unbounded session-retention failure mode described here, so I’m closing this as resolved.
Summary
The 24-hour idle-session cleanup added in #71 is working, but it is not a deterministic memory bound when ChatGPT initializes/reconnects at a high rate. In one long-running production-equivalent run, session creation substantially outpaced 24h cleanup and the DevSpace process grew into multi-GB memory usage.
This is not a claim that MCP session retention is the sole cause of every ChatGPT timeout/disconnect. It is a narrower finding: the current 24h/5m policy can still allow thousands of retained session/server graphs under high churn.
Environment / scope
Production evidence
Collector baseline:
heapUsed: 155,141,692 BAfter about 2h12m:
heapUsed: 442,791,944 BLonger advisory observation before restart:
heapUsed: 2,855,774,792 BThe 24h cleanup was actually functioning:
idle_timeoutremovals: 3,301close_completed: 3,301close_failed: 0But sustained creation continued faster than removal. Observed creation rate was roughly 2.5–3.4 sessions/minute. Session count and
heapUsedwere strongly correlated (~0.998 in the focused collector window and ~0.993 over the longer observation). Restart returned the process footprint to roughly 169 MB.Interpretation
This supports the original motivation in #71, but suggests the intentionally conservative 24h timeout is insufficient for some ChatGPT workloads. It bounds session age, not registry size or retained-memory exposure.
A paused conversation vs. memory-safety tradeoff remains: simply shortening the idle timeout can invalidate legitimate paused sessions, while leaving it long permits large accumulation.
Related upstream work
I added the detailed operational evidence to #71, and a separate bounded-stateful validation data point to #89/#209.
Independent bounded-stateful validation
To test whether a deterministic stateful bound is viable, I separately exercised a prototype with:
Retry-After: 5when no eligible idle slot existsUsing Node 24.18.0, MCP SDK 1.30.0, and Express 5.2.1 on the real server path:
This does not prove a bounded-stateful design is preferable to #209 stateless handling; it only shows it is a viable mitigation shape under this test.
Desired outcome
It would be useful to get maintainer guidance on the intended upstream direction:
Either way, I think the current 24h cleanup should be treated as an age-based mitigation rather than a deterministic memory bound for high-churn ChatGPT clients.