Skip to content

24h MCP session cleanup can be outpaced by high-churn ChatGPT reconnects #256

Description

@whitesarum

Summary

The 24-hour idle-session cleanup added in #71 is working, but it is not a deterministic memory bound when ChatGPT initializes/reconnects at a high rate. In one long-running production-equivalent run, session creation substantially outpaced 24h cleanup and the DevSpace process grew into multi-GB memory usage.

This is not a claim that MCP session retention is the sole cause of every ChatGPT timeout/disconnect. It is a narrower finding: the current 24h/5m policy can still allow thousands of retained session/server graphs under high churn.

Environment / scope

  • macOS
  • Node v24.18.0
  • ChatGPT as the MCP host
  • Observed on a production-equivalent DevSpace v1.0.7-based build with additive read-only runtime observability
  • I have not installed v1.0.8 to reproduce this exact run. However, current v1.0.8/main still appears to use the same 24h idle timeout, 5-minute cleanup sweep, and no hard session-capacity bound.

Production evidence

Collector baseline:

  • retained MCP sessions: 144
  • heapUsed: 155,141,692 B
  • physical footprint: 229,591,480 B

After about 2h12m:

  • retained MCP sessions: 586
  • heapUsed: 442,791,944 B
  • a +256 MiB heap safety threshold was reached

Longer advisory observation before restart:

  • current sessions: 3,904
  • peak sessions: 4,134
  • heapUsed: 2,855,774,792 B
  • physical footprint: 3,269,620,040 B

The 24h cleanup was actually functioning:

  • idle_timeout removals: 3,301
  • matching close_completed: 3,301
  • close_failed: 0

But sustained creation continued faster than removal. Observed creation rate was roughly 2.5–3.4 sessions/minute. Session count and heapUsed were strongly correlated (~0.998 in the focused collector window and ~0.993 over the longer observation). Restart returned the process footprint to roughly 169 MB.

Interpretation

This supports the original motivation in #71, but suggests the intentionally conservative 24h timeout is insufficient for some ChatGPT workloads. It bounds session age, not registry size or retained-memory exposure.

A paused conversation vs. memory-safety tradeoff remains: simply shortening the idle timeout can invalidate legitimate paused sessions, while leaving it long permits large accumulation.

Related upstream work

I added the detailed operational evidence to #71, and a separate bounded-stateful validation data point to #89/#209.

Independent bounded-stateful validation

To test whether a deterministic stateful bound is viable, I separately exercised a prototype with:

  • hard capacity: 256 sessions
  • active-request and live-SSE protection
  • idle-only capacity eviction
  • bounded HTTP 503 + Retry-After: 5 when no eligible idle slot exists
  • close-failure slots remaining capacity-counted

Using Node 24.18.0, MCP SDK 1.30.0, and Express 5.2.1 on the real server path:

  • 1,024 / 1,024 initialize + initialized + DELETE cycles succeeded
  • final registry/permit/protection occupancy: 0
  • in a 256-session capacity run, one live GET SSE and one disconnected-but-running tool operation remained protected
  • 254 eligible idle sessions were evicted/replaced
  • when only protected/non-eligible capacity remained, the next initialize returned the expected 503
  • maximum occupancy: 256
  • active-session evictions: 0
  • duplicate removal/terminal lifecycle events: 0
  • equal-occupancy post-GC heap delta after the eviction wave: about -2.08 MB
  • after complete cleanup, post-GC heap was about +1.12 MB above startup

This does not prove a bounded-stateful design is preferable to #209 stateless handling; it only shows it is a viable mitigation shape under this test.

Desired outcome

It would be useful to get maintainer guidance on the intended upstream direction:

  1. keep stateful sessions but add a deterministic capacity + active-work protection, or
  2. move to request-scoped/stateless Streamable HTTP as in fix(server): use stateless Streamable HTTP #209.

Either way, I think the current 24h cleanup should be treated as an age-based mitigation rather than a deterministic memory bound for high-churn ChatGPT clients.

Activity

  1. Maczuga commented on Aug 30, 2026

    @Maczuga

    Local Linux reproduction on the current desktop branch confirms the same failure mode.

    • devspace-tunnel.service uptime: ~16h
    • MCP sessions created: 963
    • MCP sessions closed: 0
    • DevSpace Node RSS/PSS: ~1.52 GB / ~1.51 GB
    • service cgroup memory peak: ~18.3 GiB
    • current policy on this branch: 24h idle timeout, 5m sweep, no capacity bound

    I am implementing the bounded-stateful mitigation here: capacity 256, active-response/SSE protection, LRU idle eviction, reservation during initialize to prevent concurrent oversubscription, 503 + Retry-After when all capacity is protected, plus focused regression coverage. I will post verification/commit details here when complete.

  2. Maczuga commented on Aug 30, 2026

    @Maczuga

    Implementation checkpoint: b572a23 (fix(server): bound MCP session retention).

    What changed:

    • hard stateful MCP capacity: 256 retained sessions
    • pending initialize requests reserve capacity immediately, preventing concurrent oversubscription
    • least-recently-used idle sessions are closed/evicted at capacity
    • active HTTP responses / live SSE sessions are protected from eviction
    • when every slot is protected, initialize returns HTTP 503 with Retry-After: 5 instead of growing memory
    • eviction close failures remain capacity-counted and are logged instead of silently freeing a slot
    • existing 24h idle cleanup remains as age cleanup, but is no longer the memory bound
    • structured logs now include session count/capacity and capacity eviction/exhaustion events

    Regression coverage includes 1,024-session churn against a 256-session cap, active-session protection, concurrent initialize reservations, close failure, idle cleanup, and shutdown cleanup.

    Verification:

    • focused MCP session tests: pass
    • npm run typecheck: pass
    • full npm test: pass
    • npm run build: pass
    • git diff --check: pass
    • generated dist/server.js contains the capacity guard used by the installed dist/cli.js serve path

    The live service still needs a restart to discard its currently retained pre-fix sessions and load this build; I am doing that after this checkpoint.

  3. Maczuga commented on Aug 30, 2026

    @Maczuga

    OOM postmortem found a second independent failure mode that the session cap does not cover:

    • 2026-08-29 15:21:51 CEST: global OOM killed python3 in devspace-tunnel.service, anon RSS ~49.5 GB
    • 2026-08-29 17:20:42 CEST: global OOM killed another python3 in the same service cgroup, anon RSS ~40.4 GB
    • both map to DevSpace workspace ws_83c8a94a1b (/home/Maczuga/Dokumenty/New World Bot) and journal entries identify the invoking tool as DevSpace bash / python3
    • current unit has MemoryHigh=infinity, MemoryMax=infinity, MemorySwapMax=infinity, OOMPolicy=stop

    So the follow-up in this same OOM hardening wave is an aggregate systemd cgroup guard for the desktop-managed background service. This is deliberately service-level (not just a per-command limit), so multiple concurrent tool children cannot collectively drive the machine into global OOM. The limit will be derived from physical RAM with a conservative ceiling and an explicit desktop env override/disable path.

  4. Maczuga commented on Aug 30, 2026

    @Maczuga

    Follow-up OOM hardening is complete in a113077 (fix(desktop): cap background service memory).

    This addresses the second failure mode from the August 29 OOMs: runaway tool children inside devspace-tunnel.service, not just retained MCP transport state.

    What changed:

    • desktop now installs a systemd drop-in for an existing devspace-tunnel.service
    • aggregate MemoryMax defaults to 20% of physical RAM, clamped to 4-16 GiB
    • MemoryHigh is 75% of the hard limit
    • MemorySwapMax is separately bounded (up to 2 GiB)
    • persistent desktop override uses OOMPolicy=continue, so a cgroup OOM does not automatically take down the whole unit
    • DEVSPACE_DESKTOP_SERVICE_MEMORY_MAX_MB allows an explicit hard-limit override
    • DEVSPACE_DESKTOP_SERVICE_MEMORY_GUARD=0 disables desktop management
    • README + focused regression coverage added

    Verification:

    • npm run desktop:check: 56/56 pass
    • npm run typecheck: pass
    • full npm test: pass
    • npm run build: pass
    • git diff --check: pass

    Live Fedora verification on the affected machine:

    • runtime MemoryHigh=12884901888 (12 GiB)
    • runtime MemoryMax=17179869184 (16 GiB)
    • runtime MemorySwapMax=2147483648 (2 GiB)
    • current service memory ~3.0 GiB while concurrent DevSpace work is active
    • memory.events: high=0, max=0, oom=0, oom_kill=0 after applying the guard

    For immediate safety I applied those three memory limits through a runtime-only systemd control drop-in, so the active service is protected without restarting the desktop/chat. I deliberately removed the persistent systemctl set-property layer because user.control has higher precedence and would defeat the desktop override/disable knobs. The live unit therefore keeps its pre-existing OOMPolicy=stop until the updated desktop app next applies its normal persistent 80-devspace-desktop-memory.conf; the committed override sets OOMPolicy=continue.

  5. Waishnav commented on Aug 31, 2026

    @Waishnav
    Owner

    @Maczuga you can raise PR against your current work, leme have a quick look while you are iterating on this one

  6. Maczuga commented on Aug 31, 2026

    @Maczuga

    Done

  7. Waishnav commented on Sep 6, 2026

    @Waishnav
    Owner

    #201 is now merged and removes the retained MCP transport-session path at the root. Modern MCP is request-scoped, and older 2025-era clients use the SDK’s stateless compatibility path.

    That removes the unbounded session-retention failure mode described here, so I’m closing this as resolved.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions