Skip to content

fix: isolate monitor polling lifecycles - #59

Open
0xdfi wants to merge 2 commits into
MiaAI-Lab:mainfrom
0xdfi:fix/monitor-poll-lifecycle
Open

fix: isolate monitor polling lifecycles#59
0xdfi wants to merge 2 commits into
MiaAI-Lab:mainfrom
0xdfi:fix/monitor-poll-lifecycle

Conversation

@0xdfi

@0xdfi 0xdfi commented Aug 24, 2026

Copy link
Copy Markdown

Summary

Monitor restarts and hot configuration changes can no longer let an older asynchronous poll overwrite the current run. GPU and CPU results now carry non-enumerable collection provenance, so consumers can distinguish a real idle sample from a fallback object produced after collection failure.

The run-generation and per-poll tokens also protect liveness, storage refresh, and shared CPU baselines. This is the freshness foundation for follow-up fleet-energy telemetry.

Validation

  • node --test server/sparks/__tests__/monitor-lifecycle.test.js server/collectors/__tests__/SystemCollector.cpuTemp.test.js - 15 passed
  • npm run typecheck - passed
  • npm run build - passed
  • Code review completed with no remaining actionable findings.

Post-Deploy Monitoring & Validation

  • Search logs for [SparkMonitor], poll error, and liveness during the first 30 minutes.
  • Change one Spark configuration while a poll is in flight, then confirm only the new target updates.
  • Healthy signal: current polls continue after the change and GPU/CPU freshness returns on the next successful sample.
  • Failure signal: metrics revert to an old target, a domain remains permanently in flight, or collection provenance stays false after successful polling. Roll back this PR if any occurs.
  • Validation window and owner: first 30 minutes after deployment; sparkDash maintainer/operator.

Compound Engineering

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant