Repository navigation
test(chat-swarm): run isolated 2-worker ChatGPT live canary #49
Description
Activity
James3014 commented
on Sep 9, 2026 OwnerAuthorMore actions2026-09-09 原生 Swarm 驗收配方:20/20 不得被 mock 或子程序替代
本 issue 的原始必做情境與 hard prerequisites 保持不變。PR #61(本次 fresh head
97fbefc041b7b60876d1dbbe45ae7c10617a61f0)已含候選 source wiring,但仍 OPEN,不能當已部署/native PASS。新的 #15 交付計畫保留本 gate,沒有縮減為單 worker 或測試程式。借鏡 ODW 的部分
借鏡 workflow + artifacts 與 essential-criteria checkpoint,把本 issue 既有情境寫成執行前凍結的測試配方及 evidence map,不安裝 ODW、不建立第二個 Swarm queue/store/scheduler。
執行前固定
- feat(chat-swarm): add durable core contract and SQLite store #46/feat(chat-swarm): add MCP peer-worker protocol and conversation identity binding #47/feat(chat-swarm): wire production MCP tools, cutover admission, config, and capability manifest #48 accepted integration、P0: Prevent live Dev MCP runtime lineage drift and capability regression #30 source/build/schema/native identity 與 Harden Dev MCP reliability across host binding, worker startup, and timeout reconciliation #15 的現行可靠性前置證據。
- exact source/tree/package/server generation/config/manifest;兩個真實 ChatGPT worker conversations 的可用 Dev 工具與 carrier binding(GitHub回報用必要 fingerprint,不洩漏 session/token)。
- 隔離的測試 realm、max workers=2;單一核對過的 Swarm state store,不能為兩個 worker各建不同DB而宣稱協作。
- 建立固定20個 bounded task attempt名冊:taskKey、worker分派、nonce、expected evidence、exact attempt。將 replay/conflict/unauthorized submit 等負向「tool calls」另外計數,不拿它們灌成20個task完成。
- 逐條連接既有15項scenario與接受条件:parallel dispatch、targeted follow-up、同輸入idempotent replay、同key changed input拒絕、wrong-worker submit拒絕、restart後unknown不重派、drain拒新派工但保留in-flight terminal result。
核對與失敗處理
完成後逐一核對20個原始attempt,結果、failed/reconcile-required都必須可追溯;unknown是已被記錄的未知,不是成功。
20/20 accounted for不等於20個都成功生成答案,也不允許忽略任一必做正向/負向驗收。任何遺失、重複執行、跨worker錯配、unauthorized submit接受、blind retry,都保持gate未通過。runtime restart/cutover由主控在隔離realm使用既有受控机制,peer workers不取得agent_start/Codex/shell/Git/repository mutation權限。工具斷線先核對同一task/attempt/store,不能把新對話或新attempt當作重新开始的免責理由。
原生的定義
兩個Node MCP clients、SDK integration tests、ODW mock/parallel agents、兩個Codex subprocesses都只能提供較低層證據,不能替代本gate的兩個真實ChatGPT conversations。若host無法維持長等待/接收互動,記錄原生interaction限制;不暗中改用OpenCLI來湊PASS,也不宣稱背景自動持續工作已成立。
本次只完成配方重規劃,沒有建立swarm、派task或執行restart。本 gate 通過後才解鎖 #51 continuity實作;#50仍可選。
James3014 commented
on Sep 11, 2026 OwnerAuthorMore actions2026-09-11 prerequisite reconciliation — source landed; native 2-conversation gate remains unchanged
Fresh GitHub state supersedes the 2026-09-09 comment that described PR #61 as OPEN:
- PR fix: bind dispatch to fresh catalogs and verified provider results #60 is MERGED (
710e5f1afb3dbb19042c17c492a2f5cb0881f10d); - PR feat: add opt-in durable ChatGPT peer swarm #61 is MERGED (
d5667c3b98d2b61228d31a24ef94d5e018d60632), carrying the feat(chat-swarm): add durable core contract and SQLite store #46–feat(chat-swarm): wire production MCP tools, cutover admission, config, and capability manifest #48 Swarm source integration line; - P0: Prevent live Dev MCP runtime lineage drift and capability regression #30 is now CLOSED / completed;
- current
James3014/devspace/mainis9f0d7712b608ff3c931ba52344455d8ed55caaf2.
These facts remove stale source/PR wording but do not themselves satisfy this native canary. Issue-level acceptance and loaded-runtime identity must still be freshly bound; source merge is not two-real-conversation evidence. #15 remains the live single-worker/host reliability owner and must satisfy the current prerequisite contract before this canary is claimable.
The native acceptance recipe remains unchanged:
- two real, independent ChatGPT conversations using the production Dev MCP surface;
- exact loaded source/build/server/config/capability identity;
- one shared verified Swarm state realm, max workers=2;
- fixed 20 bounded task attempts accounted for separately from replay/conflict/wrong-worker negative calls;
- zero lost tasks, duplicate executions, cross-worker attribution, unauthorized submit acceptance and blind retry;
- restart-unknown and drain scenarios use the same task/attempt identities;
- no
agent_start, Codex/OpenCode/Agy worker substitution, OpenCLI/browser substitution, Node-client-only proxy, or repository mutation may be used to manufacture PASS.
Keep #49 OPEN until that physical ChatGPT-host evidence exists. #51 remains downstream of this gate.
- PR fix: bind dispatch to fresh catalogs and verified provider results #60 is MERGED (
James3014 commented
on Sep 11, 2026 OwnerAuthorMore actions2026-09-11 Ultra-parity / prerequisite reconciliation
Canonical parity tracker: #104.
Fresh state update:
- P0: Prevent live Dev MCP runtime lineage drift and capability regression #30 is now closed/completed after the runtime-lineage/capability-regression repair line merged.
- Harden Dev MCP reliability across host binding, worker startup, and timeout reconciliation #15 remains open and still owns the broader single-worker / host-binding / timeout-reconciliation reliability ceiling.
- Current
mainphysically contains Chat Swarm source/tool/config/lifecycle code; therefore test(chat-swarm): run isolated 2-worker ChatGPT live canary #49 should remain the first live peer-GPT canary, not a source implementation task.
Scope remains intentionally narrow
Keep this gate focused on proving:
Main GPT -> >=2 ordinary ChatGPT peer conversations -> join -> targeted/parallel dispatch -> exact worker/task attribution -> submit/collect -> same-worker follow-up -> restart/reconcile without blind duplicateDo not add Ultra elastic runtime/Auto Compact into this canary. Those are downstream #105/#51/#106 work under #104.
Updated prerequisite interpretation
Before running #49, bind fresh evidence for:
- exact current source/build/server/capability identity;
- host-visible
chat_swarm_*tool surface; DEVSPACE_CHAT_SWARMenabled in the isolated/live test realm;- Harden Dev MCP reliability across host binding, worker startup, and timeout reconciliation #15's remaining host/direct-recipient reliability condition as it materially affects this canary.
Do not keep #30 as an unresolved blocker now that it is closed; instead consume its exact post-closure lineage/capability evidence during the fresh runtime bind.
Passing #49 proves U1 peer-swarm baseline only. It does not prove elastic worker lifecycle, worker Auto Compact, or Main proactive continuity.
#49 最新設計裁決:保留 native canary,先由 #108 修 admission / 診斷缺口
Owner 要求重新細讀 Ultra 後提供其他 agent 實作。本註記補充現行 scope,不把 source closure 重開、不簽發 live PASS、不授權重啟或修改生產服務。
研究綁定:DevSpace main
8da19eb34e2fcdd9cf36faf2064911d5cbed7020;Ultra2c61c79ffe269d31523bac37ae6ecd60c9065a5e。source 細節、API、測試與 DoD 集中於 #108,不在本 Issue 另寫第二套 admission spec。1. 修正 live 證據上限
本串既有 Main 工具記錄曾成功讀 Canary identity 並 create Swarm;Owner 貼回的新 Worker 記錄顯示同一 Canary identity 的 cutover_status 成功,join 則回「此次工具調用已被 OpenAI 的安全檢查封鎖。請再次檢查你傳送的內容。」先前另有 conversation 不支援 developer MCP 的錯誤。這是不同 failure class,不能合併成同一原因。
- 已有 native Main discovery/invocation 歷史 witness,不等於所有 Worker action 都可用。
- Worker health/read action 成功,不等於 Swarm peer
_meta已驗證,也不等於 join action permission 通過。 - 現有 evidence 支持
HOST_REPORTED_ACTION_BLOCK,不支持已證明 invite 字串是 classifier 根因,也不支持已證明 request 未達 DevSpace。 - 沒有 fresh server ingress/receipt/TTL/action-permission evidence,就不得排除 invite expiry、peer identity、schema/handler 或其他因素。
- invite 預設 TTL 900 秒;舊 invite 不應反覆被宣稱為 fresh。不得貼出或持久化原始秘密。
目前最精確 disposition:
NATIVE_PEER_ADMISSION_NOT_PROVEN / HOST_REPORTED_ACTION_BLOCK;Main 曾有 G1 read/create witness,但 G2 尚未成功,#49 維持 OPEN。G1 證據不可過度推廣為 peer write/readback 全通。2. Admission 修復的唯一 owner
#108 新增 bounded readonly peer_status / owner roster inspect、typed errors、非秘密 pending request + exact owner approval + readback。沿用既有 session-bound next/submit、SQLite 與 coordinator。
新 flow 不承諾消除 OpenAI safety,必須按正常 app/action permissions 與 host-visible schema 審查。若仍被 Host 拒絕,保留 block;不得以 renamed/encoded credential、假 readOnlyHint、不同 carrier、模擬
_meta或重複改 prompt 製造 PASS。3. 明確區分領取與喚醒
當前 source 的 next 是 immediate current-task/claim,沒有 long-wait 或 event wake;task:null 不代表 background parked。dispatch 可把 task 先標 CLAIMED/BUSY,不能當作 worker 已收到、開始生成或已完成。
本 #49 第一條 native baseline 明列
deliveryMode=MANUAL_TURN_DRIVEN:兩個真 ChatGPT Worker 在有實際使用者回合時呼叫 next,執行受控任務並 submit。每次人為觸發都記錄,不宣稱零人工或自動背景持續。原定並行派工、targeted follow-up、結果隔離仍需真實成立。自動等待/喚醒由 #105 負責。若日後測 native long-wait,它是附加 delivery evidence,不得默默取代目前 native baseline。#105 不是本 #49 admission 修復的前置依賴。
4. 更新執行順序
- fix(chat-swarm): add observable, approval-bound native peer admission #108 G0 fresh evidence rebind + G1 source/independent acceptance。
- 授權後僅在隔離 Canary install/load;記錄新的 source/build/server/manifest 与 host-visible schema,不能硬編碼舊 10-tool count。
- Main + 兩個真 ChatGPT peer:peer_status -> join_request -> Main inspect/approve -> Worker readback。
- 同一 realm 完成原 test(chat-swarm): run isolated 2-worker ChatGPT live canary #49 固定 20 個 bounded attempts;replay/conflict/wrong-worker 的負向工具呼叫另計。
- 保留全部原有 restart / RECONCILE_REQUIRED / no duplicate / drain terminal-result 場景。20/20 accounted 不代表20個全部成功,UNKNOWN 不得算成功。
#46–#48 已 CLOSED 是 source 層結案,不重寫已交付 core。#30/#15 仍依本 Issue 原有 fresh runtime prerequisite 範圍驗證,不能只看 close flag。
驗收不縮水
兩個真 ChatGPT conversations、零跨 worker 歸因、零未授權 submit、零 blind replay、零 worker repo mutation。Node clients / SDK / CLI / browser adapters 只證明各自路徑,不能替代 native claim。
完成後仍只允許 exact tested build 的
MCP_PEER_SWARM_V0_1_LIVE_CANARY_PASS,並附 delivery mode 與人工步驟。不等於 #104 Ultra parity、auto wake、Auto Compact 或 production-ready。2026-09-11 fresh native-admission G1 evidence — backend updated, ChatGPT host schema stale
Fresh test from the current Main ChatGPT conversation against
dev-cafter PR #109 merge / Canary rebuild:- loaded server identity:
sourceCommit=a3cc54c1a7fc5b64560fb4df175531063f3b9f2a buildId=devspace-1.0.7-a3cc54c1serverInstanceId=3cd81812-912d-4e8b-a922-e63268fc5a46capabilityManifestSha256=9951ecd6b724bccf099c9d15403c00069c136f5f9652d9bcd4684a45d985cfd4mode=normalreconciliationRequired=false
However, the current ChatGPT host-visible
dev-cSwarm schema in this conversation still exposes only the original 10chat_swarm_*tools. The four #108 tools are absent from this conversation catalog:chat_swarm_peer_statuschat_swarm_inspectchat_swarm_join_requestchat_swarm_approve_join
Classification:
TOOL_SCHEMA_STALEat the ChatGPT host/conversation layer. This is not evidence of 7678 runtime deployment failure: the loaded Dev MCP identity is the new merge commit and the runtime is healthy. No new swarm/admission mutation was attempted from this stale-schema Main conversation.Next gate: obtain a fresh ChatGPT conversation/app binding that actually exposes all 14 Swarm tools, then resume the #49 native admission path (
peer_status -> join_request -> inspect -> approve_join -> peer_status BOUND -> next). Keep #49 OPEN until that real host-visible 14-tool path and the remaining two-peer scenarios pass.- loaded server identity:
2026-09-11 G1 fresh ChatGPT Host schema gate — PASS
Fresh independent ChatGPT Host witness on the rebuilt 7678 Canary now sees all four #108 native admission tools in
dev-c:chat_swarm_peer_statuschat_swarm_inspectchat_swarm_join_requestchat_swarm_approve_join
The same fresh conversation called
cutover_statusexactly once and bound to:- serverInstanceId:
3cd81812-912d-4e8b-a922-e63268fc5a46 - sourceCommit:
a3cc54c1a7fc5b64560fb4df175531063f3b9f2a - buildId:
devspace-1.0.7-a3cc54c1 - capabilityManifestSha256:
9951ecd6b724bccf099c9d15403c00069c136f5f9652d9bcd4684a45d985cfd4 - mode:
normal - reconciliationRequired:
false
Disposition:
ISSUE_49_G1_HOST_SCHEMA_PASS/READY_FOR_ISSUE_49_NATIVE_ADMISSION.No swarm was created during this schema probe. This does not prove worker admission, two-peer dispatch, restart reconciliation, drain behavior, or overall #49 PASS. The next gate must create a fresh 2-worker swarm in this exact fresh Main conversation so owner identity is bound to the conversation that will later call
chat_swarm_inspect/chat_swarm_approve_join.2026-09-11 #49 G2-A native admission + targeted routing — PASS
Fresh exact runtime lineage for this gate remains:
- sourceCommit:
a3cc54c1a7fc5b64560fb4df175531063f3b9f2a - buildId:
devspace-1.0.7-a3cc54c1 - serverInstanceId:
3cd81812-912d-4e8b-a922-e63268fc5a46 - capabilityManifestSha256:
9951ecd6b724bccf099c9d15403c00069c136f5f9652d9bcd4684a45d985cfd4 - mode:
normal; reconciliationRequired:false
Worker-A evidence:
- swarmId:
swarm_75935497104a456cbe0b68b9aa30bd30 - identity.status:
RESOLVED - state:
BOUND - label:
Worker-A - workerId:
worker_223ce139aac74f54928dabef26284562
Targeted routing proof:
- taskId:
task_474735f5315140ddb2f7575b4f4f1020 - taskKey:
issue49-worker-a-routing-proof-v1 - assignedWorkerId:
worker_223ce139aac74f54928dabef26284562 - worker received exact prompt and submitted exact result
ISSUE49-WORKER-A-ROUTING-PASS - submit reached
RESULT_READY - Main collected the same exact result
- terminal lifecycleState:
COLLECTED - collectedAt:
2026-09-11T16:26:30.647Z
Disposition:
- G2-A native admission: PASS
- G2-A authenticated worker next/routing: PASS
- G2-A submit: PASS
- G2-A Main collect: PASS
- Worker-B: NOT TESTED
- two-peer parallelism / replay-conflict / wrong-worker submit / restart-reconcile / drain scenarios: NOT YET TESTED
- test(chat-swarm): run isolated 2-worker ChatGPT live canary #49 overall: OPEN
Do not infer #49 or #104 completion from this single-arm pass.
- sourceCommit:
2026-09-12 #49 G2-B native admission + targeted routing — PASS
Worker-B evidence on the same isolated Canary realm:
- swarmId:
swarm_75935497104a456cbe0b68b9aa30bd30 - workerId:
worker_cb5d246b03fb45049f8521a34a7956fa - identity.status:
RESOLVED - state:
BOUND - label:
Worker-B - workerId is distinct from Worker-A
worker_223ce139aac74f54928dabef26284562
Targeted routing proof:
- taskId:
task_98a4731f076246ed86435bf8f198e488 - taskKey:
issue49-worker-b-routing-proof-v1 - assignedWorkerId:
worker_cb5d246b03fb45049f8521a34a7956fa - result:
ISSUE49-WORKER-B-ROUTING-PASS - terminal lifecycleState:
COLLECTED - collectedAt:
2026-09-12T02:13:38.946Z
Disposition:
- G2-A single-arm native admission/routing/submit/collect: PASS
- G2-B single-arm native admission/routing/submit/collect: PASS
- two distinct real ChatGPT peer identities: PROVEN
- remaining test(chat-swarm): run isolated 2-worker ChatGPT live canary #49 acceptance still includes the fixed 20 bounded attempts, replay/conflict negative calls, Worker-B wrong-worker submit rejection, restart/RECONCILE_REQUIRED/no-blind-retry, and drain terminal-result preservation.
- test(chat-swarm): run isolated 2-worker ChatGPT live canary #49 overall remains OPEN.
- swarmId:
2026-09-12 live canary checkpoint — 6/20 accounted; owner task-ledger observability gap exposed
Current exact physical evidence:
- swarm
swarm_75935497104a456cbe0b68b9aa30bd30ACTIVE - Worker-A
worker_223ce139aac74f54928dabef26284562AVAILABLE - Worker-B
worker_cb5d246b03fb45049f8521a34a7956faAVAILABLE - A/B remain distinct bound workers
- Attempt 01 A COLLECTED
- Attempt 02 B COLLECTED
- Attempt 03 A RESULT_READY =
ISSUE49-ATTEMPT-03-WORKER-A-PASS - Attempt 04 B RESULT_READY =
ISSUE49-ATTEMPT-04-WORKER-B-PASS|WRONG_WORKER=OWNERSHIP_CONFLICT - Attempt 05 A RESULT_READY =
ISSUE49-ATTEMPT-05-WORKER-A-PASS - Attempt 06 B RESULT_READY =
ISSUE49-ATTEMPT-06-WORKER-B-PASS
Negative evidence already proven:
- Attempt 05 identical replay -> same taskId / no duplicate
- Attempt 05 changed-material replay ->
REPLAY_CONFLICT - Attempt 04 wrong-worker submit ->
OWNERSHIP_CONFLICT
Disposition: 6/20 accounted; 20/20 not yet satisfied. No restart, drain or close performed.
New observability/product finding: current owner
chat_swarm_inspectdoes not provide a durable task ledger/history. If the controller conversation loses its locally remembered taskId manifest, it cannot rediscover exact historical task IDs forchat_swarm_status/collect even though the SQLite store still owns the durable tasks. This is not a current Swarm state-machine failure, but it is a real usability/recovery gap and should be addressed downstream so normal operation never depends on a human copying task IDs between conversations.- swarm
2026-09-12 controller orchestration blocker promoted to #112
The live 20-attempt batch exposed two core control-plane gaps and #49 should not paper over them with manual taskId transport:
- owner has no durable task ledger/list API, so controller context loss yields
TASK_LEDGER_CONTEXT_LOST; - targeted
chat_swarm_dispatch(preferredWorkerId=...)rejects a BUSY preferred worker instead of queueing the targeted task for laternext, preventing one-turn batch queueing for A/B.
Canonical source inspection confirms
getTask(id),getTaskByKey(swarmId, taskKey), andlistAttempts(taskId)exist internally, but no owner-facing task enumeration exists;dispatchTaskAtomiccurrently throws when a preferred worker is unavailable/busy.These are now owned by #112: #112
Current #49 live evidence remains valid: A/B are distinct real ChatGPT peers and both passed native admission + targeted routing + submit + collect. The 20-attempt batch remains incomplete and #49 stays OPEN. Do not continue by asking the Owner to manually shuttle taskIds; fix #112, then resume the batch with durable owner recovery and queue-safe targeted dispatch.
- owner has no durable task ledger/list API, so controller context loss yields
James3014 commented
on Sep 12, 2026 OwnerAuthorMore actions2026-09-12 live canary checkpoint — 20/20 bounded task gate PASS
Fresh native ChatGPT peer evidence for swarm
swarm_75935497104a456cbe0b68b9aa30bd30now satisfies the fixed 20-attempt task gate.Exact workers:
- Worker-A:
worker_223ce139aac74f54928dabef26284562 - Worker-B:
worker_cb5d246b03fb45049f8521a34a7956fa - both remain distinct bound workers and AVAILABLE
Durable ledger final readback:
- total tasks: 20
- COLLECTED: 20
- QUEUED: 0
- CLAIMED: 0
- RESULT_READY: 0
- RECONCILE_REQUIRED: 0
- Worker-A tasks: 10
- Worker-B tasks: 10
- lost tasks: 0
- cross-worker attribution: 0
- duplicate executions: 0
- unauthorized submit acceptance: 0
Routing/result checks:
- odd attempts 01/03/05/07/09/11/13/15/17/19 -> Worker-A
- even attempts 02/04/06/08/10/12/14/16/18/20 -> Worker-B
- Attempt 01 result:
ISSUE49-WORKER-A-ROUTING-PASS - Attempt 02 result:
ISSUE49-WORKER-B-ROUTING-PASS - Attempt 04 result retained:
ISSUE49-ATTEMPT-04-WORKER-B-PASS|WRONG_WORKER=OWNERSHIP_CONFLICT - Attempt 05 identical replay -> same taskId / no duplicate execution
- Attempt 05 changed-material replay ->
REPLAY_CONFLICT - Attempt 04 wrong-worker submit ->
OWNERSHIP_CONFLICT - Attempt 19 same-logical-worker follow-up ->
ISSUE49-ATTEMPT-19-WORKER-A-FOLLOWUP-PASS
This establishes
ISSUE49_20_OF_20_TASK_GATE_PASSfor the current isolated Canary runtime generation. It does not close #49.Remaining mandatory live gates from the Issue contract:
- controlled isolated-runtime restart while one exact task is nonterminal/ambiguous -> same task identity becomes/remains
RECONCILE_REQUIRED; zero automatic second dispatch / blind retry; explicit reconciliation only after bounded evidence; - cutover drain while one already-running worker is finishing -> new dispatch rejected while permitted terminal submit/result is preserved.
No #49 overall PASS, no production 7677 change, and no swarm close is claimed by this checkpoint.
- Worker-A:
The controller reconnected to the existing Main / Worker-A / Worker-B conversations and independently read the isolated Canary ledger. The 20 original tasks remain COLLECTED. Drain probe
task_8147f8106d944c3188bb448278e0f065remains CLAIMED byworker_cb5d246b03fb45049f8521a34a7956fa, with one attempt and no result. The negative-dispatch task does not exist; this absence alone does not prove a drain rejection because the negative dispatch has not run.Restart probe
task_6bcaad7e2a5446d998c7f274e15a9f4fis now terminal FAILED with one attempt and no result/requeue. Its current terminal task row has no assigned worker; the earlier same-worker restart evidence must be read from the pre-reconciliation witness, not inferred from that terminal row.Fresh 7678 health readback reports source
d6a725227652b3b7b936b59c0248ef635d18b1c4, builddevspace-1.0.7-d6a72522, manifesta1618a3a214741e3e8f443b22265f0dfa3037f8887412b1d7744b1472db22424, server instance6368d6de-67fd-4bc9-be1f-6c55007cc0de, cutover modenormal, andsource_dirty=true. This is runtime identity evidence, not a clean-package or overall acceptance claim.The existing pending carrier pairing expired without a binding. Main separately reports that its credential-resume call was rejected by ChatGPT's platform safety layer before DevSpace execution; no detailed platform error was exposed. That reported platform rejection and DevSpace's actual
A current paired carrier is requiredresponse are distinct. There is no visible approval UI identified by Main. No alternate transport was used to perform the rejected resume, and no new pairing/swarm/probe was created.The drain gate and overall #49 acceptance remain open. Production 7677 was not operated. Independent #105 feasibility/source work can proceed under its explicit gate split, without implying that #49 passed.
James3014 commented
on Sep 13, 2026 OwnerAuthorMore actions2026-09-13 exact native resume manifest — restart + drain only
This comment supersedes any older ad-hoc resume instructions. It performs no native effect and grants no production authority. The 20/20 task gate remains accepted from the existing isolated Canary evidence; only the two original mandatory native lifecycle gates remain.
Frozen existing identities
- swarm:
swarm_75935497104a456cbe0b68b9aa30bd30 - Worker-A:
worker_223ce139aac74f54928dabef26284562 - Worker-B:
worker_cb5d246b03fb45049f8521a34a7956fa - 20 original bounded tasks: all
COLLECTED, A=10 / B=10 - existing drain probe:
task_8147f8106d944c3188bb448278e0f065, last durable evidenceCLAIMEDby Worker-B, one attempt, no result - prior restart probe:
task_6bcaad7e2a5446d998c7f274e15a9f4f, now terminalFAILED, one attempt, no result/requeue; do not infer the historical restart transition from this terminal row
G-R0 — fresh native rebind before any mutation
Use a fresh Main ChatGPT conversation that actually exposes the current
dev-ctool schema.- Call
cutover_statusread-only and record exact source/build/server/manifest/mode/reconciliation state. - Read the existing Swarm/task ledger and re-confirm the identities above.
- Fail closed if the loaded isolated runtime is dirty, stale, in reconciliation, not the intended Canary realm, or its host-visible schema lacks the required Swarm/cutover tools. The most recent historical health row had
source_dirty=true; that row is evidence only and is not an acceptance-ready starting point. - Re-establish any required Main/Worker carrier pairing only through the supported product pairing flow. Do not encode/rename credentials, invent an approval UI, switch carrier, or use an alternate transport to bypass a ChatGPT platform rejection. If a current pairing cannot be established, stop as
PAIRING_REQUIRED.
Production 7677 is out of scope.
G-R1 — restart / no-blind-retry witness
The prior restart probe is terminal and its final row is insufficient by itself. On a clean accepted isolated runtime, create one new bounded restart probe only after proving the old probe has no active attempt.
- suggested taskKey:
issue49-restart-reconcile-probe-v2 - target one already-bound peer (prefer Worker-A unless current evidence requires otherwise)
- worker claims the task so the exact task/attempt is nonterminal
- perform the normal controlled isolated Dev MCP restart/cutover path once
- reconnect Main/worker and rediscover the probe through the durable owner ledger; do not transport a remembered taskId manually as the authority source
Required PASS evidence on the same task and same attempt:
- state becomes/remains
RECONCILE_REQUIREDafter restart uncertainty; - attempt count remains exactly 1;
- no second task, second attempt, automatic redispatch, or blind retry is created;
- explicit reconciliation acts on that same durable identity only after bounded evidence;
- exact pre/post server instance and source/build/manifest lineage is captured.
If the effect/result cannot be proven, leave it unresolved; do not create v3 to escape uncertainty.
G-D1 — drain rejection + in-flight terminal preservation
First re-read
task_8147f8106d944c3188bb448278e0f065.- If it is still
CLAIMEDwith exactly one Worker-B attempt and no result, reuse it as the in-flight task; do not create a second drain probe. - If its durable state has changed, record the exact state and stop for reconciliation before creating new physical work.
Establish the exact isolated cutover lease through the currently authorized cutover flow, then enter drain with the host-visible
cutover_drain(cutoverId=...)action.While drain is active:
- issue exactly one fresh negative
chat_swarm_dispatchwith a unique taskKey; - require the dispatch to be rejected by drain admission and verify no durable task for that key was created;
- allow Worker-B to finish the already-running drain probe through the normal existing
next/submitpath; - require Main to read/collect the terminal result successfully while still preserving the same task/attempt identity.
PASS requires both halves:
new dispatch rejectedandexisting in-flight terminal submit/result preserved. Mere task absence without executing the negative dispatch is not evidence.G-CLOSE — claim ceiling
Only after G-R0 + G-R1 + G-D1 all pass may #49 claim, for the exact tested isolated runtime only:
MCP_PEER_SWARM_V0_1_LIVE_CANARY_PASSBefore closure, record:
- source/build/server instance/capability manifest before and after restart;
- exact restart task/attempt transition and no-duplicate readback;
- exact cutoverId/drain evidence;
- negative dispatch rejection and absence from ledger;
- drain probe terminal result preservation;
- confirmation of no Codex/provider/repository mutation by peer workers.
Do not infer #104 Ultra parity, #105 zero-touch wake/scale, #51 Auto Compact, production readiness, or production deployment from this gate.
- swarm:
James3014 commented
on Sep 27, 2026 OwnerAuthorMore actionsWave 4 reconciliation — Custom Dev MCP two-worker live canary retired
Classification:
SUPERSEDED/RETIRED_FOR_CURRENT_DIRECT_CONTROL_PATH.Owner architecture decision now proven by Nexus-new #1154:
- CoS is the current ChatGPT peer-worker/browser runtime for Main ChatGPT direct control.
- Desktop Commander is the direct physical host/repository tool.
- Main ChatGPT remains the sole mutation coordinator.
- Wave 3 reached
CHATGPT_DIRECT_CONTROL_PLANE_CANARY_PASSwithout Dev MCP. - DevSpace remains an optional compatibility transport when explicitly selected; it is not a required direct-control dependency.
Decision:
The custom Dev MCP peer-swarm canary is no longer required for the current direct-control path. Its idempotency/restart/reconciliation contract remains historical design evidence, but do not resume it by default.Residual boundary:
If DevSpace ChatSwarm is intentionally revived for a separate product goal, it requires a new explicit Owner contract and fresh runtime evidence rather than reopening this canary by implication.This close does not rewrite the Issue's historical implementation/evidence as wrong or completed. It records that this DevSpace-specific implementation line is no longer the selected current path.
Causal watermark:
- Issue prior
updated_at=2026-09-20T07:11:57Z - DevSpace main
3b92d6165e8531765dbaf929c3c298e0b03505bb - P1: pilot Chat On Steroids as ChatGPT peer-worker runtime #218 is now closed as completed CoS pilot
- Nexus-new Wave 4 Candidate: PR #1158 @
ba3036020c7d87c90f7cf1c4cfa84fa613b13ce2
AUTO_CHAIN=false.
Owner scope amendment — macOS only (2026-09-11)
依 Owner 最新裁決,本工作線相關功能設計、實作、交付與驗收 只處理 macOS。本裁決取代下文較早的跨平台/三平台要求。
Status
BLOCKED_BY #48, #30, #15
Purpose
Prove the first real ChatGPT Swarm path with two ChatGPT peer conversations using the production Dev MCP tool surface, while keeping the test isolated from the current production runtime and from coding-agent mutation authority.
This is the first multi-worker runtime gate. It must not start from source readiness alone.
Hard prerequisites
Do not infer these prerequisites from Issue closure alone; bind fresh runtime evidence before the canary.
Why #30 and #15 are hard blockers
The currently connected Dev MCP has historically lagged canonical
main, and #30 exists specifically because a stale/diverged runtime previously removed promoted MCP capability. #15 is the umbrella reliability contract that explicitly gates return to unattended multi-worker expansion. A Swarm live PASS on an unbound/stale runtime would not be claimable.Test realm
Use an isolated Dev MCP deployment/config for the canary where practical:
Do not mutate the current dirty source checkout.
Worker model
V0.1 workers are ordinary ChatGPT peer conversations connected to Dev MCP.
They must not invoke:
agent_startagent_continuecodex_goal_*The live canary is testing conversation coordination, identity, durability, dispatch/result isolation, and restart reconciliation — not coding ability.
Required scenario
taskKey/request; prove no duplicate execution.taskKeywith changed material input; prove conflict rejection.RECONCILE_REQUIREDand is not automatically dispatched a second time.Acceptance
Required minimum live evidence:
agent_start/ coding-provider dispatches attributable to this canaryEvidence receipt
Capture exact:
Do not persist raw ChatGPT session identifiers or auth tokens in GitHub evidence.
Failure classification
A failure must distinguish at least:
Do not retry an ambiguous task under a new attempt until reconciliation proves it is safe.
Non-goals
Claim ceiling
If all gates pass:
MCP_PEER_SWARM_V0_1_LIVE_CANARY_PASSfor the exact tested build/runtime only.Do not claim OpenCLI support, durable context rollover, production multi-worker readiness, or unrestricted autonomous engineering.
Next gate
After this live canary passes, #50 may pilot an optional OpenCLI lifecycle/runtime adapter and #51 may add durable conversation rollover/continuity.