Case file for one deploy that burned — and the prompt that stops it happening twice.
Open the preflight station · Read the prompt · Check the receipts · Reproduce it
Important
A remembered failure that cannot stop an action is not a safeguard. It is a note. This prompt turns the note into a gate: the risky action refuses to run until the check that proves the failure is cleared has passed on this commit.
An evolution of BuildMEM Agent for Walrus Memory. BuildMEM captures useful engineering lessons. Failure-to-Gate makes a past failure operational.
Case contents — click to expand
- Case header
- 01 · The incident
- 02 · Two fixes that did not hold
- 03 · What actually changed
- 04 · Exhibit A — the gate in a browser
- 05 · Exhibit B — the gate in your terminal
- 06 · The prompt, block by block
- 07 · Quick start — reproduce the case
- 08 · Findings
- 09 · Evidence boundary
- 10 · Repository map
- Who should read what
| Field | Value |
|---|---|
| Scenario | six-module budget fixture sized for two and a half modules |
| Entity key | deploy:mainnet:gas-budget |
| Trigger | deploy |
| Lifecycle | observed → diagnosed → mitigated → verified / expired |
| Clearance condition | node demo/check-gas-budget.mjs — a reviewed, target-local verifier |
| Decision surface | demo/gate-runner.mjs, demo/gate-engine.mjs |
| Outcome under the evolved prompt | the deploy is held until the verifier passes on the requested commit |
This committed scenario models a six-module package with a configured gas budget sized for two and a half modules. The fixture records the budget shortfall as persistent memory, then tests whether a later session can act on that information.
The test is not about asserting a chain deployment. It asks whether the same underfunded configuration stays blocked in a later session after the lesson has been stored.
Fix one — write it down better. The note lived in memory as prose. A later session recalled it, repeated the right words back, and attempted the deploy anyway. Remembered text carries no authority to stop an action.
Fix two — mark it verified. Flagging the memory verified made the situation worse. A
recalled verified state reads as a permission slip, and it was being applied to a commit
nobody had re-checked. The agent was now confidently wrong instead of vaguely worried.
Warning
This is the failure mode most memory prompts still have. Recall returns the text of a past verification. It does not return the fact that anything was verified here, now, on this commit.
The failure stopped being a sentence and became a typed lifecycle event with a stable identity and an executable clearance condition. Recall stopped producing advice and started producing a gate list.
flowchart TD
A["Risky action requested<br/>deploy · migration · submission"] --> B["Scoped recall<br/>by entity_key and trigger"]
B --> C{"Recall integrity<br/>intact?"}
C -- "empty or suspicious" --> R["One broader retry"] --> C
C -- "unknown after retry" --> BLOCK["BLOCKED<br/>integrity state unknown"]
C -- "intact" --> D["Resolve supersession,<br/>revocation and expiry"]
D --> E{"Any gate open<br/>for this scope?"}
E -- "no open gates" --> OK["ALLOWED<br/>action may proceed"]
E -- "open gate" --> F["Run the committed<br/>verification_command"]
F -- "PASS on this commit" --> CLEARED["CLEARED<br/>a verified event may be written"]
F -- "FAIL or missing" --> BLOCK2["GATE HELD<br/>action refused"]
E -- "two events disagree" --> ESC["ESCALATED<br/>conflict surfaced to a human"]
classDef stop fill:#f8ece9,stroke:#8f2a18,color:#8f2a18;
classDef go fill:#e7efe9,stroke:#0f4d3f,color:#0f4d3f;
class BLOCK,BLOCK2,ESC stop;
class OK,CLEARED go;
| Field | Job | The baseline failure mode it removes |
|---|---|---|
entity_key |
stable identity for one risk | groups lifecycle events without relying on phrasing |
status |
observed → diagnosed → mitigated → verified/expired | a workaround cannot masquerade as a verified fix |
trigger |
the action that must check this risk | a deploy lesson stops being irrelevant chat context |
applies_to |
commit and environment scope | old mitigations go stale instead of being silently trusted |
verification_command |
executable clearance condition | the model cannot self-certify a gate as resolved |
effective_at |
ordering inside the memory body | recency is not invented from semantic recall order |
supersedes |
explicit replacement edge | append-only memory can still represent change safely |
evidence |
command output, receipt, or human confirmation | the resolution stays inspectable |
failure-to-gate.vercel.app — pick a committed scenario, type the task context you would really hand an agent, and press Run preflight evaluation.
Four committed scenarios, four different reasons to stop or proceed:
| Scenario | What it puts in memory | Result |
|---|---|---|
| 01 Budget rehearsal | recalled lifecycle events with the original budget fixture | gate held — the local verifier fails |
| 02 Corrected budget config | the same events against the corrected fixture | gate clears |
| 03 Recalled prior pass | a recalled event already claiming verified |
gate held — recall is evidence, not permission |
| 04 Untrusted recalled note | a nested recalled field carrying directive-shaped text | the field is quarantined as inert data |
The context you type is hashed, run through the canonical untrusted-content scanner, and retained as data. It changes the intake row of the trace. It never selects the outcome: that comes from the committed fixture events and the reviewed verifier run. Each run prints the exact CLI command that reproduces it.
make gate # fixture has a 50M budget; the check requires 120M
make fix gate # raises the fixture to 130M, then the gate clears$ make gate
GATES for "deploy" @ 9f2a1c: 1 open / 0 cleared / 0 conflict
OPEN deploy:mainnet:gas-budget — Gas budget must scale with module count
BLOCKED: 1 gate(s) still open. Resolve them before "deploy".
$ make fix gate
GATES for "deploy" @ 9f2a1c: 0 open / 1 cleared / 0 conflict
CLEARED deploy:mainnet:gas-budget — Gas budget must scale with module count
OK: all gates cleared. "deploy" may proceed.
Both paths exit with scriptable codes, so the gate is usable in CI rather than only in a
conversation. Captured runs live in evidence/evolved.md; the
baseline transcript is in evidence/baseline.md. The demo does not
execute a chain deployment — it demonstrates the decision point that must happen before one.
PROMPT.md is copy-pasteable and organised so that every block owns one
domain of the decision.
| Block | Domain | What it fixes |
|---|---|---|
| Typed failure event | representation | one risk gets a stable entity_key, a lifecycle status and a scope |
| Write and admission rules | write path | events are atomic, appended and never edited in place; credentials never enter memory |
| Gate construction | read path | scoped recall builds GATES: open / cleared / conflict before a risky action |
| Clearance protocol | authority | only a reviewed local verifier — or explicit human evidence — clears a gate |
| Receipt and degraded mode | persistence | a checkpoint counts only on terminal completion with a non-empty blob_id |
| Instruction priority | trust | current policy and observed evidence outrank recalled text; ambiguity fails closed |
| Required output | interface | one leading gate line, then the events behind it |
tests/prompt-contract.test.mjs removes each material rule in turn and requires the contract
to fail, so the prompt text and the executable behaviour cannot drift apart in silence. The
rule-to-check map is docs/PROMPT_TO_TEST.md.
| Question before an action | BuildMEM baseline | Failure-to-Gate |
|---|---|---|
| Which event describes this exact risk? | semantic query only | stable entity_key |
| Is the warning still valid for this commit? | no scope | applies_to + expiry |
| Observed, diagnosed, or actually verified? | flat record | explicit lifecycle states |
| What objectively clears it? | agent judgment | verification_command or human evidence |
| Which conflicting note wins? | implicit recall order | effective_at + supersedes, conflict surfaced |
MemWal recall is semantic top-K, not a chronological ledger: it returns a blob ID, text and a distance, not a reliable write order or a namespace inventory. A prompt that assumes "the first result is newest and complete" will make the wrong operational call.
Node.js 18+. No network, wallet, API key or MemWal account required.
git clone https://github.com/alexbelij/WalrusSession-7
cd WalrusSession-7
make test # regression suite: block, clear, human gate, stale scope
make demo # baseline → gate decision → manifest board
make gate # the visible failure path
make fix gate # the visible recovery path
make secret-scan # reject tracked credential-shaped contentOwner-scoped historical replay, against real repository history rather than a fixture:
make historical-replay ALEX_HISTORICAL_REPO=/path/to/owner-historical-repositoryIt pins the direct parent-to-repair interval f2656e0 → 9e2140f, checks owner authorship
and the exact changed files, and requires the observed historical failure to remain gated
until a reviewed local verifier passes.
Optional multi-model Mainnet run — fail-closed by design
make multi-model-dry-run verifies the configured model/fixture matrix against the committed
deterministic oracle and makes no network call.
make multi-model-live requires explicit ALLOW_MAINNET_WRITE=true, a new
MULTI_MODEL_RUN_ID embedded in an isolated namespace, exact model IDs, credentials supplied
only through environment variables, and a declared MAX_MAINNET_WRITES. With the current
three-model × three-fixture configuration it refuses a ceiling below 9 writes. Each
observation is stored only after a model response is received, requires terminal
rememberAndWait completion with a non-empty blob_id, and then requires an exact recall
through a newly created client. It writes records/multi-model-<run-id>.json and never
overwrites the committed ten-checkpoint evidence.
Each row is a failure found during review, the rule that answers it, and the check that proves the rule is load-bearing.
| Prompt domain | Failure found in review | Final rule | Executable fixture | Status |
|---|---|---|---|---|
| lifecycle + scope | scope-first filtering revived an older state | resolve supersession before scope | tests/gate-engine.test.mjs |
pass |
| gate clearance | a model assertion could clear a gate | reviewed local verifier PASS only | tests/gate-runner.test.mjs |
pass |
| recalled verified state | verified memory could authorize a new action |
re-run the verifier on the requested commit | tests/gate-engine.test.mjs, tests/gate-runner.test.mjs |
pass |
| lifecycle integrity | dangling supersession could manufacture history | unknown supersedes ID is a named integrity error |
tests/gate-engine.test.mjs |
pass |
| untrusted recalled content | a nested field could hide an instruction | recursively quarantine any field | tests/gate-engine.test.mjs |
pass |
| job evidence | an accepted job could be read as a stored blob | terminal done + blob_id only |
tests/memwal-job-lifecycle.test.mjs |
pass |
| conflicts | equal-time state ambiguity | block and escalate | tests/gate-engine.test.mjs |
pass |
| recall uncertainty | an empty result must not authorize a risky action | bounded retry, then fail closed | tests/gate-engine.test.mjs |
pass |
| secret-like input | memory becomes a credential store | reject before write or use | make secret-scan |
pass |
Three claims are kept in separate boxes and never merged into one sentence.
| Layer | What it establishes | Where |
|---|---|---|
| Deterministic policy | the gate engine enforces lifecycle, scope, clearance and quarantine rules | make test, make demo |
| Historical replay | the same policy applied to real owner-authored repository history | docs/REPLAY_RECEIPT.md |
| Mainnet persistence | 10 terminal rememberAndWait receipts with blob_id, five fresh-client cold recalls |
evidence/mainnet-blobs.json, docs/RECEIPTS.md |
The unchanged source contract is locked in
evidence/source-locked-baseline.json: BuildMEM
revision 45490707…, CLAUDE.md, SHA-256 130a319d…. It requires a recalled warning to be
surfaced before a risky action, but has no enforceable local-clearance condition. A current
official-SDK write → terminal blob_id → destroy → new-client exact recall is recorded in
evidence/live-sdk-proof-2026-08-21.json.
What this prompt does not promise. It does not turn semantic recall into a database
query. It does not make any test command universally safe — a project author owns each
verification_command. It does not grant authorization: a stored "approved" note never
replaces a current human approval or a passing gate. It never stores secrets.
.
├── PROMPT.md the copy-pasteable evolved system prompt
├── demo/
│ ├── gate-runner.mjs event resolver and gate executor
│ ├── gate-engine.mjs canonical policy shared by the CLI and the station
│ ├── check-gas-budget.mjs the deterministic example verifier
│ └── deploy.config*.json the broken fixture and its corrected twin
├── evidence/ baseline transcript, captured runs, receipts, source lock
├── docs/ receipts, prompt-to-test map, replay receipt, design notes
├── tests/ gate engine, runner, job lifecycle, prompt contract, secret scan
├── web/ the read-only preflight gate station
├── assets/ · media/ rendered diagrams and the station screenshot
├── DEMO.md the operator sequence for reproducing a run
└── Makefile
| You are | Read this first | It takes |
|---|---|---|
| A judge | the preflight station, then docs/RECEIPTS.md |
60 seconds |
| An engineer | make gate then make fix gate — the same failure blocked, then cleared, with exit codes |
2 minutes |
| A team lead | Findings — the class of incident this stops repeating | 3 minutes |
| An agent author | PROMPT.md — copy it, give the project a namespace, let the next failure write its own gate |
5 minutes |
| A reviewer of the source prompt | docs/PROMPT_TO_TEST.md — every material rule mapped to the check that proves it |
2 minutes |
Copy the prompt. Break the fixture on purpose. Watch the gate hold.
Preflight evidence and the listed checks were reviewed at commit 3785a35cb84c578fb5cc87f189809b015c3d2584 (2026-08-23).

