Skip to content

fix: keep the anchoring worker alive when a transaction cannot be sent - #34

Merged
nishchal-gond merged 3 commits into
masterfrom
fix/anchor-send-guard
Sep 21, 2026
Merged

nishchal-gond merged 3 commits into
masterfrom
fix/anchor-send-guard

Conversation

@nishchal-gond

@nishchal-gond nishchal-gond commented Sep 21, 2026 •

Copy link
Copy Markdown
Owner

Requested by LPH · project thread

Before: the anchoring worker died the first time a transaction could not be submitted. anchorBatch() awaited the transaction hash and only then attached a handler to the send promise, so that handler covered the case where the send had already succeeded. A failure at submission — an address out of gas money, a nonce the node rejects, a revert the node catches before mining — rejected the awaited promise, the line attaching the handler was never reached, and the send promise's own rejection went unhandled and took the process down. Anchoring then stopped dead on a faucet running dry, which is precisely the failure a worker exists to sit through: the batch was still in the store, and the next tick would have retried it.

And a gateway that was working correctly looked broken. Receipt one document on a quiet service and nothing is anchored until the age threshold expires an hour later. That is the design, but from outside — no batches, no transactions, a worker ticking over silently — it is indistinguishable from anchoring being dead, and neither threshold could be changed.

After: the handler is attached before the await, so a send-time rejection is reported and the worker lives to retry on the next tick. OREOCHAIN_BATCH_MAX_SIZE and OREOCHAIN_BATCH_MAX_AGE_MS make both thresholds an operator's choice, GET /api/proofs/status reports the values actually in force, and a waiting worker says once per waiting spell what it is waiting for, with those numbers in the line.

How

The handler returns early while txHash is still null, because at that point the caller is about to be told by the rejection anyway and reports it as "cannot anchor batch"; logging it twice would suggest two different failures. Defaults are unchanged, and nothing is lost while a batch waits either way — the receipt already proves the service accepted the document, and the anchor upgrades that to a public fact.

Two operator-facing fixes ride along. npm run deploy-contract without --confirm now prints the deployer, its balance and the measured gas rather than refusing on an unfunded key, so a dry run is how you find out what to fund; the refusal moved to the send, where it belongs. And npm run add-exporter against an address with no contract code says so by name — "no contract code at 0x… on chain 11155111. Either the address is wrong, or OREOCHAIN_CHAIN_RPC points at a different network from the one it was deployed to" — instead of failing to decode a call to nothing.

Review notes

Reviewed from a second session. The branch's own verification, on a local node with Sepolia's chain id, was: drain the anchoring address, upload, and the worker logs cannot anchor batch on five consecutive ticks with no unhandled rejection; top it up without a restart and it sends and anchors on a later tick, with the second document's inclusion proof carrying a real transaction hash. That reproduction was a genuine submission failure rather than a simulated one — estimateGas from the zero-balance account returned 1,905,267 in 23 ms, the same as funded, so the rejection happened at send(), exactly the path this fixes.

I checked the guard by breaking it. Restoring the old ordering — the handler attached after the await — turns the anchor-worker file red:

# Error: A resource generated asynchronous activity after the test ended. This
activity created the error "Error: Returned error: insufficient funds for gas *
price + value" which triggered an uncaughtException event, caught by the test
runner.

That is the production failure exactly: in the worker there is no test runner to catch it, and the process exits. Suppressing the waiting line fails its own test by name (not ok 2 — a quiet gateway says what it is waiting for, once, rather than nothing).

One note for whoever touches this next, not a blocker. Under that mutation every named test still passes, including a transaction that fails at submission is reported, and the worker lives; only the file-level uncaught exception turns it red. The regression is caught, which is what matters, but a future reader meets "a resource generated asynchronous activity after the test ended" rather than a sentence about anchoring. An explicit process.on("unhandledRejection") assertion inside that test would name what broke.

Also checked: createProofService({...}) carries all four of verifier, batchMaxSize, batchMaxAgeMs and keyringPath after the master merge — the conflict there could have silently dropped two settings — and the waiting flag resets on both shouldFlush and pending === 0, so the line is said once per spell rather than once per process.

Merging this

This branch is a rebuild. The original fix/anchor-worker carried Co-Authored-By and Claude-Session trailers on its commits, which are not wanted in this repository's history, and could not be rewritten in place, so the work was reassembled from current master as one clean commit plus its master merges. git log --format=%B origin/master..HEAD | grep -E 'Co-Authored-By|Claude-Session' returns nothing and every commit is authored by the repository owner; I checked both rather than taking it on trust.

The first master merge had six conflicts, all read individually: the shared doc comment in server/config.mjs (the shape that has produced a SyntaxError inside an object literal on this repository before, caught there by node --check), the createProofService call above, .env.example and server/README.md purely additive, and the two generated counts. Master then moved twice more while this was in review, so it was merged forward again: that second merge was the two count lines only, and the diff against master stayed at 14 files, which is how I know nothing else came through it. The last master commit touches only js/ and the browser suite, and this branch touches no file it touches.

Duplicate-declaration scan clean, every module passes node --check, and the total is 554 — master's 549 plus this branch's 5.

Tests

554 tests pass on Node 18, 20, 22 and 24.

From a run of the whole anchoring path against a real node.

The worker died when a send failed. The guard that logs a late rejection
instead of letting it kill the process was attached after the await, so it
only ever covered a transaction that had already been accepted. A failure at
submission — an address out of gas money, a nonce the node rejects, a revert
caught before mining — rejected the awaited promise, threw past the guard,
and the send promise's own rejection reached the unhandled-rejection handler.
Anchoring stopped dead on a faucet running dry, which is precisely the
failure a loop is supposed to sit through: the batch is still in the store,
and the next tick would have retried it. The guard now goes on before the
await, and says nothing before a hash exists, because the caller is already
reporting that one as "cannot anchor batch".

Batch size and age were locked at 1000 documents or one hour. The tests
passed both; nothing in the environment could. A quiet gateway receipts one
document and then anchors nothing for an hour, which is the design working
and reads from outside as anchoring being broken — the first person to hit it
had to flush a batch by hand to find out which. They are now
OREOCHAIN_BATCH_MAX_SIZE and OREOCHAIN_BATCH_MAX_AGE_MS, defaults unchanged,
reported on the gateway's status so the worker can name the numbers actually
in force, and the worker says once per waiting spell what it is waiting for.

Both chain scripts refused to do anything at all from an unfunded address,
because connect() threw on a zero balance before either had printed a line.
That made the two things a dry run exists for — which address to fund, and
how much gas this needs — unavailable until after you had funded it. The
balance is now reported, the refusal happens at the point of sending where
--confirm already is, and both scripts print a gas estimate on a dry run.
add-exporter also checks for code at the contract address first, so the
wrong-network mistake is named instead of surfacing as web3's "Parameter
decoding error" on empty return bytes.

server/README.md takes the runbook material from that run: what a deployment
costs in gas and what to fund, keeping the anchoring address topped up, what
an unreachable endpoint looks like, and why nothing has been anchored yet.
The testnet section said everything behaves identically on a testnet, which
was wrong by exactly the first fix above.

Each of the three fixes was checked against the unfixed code and fails there.

Rebuilt on master rather than merged into it: the original branch carried
commit trailers that do not belong in this repository's history, and its
history could not be rewritten in place.
No conflicts: this branch and the receipt-claims change touch server/proofs.mjs
in different places, and server/README.md in different sections.

Every line master has that this tree does not was read: five in
scripts/chain-tools.mjs, three in add-exporter, five in deploy-contract, five
in server/chain.mjs and nine in server/README.md. All of them are lines this
branch replaces on purpose — connect() throwing on a zero balance before
either script had printed a line, the send guard attached after the await, and
the testnet section claiming a testnet behaves identically, which the first fix
here disproves.

542 = 537 on master plus this branch's five. Count regenerated, every server
and scripts module re-parsed, duplicate-declaration scan clean.
Count lines only: #32 touches js/ and this branch touches server/ and
scripts/, so the two generated totals in Readme.md and index.html were the
only conflicts. Resolved per hunk and regenerated to 554 = 549 + 5; the
diff against master is unchanged at 14 files, so nothing came through this
merge that was not already here.
@nishchal-gond
nishchal-gond merged commit 8398903 into master Sep 21, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant