This document describes what the system protects, how, and — just as importantly — what it does not protect against. Read the limitations section before describing OREOCHAIN as secure to anyone.
A reasonable instinct when building a security product is to invent a new encryption algorithm, on the theory that something nobody has seen cannot be broken. In cryptography this instinct is backwards, and it is worth being precise about why.
The security of AES is not a property of its design document. It is the accumulated result of roughly twenty-five years of public cryptanalysis: every academic group with an interest has attacked it, published what they found, and failed to break the full-round cipher. That body of failed attacks is the evidence. ChaCha20 has a comparable, somewhat shorter record.
A cipher written this month has none of that. It is not "unbroken" — it is unexamined, which is a different thing that looks identical from the inside. Historically, new ciphers from non-specialists are broken quickly once anyone looks, and the author is the person least able to find the flaw, because the same blind spot that produced it hides it.
This applies to modifying a standard cipher too. Changing AES's round count, S-box or key schedule does not produce "AES, improved". It produces a new, unanalyzed cipher wearing a trusted name — the worst of both worlds, because it invites the trust without earning it.
So OREOCHAIN uses standard primitives, from the platform's own implementation:
| Purpose | Primitive |
|---|---|
| Chunk encryption | AES-256-GCM, XChaCha20-Poly1305, or both |
| Key derivation | HKDF-SHA256 |
| Passphrase stretching | Argon2id (64 MiB, 2 passes) — PBKDF2-SHA256 read-only for legacy files |
| Content hashing | SHA-256 |
| Integrity commitment | Merkle tree, domain-separated |
The engineering novelty is in the composition, not the primitives. That is where real improvement is available, and the rest of this document is about it.
Most "encrypt the file then store it" systems use one key and one nonce for the whole object. OREOCHAIN derives an independent key and nonce for every chunk:
chunkKey(i), chunkNonce(i) = HKDF-SHA256(
ikm = fileKey,
salt = fileSalt,
info = "oreochain-envelope-v1/<suite>/chunk/<i>"
)
Two consequences:
-
Blast radius. Recovering one chunk's key — through a memory disclosure, a side channel, a bad random number on one device — reveals that chunk and nothing else. There is no single key whose loss decrypts the archive.
-
Nonce reuse becomes structurally impossible. Reusing a nonce under the same key is the failure that destroys AES-GCM completely: it leaks the authentication subkey and lets an attacker forge arbitrary messages. Because every chunk key here is used for exactly one encryption, the condition cannot arise, no matter how many chunks a file has or how many files share a passphrase.
Every chunk's AEAD tag covers additional authenticated data naming the file, the chunk's index and the total chunk count:
AAD = "oreochain-envelope-v1|<fileHash>|<index>|<totalChunks>"
The reader reconstructs this independently from the manifest. That converts a whole class of attacks from "silently succeeds" into "fails to decrypt":
| Attack | Result |
|---|---|
| Reorder two chunks | Both fail authentication |
| Duplicate a chunk | Fails at the duplicated position |
| Truncate the file and adjust the count | Every chunk fails |
| Splice in a chunk from another file | Fails |
Without position binding, an attacker who cannot read your data can still rearrange it — reordering the pages of a contract, say — and the file still decrypts cleanly. Here it does not decrypt at all.
Chunk hashes are folded into a Merkle tree whose root goes on-chain. The tree is domain-separated:
leaf(chunk) = SHA-256(0x00 || chunk)
node(left, right) = SHA-256(0x01 || left || right)
The 0x00/0x01 prefixes prevent an attacker from presenting a chunk hash as
an internal node and substituting a whole subtree for a single chunk — the
standard second-preimage attack on naive Merkle trees. A level with an odd
number of nodes promotes its last node unchanged rather than hashing it against
a copy of itself, which would make two different chunk lists collide.
This buys something a plain file hash cannot: per-chunk proof. One 256 KiB
block can be proven to belong to a registered document, on-chain, via
verifyChunk, without downloading, decrypting or disclosing the rest. Proving a
block of a 10 GB file costs the same as proving a block of a 1 MB file.
test/contract.test.js compiles the contract and runs it in a real EVM to check
the Solidity tree and the JavaScript tree produce identical roots for every tree
shape. A silent divergence there would break proofs only in production.
Three suites are available, selected per file and recorded in the manifest:
aes-256-gcm— hardware-accelerated wherever AES-NI exists, which is most machines. Default.xchacha20-poly1305— faster in pure software (older phones, embedded targets, WASM), and free of the cache-timing side channels that table-driven AES implementations exhibit. Its 192-bit nonce makes random nonces safe at any volume.cascade-aes-xchacha— inner AES-256-GCM, then outer XChaCha20-Poly1305, under independently derived keys. Plaintext stays sealed unless both ciphers are broken. Roughly double the CPU cost.
Agility is the part that matters long-term. Because the suite is recorded per file, the format can adopt a post-quantum AEAD later without invalidating anything already stored. A format that can migrate is worth more than any single cipher choice — and it is the improvement a bespoke cipher actively prevents, since a proprietary format has nowhere to migrate to.
The cascade is a genuine hedge, but be honest about the threat it addresses: a catastrophic break of AES-256, which no serious cryptographer expects. It is the right default for data that must stay secret for decades; it is overkill for a routine document.
Every other defence here assumes the attacker does not have the file key. The file key is wrapped by a key derived from a human-chosen passphrase, and the wrapped key is public — it travels in the manifest. So an attacker who fetches a manifest can guess passphrases offline, at hardware speed, with no rate limit and nobody watching.
That makes the cost of one guess the real security parameter. Not the cipher.
AES-256 is irrelevant if summer2024 takes a millisecond to test.
PBKDF2 raises that cost by iterating, but it needs almost no memory, so a GPU runs thousands of guesses in parallel and an ASIC does far better. Argon2id is memory-hard: every guess must allocate and randomly traverse tens of megabytes, which is precisely the resource parallel hardware cannot cheaply multiply. The same attacker budget buys orders of magnitude fewer guesses.
Argon2id specifically — rather than Argon2i or Argon2d — is the hybrid RFC 9106 recommends by default: Argon2d's data-dependent addressing resists time-memory trade-offs but leaks through cache side channels, Argon2i is the reverse, and Argon2id takes one pass of each.
The shipped profile is m=65536 KiB, t=2, p=1, comfortably above OWASP's
recommended settings and roughly 2.8 times as expensive to attack as the
46 MiB single-pass profile it replaced. It costs about 2.2s, which is
affordable only because it runs in a dedicated worker
(js/core/kdf-worker.js) rather than on the page's thread, so the tab stays
responsive and the cost is no longer paid in visible jank — which is what makes
raising these parameters practical. A server can afford more memory and should
use it.
On Node the same worker runs in a worker_threads thread, so a gateway keeps
serving requests while a passphrase is being stretched rather than stalling its
event loop for the duration.
Where no worker implementation exists, or where the worker script fails to load,
derivation falls back to the calling thread: slower and visibly so, but still
correct. A
worker that ran and failed, or that stopped answering, is reported instead —
retrying inline would block for the same reason and fail the same way, and a
timeout quietly followed by a second attempt is how a one-second wait becomes a
minute-long one. Implementation correctness is checked against the RFC 9106 §5.3 known-answer
vector in test/kdf.test.js — a subtly wrong KDF still produces
plausible-looking bytes and silently provides a fraction of the intended
strength.
Parameters travel with the file, and are bounded in both directions. Below the floor, whoever serves the manifest could set memory to 8 KiB and make cracking as cheap as it was before Argon2id existed, with the file still opening normally so nothing would look wrong. Above the ceiling, a manifest demanding 8 GiB turns opening a file into a denial of service against the reader.
PBKDF2 remains readable so that files sealed before this change still open. Losing them would be unrecoverable, since there is no path to the plaintext without the key.
The passphrase never encrypts data. It derives a key-encryption key that wraps a random per-file key:
fileKey = 32 random bytes
KEK = Argon2id(passphrase, kdfSalt, m=65536KiB, t=2, p=1)
wrappedFileKey = AES-256-GCM(KEK, fileKey)
Changing a passphrase rewraps ~60 bytes instead of re-encrypting the archive. Two uploads of the same document produce completely different ciphertext (fresh random file key), so identical documents are not linkable in storage — while the public file hash and Merkle root still match, so the chain recognises them as the same document.
The manifest splits into a public header and an encrypted body:
- Header (public): version, KDF parameters, wrapped key, file hash, Merkle root, chunk count, size. This is what the chain commits to.
- Body (encrypted): file name, MIME type, and the chunk table — index, storage location, plaintext hash, ciphertext hash, size.
Chunk locations and the file name are themselves sensitive. payroll-2026.pdf
leaks its contents from the name alone, and an enumerable chunk list tells an
observer exactly which blocks to target. Without the passphrase, neither is
visible.
Each downloaded chunk passes three independent checks:
- Stored-hash check — the bytes match the ciphertext hash in the manifest. Cheap, runs before any crypto, and rejects corrupted or substituted blocks without spending CPU on them.
- AEAD authentication — the chunk decrypts under its position-bound key, or it does not decrypt at all.
- Merkle reconstruction — the reassembled chunk hashes must rebuild the root recorded on-chain. This is the check that ties the bytes in your hands to the blockchain record.
Then the reassembled file's SHA-256 and length are compared to the manifest. Any mismatch raises; partial or "best effort" output is never returned.
A manifest arrives from whatever storage served its CID. Anyone who can serve those bytes — a hostile gateway, a compromised pinning service, a network attacker on a plain-HTTP gateway — controls every field in it before a single cryptographic check runs.
The interesting attacks there never reach the crypto at all:
| Hostile field | Effect without validation | Defence |
|---|---|---|
totalChunks: 4e9 |
restore loop hangs | bounded chunk count |
fileSize: 1e15 |
allocation kills the process | bounded size, consistency check |
kdf.memoryKiB: 8 |
passphrase cracking becomes cheap again | floor of 19,456 KiB enforced |
kdf.memoryKiB: 8GiB |
opening a file becomes a denial of service | ceiling enforced |
kdf.iterations: 1e12 |
client hangs in PBKDF2 | ceiling enforced |
__proto__ key |
prototype pollution | rejected outright, not sanitised |
location: "../../etc/passwd" |
path traversal / request forgery | strict alphanumeric pattern |
| duplicate chunk indices | chunk table contradicts itself | indices must be exactly 0..n-1 |
js/core/validate.js checks type, format and range on every field before it is
used, and js/core/limits.js makes every bound explicit and overridable, rather
than leaving it implied by whatever the machine happens to tolerate.
server/ sits between users and the pinning provider so the pinning credential
never reaches a browser — a browser cannot keep a secret, and any token shipped
to the page is readable by every visitor.
The important property is what the gateway is not trusted with: it never receives a passphrase, a file key or plaintext. Chunks arrive already sealed. A full compromise of that server exposes ciphertext, chunk sizes and traffic timing, not documents.
Its own defences are conventional but deliberate: constant-time API key
comparison (a naive === on a secret leaks it through timing, one character at
a time), per-key token-bucket rate limiting, request bodies capped as bytes
arrive rather than after buffering, strict CID validation before any upstream
request, and a refusal to start when misconfigured rather than silently
accepting anonymous uploads.
§2.8 treats a manifest as hostile. The same reasoning applies one layer up, to
every string the pages display: a CID and a kid from the gateway, a
transaction hash and an exporter's name from the chain, a refusal code out of
a JSON body. Those are rendered next to icons and links, which makes writing
them straight into innerHTML the natural thing to do — and that hands
whoever chose the string the ability to write markup into the page.
The content security policy the gateway serves (script-src 'self' https:,
no 'unsafe-inline') stops injected markup from running: an onerror=
attribute is inline script and the browser refuses it. It does not stop the
markup from being there, and script is not what this particular page has to
fear. The upload and retrieval pages hold a passphrase field, and a passphrase
field drawn by the service is indistinguishable from the real one to the
person typing in it.
So js/core/html.js supplies a tagged template that escapes every
interpolation, and js/App.js and js/chunked-app.js use it at every sink.
The literal parts of a template are markup because a maintainer wrote them;
the interpolations are escaped because someone else did. A fragment can only
pass through unescaped by being built with the same tag, which is visible at
the call site.
Two tests hold the line. One scans the client for an assignment to innerHTML
whose right-hand side is anything but a tagged template, a toHtml() call or
a literal with nothing in it — so the next sink somebody writes fails the
build rather than the review. The other is a browser test that puts an
<img id="pwned"> into a field the gateway is simply believed about, and
asserts no such element exists in the document afterwards.
| Threat | Mechanism |
|---|---|
| Storage provider alters stored bytes | Stored-hash check, AEAD tag, Merkle root |
| Storage provider reads documents | Client-side encryption; only ciphertext is uploaded |
| Chunks reordered, duplicated, dropped or spliced | Position-bound AAD |
| Manifest substituted | On-chain Merkle root compared against the manifest |
| Document backdated or its history disputed | On-chain block number and timestamp |
| Unauthorised registration | onlyExporter, controlled by the contract owner |
| One exporter revoking another's record | revokeDocument checks the registering address |
| Chunk hash replayed as a subtree | Domain-separated Merkle tree |
| Correlating two uploads of the same file in storage | Random per-file key |
| Hostile manifest (resource bombs, weakened KDF, prototype pollution, path traversal) | Strict validation before use |
| Pinning credential theft from the page | Credential lives in the gateway, never the browser |
| One client exhausting the pinning quota | Per-key rate limiting |
| API key recovery by timing | Constant-time comparison over hashed values |
| Gateway or chain data injected as markup into the page | Escaped at every sink; enforced by a source scan |
These are real limitations. State them plainly rather than discovering them in an audit.
- A lost passphrase means lost data. There is no recovery, no reset and no backdoor. That is the design, and it is the most common way users lose files.
- A compromised browser or device defeats everything. Encryption happens in the page; malware or a malicious extension sees plaintext and the passphrase.
- Weak passphrases fall to offline attack. PBKDF2 raises the cost per guess, it does not make a six-character passphrase safe. The wrapped key is public.
- Argon2id still cannot rescue a genuinely weak passphrase. It raises the
cost per guess by orders of magnitude; it does not make
password1safe. The wrapped key is public, so a short passphrase remains the most likely way an attacker gets in. - Key derivation takes about 2.2 seconds, and longer on a low-end phone. The worker keeps the page responsive, but the user is still waiting. That wait is the price of the attacker's cost; it cannot be reduced without reducing theirs.
- A worker does not isolate key material from the page. It is a separate thread in the same origin and the same process, not a security boundary. Anything that can run script on the page can still reach the derived key; the worker is a responsiveness measure, not a containment one.
- Revocation does not delete anything.
revokeDocumentclears the on-chain record. Chunks already pinned remain wherever they were pinned. Treat any upload as permanent. - Availability depends on pinning. If nothing pins the chunks, the chain still says "registered" while the bytes are gone. Chunked storage helps (chunks can be re-pinned and replicated independently) but guarantees nothing on its own.
- Chunk sizes and counts leak metadata. An observer sees roughly how large a file is and when it was registered. There is no padding.
- The gateway bounds rate, not total volume. A client within its rate limit can still pin indefinitely. Per-user quotas are not implemented.
- Gateway rate limits are per process and in memory. Behind several instances, each enforces its own share, and all buckets reset on restart.
- Not formally audited. The composition is built from standard primitives and is tested, but it has not been reviewed by a professional cryptographer. Do not describe it as audited.
- The contract has not been audited either. It is deliberately small and uses no external calls, but the same caveat applies before it holds anything valuable.
- Never put credentials in client-side JavaScript. An earlier version of
this project embedded a live Pinata API key and secret in
js/App.js, where every visitor could read them. Credentials belong on a server, reached through thebackendstorage mode. Anything committed to git must be treated as public forever — rotate it, don't just delete it. - Keep
js/config.jsout of version control. It is gitignored. Onlyjs/config.example.jsis committed. - Do not lower the Argon2id memory below 19,456 KiB. It is the OWASP floor and the main thing standing between a weak passphrase and an offline attacker. Raise it server-side, where a slower derivation costs nobody a frozen tab.
- Deploy the contract yourself and set the owner to a wallet you control. The owner controls the exporter allowlist.
- Serve over HTTPS. WebCrypto is unavailable on insecure origins, and the page's integrity is what everything else rests on.
- Allow
worker-src 'self'in your Content-Security-Policy. Browsers that honourworker-srcdo not fall back toscript-src, so omitting it blocks the derivation worker and silently pushes every derivation back onto the page's thread. The bundled gateway already sets it.
- A backend pinning proxy so credentials leave the browser entirely, plus rate limiting and per-user quotas.
- Multi-recipient key wrapping — wrap the file key to several public keys so a document can be shared without sharing a passphrase.
- Hybrid post-quantum key wrapping (ML-KEM alongside the classical wrap) for documents that must stay confidential for decades. "Harvest now, decrypt later" is a real concern for long-lived records.
- Replication across independent pinning providers, with on-chain challenges
using
verifyChunkto prove a provider still holds a given block. - A professional cryptographic review before this protects anything that matters.