Blobs
Documents are JSON plus Yjs sub-documents, and both are capped at a size that rules out a photo. Files need a third kind of content: blobs — immutable byte objects held in an object store (R2 on Cloudflare, S3-compatible elsewhere, filesystem or memory for development) and referenced from document JSON by handle. The bytes never touch the sync protocol; the handle syncs like any other field.
Why blobs are core, not an app feature
Section titled “Why blobs are core, not an app feature”The tempting version — an upload route in the app and the object key in a string field — cannot be made correct, because four of its obligations are only reachable from inside the engine:
- Lifecycle is transactional with writes. Knowing when a blob becomes referenced or unreferenced means seeing every accepted patch at the write boundary, atomically with the document row. An app observing changes after the fact races concurrent writes: bytes leak forever, or a cleanup deletes a blob just as a write re-references it. This is exactly why Yjs sub-documents and their reference counting live in core; blobs are the same problem with different bytes.
- Authorization parity. Read filtering promises that a hidden document is indistinguishable from a nonexistent one. An app-owned download route with its own checks is a permanent side door: a private document’s image gets a weaker, drifting access model than the document itself.
- Trash and purge semantics. Delete → restore must keep the bytes; purge must destroy them. Only the engine sees those transitions.
- Schema integrity. “A synced document never points at missing bytes” needs the write boundary: the schema-derived set of handles is verified against the catalog inside the write’s transaction. A plain string field lets every client sync a document whose image 404s.
The answer to the engine-size objection is the shape of the integration, not exclusion: core owns the semantics, the edges own the bytes. Hosts own the transport (the Worker or Node process streams bytes; core never runs an HTTP server), adapters own the store, apps own presentation, and the whole lane is opt-in — a server without an object storage adapter pays nothing, and refuses a schema that declares a blob field.
The one insight that makes this cheap
Section titled “The one insight that makes this cheap”Blobs are immutable. A “changed” image is a new blob and a JSON patch that swaps the handle. From that:
- Blobs need no realtime sync at all. The engine syncs the handle over the ordinary JSON lane — ordinary optimistic update, ordinary guards, ordinary conflict rules — and the bytes move over plain HTTP, out of band, exactly once per blob.
- No blob sequence numbers, no blob event log, no blob replay. The event log stays byte-free; history replay reproduces handle changes. The bytes behind a historical handle survive until the sweep reclaims them, but the read gate follows the current references: a reader can fetch them only while a document they can read still holds the handle, so a history view that shows old attachments needs a copy under system authority, not the by-id URL.
- Downloads are cacheable forever in principle, since the bytes under a handle never change. How far a deployment cashes that in is the host’s choice, below.
So the design is deliberately asymmetric with the Yjs lane: Yjs bytes are mutable CRDT state and ride the sync protocol; blob bytes are immutable content and never touch it.
A blobRef field
Section titled “A blobRef field”A schema declares a blob field as its own
value kind, mirroring yjsRef: the JSON value is a bare, server-minted,
opaque id.
{ "type": "blobRef", "accept": ["image/*", "application/pdf"], "maxBytes": 5000000 }- Membership is schema-derived. The write boundary collects the document’s live blob ids from the schema, the same walk that finds its Yjs sub-documents. A write introducing a handle to a blob that doesn’t exist, was never finalized, or has been swept is rejected as a precondition failure — the same benign category as a reclaimed sub-document.
- The catalog owns the facts the platform needs — size, content type,
SHA-256, timestamps, attribution. Display metadata (alt text, dimensions,
crop) is the app’s business, in sibling JSON fields. A structured value
duplicating catalog facts into every snapshot was considered and rejected:
one source of truth, and a
HEADon the download URL answers the rest. acceptandmaxBytesconstrain what the field takes: media-type patterns (exact, or atype/*wildcard) and an inclusive size ceiling under the server-wide limit. Both are checked against the catalog’s facts — the type declared at upload, the size recorded at finalize — never the bytes, soacceptis a schema and UX constraint, not a security boundary. The read side is what protects readers.
A blob is also a fourth kind of reference: a handle to bytes rather than a pointer into JSON, reference-counted like a Yjs reference, but with the target living outside the document store.
Upload, then write
Section titled “Upload, then write”Upload is two-phase HTTP followed by an ordinary document write:
sequenceDiagram
accTitle: Upload a blob, then commit its handle
accDescr {
The client POSTs the bytes to the host's blob route with the intended
write named in the query. The host asks the server to stage the upload,
which runs the ordinary write gate on that intent and the field
constraints, then streams the body into the object store and asks the
server to finalize the catalog row. The client receives a handle and
commits it in a normal document write; the write boundary verifies the
handle and records the reference edge in the same transaction, flipping
the blob from staged to live.
}
participant C as client
participant H as host (Worker / Node)
participant S as server
participant O as object store
C->>H: POST …/blob?kind=&docId=&docType= (bytes)
H->>S: stage upload — authorized as that write
S-->>H: blobId + key
H->>O: stream body to key
H->>S: finalize (size, sha256)
H-->>C: handle
C->>S: ordinary doc write, data.image = blobId
Note over S: verify handle + record edge<br/>in the write's transaction
POSTthe bytes to the folder’s blob route, naming the write the handle is intended for:kind(create or update),docId,docType, and optionally the fieldpath. The server runs the ordinary write gate on exactly that operation — the docType’s rule, or the per-documentsys:accessentry where one exists — so a refusal happens before any byte moves, with the same verdict the committing write will get. The field constraints are checked here too, over everyblobRefleaf the intent could land in.Content-Lengthis required: only a declared length lets the size limit run before bytes move and lets the store stream without buffering.- The host streams the body into the object store and finalizes the catalog row with the size and hash it observed. On Cloudflare the Worker holds the R2 binding and streams; the Durable Object sees only metadata calls, so no large body passes through its memory.
- The client writes the document — a normal optimistic write with the handle in the field. The write boundary collects the document’s blob membership, the storage adapter re-verifies every handle inside the content write’s transaction and records the reference edges there, and the blob flips from staged to live.
Ordering is enforced by the integrity gate rather than by convention: a handle can only be committed after finalize, so a synced document never points at bytes that aren’t there. The intent is a hint for authorization, not a binding — the handle may be committed wherever the write boundary allows.
On the client this is client.uploadBlob(bytes, { contentType, kind, docId, docType })
resolving to a handle, and client.blobUrl(blobId) for the download URL, both
configured through ClientConfig.blobs — the folder’s blob route as a bare
path. A refused upload maps back to the write vocabulary (unauthorized,
sizeLimitExceeded, …), so apps handle it like a denied write. The React
integration wraps the same calls as a useBlobUpload hook and a render-safe
blobUrl helper.
Download and the read gate
Section titled “Download and the read gate”GET …/blob/:id asks the server one question: may this principal read this
blob? A blob is readable iff the principal can read at least one document
that references it — the existing read filtering
evaluated per referencing document — or, while it is still staged, iff the
principal is its uploader, so previews work before commit. “Uploader” is
matched by subject: an anonymous upload has no subject to match, so it
becomes readable only once a readable document references it. Anything else is a
404: a hidden blob is indistinguishable from a nonexistent one, exactly as for
documents. No new authorization vocabulary was added.
The response carries the content type, a strong ETag (the SHA-256), honors
Range, and a Cache-Control the host states rather than core defaults,
because only the host knows how its read gate is wired:
- Revalidate — the gate runs on every request (identity rides cookies, as
it must for a URL that lands in
<img src>). Success isprivate, no-cache: every reuse re-runs the read gate, and the strong ETag keeps that a byte-free 304. This is what both hosts ship today. - Immutable — the URL carries its own authorization (signed) or none (a
secret or public URL; the id is a random server-minted capability).
Success is
public, max-age, immutable, cacheable by browsers and shared caches alike, with no round-trip to the server at all.
Denials are never cacheable under either policy: a staged blob goes live, a trashed referrer is restored, and a cached 404 must not outlive the grant.
Identity never rides the URL. blobUrl answers a bare by-id path and the
client refuses an endpoint with a query string, because those URLs end up in
<img> tags, target="_blank" links, browser history and referrers. Cookies
are the channel; short-lived signed URLs are the planned addition for CDN
offload — a capability minted by a separate step, not identity on the by-id
URL.
Serving uploads safely
Section titled “Serving uploads safely”Serving user uploads from the app’s origin, with cookie auth, is the classic stored-XSS setup: the content type is attacker-supplied at upload, and SVG or HTML uploads would otherwise run script as the site. The download route treats it as one, and the rules are pinned by core’s own tests because the HTTP handlers live in core and every host mounts the same ones:
X-Content-Type-Options: nosniffon every blob response.- An inline-rendering allowlist — raster images, video, audio, PDF.
Everything else,
image/svg+xmland anythingtext/*or*+xmlincluded, is served asContent-Disposition: attachment, so the browser downloads it instead of rendering it in the origin. Content-Security-Policy: sandboxon blob responses as defense in depth.
The production-grade mitigation — a separate, cookie-less origin for user content — is what signed URLs will provide. The header rules are what make same-origin serving safe until then.
Lifecycle and garbage collection
Section titled “Lifecycle and garbage collection”stateDiagram-v2
accTitle: Blob lifecycle
accDescr {
An upload creates a staged blob. A committed write referencing it makes
it live. A live blob whose last reference is removed, or whose
referencing documents are all purged, waits out the retention period
and is then deleted. A staged blob never referenced is deleted once it is
older than the retention period.
Deleted means bytes destroyed and the catalog row gone.
}
[*] --> staged: upload + finalize
staged --> live: referenced by a committed write
staged --> deleted: never referenced, retention elapsed
live --> deleted: last reference removed or referrers purged, retention elapsed
deleted --> [*]
- Staged — finalized, not yet referenced. Swept once older than the host-chosen retention period, so abandoned uploads don’t leak.
- Live — referenced by at least one non-purged document. Documents in the trash still count: delete → restore must round-trip, so bytes survive until purge.
- Deleted — bytes destroyed, then the catalog row. The retention period between “unreferenced” and deletion is policy, not a correctness mechanism: it keeps sweeps unhurried and leaves a write that re-references the blob (or an export under system authority) a window. It is not a read window: the read gate follows the references, not the retention clock.
A document losing a handle starts that blob’s retention clock whatever removed it — an ordinary write, a purge, or a schema migration that drops the field.
The sweep is where the correctness lives, and its ordering is the point. The naive order — find candidates, delete bytes — races a concurrent write that commits a fresh reference in between, leaving a synced document pointing at destroyed bytes. So the sweep tombstones first, transactionally: for each candidate, one catalog transaction re-verifies the sweep condition and marks the row deleted. Reference edges commit in the same transactional store, so a concurrent write either lands its edge first (the re-check sees it and skips the blob) or runs after (the write’s in-transaction verification now rejects the handle). Only then are the bytes deleted, then the tombstone. Object-store deletion isn’t transactional with the catalog, so those last two steps are at-least-once: a failure leaves a tombstone that retries on the next sweep.
Retention stays with the host: sweepBlobs takes a retentionMs, the way
purgeDeletedDocuments takes deletedBefore, so the server holds no retention
configuration. One retention covers both kinds of candidate — a staged upload
older than it is abandoned, a blob unreferenced for longer than it is
reclaimable. Around a day is the sane policy, kept longer than any plausible
upload. DEFAULT_BLOB_SWEEP_POLICY is that day with a batch of 100, and
assertBlobSweepPolicy is the check the sweep itself applies: a retention
from zero to ten years, a batch limit of one or more. A host runs the same
check on its configured policy at startup.
When the sweep runs
Section titled “When the sweep runs”The host owns the schedule, and the sweep tells it when the next one is due:
its result’s nextSweepAt is already past when candidates remain (the batch
limit was hit), the moment the oldest pending upload or unreferenced blob
outlives the retention, or null when nothing is pending. Between sweeps, the
server’s onBlobSweepCandidate hook fires whenever a candidate can appear — a
blob staged, a write that replaces a document’s blob references, a
schema change releasing them on the next read or
validation pass, a purge — so a host arms its schedule from one hook.
- On Cloudflare the Durable Object’s alarm runs the sweep, and the alarm
is armed by activity rather than by a clock: the sweep-candidate hook, or a
wake with no alarm, arms it one retention period out, and after a run the
sweep’s
nextSweepAtre-arms it for when the next candidate comes due, or not at all. No sweep alarm is set less than a second out, so a burst of activity coalesces into one run. A folder with nothing pending never wakes for a sweep, which keeps idle folders free. The policy is theblobSweepPolicy()override, read once when the Durable Object starts. - On Node a timer sweeps every folder in the database that has something to reclaim, which the catalog answers in one query — so a folder nobody has opened since the process started is swept too, and a blob unreferenced before a restart is reclaimed by the first tick after it. Folders then take turns, one batch each per round, and the catalog is asked again every round, so neither one folder’s backlog nor garbage that keeps coming due can hold up the rest.
Where the pieces live
Section titled “Where the pieces live”- Core — the
blobRefkind, the catalog (tables in the same storage adapter as the documents, so reference edges are transactional with writes), the server API, the dumb object storage adapter interface (put/get/head/delete by opaque key), and the shared HTTP handlers every host mounts. - Object storage adapters — R2 (a bucket binding, one bucket shared by
many Durable Objects through per-folder key prefixes), S3-compatible (plain
fetchand SigV4; AWS S3, R2’s S3 API, MinIO), filesystem (the Node harness’s dependency-free default), and memory for tests. - Hosts — the Cloudflare Worker mounts the handlers next to the WebSocket route and streams to R2; the Node harness mounts the same handlers under the same principal resolver as its WebSocket upgrade.
- Export and import — a document export ships its blobs’ metadata and bytes alongside both event streams, and an import re-uploads them. Because blobs are immutable, per-document restores compose: an import reuses an already-present blob whose facts match, and rejects a mismatch as a genuine id collision.
- Apps — presentation. Thumbnails, resizing, pixel dimensions, alt text, which document types carry files. datadata never inspects image bytes; the playground’s image lane (dimensions committed next to the handle in the same write, resized variants derived on read behind the same read gate) is about a hundred lines of app code over two existing boundaries, which is the evidence that settled this as a decision rather than a default. The generic editor and viewer components render a blob as a link, never a preview: a preview is presentation, and a bare id has no content type.
The conformance suite pins the whole contract — upload, integrity gate, reference-then-sweep, delete → restore → purge, the read gate, size limits, export round-trips, sweep-versus-concurrent-write, and the response headers — against all three storage runners, so memory, Durable Object + R2, and Postgres + filesystem or S3 behave identically.