This is the abridged developer documentation for datadata # datadata > A sync engine where everything — including the schemas — is a document. Documents all the way down User data, schemas, and even the engine’s own system state are all documents, read through one small API — like the Unix filesystem. [Read the core idea](/introduction/documents-all-the-way-down/). Optimistic by default Changes apply locally the moment they’re made and are confirmed — or cleanly rejected — by an authoritative server. [How sync works](/concepts/sync-and-optimistic-updates/). Keeps working offline An opt-in IndexedDB layer makes the write queue and document cache durable: a reload becomes a reconnect, cached documents render immediately, and two offline tabs converge through the shared store with no server round-trip. [Offline persistence](/concepts/offline-persistence/). Structured data and rich text A **hybrid** model by design: structured data as guarded, server-ordered JSON Patch, rich text as embedded **Yjs** (CRDT) on [its own sync lane](/concepts/two-sync-lanes/) — each with the merge semantics it needs. [Rich text with Yjs](/concepts/rich-text-with-yjs/). Files without a side door Images and attachments are immutable **blobs** referenced by handle: the handle syncs like any field, the bytes move over HTTP, and the engine owns their integrity, lifecycle and read authorization — no app-level upload route with its own rules. [Blobs](/concepts/blobs/). A staging area for changes A staged session lets humans and AI agents gather edits, preview conflicts three-way, and commit them atomically. [Staged sessions](/staged-sessions/overview/). Built with agents in mind The surface is small and uniform enough to wrap cleanly for an LLM — and because schemas are themselves documents, an agent can define new document types at runtime, not just fill in existing ones. [datadata and AI agents](/ai-agents/why-agents-like-datadata/). Mind the gap This is early-stage software. Known issues — including problems we’ve found with the JSON Patch RFC — live in the open. [Known issues & open questions](/known-issues/). Under the hood Wire protocol, storage model, and the reasoning behind the design. [Design decisions](/comparison/design-decisions/). # What is datadata? > A sync engine where everything — including the schemas — is a document. Server-authoritative sync with optimistic updates, live and staged sessions, and Yjs integration. datadata is a **sync engine where everything — including the schemas — is a document**. In sync-engine terms, that makes it a server-authoritative document sync engine: clients subscribe to documents, edit them optimistically, and an authoritative server orders, validates, and broadcasts every change in real time. Offline edits replay on reconnect, and an opt-in [persistence adapter](/concepts/offline-persistence/) carries both writes and reads across a reload — but there is no distributed merge for structured data: the server stays in charge, by design. The name is a working title. ## The pieces datadata is a TypeScript library, not a hosted service. It currently consists of: * **The core library** — client, server, live and staged sessions, the schema and presence systems, and an in-process transport, with no opinion about where each side runs. * **A Cloudflare backend** — runs the server inside a Durable Object with SQLite storage and WebSocket transport. This is the production deployment shape; see [Architecture](/architecture/system-overview/). * **A Node + Postgres backend** — the same server behind a Postgres storage adapter and a Node WebSocket host, serving many folders from one database. It passes the same conformance suites as the Cloudflare backend but has not run in production, and its only Postgres driver so far is the embedded PGlite; see [Node + Postgres deployment](/architecture/node-deployment/). * **Object stores for files** — R2, S3-compatible, and filesystem adapters for the [blob lane](/concepts/blobs/), all conformance-tested. * **React bindings** — a Jotai-based integration for subscribing to documents from React components. * **React components** — utility editor and viewer components for exercising a session in an app, kept apart from the bindings. * **Devtools** — a panel for inspecting live documents, subscriptions, and pending optimistic updates. ## The model in one paragraph Everything is a **document**: a typed JSON value with a server-assigned sequence number. Structured changes travel as JSON Patch; collaborative text lives in embedded Yjs documents. Schemas are themselves documents, editable through the same API as user data. On top of the live client sits the **staged session**: a staging area where a human or an AI agent accumulates changes, previews conflicts against the live head, and commits atomically. ## Who this site is for * **Application developers** evaluating sync engines — start with [Documents all the way down](/introduction/documents-all-the-way-down/) and [When to use it](/introduction/when-to-use/). * **Sync-engine developers** comparing notes — start with [Design decisions](/comparison/design-decisions/) and the [wire protocol](/architecture/wire-protocol/). We’d genuinely like to [compare notes](/contact/). # Documents all the way down > The core idea — user data, schemas, system state, and presence are all documents behind one small API. Most sync engines juggle several kinds of things, each with its own API: relational engines sync **rows** but evolve schema through out-of-band **migrations**; document engines separate **data** from the **schemas** that type it; nearly all keep the engine’s **own sync state** on a side channel apart from your data. datadata has one kind of thing. **Everything is a document**, and there is one small API for all of them. ## Four kinds of documents, one API **User documents** hold application data. A document has an id, a type, a server-assigned sequence number, and JSON data. The whole API is a handful of calls — `createDocument` and `updateDocument` to write, `getDocument` to read, and an `onDocumentChange` subscription to follow along — and you use the same ones for every kind of document below. **Schemas are documents.** A document type — `task`, say — gets its schema from a document at `sys:schema:task`, itself a document of type `sys:schema`. Defining a new type means writing that document, through the same `createDocument`/`updateDocument` calls as any other. Schemas validate against a meta-schema, carry their own migration history, and sync to clients like any other document. See [Schemas as documents](/concepts/schemas-as-documents/). **System state is documents.** The engine exposes its own state as read-only documents you can subscribe to — the folder index (`sys:index`), the trash (`sys:trash`), a live view of [staged work](/staged-sessions/overview/) (`sys:session`), and more. These are synthesized read-models, never written directly — but to a client they’re just documents. [System documents](/concepts/system-documents/) enumerates the full set. **Presence is documents too.** Ephemeral per-participant state — who’s here, where their cursor is, what they’re streaming — lives in a companion `sys:presence::` document that rides the same subscribe/read/update path as everything else. What sets it apart is that it’s never stored durably: the server keeps only an in-memory aggregate, and each client is the source of truth for its own cells. See [Presence](/concepts/presence/). ## Why this matters **The API stays tiny.** Subscribe, read, create, update. Learn it once and you can manipulate data, evolve schemas, and introspect the engine. There is no separate admin API, migration tool, or metadata endpoint to learn. **Everything is referential.** Documents point at other documents by id — including schemas (`sys:schema:`) and staged work (`sys:stage:`). A small set of concepts composes instead of multiplying. **UI and agents get introspection for free.** A devtools panel, a “pending changes” sidebar, or an AI agent inspecting its own staged edits all work the same way: subscribe to a system document and render it. **The data model is extensible at runtime.** Because schemas are documents, anything connected to a folder — application code, a migration, or an [AI agent](/ai-agents/why-agents-like-datadata/) — can design a new document type, write its schema, and start creating documents of that type, with no out-of-band tooling or redeploy. And since a schema is just a document, a new type can be staged in a [session](/staged-sessions/overview/) alongside the very documents that use it and committed atomically. # One timeline per document > The experience behind the design — meaning splits across histories that move at different speeds, so datadata keeps each document's history in one place. No lesson shaped datadata’s design more than this one. It comes from years of building systems with heavy schema evolution — [a headless CMS](/introduction/where-datadata-comes-from/) and [Shiftic](/introduction/where-datadata-comes-from/#shiftic-2023): **a document’s meaning should never be split across histories that move at different speeds.** It matters because data and the schema that gives it meaning only make sense read *together*. Let their histories drift onto separate timelines and the past gets harder to read — to say what a document meant when it was written, you have to reconstruct which schema was in force, rather than have the data carry that answer itself. ## How meaning gets split In a conventional stack, what a stored record *means* lives in at least three places: the data itself, the database schema, and the source code that reads and writes it. Each has its own history — the data’s, the migration log’s, the repository’s — and they are correlated only by deploy timestamps. Answering “what did this document mean when it was written?” becomes archaeology across git history, deploy logs, and migration tables. Schema evolution widens the cracks. Database changes and code changes never land atomically — there is always a window where new code meets old shapes or old code meets new ones. And some changes leave no trace where the data lives at all: when the schema is defined in application code — a Zod schema, say — rather than in a migration table, remove a field and the only record of the removal is a commit. ## Distribution splits it again With one central database, you can at least pretend a change happened at a single moment. In a distributed deployment there isn’t even that. datadata runs [one isolated store per folder](/architecture/cloudflare-deployment/), and a rollout reaches each folder at its own moment, as instances restart. “When did the schema change?” has no global answer — only a per-folder one. Any design that assumes one deploy moment is telling a fiction somewhere. ## datadata’s response Make the per-document event log the only timeline that matters. Three decisions follow directly: * **[Schemas are documents.](/concepts/schemas-as-documents/)** Schema history lives in the same event log as the data it governs. Even TypeScript-authored schemas are upserted into documents at startup — so a code rollout becomes a document write *in each folder*, and each folder’s log records when the change reached *it*: the only answer an eventually-consistent fleet can give. * **[Changes are data, not domain events.](/comparison/design-decisions/#changes-are-data-not-domain-events)** A JSON Patch is self-applying. Domain events need a reducer to interpret them — and that reducer is a second timeline of its own: fix a bug in it and last year’s events quietly mean something new on replay. A patch has no interpreter to drift from, so history replays identically forever. * **[Every event is tagged with its schema version.](/architecture/portable-event-streams/)** Any historic version of any document is interpretable with the exact schema it conformed to — in place, or exported to somewhere that has never seen the source code. What this buys: point-in-time introspection has exactly one place to look, and a folder’s history is portable because it carries its own interpretation with it. ## Staging without a second timeline You might expect a review-before-it-lands layer to work like a branch: fork the document, edit the copy, merge later. That spins up a divergent history you have to reconcile on the way back. [Staged sessions](/staged-sessions/overview/) do the opposite: staged changes ride as an overlay on the live document, which keeps advancing underneath them. The session tracks the moving head in real time and flags the moment an upstream change conflicts with staged work — it stays caught up with the timeline rather than drifting away from it. Commit lands the staged changes on the current head atomically, or discard drops them and nothing reaches the document. Either way there is only ever one committed timeline, with a provisional layer kept separate on top. # Where datadata comes from > The two systems datadata draws on — an open-source CMS and an AI platform for organizational change — and the lessons that became its design decisions. datadata is a new take, but not from a blank page. It draws on two systems I’ve spent years building — [Dossier](https://www.dossierhq.dev), an open-source headless CMS, and [Shiftic](https://shiftic.com), an AI platform for designing and measuring lasting organizational change. Most of datadata’s design decisions are answers to something one or the other taught me. ## Dossier (2020–) From October 2020 I built Dossier as a solo project — an open-source headless CMS where you bring your own auth, database, and backend, with a schema-driven admin UI on top. It never gained traction, but it’s where most of datadata’s instincts were formed: one idea that survived intact, and several that became problems datadata had to solve. **The complete event stream — the idea that survived.** Dossier captured every mutation, schema changes included, as one ordered stream of sync events; replaying it on an empty database reproduced everything, schema and content together, across backends. That idea came through and deepened: it’s datadata’s [one timeline](/introduction/one-timeline/) principle and the basis of [portable event streams](/architecture/portable-event-streams/). Three bets aged badly, and each became a problem the build below had to answer: * **Betting on the generic admin UI.** A schema-driven admin interface was a pillar of the CMS value proposition — and a large share of the code. AI agents eroded that pillar mid-project: bespoke, task-specific UIs became cheap to build, often cheaper than adapting a generic one. * **Owning a rich text format.** Dossier’s rich text format changed twice (Editor.js blocks, then Lexical’s editor state), and each switch meant migrating stored content and rebuilding editor UI. * **A migration rule matrix.** Schema changes were governed by a matrix of allowed and unsupported operations, some triggering re-validation or re-indexing duties — expressive, but heavy and easy to get wrong. One more lesson was structural rather than a pivot. Dossier’s database adapter abstracted over low-level primitives like transactions, across Postgres, SQLite, and D1 — and D1, which has no interactive transactions, proved that boundary wrong rather than just awkward. datadata draws its two extension points higher up: a [storage adapter and event bus](/architecture/system-overview/) that own whole operations end to end. Those boundaries were deliberately drawn around a single production backend, and they held when the second arrived: the [Node + Postgres backend](/architecture/node-deployment/) reuses the engine untouched, with shared conformance suites rather than a shared SQL layer keeping the two adapters in step. ## Shiftic (2023–) Dossier is the longer arc, but it isn’t the only system feeding datadata. Since 2023 I’ve been CTO of Shiftic, where we built a comparable engine to power structured UI and AI agents. Shiftic settled on the same core move — JSON Patch as the unit of change, event-sourced — which is evidence, not coincidence, for datadata’s [changes-as-data](/comparison/design-decisions/#changes-are-data-not-domain-events) bet. It also brought the half Dossier never had: realtime sync. Shiftic needed it for the obvious reason — multiplayer humans editing together — and then for a less obvious one: the moment you add AI agents, the UI is multiplayer by default. An agent and a person working the same document are just two participants, so a sync engine stops being a collaboration nicety and becomes the substrate agents run on. Shiftic had optimistic updates too, but bespoke — hand-rolled for the use cases that needed them rather than a property of the engine, which is the generalization datadata took on. But Shiftic keeps schema in source code, so data, schema and migrations still live on separate timelines — the exact split that datadata’s [one timeline](/introduction/one-timeline/) principle and [schemas as documents](/concepts/schemas-as-documents/) set out to close. Both a positive influence and a working example of the pain it answers. ## How datadata came together (2025–) datadata didn’t start as a library. It started app-first — real applications built directly on Yjs — and the library grew out of what those apps kept reaching for. In **August 2025** that became its own thing: the datadata library, with [JSON Patch](/concepts/changes-as-json-patch/) as the unit of change. Working in it, I immediately missed two things from Shiftic — the JSON-based programming model and having schemas — so schemas went in from the start, defined with Zod. By **October 2025** the cost of JSON Patch for rich text was clear — the same rich-text pain Dossier had already taught, now from the patch side. I explored several alternatives and landed on a hybrid: JSON for structure, [Yjs for rich-text fields](/concepts/rich-text-with-yjs/), so the format question belongs to a project whose whole job it is. More demo apps followed as stress tests, and the library refined and settled. In **May 2026** the Zod schemas gave way to datadata’s own [declarative schemas](/concepts/schemas-as-documents/), which also made it possible to change a schema at runtime. In place of Dossier’s rule matrix, the answer is deliberately smaller: [three explicit, append-only migration operations](/concepts/schema-evolution/), with documents that no longer fit flagged rather than blocked or dropped. In **June 2026** came [staged sessions](/staged-sessions/overview/): a way to see and approve changes to a document before they land — an AI agent’s work that a person signs off on, or human edits held back for review the same way. Dossier had a coarse precursor — a published/draft split on every document — but staged sessions are finer-grained. Promoted to a library primitive, they’re the review surface that replaces Dossier’s generic admin UI, now that the API is [small enough for an agent’s tool definitions](/ai-agents/why-agents-like-datadata/). The same release reworked how Yjs content is stored and synced, and added a [presence](/concepts/presence/) system. In **July 2026** three things landed, and a smaller fourth closed a loop. Offline support became real: an [opt-in persistence adapter](/concepts/offline-persistence/) carries both writes and reads across going offline — queued writes replay on reconnect, cached documents render immediately, and a reload is just a slow reconnect. It stays a property of [sync and optimistic updates](/concepts/sync-and-optimistic-updates/) rather than a separate mode. And [authorization](/concepts/authorization/) moved into the library: one principal model whether the writer is a person’s frontend, trusted server code, or an AI agent, with declarative access rules that live in the schema and per-document read filtering. Dossier left auth entirely to the host app; agents as ordinary participants — the multiplayer-by-default point above — are what pushed access control into the substrate: an agent should get a narrower grant than the person it works alongside. And the engine was rewritten so a storage adapter no longer has to be synchronous: every operation became a generator that yields its storage requests, run by either a synchronous or an asynchronous interpreter, and the first [Postgres adapter with a Node host](/architecture/node-deployment/) followed the same week — the second backend the Dossier lesson above had been waiting for, held to the Durable Object’s behavior by shared conformance suites rather than shared SQL. The same month closed the loop on deletion with purge: a soft-deleted document can be hard-deleted and its id released. In **August 2026** two of those threads got their endings. Authorization went fully declarative: per-document access entries joined the schema-level rules, and the last imperative escape hatch — app-supplied callbacks that could veto a read or a write — was removed, so policy is entirely data, which is what lets a client evaluate the same rules locally and predict the server’s verdict. And the wire stopped spelling Yjs updates as JSON arrays of numbers: every event is now one [binary frame](/architecture/wire-protocol/) with a JSON header and a table of byte payloads, so a rich-text delta costs what it weighs. In **August and September 2026** the last kind of data without a home got one: files. [Blobs](/concepts/blobs/) are immutable byte objects referenced from documents by handle — the handle syncs like any field, the bytes move over HTTP, and the engine owns integrity, lifecycle, and read authorization. An upload names the write it is for, so authorization runs at upload time, the same gate whether the uploader is a person or an agent. Stores for R2, S3-compatible services, and the filesystem shipped together, conformance-tested on every backend. ## The thread The thread runs through both. Dossier was content infrastructure with sync underneath; Shiftic put structured UI and AI agents on a realtime engine but kept schema in code. datadata promotes what each got right to the product itself — Dossier’s self-contained, schema-aware event log and Shiftic’s changes-as-data over realtime sync — and sheds the layers that aged worst. # When to use it (and when not) > A fit guide for datadata in its current state. This page is the short version. For the per-feature detail, see [Where datadata sits](/comparison/where-datadata-sits/) and [Limitations](/known-issues/limitations/). ## A good fit today * **Collaborative document apps** — multiple people editing structured documents and rich text in real time, with instant local feedback. * **Human + AI co-editing** — an agent stages changes in a [session](/staged-sessions/overview/), a human reviews a three-way preview and commits. This is the workflow datadata is being honed against. * **Evolving domains** — schemas are documents with their own migration history, so the data model can change while the app is running. * **Apps already on Cloudflare** — the production backend is a Durable Object per folder with SQLite storage; if that’s your platform, deployment is natural. A [Node + Postgres backend](/architecture/node-deployment/) exists and passes the same conformance suites, but hasn’t run in production yet. * **Offline-capable apps** — an opt-in [persistence adapter](/concepts/offline-persistence/) makes the write queue and document cache durable, so a reload becomes a reconnect and cached documents render while disconnected. Browser storage makes that durability best-effort, and long-offline unguarded writes replay last-writer-wins — offline work that deserves review belongs in a [staged session](/staged-sessions/overview/). ## Not a fit (today) * **Apps needing per-field permissions.** [Authorization](/concepts/authorization/) covers role rules per document type, per-document access entries, read filtering that hides documents, and scope caps for agents — but visibility is whole-document: a user sees a document entirely or not at all. * **Large datasets.** A folder is the authority unit and lives in one Durable Object; clients subscribe per document, but a document syncs whole (no partial replication within a document), individual documents have a size cap, and there is no event-log compaction yet. * **Tabular, query-shaped data.** Sync is whole-document — you can’t select or project parts of documents, and there is no query language. The supported pattern is a derived store (listen to document changes, project into another document or database, query that), but if your data is fundamentally tabular, datadata is probably not the right fit. ## Not a fit (by design) Some of the above will change; the server-authoritative model won’t. If you want decentralized authority, peer-to-peer sync, CRDTs for all data (not just text), or end-to-end encryption (which is incompatible with server-side validation), datadata is the wrong tool. ## Sounds like a fit? If your app lands in the good-fit column, [get in touch](/contact/) — the library isn’t open source yet, but early access ahead of open-sourcing is a possibility. # Status > What's solid, what's experimental, and what "not open source yet" means. datadata is **early-stage software**. It runs real applications, but it is being actively reshaped and backwards compatibility is explicitly not a goal yet. ## Maturity at a glance This table rates how solid the things datadata *does* are, with a last row for the headline gaps. The fuller catalogue of gaps and deliberate non-goals lives in [Limitations](/known-issues/limitations/). | Area | State | | --------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Core client/server sync, optimistic updates | Solid, exercised daily | | JSON Patch changes with guarded writes | Solid, with [known RFC issues](/known-issues/json-patch-rfc/) | | Domain commands (named mutators, `doc:command`) | New; registry shared by client and server, re-executed prediction, per-command access rules, refusals with codes; not yet journaled offline; see [Domain commands](/concepts/domain-commands/) | | Yjs rich-text integration | Solid; [two independent sync lanes](/concepts/two-sync-lanes/), Yjs updates cross the wire as [binary frames](/architecture/wire-protocol/) | | Staged session (staging, conflict preview, commit) | Converging; bindable staged editors, crash-safe commits, bases resolved from the event log | | Schemas as documents, migrations | Solid; schema evolution and migrations are explicit | | Document lifecycle (delete, restore, rename, purge) | Solid; soft delete with a live trash listing, restore, rename, and purge | | Cloudflare Durable Object backend | Production shape for our apps | | Node + Postgres backend | New; the same engine over a Postgres storage adapter (PGlite driver so far) and a Node WebSocket host, passing every conformance suite; not yet in production; see [Node + Postgres deployment](/architecture/node-deployment/) | | Projection layer (per-document views) | Reads work over the full synced document; writable lenses very alpha | | Presence + Yjs awareness (live and staged) | Solid after several hardening rounds | | Offline reconnect replay | Both lanes replay (JSON last-writer-wins, Yjs by CRDT merge); an opt-in [persisted write queue](/concepts/offline-persistence/) carries the buffer across a page reload | | Authorization (write rules, read filtering, capabilities) | Fully declarative — schema rules, per-document `sys:access` entries, read filtering, scope caps; see [Authorization](/concepts/authorization/) | | Durable offline write queue | New; IndexedDB-backed, every tab journals and a dead tab’s queue is adopted; see [Offline persistence](/concepts/offline-persistence/) | | Offline reads (persisted doc cache) | New; offline reloads render cached documents with pending writes re-projected, `offline-cached` sync status; see [Offline persistence](/concepts/offline-persistence/) | | Offline tab-to-tab sync | New; two offline tabs converge through the shared store (Yjs content CRDT-merges via the cache; pending JSON writes mirror read-only, offline creates listed); see [Offline persistence](/concepts/offline-persistence/) | | Blobs (files as immutable handles) | New; R2, S3-compatible and filesystem stores, integrity gate and read gate conformance-tested on every backend, garbage collection scheduled by both hosts; see [Blobs](/concepts/blobs/) | | Log compaction | Not yet — see [Limitations](/known-issues/limitations/) | ## Not open source (yet) The source isn’t public yet — partly because the API is still moving, and partly because good docs felt like the right first artifact. This site documents how datadata works and why, so that: * application developers can evaluate the model before betting on it, and * sync-engine developers can compare notes on the design. If you want to look closer, try it, or argue with a design decision, [get in touch](/contact/). ## How gaps are documented Every page on this site flags its own gaps inline, and the [Known issues & open questions](/known-issues/) section aggregates them with full write-ups. If something reads as more finished than it is, that’s a bug in the docs — [tell us](/contact/). # Live demos > The full engine running in your browser — real clients, real server, no backend. Each demo runs the complete engine [in-memory, in the page](/architecture/running-in-memory/): real clients, a real server, and a latency slider standing in for the network. ## Contingency [**Contingency**](/demos/contingency/) is a co-op tower defense game about [staged sessions](/staged-sessions/overview/). You draw plans; a scripted commander builds live under them. Each plan is a staged session over the live board, so it turns red when the board moves under it, and a deploy is one atomic commit. Presence carries the pointers and the enemies, and a [shared clock](/concepts/shared-clock/) moves the enemies on every screen without sending where they are. ## Kanban [**Kanban**](/demos/kanban/) is a board you share with three simulated teammates: a product owner, a developer and QA. It shows the everyday collaboration features: live cursors and drags as [presence](/concepts/presence/), card descriptions as [Yjs text](/concepts/rich-text-with-yjs/) with everyone’s caret, guarded moves that snap back when two people move one card, and card fields added at runtime as writes to the board’s [schema document](/concepts/schemas-as-documents/). # Contingency: plans as staged sessions > A co-op tower defense game where every plan is a staged session over the live board, running the whole engine in the page. You plan; Bo, a scripted commander, builds live. Every plan you draw is a [staged session](/staged-sessions/overview/) over the board Bo is building on, so you watch it go red the moment Bo builds where it wanted to. Everything runs in this tab: two clients, one server and a simulation, on the [in-memory engine](/architecture/running-in-memory/), with no backend. [Play it full window](/play/contingency/), which is the better way on a phone. Under the game are its documents, as one client holds them. Switch between your client and Bo’s, raise a latency, and watch his writes sit unconfirmed and the two views disagree until the server catches up. A plan’s card lists its staged operations, each write behind the `test` that guards it, and the guards that no longer hold turn red. ## How to play * **Start.** Nothing moves until you press Start, and you can plan before you do: close the intro card and Start waits in the wave forecast at the top. The first wave comes 30 seconds after. * **Plan.** Click or tap a cell and pick what to build there, or what to do with the tower on it. The first change starts a plan. Its towers show as blue holograms, and every other plan’s cells carry that plan’s badge. * **Deploy.** “Deploy” puts the whole plan on the board at once, for one of your two command points. * **Watch it go red.** Bo builds live and doesn’t look at your plans. When he builds on a cell your plan needs, the plan turns red. “Rebase on the board” takes the board as it is now and keeps the rest of the plan. * **Spend your own gold.** You and Bo each have a wallet, and every kill pays into both. Bo spends his on live towers; your plans spend yours. * **Break the network.** Your latency slider and unplug switch sit at the foot of the plan panel; Bo’s are in his view. Unplugged, you can keep planning, but you can’t deploy. The castle has 20 lives; hold it for eight waves to win. ## What datadata is doing **A plan is a staged session with a host document.** Each plan is a `plan:` document that [holds its session’s changeset](/staged-sessions/overview/#the-changeset-can-live-in-a-document). Because the changeset lives in a document, every pane can open every plan, and a plan survives an unplugged pane. The holograms are the session’s staged view: the live board with the plan laid over it, which keeps moving as Bo builds. **A red plan is a derived conflict.** Every buildable cell exists from the start as `{ tower: null }`, so a staged placement is a `replace` guarded by a `test` that the cell is still empty. When Bo builds on that cell, the guard no longer holds and the plan’s stage is blocked. Plans use the default `autoMerge` policy: Bo building elsewhere leaves them clean. Rebase is the [resolution verb](/staged-sessions/conflict-preview-and-resolution/#the-resolution-verbs) that accepts the live board. **A deploy is one atomic commit.** Deploying stages a ledger entry for the plan’s cost on your wallet, the `bank:planner` document, then commits: the towers on `board` and the cost in your wallet land together or not at all. **Live building is a guarded write.** Bo’s towers go straight to the board as guarded patches, applied [optimistically](/concepts/sync-and-optimistic-updates/) and confirmed or rejected by the server. Raise his latency and watch his towers appear late in your view. **Gold is a ledger, not a counter.** You and Bo each have a wallet, a `bank:planner` and a `bank:commander` document. Every purchase, kill and deploy adds an entry under a fresh key, so spending never conflicts with anything; a balance is the sum. Every kill pays into both wallets. datadata’s authorization doesn’t read document contents, so nothing in the engine stops two deploys from overspending. The game allows debt instead: while a wallet is negative, the towers it paid for fire at half rate. **Pointers and enemies are presence.** Each player’s pointer, and the plan they’re viewing, are [presence](/concepts/presence/) cells. So are the enemies: the simulation publishes their positions ten times a second, and they never enter the event log. ## Gaps * **The second player is a bot, in the same tab.** Two people in two browsers would need a server; the demo shows the engine, not a deployment. * **One changeset per host document.** Each plan needs its own document, because a host holds one changeset. [Multiplexed changesets](/known-issues/open-questions/#multiplexed-changesets) are an open question. * **No overspend guard in the engine.** Debt is a game rule standing in for an integer decrement the engine doesn’t have. * **Unmeasured on phones.** The game lowers its resolution and shadow detail on touch devices, but its frame rate on a real phone hasn’t been measured. # Kanban: a board shared with three teammates > A kanban board where three simulated teammates work alongside you, with live cursors, collaborative card descriptions and card fields added at runtime, running the whole engine in the page. You share this board with three simulated teammates. Ada, the product owner, writes new cards into the Inbox. Bo, the developer, pulls them through To do and Doing. Cleo, QA, tests what lands in Review and passes it or sends it back. Each of them works through a client of their own, the same way you do, so every pointer, drag and keystroke you see has gone through datadata. Everything runs in this tab on the [in-memory engine](/architecture/running-in-memory/), with no backend. They wait until you press Start, so you can look around first. [Open it full window](/play/kanban/), which is the better way on a phone. ## What to try * **Triage the Inbox.** Drag cards into To do. If you leave the Inbox alone, Bo starts pulling from it himself. * **Watch the cursors.** Every pointer sits on the card it’s over, on every screen. A card someone is carrying fades in its column and travels as a ghost under their pointer. * **Type together.** Open a card Ada is writing and you see her caret in the text. She sometimes joins a card you have open, too. * **Race for a card.** Raise your latency and move a card Bo or Cleo is about to move. One move wins; the other snaps back and says who got there first. * **Add a field.** “Card fields” adds Priority, Estimate or a field of your own to every card. Ada starts filling it in. * **Go offline.** Keep moving and editing cards; it all lands when you come back. ## What datadata is doing **The board is one document, with the cards in a record.** A move rewrites one card’s `column` and `index`, so you and Bo moving different cards touch different paths, and both land. The `index` is a [fractional index](/concepts/changes-as-json-patch/#ordered-collections-without-arrays), healed per column: when two cards are dropped into the same slot at once, the schema re-spreads them instead of letting them tie. **A move is a guarded write.** It holds only if the card is still where the mover last saw it. When two people move the same card at once, the server takes the first and rejects the second, which rolls back on its own [optimistic](/concepts/sync-and-optimistic-updates/) copy. That is the snap-back. See [Conflicts](/concepts/conflicts/#live-editing--resolved-as-it-happens). **Descriptions are Yjs.** Each card’s description is a Y.Doc [referenced from the card](/concepts/rich-text-with-yjs/), so two people typing in it both keep their words. The carets are Yjs awareness, carried on a [presence](/concepts/presence/) channel. Deleting a card deletes its description with it. **Pointers are presence.** Each person’s pointer, the card they’re carrying and the card they have open are a presence cell on the board. A pointer is anchored to a card or a column rather than to screen pixels, so it lands on the same card whatever your layout. The teammates read the same presence: they don’t pick up a card someone else is holding. **The schema is a document too.** Adding a field writes to the board’s schema document, `sys:schema:board`. Every client runs with [dynamic schemas](/concepts/schemas-as-documents/#where-the-schema-truth-lives), so it starts validating cards against the new schema the moment the write arrives. Before that, a card carrying the field is rejected as an unknown field. ## Gaps * **The teammates are bots, in the same tab.** Three people in three browsers would need a server; the demo shows the engine, not a deployment. * **Fields can’t be removed yet.** Removing a field is a [schema migration](/concepts/schema-evolution/); the demo only adds them. * **One schema per document type.** Every board of a type shares its schema, so per-board fields would mean a type per board. # Documents & folders > The unit of data and the unit of sync. ## Documents A document is the unit of data: * **`docId`** — a string id, unique within its folder: 1–128 letters, digits, `-` and `_`, which fits UUIDs, nanoids and prefixed ids like `project_42`. The one exclusion is the name of an `Object.prototype` member (`__proto__`, `constructor`, `toString`, …): ids end up as keys of other documents, such as the folder’s index, and those names can’t safely be one. The same rule covers the ids of embedded Yjs documents, blobs and presence cells, and creating or importing a document under any other id is rejected. * **`type`** — a string naming the document type. Every type has a schema document (`sys:schema:`) that writes are validated against. * **`sequence`** — a monotonically increasing number assigned by the server on every accepted change. It orders the document’s log and is what a [sequence guard](/concepts/changes-as-json-patch/) anchors to. It counts JSON-lane changes only — edits to embedded Yjs documents travel on a [separate lane](/concepts/two-sync-lanes/) and don’t advance it. * **`generation`** — an opaque token telling this document apart from anything else that has ever carried its `docId`. It is minted when the document is created — by the creating client, so its not-yet-confirmed copy already has one — or imported, and never changes afterwards — not on edits, delete or restore. Only [purge](#deleting-documents) followed by a new create under the same id produces a new one. * **`data`** — the JSON payload. Rich text fields hold references to embedded [Yjs documents](/concepts/rich-text-with-yjs/) rather than raw text. A document can also have a user-facing **name** — an optional display label that is *index-level metadata*, not part of `data`. Set it at create time or change it later with `renameDocument`; renaming never touches the document’s `data`, `sequence` or event log. A name is capped at 1024 characters; an empty or whitespace-only name is treated as unnamed (`null`). By convention a `/` in the name nests documents into display-only subfolders — a UI grouping over the flat label, not an engine-level hierarchy (the folder below is still one flat namespace). ## Folders A **folder** is the unit of synchronization and authority: one server instance owns one folder, orders all of its events, and broadcasts to all of its subscribers. In production a folder maps to one Cloudflare [Durable Object](/architecture/cloudflare-deployment/). On the [Node backend](/architecture/node-deployment/) one process and one Postgres database serve many folders, each still with its own server instance and its own ordering. Clients subscribe per document, not per folder — a client only receives changes for documents it has subscribed to. Schemas are **per folder**, too: a document type is defined by the `sys:schema:` document living inside the folder, and a document’s `schema_sequence` is a folder-local coordinate. Two folders can carry the same type name at independently evolved schemas. Admission is **folder-granular** — a connection presents a host-minted token to enter the folder — but inside it the engine authorizes every write and filters every read across three declarative grains: folder roles, access rules per document type, and per-document entries in the folder’s `sys:access` document — with hard scope caps on the principal above them all. See [Authorization](/concepts/authorization/). ## Deleting documents Deletion is **soft**: a deleted document leaves the index and every read path but is listed in [`sys:trash`](/concepts/system-documents/#folder-state), with its row, event log, and embedded Yjs state retained so it can be restored. Delete is a *lifecycle* change, recorded in the folder’s membership log rather than the document’s own event log — restoring a document returns it at the exact sequence it left, history intact. Deletion is also **observable**. Live subscribers are pushed a `doc:deleted` event the moment a delete commits (an open editor can show “this document was deleted” instead of silently going stale), and subscribing to an already-deleted document answers `doc:deleted` too — deliberately distinct from “never existed”. Subscriptions survive the deletion, so a restore pushes the document straight back to everyone who was watching. A deleted document still **owns its id**: creating a new document under a deleted docId is rejected with a distinct `deleted` error (restore it, or pick a new id) rather than the benign already-exists signal. Deletes are **optimistic and stackable**: the document reads as gone locally the moment you call `deleteDocument`, and you can delete a document whose own create is still in flight — the pair goes to the server in author order rather than being cancelled against each other locally, so both writes report their real outcome. If the create is rejected, the trailing delete is dropped along with it and the failure surfaces once. If the create lost the id to another client’s create, the delete is refused too — it names the document this client created, never the one that won the id. **Purge** is the second stage — hard deletion, for storage reclamation and right-to-erasure. It takes a *deleted* document, destroys its event logs, Yjs state and snapshot — and, once a later sweep finds them unreferenced for longer than the retention period, the bytes of any [blob](/concepts/blobs/) that no other document still references — and **releases its id**: afterwards the id behaves exactly like one that never existed — a subscribe answers `notfound`, restore has nothing to resurface, and a create under it simply starts a fresh document: a new **generation**, back at sequence 1. A client still holding a copy of the purged document — offline through the whole purge, say — is not fooled by a matching sequence: on reconnect its copy (embedded Yjs content included) is replaced by the new document, and edits, deletes and renames it made to the old one are rejected rather than applied. So is a restore it queued for the old one, even once the new document has been deleted too: a restore names the generation the client saw deleted, so it never brings back a different document. That includes the purged document’s own create: replayed by a client that never saw it acknowledged, it is rejected rather than bringing the purged generation back. That release is what makes purge safe for deterministic ids (one document per domain entity, a `settings` singleton): an id is never permanently burned. One erasure boundary to know: purge destroys a document’s *content*, never its *identifier* — the original id is kept in a server-side audit record, so don’t encode data in a docId that would itself need erasing. Purge has two entry points, and both are **batch-shaped and all-or-nothing**: one storage transaction, one `sys:trash` sequence step for the whole batch, and a failing member anywhere destroys nothing. Clients purge over the wire with `purgeDocuments` (the “delete forever” button) — but only for document types whose schema **explicitly opts in** with an `access.purge` rule; the default is *nobody*, admins included (see [Authorization](/concepts/authorization/)). It is deliberately never optimistic and never queued offline: irreversible destruction doesn’t sit in a replay buffer. On the server, `purgeDocuments` (system authority) empties specific tombstones, and `purgeDeletedDocuments` sweeps everything deleted before a cutoff, oldest first and optionally bounded — the mechanism behind “trash empties after 30 days”. The retention window and the trigger (a Durable Object alarm, a cron, opportunistically on wake) belong to the application; datadata deliberately owns no scheduler. One carve-out: `sys:` documents — including schema documents — can’t be deleted, and therefore can’t be purged; see [Limitations](/known-issues/limitations/). ## System documents The engine maintains a handful of documents itself — the folder index, the trash, schemas, staged-work views, presence — each readable and subscribable like any other document. [System documents](/concepts/system-documents/) enumerates them. Limitations A folder is also the *scaling* unit: there is no cross-folder sync or federation, and a document syncs whole — no partial replication within a document. See [Limitations](/known-issues/limitations/). # Schemas as documents > Document types are defined by schema documents, edited through the same API as data. Take any document type — say `task`. Its schema is itself a document: `sys:schema:task`, of type `sys:schema`. It is created and updated through the same API as user data, synced to clients like user data, and validated like user data — against a meta-schema. ## What a schema describes A schema document declares a document type’s shape, plus the rules that govern it: * **Field types** — objects (fixed, named keys), records (open string-keyed maps), arrays, strings, numbers, booleans, and enums. And `json` — a field that holds *any* JSON value, the escape hatch from a declared shape. * **References** — fields that point at other documents (or at records within the same document), with integrity rules for what happens when the target disappears. See [References & integrity](/concepts/references-and-integrity/). * **Defs** — reusable type definitions, referenced from fields so a shape can be declared once and shared across the schema. * **Access rules** — the declarative per-type role rules the engine evaluates on every write and read. See [Authorization](/concepts/authorization/). * **Migrations** — an append-only log of schema evolutions (rename, remove, remap). A document written under an old schema version is brought forward on its next read and the migrated data is persisted back — the full story is on [Schema evolution](/concepts/schema-evolution/). Every document type has a schema — there is no schemaless document. When you need loose or open-ended data, reach for a `json` field: it rides inside a validated type and syncs and patches like any other field — accepting any JSON-encodable value, just without a declared shape. One rule holds for every document, typed fields and `json` alike: strings — values and object keys — are well-formed Unicode, as [I-JSON](https://datatracker.ietf.org/doc/html/rfc7493) requires. A lone surrogate (half an emoji, left by a careless `.slice`) is rejected at the write like any other invalid data. ## Why schemas want to be documents **One API.** Evolving the data model is a document write, not a deploy. The same subscription machinery that gives you live user data gives you live schema changes. **The model is editable at runtime.** Introducing a document type is a document write — no deploy, no code generation, no out-of-band tooling. A user building types from a UI and an [AI agent](/ai-agents/why-agents-like-datadata/) that decides it needs a new type take the very same path. **[One timeline, not two.](/introduction/one-timeline/)** As documents, schemas have their history in the same event log as the data they govern — introspecting a document as it was at any point in time means looking in exactly one place, not correlating deploy times against data versions. It’s also what makes the data [portable](/architecture/portable-event-streams/): a folder’s history carries its own interpretation with it. **Schemas version like data.** Sequence numbers, event history, and guarded writes apply to schema changes exactly as they do to data changes. ## Where the schema truth lives Schemas being documents doesn’t mean giving up TypeScript. There are three ways to source them, and they compose: **Authored in TypeScript.** A schema can be defined in code, and the document data types are *inferred from the schema literal* — compile-time validation and typed reads/writes with no code-generation step. The defined value is a plain schema document’s data; the types are phantom. **Upserted at startup.** The host application writes its TypeScript-defined schemas into the `sys:schema:*` documents when a folder boots — a no-op when nothing changed. Rolling out a schema change in code becomes a document update in each folder, still subject to the server’s append-only [migration rules](/concepts/schema-evolution/). This is what makes TypeScript authoring compatible with the one-timeline argument above: the upsert turns every code change back into document history, so even when git is where schemas are *written*, the schema documents remain where their history *lives* — the repository never becomes a second timeline you’d have to consult. **Document-driven only.** Types created at runtime — by a user, or by an [agent](/ai-agents/why-agents-like-datadata/) — exist purely as schema documents, with no TypeScript counterpart. Whatever the source, **the server’s write path always validates against the schema documents** — the TypeScript definitions are never a second, shadow truth. Clients then choose per app: bundle the static TypeScript schemas (typed, instant optimistic validation, no waiting for schemas to sync) or, with dynamic schemas enabled, subscribe to `sys:schema:*` and validate against whatever the folder currently declares — required when document types are born at runtime. ### Waiting for a type’s schema A dynamic-schema client doesn’t write a type until it knows that type’s schema. `createDocumentAndWait` and `updateDocumentAndWait` wait for it before they stage the write. The fire-and-forget `createDocument` and `updateDocument` throw if the schema isn’t resolved yet, so wait first: ```ts const schema = await session.waitForSchema("chatConversation", { signal }); session.createDocument({ docId, type: "chatConversation", data }); ``` `waitForSchema` asks about the type alone. No document of the type has to exist, and nothing is subscribed on your behalf. It resolves with where the schema comes from: | `source` | Meaning | | --------- | -------------------------------------------------------------------------------------------- | | `dynamic` | The client holds the type’s `sys:schema:` document. Carries `sequence` and `confirmed` | | `static` | The folder has no schema document for the type; the bundled schema validates it | | `none` | No schema document and no bundled schema. Nothing validates the type, so writes are refused | | `staged` | A staging session only: the schema is staged in that session and not committed yet | A schema held in the offline cache counts. `confirmed` is `false` until the server has answered, and the write is still accepted, so an app that boots offline with its schemas cached can author. A client that boots offline with nothing cached has no way to know, and the wait stays pending until it connects. Pass a `signal` to bound it. The answer holds for a write you issue right away. It can change later: when someone creates a schema document for a type that had none, the type is unresolved until that document has synced. Wait again before a later plain write, or use the awaited writes, which hold through it. If the schema could not be read (the server answered an error for `sys:index` or for the schema document), the wait rejects with a `SchemaUnavailableError` and writes of the type are refused with the same error. An absent schema is not an error: that is the `static` or `none` answer. Before an update of a document you know exists, await the prepared document’s `available`. It includes the schema wait and also asserts the document is readable, so it rejects for a missing or deleted document. For a create-or-update, use `waitForSchema`. # Schema evolution > Versioned schema documents, explicit append-only migrations, and what happens to documents that no longer fit. [Schemas are documents](/concepts/schemas-as-documents/), so schema evolution is document editing — but evolution has rules of its own. datadata’s stance is **evolution-first**: schemas change freely while the app runs, migrations are explicit, and documents that no longer fit are flagged, never dropped. ## Versioning A schema’s version is its document’s **sequence number** — there is no separate version field. Every accepted edit to `sys:schema:` advances it, and every document tracks which schema sequence its data conforms to. Every event in a document’s history also records the schema sequence in effect when it was written — which is what keeps historic versions [interpretable anywhere](/architecture/portable-event-streams/). Because that number counts *this folder’s* accepted edits to the schema document, it is a [folder-local coordinate](/concepts/documents-and-folders/): the same type can sit at different sequence numbers in two folders depending on when each folder’s schema was last edited. A version number only means something inside its own folder — there is no global schema version to compare across folders. ## Migrations are explicit and append-only Data-shape changes are declared as **migrations** in the schema document. Three operations exist, deliberately minimal: * **`rename`** — move a field to a new name. * **`remove`** — delete a field. Removal is *never inferred*: dropping a field from the schema without a `remove` migration makes documents still holding it invalid, rather than silently discarding data. * **`remap`** — rewrite scalar values (old → new pairs). Collapsing several old values into one is allowed; one-to-many is not. Remaps are pure — the new value depends only on the old one. `migrations` is a record — each migration sits under a **key** its author picks: ```ts // As stored: `sequence` is the server's stamp, whatever the author wrote. migrations: { "001-rename-title": { op: "rename", sequence: 2, from: "title", to: "heading" }, "002-remove-body": { op: "remove", sequence: 4, path: "body" }, } ``` The log is **append-only, enforced by the server**: a schema write may add keys but never modify or drop a committed one, and the server stamps each new migration with the schema sequence it took effect at — authors don’t control the stamps. A **rename never replaces a field unasked**. Where the target name may already hold a value, the rename has to say what happens then, with `onConflict`: * `keep` — nothing: both keys stay, and validation reports the one the schema doesn’t declare. * `replace` — the renamed value takes the name; the old occupant is dropped. * `discard` — the occupant keeps the name; the renamed field is dropped. ```ts // Merge `nickname` into `displayName`, keeping displayName where both exist. "003-merge-nickname": { op: "rename", sequence: 5, from: "nickname", to: "displayName", onConflict: "discard", }, ``` The server requires `onConflict` exactly where a collision is possible and refuses it everywhere else, so it always means something. It is required when the target is declared in that container (in the schema being replaced, in any variant of a union, or because an earlier migration of the same write renamed something there), and when the container’s keys are data: a record’s keys, an object that keeps undeclared members, a `json` value. A rename of a field onto itself is always refused. Instead of a policy you can `remove` the target under an earlier key; to swap two fields, go through a temporary name (`a → tmp`, `b → a`, `tmp → b`). Merging two fields into one is the second rename with `discard`: `a → c`, then `b → c` leaves `c` holding whichever of the two was present, `a` first. Whether a rename with a policy actually happened depends on each document, so an offline edit written against the old shape that touches either of its keys can’t be carried forward faithfully: it is rejected as a precondition failure to refetch and rebase. Edits elsewhere in the document carry over as usual. Migrations reach nested data through dotted paths, with `*` standing for every member at that level (a rename’s `at: "items.*"` renames a field in every element of `items`). **An array is entered through `*` only**: the server refuses a migration whose path steps into an array by index (`items.0`). A position means a different element as soon as something is inserted or removed before it, so a migration aimed at one element could not be carried faithfully across the edits written against the old shape. A record is keyed by name, not position, so `labels.en` stays a valid path. What authors *do* control is the key, and the keys’ lexicographic order is the replay order (English collation, case- and accent-sensitive — not numeric, so pad numbers: `010` sorts after `009`, `10` sorts before `9`). A new migration’s key must sort after every committed one — the server rejects a key that would land inside the committed range, since that would silently reorder replay for older documents, and rejects an empty key. Keys can be written by hand or generated by tooling; a convention like `002-rename-title` sorts correctly and stays readable. Two authors adding migrations at once merge as long as their keys differ and the later write’s key sorts last; otherwise the later write is refused and needs a new key. Where several people author migrations concurrently, a prefix that sorts by authoring time — a date (`2026-09-21-rename-title`) or a ULID — makes the out-of-order case rare and a shared key all but impossible. A [staged session](/staged-sessions/conflict-preview-and-resolution/#when-the-schema-moves-under-you) reports a staged key the server would refuse before commit, and rekeys it in one operation. Tooling that writes schema documents on an author’s behalf can reach the same verdict itself: `@repo/datadata/schema` exports `compareMigrationKeys` (the key order), `orderedMigrations` (a log’s entries in replay order) and `migrationLogRefusal` (the append-only verdict on a submitted log against the committed one, with the key it is about), so a host mints a key that sorts last and checks its log with the rule the server enforces, rather than re-implementing the collation. `migrationShapeRefusal` is the server’s other verdict, on the migrations a write adds, judged against the schema it replaces. When the author can’t see the committed log at all — an LLM evolving a type through a tool, or a form that collects “the migrations for this change” — `appendMigrations` mints the keys: ```ts import { appendMigrations } from "@repo/datadata/schema"; const migrations = appendMigrations(committedSchema.migrations, [ { label: "rename-title", migration: { op: "rename", from: "title", to: "heading" } }, { migration: { op: "remove", path: "body" } }, ]); // { …committed, "0003-rename-title": …, "0004": … } after a log ending in "0002-…" ``` It appends in the order given, for any committed log: it continues a counter the last key ends in (`0002-…` → `0003-…`) and otherwise extends that key (`zebra` → `zebra~0001-…`), checking each key against the collation — under which simply appending characters to a key doesn’t always sort it later. The migrations it takes have no `sequence`, since the server stamps every new one. A label can’t start with a digit or contain `~`, since the key would then read as something else. `parseMigrationKey` reads a minted key back into its stem, counter and label. When a staged session reports that the committed log moved on under its migrations, re-mint them after the new log and keep their labels: ```ts import { appendMigrations, parseMigrationKey } from "@repo/datadata/schema"; // staged: the [key, migration] entries the new log doesn't have yet const reminted = appendMigrations( movedSchema.migrations, staged.map(([key, { sequence, ...migration }]) => ({ label: parseMigrationKey(key)?.label, migration, })), ); ``` ## Migration runs on read — and writes back When a document is read whose conformed sequence is behind the schema, the server brings it forward: it replays the migrations stamped after that sequence — in key order — validates the result (backfilling declared defaults for added fields), and — if anything changed — **persists the migrated data back** as a normal change with a new sequence number. Each document pays the migration cost once, on its next read, not on every read. The server’s `onDocumentChange` doesn’t fire for the write-back — the schema write that caused it is the change that gets announced. That change’s patch says what the migrations did: a `rename` is a JSON Patch `move` from the old name to the new one — not a `remove` and an `add` — a `remove` is a `remove`, and a `remap` a `replace`. Backfilled defaults follow as ordinary `add`s. The write-back doesn’t wait for the result to be valid. If the migrations rewrote the document but it no longer fits the schema, the migrated shape is still what’s persisted — it’s what every reader is handed, so it’s what the event log has to hold too — and the document is flagged alongside it. Only the flag says whether it conforms; the conformed sequence says how far the stored data has been carried. **Subscribed documents don’t wait for a read.** A client that already holds a document never reads it again, so the schema write itself brings forward every document of that type with a live subscriber, right after it commits. Each subscriber receives the migration as an ordinary patch or, for a document the new schema flags invalid, a fresh `doc:init` carrying the flag. Either way, the migrated state reaches open clients as soon as the schema change lands. Beyond subscriptions there is no proactive bulk sweep: a document nobody reads keeps its old shape, and its pending migrations simply accumulate until it’s next loaded. **A copy says how far it has been carried.** The schema change and a document’s migration reach a subscriber as two frames, the schema first, so for a moment a client holds the new schema and a copy still in the old shape. Every frame that carries a document’s data — `doc:init`, the `doc:patch` of each write, and `doc:resume` — names the document’s conformed sequence, and the client keeps it on its copy as `conformedSchemaSequence` (cached with the copy, so a reload in that moment still knows). A copy whose sequence trails its schema’s is one the server still owes a migration, and everything on the client that has to read it in the current shape replays from exactly there instead of guessing. It is “carried to”, not “conforms at”: it moves for a document the new schema flags, too. Documents of a [static-schema](/concepts/schemas-as-documents/) type carry none — there is no migration log for them to be behind. **A client that was offline is told when it reconnects.** Its resubscribe is a read like any other, and a document that read flags is always answered with a full `doc:init` carrying the flag — even when the schema change reshaped nothing and the client’s copy is current byte for byte. Validity isn’t a function of the sequence, so a matching one doesn’t earn the usual no-data `doc:resume`; the document is resent on each reconnect until it’s repaired. A `doc:resume`, in turn, tells the client the document is valid and clears a flag it held. ``` flowchart TB accTitle: What happens to a document on read accDescr { When a document is read, the server compares the sequence its data conforms to against the schema's current sequence. If it is current, the document is served as stored. If it is behind, the migrations stamped after that sequence are replayed in order and the result validated, backfilling declared defaults. A result that changed is persisted back as a normal change at a new sequence, whether or not it validates. A result that no longer fits the schema is delivered anyway, flagged with its violations. } read["document read"] --> behind{"conformed sequence
behind the schema?"} behind -->|"no"| serve["deliver to the reader"] behind -->|"yes"| replay["replay the migrations stamped after
that sequence, in key order"] replay --> changed{"did anything
change?"} changed -->|"yes"| persist["persist back as a normal change
at a new sequence"] changed -->|"no"| validate persist --> validate{"valid under
the current schema?"} validate -->|"yes"| serve validate -->|"no"| flag["flag with located violations
data intact · migrated shape persisted"] flag --> serve ``` ## Pending edits survive a migration A client’s pending write — an edit still unsent, or one made offline and replayed on reconnect — is authored against the shape the client held. If a migration moves that shape before the write lands, the patch as written no longer fits: `replace /title` finds no `title` after `title` was renamed to `name`. Rejecting that as ordinary drift would silently discard offline work the user could not have prevented. So every write names the **schema sequence it was authored against** (the `schemaSequence` on `doc:create` and `doc:update` — the sequence of the `sys:schema:` document the client validated the write with), and the server **brings the write forward** through the migrations stamped since, exactly as it brings a stored document forward on read: * A `rename` moves the paths the patch targets (and the keys inside any value it writes). One with an `onConflict` policy may not have happened in a given document, so an edit touching either of its keys is rejected as a precondition failure to refetch and rebase; a value that contains its container is still carried, with the policy applied to it. * A `remove` drops the operations that edit the removed field — the schema author’s stated intent, not a loss — and strips the field from any value the patch writes. A patch emptied this way is acknowledged as a no-op, never rejected. A remove whose path lands on an array slot removes nothing (arrays are never left with holes), so an edit there is kept. * A `remap` rewrites the values the patch sets at the remapped path. Guard `test` ops are rewritten with the rest, so a `patch`-guarded edit tests the migrated document for the migrated value. A create’s `data` is carried the same way. A raw `move` or `copy` is carried too, including one whose transferred value a migration rewrites inside of, as long as the migration rewrites it alike at both ends (reordering the elements of an array whose every element gets a field renamed, say). What cannot be carried, besides an edit to a policy rename’s keys, is a transfer the migration rewrites at one end but not the other, or one that a `remove` deletes at either end: the author moved the pre-migration value, which no longer exists, so such a write is rejected as a precondition failure to refetch and rebase, rather than applied as something the author did not write. datadata’s own clients never emit `move` or `copy`. Only a client validating against a bundled [static schema](/concepts/schemas-as-documents/) sends no sequence — it has no migration log to name — and its writes apply as authored. A dynamic-schema client never falls into that case by accident: it refuses a write until the type’s schema document has synced, so every write it sends names a sequence. The client carries its own pending copy of the write the same way. When its synced `sys:schema:` document moves past the sequence a pending write names, it rewrites the pending patch (or a pending create’s data) with the same migrations, re-stamps it with the new sequence, and re-journals it — so the edit keeps rendering over the migrated copy while it waits for its acknowledgement, including a whole offline queue replayed after a migration, and a reload or another tab sees the rewritten write. The client never migrates its confirmed copy itself: that arrives from the server as an ordinary patch, and a pending update is rewritten once its copy’s `conformedSchemaSequence` says the copy is at the new shape (or the migrations it is still owed don’t touch it). An edit made in the moment between the schema frame and the migration frame is authored in the shape the copy is still in: the client checks it as the server will land it — carried through the owed migrations, then validated — stamps it with the copy’s sequence rather than the schema’s, and rewrites it with the rest when the migration arrives. Re-sending a rewritten write that was already on the wire is safe, because writes are deduplicated by event id. The server’s bring-forward remains the authority: any pending write the client cannot rewrite — it has no local copy of the document, say — goes out as authored and is brought forward on arrival. Work held in a [staged session](/staged-sessions/overview/) is carried the same way, without being rewritten: each staged change records the sequence it was authored at, and the session [reads the stack brought forward](/staged-sessions/conflict-preview-and-resolution/#when-the-schema-moves-under-you), so a migration never surfaces as a conflict. ## Presence schemas don’t migrate The one exception is a [presence](/concepts/presence/) schema (`sys:schema:presence:`). Presence data is ephemeral — never stored, so never brought forward — which means a migration could never fire. So a presence schema carries **no migration log at all**: a write that declares migrations is rejected. The field shape still evolves in any way you like (add, remove, retype, restructure) purely by editing the schema, since there is no stored presence for the change to strand; live cells simply re-validate against the new shape on their next write. ## Invalid documents are flagged, not dropped Schema edits are **not checked against existing documents** — you can tighten a type or add a required field freely, and documents that no longer fit become invalid. What happens then is asymmetric on purpose: * **Writes are strict.** A change that would leave a document invalid under the current schema is rejected (a `schemaValidation` error on the [wire](/architecture/wire-protocol/)). * **Reads are relaxed.** An invalid document is still delivered — flagged with the violation and its data intact, with the flag recorded so the document can be enumerated and repaired later. The app (or an [agent](/ai-agents/why-agents-like-datadata/)) decides how to repair it. Invalidity is discovered on read — at schema-write time only the documents with a live subscriber are checked, and their subscribers see the flag appear, or disappear when a later edit loosens the schema again — so a document nobody has loaded since the tightening isn’t known to be broken yet. The **discovered invalid documents** are enumerable, so “what broke when we tightened the schema?” is a query over what reads have surfaced so far. When you need the full audit, a **budgeted validation sweep** forces that discovery: it reads every document whose conformed sequence trails the current schema (in host-sized batches, off the hot path), migrating the ones it can and flagging the rest — after which the enumeration is complete, and “is the migration done?” is answerable. Each flagged document carries its violations as located paths, attributed to the schema shape or to a broken cross-document reference. ## Unknown fields Object types and discriminated unions choose how to treat keys the schema doesn’t declare: * **`reject`** — the default. An unknown field is a validation issue. * **`strip`** — accept the value but drop unknown fields from the validated output, for open-by-design shapes. * **`keep`** — accept and retain unknown fields, as long as the retained values are still JSON-compatible. Gaps & open questions Nothing warns at schema-write time that an edit will strand existing documents — you find out as flagged-invalid reads, or by running the validation sweep afterwards. Repair workflows for flagged documents are still being shaped in the apps. See [Open questions](/known-issues/open-questions/). # Sync & optimistic updates > How a change travels from a client through the server and back. datadata uses **optimistic updates**: every change applies to the local view immediately, while the server remains the single source of truth. Editing keeps working while you’re offline — changes replay on reconnect — but a change isn’t *durably* committed until the server accepts it. ## The life of a change Numbered for reading, not for timing: step 1 finishes before the server has heard anything, and the application spends steps 2 and 3 already rendering the change. 1. A client calls `updateDocument`. The change is applied to the local overlay instantly and sent to the server tagged with a client event id. 2. The server validates the change (schema, guards, limits), assigns the document’s next **sequence number**, persists it, and appends it to the document’s event log. (The sequence versions the JSON snapshot only — Yjs deltas travel in [their own lane](/concepts/two-sync-lanes/), under their own cursor.) 3. The server broadcasts the change to every subscriber of that document. The originating client recognizes its own event id in the broadcast and retires the optimistic entry — the local view and the confirmed view now agree. ``` sequenceDiagram accTitle: The life of an optimistic change accDescr { The app calls updateDocument. The client applies the change to its local overlay at once, so reads reflect it immediately, and sends it to the server while the app is already rendering it. If the server accepts, it validates the change, assigns a sequence, persists it and broadcasts it to every subscriber; the originating client recognises its own event id and retires the optimistic entry. If the server rejects it, the client discards the optimistic entry and reads fall back to confirmed state. } participant app as App participant client as Client participant server as Server participant peers as Other subscribers app->>client: updateDocument(...) client-->>app: local reads reflect the change at once client->>server: doc:update — patch + client event id Note over app,client: the app renders the change
while the round-trip is still in flight alt server accepts server->>server: validate · assign sequence
persist · append to event log server-->>peers: doc:patch server-->>client: doc:patch — same client event id client-->>app: optimistic entry retired —
local and confirmed now agree else server rejects — failed guard, validation, limit server-->>client: doc:error — tied to the client event id client-->>app: optimistic entry discarded —
reads fall back to confirmed state end ``` If the server rejects the change — a failed [guard](/concepts/changes-as-json-patch/), a validation error — the client discards the optimistic entry and the local view falls back to the confirmed server state. The UI sees a clean signal, not corruption. Every rejection carries a category, such as `preconditionFailed`, `schemaValidation` or `timeout`, and `DOCUMENT_ERROR_CATEGORY_TRAITS` says how to handle each one: whether trying again later can succeed (`retryable`), whether the write may have committed after all, so the document should be read again first (`outcomeUnknown`), and whether it points at a bug in the app rather than anything the user did (`appFault`). An app branches on the traits and keys only its wording by category, so a category added later needs no change to its handling. ## Reading `getDocument` returns the merged view: confirmed server state with any pending optimistic changes overlaid. Subscriptions deliver an initial full snapshot (`doc:init`) followed by incremental patches (`doc:patch`) — see the [wire protocol](/architecture/wire-protocol/). The initial snapshot doesn’t have to travel over the socket: a server-rendered page can fetch documents over plain HTTP and hand them to the client as preloaded state. The subscription then carries the preloaded sequence, and the server confirms with no data transfer — unless changes arrived between the HTTP fetch and the subscribe, in which case it sends a fresh snapshot. Either way, embedded Yjs documents catch up with exact [state-vector diffs](/concepts/two-sync-lanes/), not re-sent states. Some of this state lives in **system documents** — subscribable documents the client, server, or session maintains rather than user data, such as the client’s own sync bookkeeping (subscription states, pending optimistic counts, errors) that the devtools panel renders. They’re read like any other document; see [System documents](/concepts/system-documents/). ## What “optimistic” does not mean ### It hides latency and disconnects — but not durability Optimistic updates mask network latency, and they carry over a disconnection: writes made while the socket is down are held locally and replayed on reconnect — structured (JSON Patch) writes last-writer-wins, the Yjs lane by [CRDT self-heal](/concepts/two-sync-lanes/) (deltas merge cleanly no matter how late they arrive). The boundary is durability, not connectivity: by default that replay buffer lives only in memory, so a page reload drops anything still unsent. Opt-in [offline persistence](/concepts/offline-persistence/) is what extends the buffer to disk — a journaled write survives the reload and replays on the next boot. ### Replay is de-duplicated, not silently lossy Re-sending a buffered write is safe to repeat: the server dedups by client event id, so a write that committed but whose acknowledgement was lost can be re-sent without applying twice. Dedup is time-bounded, though — a write left unconfirmed too long is dropped and reported as unconfirmed (an awaited write rejects, a fire-and-forget one hits `onWriteError`) rather than risking a silent double-apply. The exact rules — the replay horizon, sequence-guard exemptions, and the sweep that recovers a write left unanswered on a live connection — are in [Reconnect & replay](/architecture/reconnect-and-replay/). Gaps & open questions Offline writes replay on reconnect in both lanes, and with opt-in [offline persistence](/concepts/offline-persistence/) the queue is durable — unsent writes survive a page reload. Unguarded JSON writes still reconcile last-writer-wins. See [Limitations](/known-issues/limitations/). # Changes as JSON Patch > Structured changes travel as RFC 6902 patches, with guards for conflict detection. Structured changes to a document’s JSON data travel as **JSON Patch** ([RFC 6902](https://datatracker.ietf.org/doc/html/rfc6902)) — ordered lists of `add`, `remove`, `replace`, `move`, `copy`, and `test` operations. Patches are compact, human- and agent-readable, and double as the document’s audit trail in the [event log](/architecture/storage-and-event-log/). Heads up We’ve hit real problems applying RFC 6902 in a collaborative setting. Some are fully worked around — array addressing, defused by storing collections as ordered records instead of arrays. The sharpest one is still open: `add` has no precondition, so guards can’t catch two writers creating the same new key, and the later write silently wins. These and more are written up in [JSON Patch RFC issues](/known-issues/json-patch-rfc/). ## Guards A patch that was computed against one version of a document may not be safe to apply to a newer version. Writers choose how strict to be, per update: * **Sequence guard** — “apply only if the document is still at sequence N.” The strictest mode: any concurrent change rejects the write. * **Patch guard** — the patch carries RFC 6902 `test` operations asserting the values it’s about to change. Concurrent changes to *other* parts of the document are tolerated; a failed `test` rejects the write. When generating a patch by diffing, `test` operations are emitted automatically for `replace` and `remove` operations, so the patch defends itself by construction. Additions are the gap — an `add` has no prior value to assert, so guarded writes catch same-path *changes and deletions* but not two writers creating the same new key (the [`add` precondition issue](/known-issues/json-patch-rfc/)). A rejected guard is **expected, not an error** — nothing to log and page on — but the write itself is genuinely **dropped**, not applied. The optimistic entry is rolled back, the client snaps to the winning state, and it’s on the caller to re-derive against that state and retry; nothing replays the lost write automatically. The [staged session](/staged-sessions/conflict-preview-and-resolution/) builds its conflict detection on the same primitive, turning that drop into a reviewable conflict instead. Even an **unguarded** write can miss this way: if a concurrent write removed the very path the patch targets (a `replace` whose entry is gone, an array index past the end), the patch no longer resolves against the stored document. The server applies patches with strict RFC 6902 validation and rejects that drift with the same benign precondition-failed signal as a guard miss — rolled back, document healthy, caller rebases — rather than surfacing an opaque engine error. ## Moving a value An update is written as the difference it made, so a value you take from one place to another is sent as a `remove` and an `add` that carries your copy of it. If someone else edited the value a moment earlier, your copy overwrites their edit. To move a value and keep what others changed in it, say that it is a move. The update callback’s third argument does that: ```ts session.updateDocument({ docId, type: "board", callback: (doc, _yjsDocs, ops) => { ops.move("/todo/t1", "/done/t1"); doc.data.done.t1.title = "Shipped"; }, }); ``` `move` takes two [JSON Pointers](https://datatracker.ietf.org/doc/html/rfc6901) into `doc.data`. The draft changes at once, so the callback can go on editing the value at its new place. The write is sent, stored and broadcast as a `move` followed by the edits, and the server moves the value it holds. * **Move first, edit after.** A value you changed before moving it is sent as the plain difference. * **Object members only.** `move` takes a member of one object to a member of another, replacing what is there. It does not move array elements: that shifts the positions behind them, which is the problem the next section is about. * **Your own unconfirmed edits come first.** While an earlier edit of yours to the same value is still on its way to the server, the move is sent as the plain difference, which carries that edit with it. * **A move stays a move.** Made offline, or sent again after a dropped connection, it still goes out as a `move`. There are two exceptions, and in both the move is sent as the plain difference. One is when a later edit of yours to the same value reached the server ahead of it, so that edit is not overwritten. The other is when the page reloaded, or the tab closed, before the move was sent. * **A patch guard still guards.** With `guard: "patch"` the move carries a `test` of the value it moves, so a concurrent edit inside it rejects the write instead of being carried along. * **On the server too.** The callback of the server’s own `updateDocument` takes the same third argument, and its move is stored and broadcast as a `move`. * **Staged sessions** stage the plain difference. `ops.move` moves the value in the draft there and nothing more. Most of the time you do not need it. If a card’s column is a field of the card and its position a fractional index, moving the card is one small `replace`, and an edit to its title touches a different path. Reach for `move` when where a value sits in the document is part of what it means. ## Ordered collections without arrays The array-addressing problem shapes how you model collections. A collection inside a document is stored as **records keyed by id**, each carrying a fractional index for ordering, rather than as a JSON array. This is deliberate: JSON Patch addresses array elements by position, which is unstable when several writers insert and remove concurrently. Fractional indices make “insert between A and B” a single-key write that doesn’t disturb neighbors. (This is one of the workarounds discussed in [JSON Patch RFC issues](/known-issues/json-patch-rfc/).) The order of an object’s keys is not part of the document, so don’t sort by it. A document is compared as a value everywhere, and nothing keeps its keys in the order they were added: the server and each client apply writes in their own order, and a reload hands a client the server’s copy. In a chat where two people post at once, one screen can list the messages in the order it saw them while the server’s copy has them the other way round. To show records in order, sort them by their fractional index, or by a timestamp or another field of their own. ## Why patches — not CRDT operations, not domain events? For structured data, datadata deliberately uses server-ordered patches with guards rather than CRDTs — and rather than an application-defined event vocabulary with reducers: a patch is self-applying data, so history can be read and replayed in any context without the application’s code. The trade-offs of both choices are discussed in [Design decisions](/comparison/design-decisions/). For collaborative text, where convergence-per-keystroke matters, it uses [Yjs](/concepts/rich-text-with-yjs/) instead. Patches carry a real cost here worth naming: a diffed edit to a text field is a whole-value `replace`, so each change ships and stores the **entire** string — the [event log](/architecture/storage-and-event-log/) and the wire bloat fast for large, frequently edited prose. Yjs sends compact per-edit deltas instead, so for that shape of data it’s the better field type. The two travel together in the same update. # Domain commands > Named operations with server-side mutators — an update by name instead of by patch, predicted locally and decided on the server. A [JSON Patch](/concepts/changes-as-json-patch/) says “put this value here”. A **domain command** says “perform this operation against the state you find”: `postMessage`, `editMessage`, `increment`. The app registers a **mutator** per command — a pure function from the document, the command’s arguments and a server-stamped context to the document’s new data, or a refusal — and gives the same registry to the client and the server. The client runs the mutator over its local copy for an optimistic prediction; the server runs it over the document it holds, and that run is the one that counts. The mutator gets its own copy of the document either way, so it may edit the data in place and return it, as an update callback does. ```ts import { defineCommands } from "@repo/datadata/commands"; export const COMMANDS = defineCommands(SCHEMAS, { chat: (command, schema) => { // The message's own shapes, minus what the server stamps. const { authorId, postedAt, ...authored } = schema.root.fields.messages.values.fields; return { postMessage: command({ args: { type: "object", fields: { id: { type: "string" }, ...authored } }, contract: "relative", run: (doc, args, { principal, now }) => { if (args.id in doc.data.messages) { return { refuse: { code: "duplicateId", message: `${args.id} is taken` } }; } const message = { ...args, authorId: principal.subject, postedAt: now }; return { data: { ...doc.data, messages: { ...doc.data.messages, [args.id]: message } } }; }, }), }; }, }); client.live.command({ docId, type: "chat", name: "postMessage", args: { id, body } }); await client.live.commandAndWait({ docId, type: "chat", name: "editMessage", args: { id, body } }); ``` The registry goes to `createClient({ commands })` and `createServer({ commands })` (the Durable Object and the Node server take it the same way). The client’s sessions and the server’s `commandDocument` are typed by it: `type` names the document’s type, `name` must be a command defined for it and `args` must fit that command’s shape, all checked at compile time, and a `type` that is not the document’s own is refused. It is keyed by document type first: a command belongs to the type it is defined under, so two types may each have a `rename` with arguments of their own. Each callback also receives the type’s schema, so an arguments shape is picked from the document’s own field shapes, constraints included, rather than written again; the type and the validation of the arguments then come from the one definition. Running a command on a type that does not define it, or with arguments that fail their shape, is refused before any mutator runs. ## Why not a patch The mutator reads facts the client cannot be trusted with — the server’s clock, the connection’s principal — and encodes rules a patch cannot carry: only the author may edit their message, a counter increments whatever it finds. Two `increment`s from two users make `+2`, because each runs against the count it meets on the server; two patches to the same path would make `+1`. And the server validates the mutator’s result against the document type’s schema like any other write, so a command cannot produce a document a patch could not. ## Contracts and guards A command declares what it promises under concurrency, and the client derives the guard from it: * **`relative`** — the operation applies to whatever the server holds (`increment`). No guard. * **`overwrite`** — the later command wins, as two patches on one path would (`editMessage`). No guard; a mutator that wants more compares a revision it is handed in the arguments with the one it finds, and refuses. * **`exact`** — compare-and-set over the whole document the caller saw. The command carries a `"sequence"` guard and is refused `preconditionFailed` if the document moved. ## Prediction The client applies the mutator’s result to its optimistic view at once, like an update. The difference is that the prediction is **re-executed**: each render runs the pending command again over the document as that render meets it, so a command queued behind another pending write renders its effect over that write, as the server will run it. An `exact` command is the exception: it renders its effect only while the document is the one the caller saw, and nothing once another write lands under it, since the server will refuse it. A prediction the client cannot make — no principal is known yet, the mutator refuses on this copy — renders nothing, and the server’s answer decides. Writes on one document are sent one command-bounded step at a time: a command goes once every earlier write on the document is acknowledged, and an update staged behind a pending command waits for the command’s answer, because the command’s effect over those writes is decided on the server. With [offline persistence](/concepts/offline-persistence/) a pending command is journaled like any other write, in its place among the document’s updates, so a reload replays it under its original id ahead of the updates staged behind it, and renders it re-executed over the document as before. No `sys:` document takes a command: one sent at a system document is refused `unknownCommand`. ## Refusals A mutator refuses with a **code** of its own — `"notAuthor"`, `"duplicateId"` — and the server answers a `doc:error` of category `commandRefused` carrying it. The originator rolls its prediction back, `commandAndWait` rejects with a `CommandRefusedError` exposing the code, and the fire-and-forget form reports it through `onWriteError`. Three codes are the library’s: `unknownCommand`, `invalidArgs` and `mutatorError` (the mutator threw). A refusal is a benign outcome, like a guard miss: the document is intact, and the caller decides what to do with the code. ## Authorization Each command is its own [access kind](/concepts/authorization/): a schema’s `access` block grants it under `commands`, by name, with the same rule vocabulary as the write kinds. ```ts access: { update: { role: "admin" }, commands: { postMessage: { role: "editor" }, editMessage: { role: "editor" } }, }, ``` A command’s default is **nobody** — it reads only its own rule, never the `update` rule or the `write` fallback — so registering a command grants nothing until the schema names it. That is what makes commands a narrower lane than raw updates: the chat above lets editors post and edit through the mutators, while a raw `doc:update` to the chat needs an admin. The client’s `can()` answers for a command kind like any other, so a UI can hide what the principal cannot run. A command needs read access like an update does, and the server decides authorization before it consults the registry, so a refusal never reveals which commands exist. ## Server agents Code running beside the server, such as an AI agent in a Durable Object, runs a command with `server.commandDocument(context, { docId, type, name, args })`, the direct counterpart of `server.updateDocument`. It runs as the context’s principal and meets the same gates as a client’s command, so an agent can be granted the few commands it needs instead of `update` on the whole document. The mutator’s context carries that principal and the server’s clock. A refusal throws a `RefusedCommandError` carrying the mutator’s code, the library’s refusals throw the same error with their own codes, an authorization failure throws too, and nothing is stored. The call runs against the document as the server holds it, with no guard, whatever the command’s contract. # Rich text with Yjs > The hybrid model — JSON Patch for structure, Yjs CRDTs for collaborative text. Character-level collaborative editing is the one place where last-write-wins or guarded patches are clearly the wrong tool: two people typing in the same paragraph should both win. For that, datadata embeds [Yjs](https://yjs.dev/) — a mature CRDT implementation — rather than reinventing text merging. ## The hybrid model A document’s structured data is JSON, changed by [guarded patches](/concepts/changes-as-json-patch/). Where a field needs collaborative text, it holds a **reference** to an embedded Yjs document: ```json { "title": "Roadmap", "body": "yjs_a1b2c3" } ``` The Yjs document is created alongside its parent and travels with it: subscribing to the parent delivers the Yjs state too, and a single [wire protocol](/architecture/wire-protocol/) update can carry a JSON patch and Yjs deltas together. But the two are versioned independently — the document’s sequence tracks the JSON snapshot only, and Yjs deltas advance their own per-document cursor. That split, and what it fixes, has [its own page](/concepts/two-sync-lanes/). ## Working with Y.Docs Sessions work with real `Y.Doc` instances, handed out by the lease that keeps the document subscribed: `session.prepareDocument(docId).getYDoc(yjsId)` returns the live one — a single shared instance per target, so every editor bound to it sees the others’ keystrokes synchronously. Local edits are captured and auto-synced (debounced), remote deltas apply as they arrive, and the whole Yjs editor ecosystem (ProseMirror/Tiptap bindings, etc.) takes the instance as-is. For cursors, the same pair gets a real y-protocols `Awareness` — bridged over [presence](/concepts/presence/), with no second wire protocol. A `Y.Doc` lives as long as its document is subscribed. Releasing the last lease destroys it, and a later prepare hands out a **new** instance for the same id, so keep the handle for as long as an editor is bound to its doc and take the doc again after re-preparing. A released handle throws rather than hand out a doc that may be gone, and an edit to a destroyed doc logs a warning, since it never syncs. In React, prepare and read the doc at commit time (an effect or a `useSyncExternalStore` subscription that releases on cleanup), never during render: StrictMode’s mount, cleanup, mount would otherwise leave the component bound to the destroyed instance. `useSessionYDoc(session, { docId, yjsId })` from `@repo/datadata-react-components` does exactly that, and also rebinds when a staging session’s copy is dropped. The lifecycle follows the reference. A Y.Doc is authored in the **same write** that adds its reference field (the create/update callback’s accessor), so reference and content are never transiently inconsistent — and a reference whose Y.Doc was never written simply initializes empty on first use. Removing the referencing field deletes the orphaned Y.Doc from storage (a tombstone in the [Yjs lane’s log](/architecture/storage-and-event-log/)). Reintroducing the same reference id later does **not** recover the deleted content — the tombstone stands; the id merely becomes reachable again and initializes empty on first use, exactly like a never-written reference. Loading granularity is the document: subscribing delivers **all** of a document’s embedded Y.Docs — full states on first contact, exact [state-vector diffs](/concepts/two-sync-lanes/) on every reconnect after — so a document with many large texts should be modeled as more documents. The same model holds inside a staging session: [the prepared document’s `getYDoc`](/staged-sessions/staged-rich-text/) hands an editor a staged Y.Doc to bind to, so rich text can be drafted and reviewed before it commits. These Y.Docs are all **durable** — stored, versioned, replayable. The opposite case — transient text needed only while it is produced, like an LLM response streamed token by token — instead rides an ephemeral Y.Doc owned by a [presence](/concepts/presence/) cell, never stored and reaped when the cell goes; see [Presence](/concepts/presence/) for that lane. ## Two merge semantics, on purpose | | Structured data | Rich text | | ---------------- | --------------------------------------------------------------------------------------------------------- | --------------------------- | | Format | JSON + RFC 6902 patches | Yjs binary updates | | Concurrency | Server-ordered, guards detect conflicts | CRDT, always converges | | Conflict surface | Explicit — rejected guards, [staged conflict previews](/staged-sessions/conflict-preview-and-resolution/) | None — merging is automatic | This split is a deliberate [design decision](/comparison/design-decisions/): structure benefits from explicit, reviewable conflicts; prose benefits from silent convergence. The boundary extends to validation: the [schema’s](/concepts/schemas-as-documents/) jurisdiction deliberately ends where Yjs begins. Validation governs the JSON structure (including that a field holds a Yjs reference); constraints on the collaborative text *content* — length, structure, formatting rules — are intentionally left to the application and its editor. Gaps & open questions Yjs update logs are append-only today, with no compaction, and Yjs garbage collection is not supported yet. See [Limitations](/known-issues/limitations/). The cursor and ephemeral-streaming machinery rides [presence](/concepts/presence/). # Presence > Ephemeral per-document participant state — who's here, where their cursor is, what they're streaming — as an ordinary document. Some state belongs to a moment, not to history: who has a document open, where their cursor sits, the display name to show beside it, a half-generated LLM response still arriving. None of it should be versioned, replayed, or kept once the participant leaves. datadata models all of it as **presence** — and, true to [the rest of the system](/introduction/documents-all-the-way-down/), presence is just another document. State a client could compute from an agreed time — an enemy on a fixed route, a countdown — needs no presence at all: see the [shared clock](/concepts/shared-clock/). ## An ephemeral document of cells Presence lives in **channels**: a presence document `sys:presence::` is a flat record set of **cells** keyed by `presenceId`, one per participant, each `{ subject, state }`. The `` names the channel’s schema, and `` is just a document id — the app gives it meaning (typically “the document this channel is about”); the server never resolves it, only caps it at 256 characters. That independence is deliberate: a channel works on a document that doesn’t exist (yet) server-side — say, a staged create two sessions are collaborating on — and one channel type (a generic editor’s `cursors`, say) serves any document. Which channels an app opens on which documents is convention in app code, exactly like deterministic docIds; a common convention is one channel per docType. The `state` is an application-defined JSON object — cursor, selection, display name, whatever the app puts there — [schema-validated](/concepts/schemas-as-documents/) against the channel’s **required** schema at `sys:schema:presence:`. A channel whose schema isn’t declared doesn’t exist: subscribing or publishing to it answers an explicit `unknownType` error — a loud configuration signal, never a silent “not found”. The presence schema is an ordinary schema document with one twist: because presence is never stored, it [carries no migration log](/concepts/schema-evolution/#presence-schemas-dont-migrate) — its shape changes freely and each cell just re-validates on its next write. The server stamps `subject` from the authenticated connection, so a cell can’t lie about who owns it. Presence needs almost no protocol of its own. The presence document rides the same [wire](/architecture/wire-protocol/) as any other — the ordinary subscribe/init/patch/update path, plus a single presence-only event (`doc:resync`, below) — so presence reuses the engine’s existing sync, fan-out, and optimistic-update machinery rather than bolting on a parallel channel. What makes it presence rather than a document is its **lifetime**: presence is never stored durably. The server keeps only an in-memory aggregate — each client is the source of truth for its own cells, bounded by the connection that published them and cleared on disconnect or explicit removal. Nothing is logged or replayable. So when a host restarts or wakes from hibernation it loses that aggregate outright, and rebuilds it by sending each still-subscribed presence document a **`doc:resync`** — the one event unique to presence — asking every subscriber to republish its cells. The nudge deliberately isn’t an empty `doc:init`, which would clear the peers a client is already showing; instead each client marks its peers [stale](#liveness-and-the-grace-window), republishes its own cell, and the roster refreshes as everyone’s republishes arrive, without flickering empty. ## Read and written per-session A client never touches the raw aggregate. It works through a synthetic **view**, `sys:presence-view::`, that projects the cells into `{ self, peers }` and is read and written **per-session** with the same `getDocument` / `updateDocument` / `onDocumentChange` a session uses for any document. A session writes only its own `self` cell; the other participants arrive as `peers`. The view is computed on read — it is never itself persisted. Because presence is per-session, two sessions on the same client (the [ambient `client.live`](/staged-sessions/overview/) and a scoped one, say) each own a distinct cell, and a [staging session](/staged-sessions/overview/) publishes its presence live even while its document edits stay staged. Each session mints its own `presenceId`. Pass `presenceId` to `createLiveSession` to choose it instead — for deterministic ids in tests, or to correlate with another system. The id is the cell’s key in the channel document, so it follows [the same rules as a document id](/concepts/documents-and-folders/#documents). Creating a session with any other id throws, and the server rejects a cell written under one. ### Liveness and the grace window When a peer’s connection closes, the server removes that peer’s cells at once: the departure reaches everyone else as an ordinary `doc:patch` removing the cell, and the peer leaves the view. There is no disconnect event and no timeout on the server. The **`stale`** flag covers the other case: when **your own** view of the channel can no longer be trusted. That happens when your client reconnects, or when the server wakes from hibernation and sends `doc:resync`. Your client can’t tell yet which peers are still there, so it keeps showing them, marked `stale`, instead of blanking the roster. Each peer’s republish (or the fresh `doc:init` after a reconnect) confirms that peer and clears its flag. A peer nobody confirms within the grace window, `presenceGraceMs` on the client (10 seconds by default), is hidden from the view. The UI can dim stale collaborators while the roster settles instead of flickering empty. A client that changes its own cells while offline sends nothing: the reconnect republishes each cell’s latest state, so peers see only where it ended up. ``` stateDiagram-v2 accTitle: A peer's cell as one client sees it accDescr { A peer's cell is live once published. If the peer's connection closes, the server removes the cell and it leaves the view. If this client reconnects, or the server wakes and sends doc:resync, the cell is marked stale. The peer's republish, or a fresh snapshot, makes it live again. If the grace window elapses first, the cell is hidden. } [*] --> Live: peer publishes its cell Live --> [*]: peer's connection closes
server sends a cell remove Live --> Stale: this client reconnects,
or the server sends doc:resync Stale --> Live: peer republishes,
or a fresh doc:init lists it Stale --> [*]: grace window elapses
presenceGraceMs, 10 seconds by default ``` ### Bounded by construction Presence is a hot, fan-out-heavy channel, so it is capped: each cell’s JSON `state` is bounded (`max_presence_state_bytes`, 4 KiB), keeping the per-keystroke traffic small. Larger transient payloads don’t go in the cell at all — they use the ephemeral Yjs lane below. ## Two reserved lanes Most of a cell’s state is opaque app data, but two reserved fields turn presence into transport for richer collaboration. ### Awareness — collaborative cursors The `sys:awareness` field carries [y-protocols](https://github.com/yjs/y-protocols) awareness states, keyed by `yjsId`. A bridge publishes each Y.Doc’s local awareness there (merged with the app’s own fields) and injects peers’ entries into a real `Awareness` instance, so editor bindings like y-prosemirror’s cursor plugin plug in directly — no second wire protocol. The numeric awareness client ids editors see are allocated **client-locally** per `presenceId` and never cross the wire. This is how [rich-text editors](/concepts/rich-text-with-yjs/) get live carets, including [staged rich text](/staged-sessions/staged-rich-text/) whose cursors ride each session’s own cell. You never declare `sys:awareness` in your presence schema — datadata composes it in (an optional map of opaque cursor state) wherever your schemas are registered, so the whole cell validates against one schema and the field shows up in schema introspection exactly as the server enforces it. ### Ephemeral Y.Docs — streaming transient text A presence schema may declare a Yjs reference field, letting a cell own an **ephemeral Y.Doc** that lives only in server memory — never stored, never logged, reaped when the cell goes. The motivating case is an LLM response streamed token by token: peers watch it arrive live, but the half-generated text is worthless once the turn ends. Folding each token as a Yjs delta is a few bytes, where re-sending the whole growing string as a JSON patch every token is O(n²) on the wire and would blow the cell’s size cap. The lane is single-writer (a cell may only stream into Y.Docs it references), bounded by `max_presence_yjs_update_bytes` (1 MiB) per delta, and survives the author’s reconnect — on a hibernation wake the owner re-sends the Y.Doc’s full state. When the stream finishes, the author writes the final text to a **durable** document and clears the cell. See [Rich text with Yjs](/concepts/rich-text-with-yjs/) for the durable counterpart. # Shared clock > An opt-in estimate of the server's clock, so clients agree on time and compute what they would otherwise have to send. Some state needs no syncing at all once everyone agrees on the time. An enemy walking a fixed route at a fixed speed is in the same place on every screen if every client knows when it set off; a countdown ends together if every client knows when it ends. Sending the position ten times a second over [presence](/concepts/presence/) works, but it is bandwidth spent on a fact each client could compute. What those clients lack is a clock they agree on. The **connection clock** is that clock. It is off until the app asks for it: ```ts client.clock.start(); // Later, on any client of the same server: await client.clock.ready(); // synced: now() reads server time const waveStartsAt = client.clock.now() + 3_000; // write this into a document // …and each frame, on every client: const elapsed = client.clock.now() - game.waveStartsAt; ``` `client.clock.now()` is server time, in milliseconds since the Unix epoch. Write it into documents, compare it across clients, compute from it on your own frames. Until the first answer it is the local clock, so wait on `client.clock.ready()` before writing a time other clients will compute from; it resolves once the clock has synced, at once if it already has. ## How it estimates A started clock sends `clock:ping` over the connection and the server answers `clock:pong` with its own time ([wire protocol](/architecture/wire-protocol/)). From one round trip the client knows the server read its clock somewhere between sending and receiving, and takes the midpoint, so a single sample is off by at most half its round trip. It takes five samples back to back on every connect, then one a minute, and keeps the offset from the **fastest** round trip among its recent samples (NTP’s minimum-delay filter): the shortest trip leaves the least room for the two legs to differ. `client.clock.status` reports `synced`, the `offsetMs` and the `rttMs` of the sample it came from; `client.clock.subscribe` hears every change. `now()` is the local monotonic clock plus that offset, and a better sample does not make it jump: it **slews** to the new offset, running up to 10% fast or slow until it has caught up, the way NTP and game clocks do. So `now()` never runs backwards, and a 20 ms correction is spread over 200 ms. Only the first sync, and a correction over 250 ms, step at once; `status.offsetMs` is always the estimate itself, which `now()` may still be catching up with. Offline `now()` keeps counting on the last offset. A reconnect samples afresh, since a new connection may take a different path, but the old estimate stays in the running. Each sample is off by at most half its round trip, so a slow first answer whose range overlaps the old estimate’s cannot replace it. One whose range does not proves the old estimate wrong (the local clock paused while the laptop slept, say), and wins at once. ## Opt-in, per connection A client that never calls `start()` sends no clock traffic, which matters on a [Durable Object](/architecture/cloudflare-deployment/): the heartbeat is answered at the edge without waking the object, but a clock ping is a real message the object handles. `start({ burstSamples, resyncIntervalMs, sampleTimeoutMs })` tunes the cadence; `resyncIntervalMs: null` samples on connect only. The clock belongs to the connection, not to presence or to a document: one clock serves every channel and every document the client touches. ## In one page In the [in-memory demos](/architecture/running-in-memory/) every client shares the page’s clock, so the offset is near zero by construction; a `LatencyLink` delays both directions alike, which the midpoint cancels. To see the clock correct something, give a client a skewed local clock: ```ts connectInMemoryClient(connection, schemas, { localClock: () => performance.timeOrigin + performance.now() + 2_000, }); ``` ## Gaps & open questions * **Agreeing on time is not agreeing on outcomes.** A computed enemy still dies where the authority says it does, and that news arrives one-way latency late. Draw slightly in the past, as games do: Contingency draws half the clock’s `rttMs` plus a tick behind `now()`, so a kill has arrived before the enemy walks past it. * **Precision is the network’s.** The estimate is as good as the most symmetric recent round trip; a path whose legs are consistently lopsided biases it by half the difference, and nothing on the client can see that. # Two sync lanes > Structure and rich text version independently — one sequence for the JSON snapshot, one for the Yjs lane. The [hybrid model](/concepts/rich-text-with-yjs/) puts two kinds of change in one document: guarded JSON patches for structure, Yjs deltas for rich text. They have different merge semantics on purpose — and it turns out they want different sync mechanics too. datadata syncs them as **two independent lanes**. ## One version each A document’s **`sequence`** versions the JSON snapshot only: it advances by one per confirmed JSON change, and it is what resume checks, sequence [guards](/concepts/changes-as-json-patch/), and the staged session’s [conflict gate](/staged-sessions/conflict-preview-and-resolution/) compare against. The Yjs lane has its own per-document version, **`yjs_sequence`**, and its own [stored log](/architecture/storage-and-event-log/). A Yjs-only update advances `yjs_sequence`, leaves the current `sequence` unchanged, appends nothing to the JSON event log, and never rewrites the snapshot. The two versions never read each other. `yjs_sequence` never leaves the server. Clients are sent no Yjs version at all, because they have nothing to do with one: the lane is a CRDT, so updates merge in any order. The single place ordering matters — deleting an embedded Y.Doc — is ordered by `sequence`, since a Y.Doc lives exactly as long as a `yjsRef` field names it, and removing that reference is a JSON change. ``` flowchart TB accTitle: One update, two independently versioned lanes accDescr { A single doc:update carries JSON patches and Yjs deltas together. The JSON half is validated against guards, advances the document sequence by one, and is written to the snapshot and the JSON event log. The Yjs half merges as a CRDT in any order, advances a separate yjs_sequence, and is written to the Yjs state and its own update log. The two versions never read each other, and yjs_sequence never leaves the server. Both halves leave as one doc:patch. } update["one doc:update
JSON patch and/or Yjs deltas"] subgraph jsonLane["JSON lane — order-sensitive"] direction TB j1["validate · check guards"] j2["sequence + 1"] j3[("snapshot + JSON event log")] j1 --> j2 --> j3 end subgraph yjsLane["Yjs lane — CRDT, merges in any order"] direction TB y1["merge delta"] y2["yjs_sequence + 1
never sent to clients"] y3[("Yjs state + update log")] y1 --> y2 --> y3 end out["one doc:patch
a Yjs-only update repeats sequence unchanged"] update --> j1 update --> y1 j3 --> out y3 --> out ``` Why bother? A single sequence number shared by both lanes makes the CRDT lane trip machinery built for the other one: * a sequence guard fails because someone *typed* — a concurrent change that, being CRDT, conflicts with nothing; * a reconnecting client whose JSON is current gets a full snapshot re-send because only rich text has moved; * a [session](/staged-sessions/overview/) watching for head movement surfaces conflicts for edits that merge unconditionally. All three are the same category error: describing an order-insensitive lane with the snapshot’s version number. Splitting the versions fixes them structurally instead of special-casing each one. ## What rides neither lane [Blobs](/concepts/blobs/) — files referenced from document JSON — are not a third sync lane. Their bytes are immutable, so there is nothing to sync: the handle travels over the JSON lane as an ordinary field, and the bytes move over plain HTTP, exactly once per blob, without a version, a log, or a replay of their own. ## State-vector catch-up The Yjs lane doesn’t catch up with snapshots. A subscribe carries a **state vector** per embedded Y.Doc — Yjs’s compact summary of “what I already hold” — and the server answers with **exact diffs**: precisely the missing content, nothing for a Y.Doc that’s current, a tombstone notice for one that’s been deleted. The diffs ride whichever answer the JSON lane earns — `doc:init` or `doc:resume`, see the [wire protocol](/architecture/wire-protocol/) — so a reconnect after an hour of pure typing is a tiny resume plus a diff, never a full re-send. A subscribe without vectors falls back to full Yjs states. ## Catch-up runs both ways The exchange is symmetric. The server’s answer also carries **its own** state vectors, and a client holding Yjs content the server lacks pushes exactly the missing diffs back as one ordinary update. Three situations produce that divergence: edits made while disconnected (CRDT deltas merge cleanly later, unlike guarded patches), content held across a [soft delete](/architecture/reconnect-and-replay/) and released by the restore that undid it, and the [write-behind window](/architecture/storage-and-event-log/) — the bounded period where a streamed Yjs burst has been broadcast but not yet durably stored. If the server crashes inside that window, the next reconnect of any client that saw the broadcasts restores the burst. Gaps & open questions The lane split keeps conflicts and re-inits scoped to real JSON changes, but both logs remain append-only with [no compaction](/known-issues/limitations/). Both lanes replay writes buffered across a disconnect — the JSON lane last-writer-wins, the Yjs lane by CRDT merge — but that buffer is in-memory only: it does not survive a page reload. See [Open questions](/known-issues/open-questions/). # Blobs > Files — images, PDFs, attachments — as immutable blobs referenced from document JSON, with lifecycle, authorization and integrity owned by the engine. Documents are JSON plus Yjs sub-documents, and both are capped at a size that rules out a photo. Files need a third kind of content: **blobs** — immutable byte objects held in an object store (R2 on Cloudflare, S3-compatible elsewhere, filesystem or memory for development) and referenced from document JSON by handle. The bytes never touch the sync protocol; the handle syncs like any other field. ## Why blobs are core, not an app feature The tempting version — an upload route in the app and the object key in a string field — cannot be made correct, because four of its obligations are only reachable from inside the engine: * **Lifecycle is transactional with writes.** Knowing when a blob becomes referenced or unreferenced means seeing every accepted patch at the write boundary, atomically with the document row. An app observing changes after the fact races concurrent writes: bytes leak forever, or a cleanup deletes a blob just as a write re-references it. This is exactly why Yjs sub-documents and their reference counting live in core; blobs are the same problem with different bytes. * **Authorization parity.** [Read filtering](/concepts/authorization/) promises that a hidden document is indistinguishable from a nonexistent one. An app-owned download route with its own checks is a permanent side door: a private document’s image gets a weaker, drifting access model than the document itself. * **Trash and purge semantics.** Delete → restore must keep the bytes; purge must destroy them. Only the engine sees those transitions. * **Schema integrity.** “A synced document never points at missing bytes” needs the write boundary: the schema-derived set of handles is verified against the catalog inside the write’s transaction. A plain string field lets every client sync a document whose image 404s. The answer to the engine-size objection is the shape of the integration, not exclusion: **core owns the semantics, the edges own the bytes**. Hosts own the transport (the Worker or Node process streams bytes; core never runs an HTTP server), adapters own the store, apps own presentation, and the whole lane is opt-in — a server without an object storage adapter pays nothing, and refuses a schema that declares a blob field. ## The one insight that makes this cheap **Blobs are immutable.** A “changed” image is a new blob and a JSON patch that swaps the handle. From that: * Blobs need **no realtime sync at all**. The engine syncs the handle over the ordinary [JSON lane](/concepts/two-sync-lanes/) — ordinary optimistic update, ordinary guards, ordinary conflict rules — and the bytes move over plain HTTP, out of band, exactly once per blob. * No blob sequence numbers, no blob event log, no blob replay. The [event log](/architecture/storage-and-event-log/) stays byte-free; history replay reproduces handle changes. The bytes behind a historical handle survive until the sweep reclaims them, but the read gate follows the *current* references: a reader can fetch them only while a document they can read still holds the handle, so a history view that shows old attachments needs a copy under system authority, not the by-id URL. * Downloads are cacheable forever in principle, since the bytes under a handle never change. How far a deployment cashes that in is the host’s choice, below. So the design is deliberately asymmetric with the Yjs lane: Yjs bytes are *mutable CRDT state* and ride the sync protocol; blob bytes are *immutable content* and never touch it. ## A `blobRef` field A [schema](/concepts/schemas-as-documents/) declares a blob field as its own value kind, mirroring `yjsRef`: the JSON value is a bare, server-minted, opaque id. ```json { "type": "blobRef", "accept": ["image/*", "application/pdf"], "maxBytes": 5000000 } ``` * **Membership is schema-derived.** The write boundary collects the document’s live blob ids from the schema, the same walk that finds its Yjs sub-documents. A write introducing a handle to a blob that doesn’t exist, was never finalized, or has been swept is rejected as a precondition failure — the same benign category as a reclaimed sub-document. * **The catalog owns the facts the platform needs** — size, content type, SHA-256, timestamps, attribution. Display metadata (alt text, dimensions, crop) is the app’s business, in sibling JSON fields. A structured value duplicating catalog facts into every snapshot was considered and rejected: one source of truth, and a `HEAD` on the download URL answers the rest. * **`accept` and `maxBytes`** constrain what the field takes: media-type patterns (exact, or a `type/*` wildcard) and an inclusive size ceiling under the server-wide limit. Both are checked against the catalog’s facts — the type declared at upload, the size recorded at finalize — never the bytes, so `accept` is a schema and UX constraint, not a security boundary. The [read side](#serving-uploads-safely) is what protects readers. A blob is also a fourth kind of [reference](/concepts/references-and-integrity/): a handle to bytes rather than a pointer into JSON, reference-counted like a Yjs reference, but with the target living outside the document store. ## Upload, then write Upload is two-phase HTTP followed by an ordinary document write: ``` sequenceDiagram accTitle: Upload a blob, then commit its handle accDescr { The client POSTs the bytes to the host's blob route with the intended write named in the query. The host asks the server to stage the upload, which runs the ordinary write gate on that intent and the field constraints, then streams the body into the object store and asks the server to finalize the catalog row. The client receives a handle and commits it in a normal document write; the write boundary verifies the handle and records the reference edge in the same transaction, flipping the blob from staged to live. } participant C as client participant H as host (Worker / Node) participant S as server participant O as object store C->>H: POST …/blob?kind=&docId=&docType= (bytes) H->>S: stage upload — authorized as that write S-->>H: blobId + key H->>O: stream body to key H->>S: finalize (size, sha256) H-->>C: handle C->>S: ordinary doc write, data.image = blobId Note over S: verify handle + record edge
in the write's transaction ``` 1. **`POST` the bytes** to the folder’s blob route, naming the write the handle is intended for: `kind` (create or update), `docId`, `docType`, and optionally the field `path`. The server runs the **ordinary write gate on exactly that operation** — the docType’s rule, or the per-document `sys:access` entry where one exists — so a refusal happens before any byte moves, with the same verdict the committing write will get. The field constraints are checked here too, over every `blobRef` leaf the intent could land in. `Content-Length` is required: only a declared length lets the size limit run before bytes move and lets the store stream without buffering. 2. **The host streams the body** into the object store and finalizes the catalog row with the size and hash it observed. On Cloudflare the Worker holds the R2 binding and streams; the Durable Object sees only metadata calls, so no large body passes through its memory. 3. **The client writes the document** — a normal optimistic write with the handle in the field. The write boundary collects the document’s blob membership, the storage adapter re-verifies every handle inside the content write’s transaction and records the reference edges there, and the blob flips from *staged* to *live*. Ordering is enforced by the integrity gate rather than by convention: a handle can only be committed after finalize, so a synced document never points at bytes that aren’t there. The intent is a hint for authorization, not a binding — the handle may be committed wherever the write boundary allows. On the client this is `client.uploadBlob(bytes, { contentType, kind, docId, docType })` resolving to a handle, and `client.blobUrl(blobId)` for the download URL, both configured through `ClientConfig.blobs` — the folder’s blob route as a bare path. A refused upload maps back to the write vocabulary (`unauthorized`, `sizeLimitExceeded`, …), so apps handle it like a denied write. The React integration wraps the same calls as a `useBlobUpload` hook and a render-safe `blobUrl` helper. ## Download and the read gate `GET …/blob/:id` asks the server one question: may this principal read this blob? A blob is readable iff the principal can read **at least one document that references it** — the existing [read filtering](/concepts/authorization/) evaluated per referencing document — or, while it is still staged, iff the principal is its uploader, so previews work before commit. “Uploader” is matched by subject: an anonymous upload has no subject to match, so it becomes readable only once a readable document references it. Anything else is a 404: a hidden blob is indistinguishable from a nonexistent one, exactly as for documents. No new authorization vocabulary was added. The response carries the content type, a strong ETag (the SHA-256), honors `Range`, and a `Cache-Control` the **host** states rather than core defaults, because only the host knows how its read gate is wired: * **Revalidate** — the gate runs on every request (identity rides cookies, as it must for a URL that lands in ``). Success is `private, no-cache`: every reuse re-runs the read gate, and the strong ETag keeps that a byte-free 304. This is what both hosts ship today. * **Immutable** — the URL carries its own authorization (signed) or none (a secret or public URL; the id is a random server-minted capability). Success is `public, max-age, immutable`, cacheable by browsers and shared caches alike, with no round-trip to the server at all. Denials are never cacheable under either policy: a staged blob goes live, a trashed referrer is restored, and a cached 404 must not outlive the grant. Identity never rides the URL. `blobUrl` answers a bare by-id path and the client refuses an endpoint with a query string, because those URLs end up in `` tags, `target="_blank"` links, browser history and referrers. Cookies are the channel; short-lived signed URLs are the planned addition for CDN offload — a capability minted by a separate step, not identity on the by-id URL. ## Serving uploads safely Serving user uploads from the app’s origin, with cookie auth, is the classic stored-XSS setup: the content type is attacker-supplied at upload, and SVG or HTML uploads would otherwise run script as the site. The download route treats it as one, and the rules are pinned by core’s own tests because the HTTP handlers live in core and every host mounts the same ones: * `X-Content-Type-Options: nosniff` on every blob response. * An **inline-rendering allowlist** — raster images, video, audio, PDF. Everything else, `image/svg+xml` and anything `text/*` or `*+xml` included, is served as `Content-Disposition: attachment`, so the browser downloads it instead of rendering it in the origin. * `Content-Security-Policy: sandbox` on blob responses as defense in depth. The production-grade mitigation — a separate, cookie-less origin for user content — is what signed URLs will provide. The header rules are what make same-origin serving safe until then. ## Lifecycle and garbage collection ``` stateDiagram-v2 accTitle: Blob lifecycle accDescr { An upload creates a staged blob. A committed write referencing it makes it live. A live blob whose last reference is removed, or whose referencing documents are all purged, waits out the retention period and is then deleted. A staged blob never referenced is deleted once it is older than the retention period. Deleted means bytes destroyed and the catalog row gone. } [*] --> staged: upload + finalize staged --> live: referenced by a committed write staged --> deleted: never referenced, retention elapsed live --> deleted: last reference removed or referrers purged, retention elapsed deleted --> [*] ``` * **Staged** — finalized, not yet referenced. Swept once older than the host-chosen retention period, so abandoned uploads don’t leak. * **Live** — referenced by at least one non-purged document. Documents in the [trash](/concepts/documents-and-folders/) still count: delete → restore must round-trip, so bytes survive until *purge*. * **Deleted** — bytes destroyed, then the catalog row. The retention period between “unreferenced” and deletion is policy, not a correctness mechanism: it keeps sweeps unhurried and leaves a write that re-references the blob (or an export under system authority) a window. It is not a read window: the read gate follows the references, not the retention clock. A document losing a handle starts that blob’s retention clock whatever removed it — an ordinary write, a purge, or a [schema migration](/concepts/schema-evolution/) that drops the field. The sweep is where the correctness lives, and its ordering is the point. The naive order — find candidates, delete bytes — races a concurrent write that commits a fresh reference in between, leaving a synced document pointing at destroyed bytes. So the sweep **tombstones first, transactionally**: for each candidate, one catalog transaction re-verifies the sweep condition and marks the row deleted. Reference edges commit in the same transactional store, so a concurrent write either lands its edge first (the re-check sees it and skips the blob) or runs after (the write’s in-transaction verification now rejects the handle). Only then are the bytes deleted, then the tombstone. Object-store deletion isn’t transactional with the catalog, so those last two steps are at-least-once: a failure leaves a tombstone that retries on the next sweep. Retention stays with the host: `sweepBlobs` takes a `retentionMs`, the way `purgeDeletedDocuments` takes `deletedBefore`, so the server holds no retention configuration. One retention covers both kinds of candidate — a staged upload older than it is abandoned, a blob unreferenced for longer than it is reclaimable. Around a day is the sane policy, kept longer than any plausible upload. `DEFAULT_BLOB_SWEEP_POLICY` is that day with a batch of 100, and `assertBlobSweepPolicy` is the check the sweep itself applies: a retention from zero to ten years, a batch limit of one or more. A host runs the same check on its configured policy at startup. ### When the sweep runs The host owns the schedule, and the sweep tells it when the next one is due: its result’s `nextSweepAt` is already past when candidates remain (the batch limit was hit), the moment the oldest pending upload or unreferenced blob outlives the retention, or `null` when nothing is pending. Between sweeps, the server’s `onBlobSweepCandidate` hook fires whenever a candidate can appear — a blob staged, a write that replaces a document’s blob references, a [schema change](/concepts/schema-evolution/) releasing them on the next read or validation pass, a purge — so a host arms its schedule from one hook. * **On Cloudflare** the Durable Object’s alarm runs the sweep, and the alarm is armed by activity rather than by a clock: the sweep-candidate hook, or a wake with no alarm, arms it one retention period out, and after a run the sweep’s `nextSweepAt` re-arms it for when the next candidate comes due, or not at all. No sweep alarm is set less than a second out, so a burst of activity coalesces into one run. A folder with nothing pending never wakes for a sweep, which keeps [idle folders free](/architecture/cloudflare-deployment/). The policy is the `blobSweepPolicy()` override, read once when the Durable Object starts. * **On [Node](/architecture/node-deployment/)** a timer sweeps every folder in the database that has something to reclaim, which the catalog answers in one query — so a folder nobody has opened since the process started is swept too, and a blob unreferenced before a restart is reclaimed by the first tick after it. Folders then take turns, one batch each per round, and the catalog is asked again every round, so neither one folder’s backlog nor garbage that keeps coming due can hold up the rest. ## Where the pieces live * **Core** — the `blobRef` kind, the catalog (tables in the same storage adapter as the documents, so reference edges are transactional with writes), the server API, the dumb object storage adapter interface (put/get/head/delete by opaque key), and the shared HTTP handlers every host mounts. * **Object storage adapters** — R2 (a bucket binding, one bucket shared by many Durable Objects through per-folder key prefixes), S3-compatible (plain `fetch` and SigV4; AWS S3, R2’s S3 API, MinIO), filesystem (the Node harness’s dependency-free default), and memory for tests. * **Hosts** — the Cloudflare Worker mounts the handlers next to the WebSocket route and streams to R2; the Node harness mounts the same handlers under the same principal resolver as its WebSocket upgrade. * **Export and import** — a document export ships its blobs’ metadata and bytes alongside both [event streams](/architecture/portable-event-streams/), and an import re-uploads them. Because blobs are immutable, per-document restores compose: an import reuses an already-present blob whose facts match, and rejects a mismatch as a genuine id collision. * **Apps** — presentation. Thumbnails, resizing, pixel dimensions, alt text, which document types carry files. datadata never inspects image bytes; the playground’s image lane (dimensions committed next to the handle in the same write, resized variants derived on read behind the same read gate) is about a hundred lines of app code over two existing boundaries, which is the evidence that settled this as a decision rather than a default. The generic editor and viewer components render a blob as a link, never a preview: a preview is presentation, and a bare id has no content type. The conformance suite pins the whole contract — upload, integrity gate, reference-then-sweep, delete → restore → purge, the read gate, size limits, export round-trips, sweep-versus-concurrent-write, and the response headers — against all three storage runners, so memory, Durable Object + R2, and Postgres + filesystem or S3 behave identically. Gaps & open questions Blobs are new (September 2026). Uploads stream through the host, capped by its request body limit; presigned direct-to-store uploads and multipart are future work behind the same client contract. Blobs default to a 16 MiB ceiling so every accepted blob stays exportable inline; raising it trades away export of the larger ones until a streaming export exists. Signed URLs and a cookie-less serving origin are planned, not built. The backendless [in-memory client](/architecture/running-in-memory/) has no blob lane at all. See [Limitations](/known-issues/limitations/). # Offline persistence > The opt-in durable write queue and document cache — what survives a reload, how long each kind of write survives offline, and the caveats. The client takes an **opt-in persistence adapter**. With one configured, every pending write the client tracks — create, update, command, delete, rename, restore — is journaled write-through to local storage, and the next client on the same namespace (a reloaded page) rehydrates the queue and replays it. **A reload becomes a reconnect.** Without an adapter the client is memory-only: disconnected writes replay on reconnect, but a reload loses anything still unsent. ## A reload is a reconnect — the same replay rules There are no separate cold-start semantics. Rehydrated writes replay through the exact [reconnect contract](/architecture/reconnect-and-replay/): writes that never reached the server replay however long you were away; a write that was *sent* and sat unconfirmed past the replay horizon (\~30 minutes) surfaces as unconfirmed rather than risking a double-apply; sequence-guarded writes keep compare-and-set semantics at any age. The journal stores each write verbatim — the exact wire event under its original event id — so the server’s de-duplication cannot tell a replay-after-reload from a replay-after-reconnect. One write kind gets extra care: a **sent create** (creates aren’t de-duplicated). When the hydrated document cache already proves the document exists — say the reload raced the create’s already-exists rejection — the merge drops the stale create on the spot instead of rehydrating it, so it can never mask or clobber the newer cached state, offline included. Otherwise the re-subscribe’s authoritative answer reconciles it as usual. One thing deliberately does **not** survive a reload: **promises**. An awaited write can’t resolve on a page that no longer exists, so rehydrated writes are fire-and-forget, and their outcomes report through the client’s write-error callback and the pending counts on [`sys:client-docs-status`](/concepts/system-documents/). ## Discarded writes are observable The two paths that **discard** queued writes — the retention bound evicting a stale record at load, and the replay horizon dropping a sent-but-unconfirmed write — never do so silently. Each lands on `sys:client-docs-status` as a per-document `discardedWrites` entry: how many writes were dropped, why (`retentionExpired` or `replayHorizon`), when the oldest of them was **made** (`oldestStagedAt` — every queued write carries its authoring time, so the loss is datable even when nothing was ever sent), and the oldest first-send time when any of them ever reached the transport. An awaited write still gets its rejection, but a rehydrated fire-and-forget write has no promise left to reject — the status document is its witness, and the report is what lets an app tell the user “some offline changes from last week couldn’t be recovered” instead of saying nothing. Entries accumulate for the session and the document is reactive, so a badge or toast hangs off the ordinary subscription machinery. ## Pending writes are datable The pending counts on `sys:client-docs-status` say *that* unsent work exists; each document’s entry also says *since when*. `oldestStagedAt` is the authoring time of the document’s oldest pending write — stable across restamps, reloads (it persists with the journal) and adoption, and spanning the namespace like the count itself, so a sibling tab’s queued write dates it too. It is what lets a UI say “unsent changes from Tuesday” rather than merely “3 unsent changes”, and it clears with the count as the writes ack. Rich text dates too, its own way. Yjs edits don’t live in the write queue — they ride the document cache’s snapshot — so the client stamps the moment a document’s Yjs lane first goes **ahead** of the server (a dirty episode) and feeds that stamp into `oldestStagedAt`. The stamp persists on the cache record (`localYjsAheadSince`), so after a reload a document whose only unsent work is rich text still counts as pending, still carries its date, and the sync status truthfully reports **disconnected-saving** (“changes pending”) rather than offline-cached (“showing saved data”). The claim clears only when the lane verifiably settles — a flush’s acknowledgment, or the reconnect state-vector exchange confirming the server holds everything — and briefly over-claiming (a sibling tab may have pushed the same content already) errs toward caution by design. The cache is shared, so the marker merges across tabs: two claims keep the earliest stamp, and a settled tab’s write clears another tab’s claim only when its snapshot verifiably contains that tab’s content, so an offline tab’s unsent edits stay marked even while a sibling keeps editing online. ## Offline reads: the persisted document cache The same adapter caches **server truth** write-through: every answer the server gives about a document — initial load, resume, patches, and the negative outcomes, so a deleted or never-existing document renders truthfully. Each document’s Yjs state is re-encoded from the **live** documents at cache time, which is what carries local unsent edits across a reload: server state plus local ahead-content in one snapshot, self-healed both ways by the reconnect state-vector exchange — and stamped with the dirty-episode marker described above whenever the snapshot carries such unsent content, so the reload knows it is rendering more than saved data. On boot the cache hydrates before anything reaches the wire, so an offline reload **renders**: cached documents appear immediately, pending writes re-project on top, and the sync status reports **offline-cached** — “showing saved data” — until the server re-confirms each one. Online, the re-subscribe carries what the cache knows (each document’s sequence number, generation and Yjs state vector), so the server answers with a resume or deltas instead of re-shipping documents the client already has. A cached copy of a document that was purged while the client was away — its id since reused by a new document — names the old generation, so it is replaced with the new document rather than resumed, and the edits queued against it are rejected instead of applied. So is a restore queued for the old document, even when the new document has since been deleted too — the cached tombstone keeps the generation it restores. “Showing saved data” is answerable **per document** as well as app-wide. Each hydrated document’s entry on `sys:client-docs-status` carries `cacheHydrated` along with the `cachedAt` of the snapshot being rendered, so a badge can say *when* the data was saved and not merely that it was. The flag is removed the moment that document’s own authoritative answer lands, which means the badge clears **per document** on reconnect rather than all at once when the connection returns — a document whose answer is still in flight keeps saying “saved data”, because that is still what is on screen. An offline edit does not clear it either; the document is being edited, but what it renders is still the saved copy. Cached documents render, but apps still subscribe to everything they read, connection or not — the cache never bypasses that contract. The separate `subscribed` field answers the narrower question of whether this session’s subscription has been answered: it flips only on a **real** authoritative answer, never from the cache and never from an optimistic write. Housekeeping is automatic: the cache is shared across tabs (the newest cursors win; a tie merges the Yjs snapshots), writes coalesce per document, and at boot the oldest records are evicted past the caps (1000 documents / 64 MiB by default), with `sys:` documents pinned — schemas and the index are what offline validation and rendering hang off, so a warm offline boot resolves a document’s schema without a round-trip. A document holding Y.Doc edits not yet sent to the server is pinned too: the cached record is the only place those edits live until they are sent. Boot evictions are counted on `sys:client-docs-status` (`evictedCachedDocuments`) — benign next to discarded writes, since the data is refetchable once online, but on an offline boot an evicted document simply isn’t there, and the count is the app’s explanation. ## Knowing when you can edit Truthful state needs truthful **waits**, so there are two. A document is *settled* only once the server has answered for it, a wait that correctly stays pending through an offline boot. Editing needs less: the document only has to be *available* — readable and backed by a resolved schema, which a cache-hydrated document already is — and that wait resolves as soon as the document materializes, rejecting if it turns out not to exist. Offline-capable flows wait for available before editing, and for settled only where they need server-confirmed truth. ## What survives a long offline period Three kinds of write, three answers: | Write type | Long-offline behavior | | ------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Yjs (rich text)** | CRDT deltas merge at any age — no horizon. Offline content editing is the CRDT lane’s home game. | | **Sequence-guarded JSON** | Exactly-once at any age: if the server moved, the replay draws a benign rejection and the caller rebases. Right for discrete transactional ops, wrong for long editing sessions. | | **Unguarded JSON** | Replays last-writer-wins if never sent; surfaces as “unconfirmed, refetch” if sent and past the horizon. With a persisted queue this is the common case for long offline periods. | That last row is a genuine semantic caveat: a silent last-writer-wins replay of week-old writes is protocol-legal but can surprise the humans involved. For offline work that deserves review before it lands, use a [staged session](/staged-sessions/overview/) — its changesets live in the host document and sync through the server, so they are *already* durable, and a reconnect produces a **changeset to review** instead of silent writes. ## Staged sessions offline Staging and the journal compose with no extra wiring: a changeset is ordinary writes to its host document, so staging journals like any other write. Staged work survives a reload, renders from the cache while offline, and a fresh session over the same host resumes it. Commit awaits the server’s acknowledgment of every stage, so it is inherently an online act — starting one offline leaves it waiting for the reconnect. Commit is also **crash-safe**: a page that dies mid-commit leaves journaled writes that replay on the next boot, and a session resumed over the same host recognizes them as its own commit already in flight rather than re-applying the staged work. The [commit intent](/staged-sessions/changesets-and-lanes/#commit-intents) that makes this work is a staged-session mechanism, not an offline one. ## Several tabs, one store Tabs on the same folder and user share the journal and the cache, and all of them stay durable. Each journals its own writes, a dead tab’s unsent writes are adopted by a surviving one, and offline rich-text edits reach sibling tabs through the cache with no server in the loop. Pending JSON writes are visible in siblings but never merge, so a “changes pending” indicator is truthful across tabs — at the cost of two tabs editing the same field rendering their own value until a reconnect settles the order. [Offline across tabs](/architecture/offline-multi-tab/) is the contract. ## Scope it to the principal, and other caveats * **The namespace must be scoped to both folder and principal.** The journal carries document data: key it by folder alone and two accounts on one origin read each other’s pending writes. Clearing the store belongs in logout. * **Durability is best-effort.** Browsers evict IndexedDB under storage pressure; Safari prunes stale origins; private browsing is ephemeral. The guarantee is “survives reload in practice”, never “survives anything” — and an evicted store degrades to exactly the memory-only behavior, never corruption. * **A deletion wins over unsent Y.Doc edits.** A direct Y.Doc edit made to a document that has been deleted — seen by this tab, or by another tab while this one was offline — is kept in memory and sent if the document is restored while the tab is open. It does not survive a reload: the cached record is the deletion, and a tab’s copy cached before the deletion was observed never replaces it. Yjs content written through `updateDocument` is journaled with the write and is not affected. * **Updates need a subscription; lifecycle writes do not.** An update is tracked (and so journaled) only for a document this session subscribes to, because its acknowledgment arrives over that subscription. A create, delete, rename or restore is tracked whether or not the document is subscribed: the server answers the originator directly, so the write is queued, journaled and replayed like any other. What the missing subscription withholds is the local copy — the caller gets the outcome, not a readable document. * **The queue is capped.** At the cap (1000 writes by default), a further write is refused with a distinct queue-full error: an awaited write rejects, a fire-and-forget one reports through the write-error callback. Refusing new work beats silently dropping journaled writes the user already made. * **Apps can tighten, never loosen.** An app-level retention bound evicts stale queue records on boot below the protocol’s horizon; nothing an app configures can make a sent unguarded write replay past it. What the bound evicts reports through `discardedWrites` like every other discard. * **The cache outlives authorization, briefly.** Cached documents stay readable offline after an access revocation — the client already saw the data — until the next connection replaces them with a not-found. It is the second reason the namespace must be principal-scoped. # Conflicts > How datadata resolves conflicts — automatically during live editing, deferred for review in staged sessions. “Conflict handling” in datadata is not one mechanism but several — and the useful way to organize them is by **when** a conflict gets resolved. The server keeps one authoritative history underneath all of it. Live editing settles conflicts the moment a write lands, with no human in the loop. [Staged sessions](/staged-sessions/overview/) defer resolution: the few conflicts that genuinely need a decision are surfaced for review rather than merged behind your back. ## Live editing — resolved as it happens Outside a staged session, a conflict is settled the instant the write is processed. You never accumulate one to look at later. 1. **Server ordering.** Every accepted change gets the document’s next sequence number. There is no distributed disagreement about history — the folder’s server is the authority. 2. **Last-write-wins by default.** Two writers racing on the same field both land, in server sequence order; the later one overwrites the overlapping values. Nothing is dropped — the write with the higher sequence number simply wins. 3. **Guards, when you want detection.** A write can demand that the document (or just the values it touches) hasn’t changed since the writer last saw it. A failed [guard](/concepts/changes-as-json-patch/) rejects the *whole* write cleanly — conflict detection without locking. The rejected write is dropped, not replayed; the caller refetches and retries. ## Staged sessions — resolution deferred to review A [staged session](/staged-sessions/overview/) defers the decision, not the detection — and it reuses the same guard mechanism. Every staged patch embeds `test` guards asserting the values its author saw; folding the staged patches, guards included, onto the live head either merges cleanly (disjoint upstream changes need no attention) or blocks the stage. Conflicts are *derived* from head and the patches — no per-document snapshot is stored. Anything that actually collides is surfaced and held until you resolve it (amend, drop, or rebase onto head) before commit. Nothing that collides is auto-merged behind your back. ## Rich text is conflict-free in either mode Inside [Yjs fields](/concepts/rich-text-with-yjs/), concurrent edits always converge; there is nothing to detect or resolve — live or staged. This is structural, not just policy: Yjs deltas travel in [their own sync lane](/concepts/two-sync-lanes/) and never advance the JSON sequence number that guards and conflict gates compare against — so someone typing in a document can’t fail your guard or block your stage. ## What this is not datadata does **not** do operational transformation, and does not use CRDTs for structured data. As the live rules above spell out, it favors *explicit, reviewable* conflict handling — server-ordered last-write-wins with opt-in guards — over silently merging structured edits. Why we landed here, and what it costs, is covered in [Design decisions](/comparison/design-decisions/). Gaps & open questions Conflict *presentation* is still evolving: the three-way preview gives you the data, but patterns for resolving (field-level pick-and-choose vs. restage-from-head) are being worked out in the apps. See [Open questions](/known-issues/open-questions/). # References & integrity > Documents point at documents — and the schema knows what should happen when targets disappear. datadata’s model is referential: schemas mark fields as references — to a record in the same document, to another document, or to a collaborative sub-document. Declaring them in the schema is what lets the engine enforce integrity (and, where it can, repair it) instead of leaving dangling ids to the application. ## Declaring references A [schema document](/concepts/schemas-as-documents/) marks fields as references. There are four kinds, each declared as such in the schema: * **Entity references** point at a record *inside the same document* — for example, an edge in a diagram pointing at two nodes. * **Document references** point at *another document* of a given type, within the same [folder](/concepts/documents-and-folders/) — the reference can’t reach across a folder boundary. * **Yjs references** hold a handle to one of the document’s own collaborative sub-documents — a [Yjs](/concepts/two-sync-lanes/) Y.Doc with its own CRDT sync lane and lifecycle. * **Blob references** hold a handle to an immutable file — an image, a PDF — in the folder’s object store. See [Blobs](/concepts/blobs/). The first two are constraints over an ordinary id; a yjs reference and a blob reference are their own value kinds — handles to content held outside the JSON, not pointers into it. ### Rules for entity references An entity reference also declares what should happen when its target goes away. Because the referrer and the target live in the same document, these rules are enforced within that document, at each write: * **restrict** (the default) — the write is rejected while a referrer still points at the target, so you can’t delete a node while an edge needs it. * **cascade** — deleting the target also drops the referrers that pointed at it. * **unreferenced** — a target that loses its last referrer is reclaimed (reference-counted: the target lives only as long as something points at it). An entity reference can also bound how many referrers a target must carry (for example, “at least one”). These rules make integrity a property of the schema rather than something each write path re-implements — a raw update gets the same cleanup as any other. ## Checking and repairing How a reference is enforced depends on its kind — and the guarantees differ because the target lives in a different place each time: * **Entity references** are both checked and repaired at every write. Because referrer and target live in the same document, this is atomic: the declared rules run first (cascade drops dangling referrers, unreferenced reclaims targets nothing points at, iterated to a fixpoint), then restrict and cardinality violations reject the write. The repaired state is the same whether the client applied it optimistically or the server applied it authoritatively, so sync converges. * **Yjs references** are reference-counted. When a write leaves one of the document’s sub-documents with no referring field, that Y.Doc is deleted in the same write — no separate cleanup step. A reference to a sub-document that doesn’t exist yet is fine: it’s created lazily on first use. * **Blob references** are reference-counted too, with the opposite creation rule: the bytes must exist first. A write naming a blob that was never uploaded, never finalized, or already swept is rejected, and the reference edge is recorded in the same transaction as the write — so a synced document never points at missing bytes. A blob whose last reference is removed is reclaimed by a later sweep, not inline; the [lifecycle](/concepts/blobs/#lifecycle-and-garbage-collection) explains why. * **Document references** point at another document *within the same [folder](/concepts/documents-and-folders/)* — the existence check resolves against the folder’s own registry, so a document reference can’t reach outside it. They are checked at write time only: the target must exist when the write lands, but there is no cross-document repair — a target deleted afterward leaves a dangling pointer, which surfaces the *referrer* as an invalid document rather than erroring. There’s no transaction spanning the two documents. Every validated write also records the document’s outbound references in a folder-level index, so the invalid-document enumeration **derives** dangling-reference invalidity from that index against the live document set — deleting a target lists its referrers instantly, and restoring it heals the listing just as instantly, with nothing persisted into the referrers that could go stale. Clients subscribed to a referrer see the same flip: deleting, restoring or creating the target sends them the referrer’s current state, flagged or cleared. The flag lists every broken reference, not just the first, each located at the field that holds the id and naming the rule and the missing id, so a repair can fix them all in one write. This matters doubly for [AI agents](/ai-agents/why-agents-like-datadata/): an agent can validate the referential consistency of its staged changes before committing, and the rules it must respect are readable from the schema documents themselves. Gaps & open questions The cross-document guarantee is the weak one: document references are checked at write time but not enforced transactionally across documents, so concurrent writers can still race a reference into dangling between check and commit. The reference index makes such breaks enumerable the moment they happen, but the repair workflows over that listing are still being shaped by the apps. See [Open questions](/known-issues/open-questions/). # Authorization > Principals, scopes, declarative access rules, read filtering, and capability-aware clients — one identity model for users, system code, and AI agents. Every access in datadata runs as a **principal** — one identity model that covers a user’s frontend actions, server-side system code, and AI-generated edits alike. The host attaches a principal to each connection (and to in-process clients), and the engine gates it on both sides: it authorizes **every stored-document write** — create, update, delete, rename, restore, purge — and filters **every read**. A denied write is rejected like any other rejected write, so the writer’s optimistic update simply rolls back; a document a principal may not read simply does not exist for it (see [Read filtering](#read-filtering)). ## The principal A principal carries four things: * **Subject** — *who* the access is on behalf of: a user or service identity, or nobody for anonymous. * **Actor** — *what kind of code* is exercising it: a user’s frontend, a named AI agent (`agent:chat`), or server-side system code (`sys:schema-seed`). Actors are provenance and policy-targetable labels, carried so edits stay attributable — never gates of their own. * **Grants** — facts the app resolved at connect time, like a role or a membership. * **Scopes** — an optional hard cap the engine enforces no matter what any rule would otherwise allow. Subject and actor are independent: the subject is *whose behalf* the access is on, the actor only *what kind of code* is exercising it. A named AI agent therefore acts as its subject, not as a separate identity — which is what keeps its edits attributable to the person behind them. ## The composition contract Reads and writes run the same gate, and the whole policy is **declarative**: every verdict is computed from the governance documents and the principal, so both sides of the wire can evaluate the same rules. At each choke point — one for writes, a parallel one for reads — two checks compose, the second able only to *further restrict* the first: 1. **Scope attenuation.** If the principal carries scopes, the access must fall inside them — for **every** actor, before any rule is consulted. This is the hard cap for AI agents: an agent scoped to creating and updating notes can never delete anything or touch another document type, no matter what the rules say. Scopes attenuate **visibility too** — read is a scope kind alongside the write kinds, so that same agent sees nothing outside notes unless it is also scoped to read more broadly. (On the read path the system-document exemption is applied even before this cap, so a scoped agent keeps the index, schemas, and its own identity — see [Read filtering](#read-filtering).) 2. **Declarative rules**, most specific grain first: the folder’s [per-document entry](#per-document-entries) for the target id if it has a rule for the operation, else the document type’s access rules (below) — evaluated against the folder’s roles. The write-kind rules run on the write path, the read rule on the read path. Facts from external sources (a database, an API) enter as **connect-time grants** — the app resolves them at its trust boundary and the engine consumes them as rule conditions; nothing in the decision path does I/O. Presence is governed like any document: its presence schema’s `access` block decides who may read and write it (a presence write is an `update` on the `presence:` docType, resolved from the channel’s id), and per-document entries reach presence channels by their ids too. The channel is governed independently of the document it is about — that document’s read rules never gate presence — and the library adds one structural rule on top, because the schema can’t express it: per-cell ownership. ## System authority Some host operations must run **above** the folder’s own policy: seeding schemas at boot, creating the governance document itself, applying an operator’s schema migration. That privilege is a property of the **call**, never of any principal: a direct server call can carry an explicit system authority marker that skips the policy layers — the governance meta-gate, the access rules, and the per-document entries — while **scope attenuation still binds**, so a host can self-attenuate a system call. **Identity never implies authority.** No actor value grants anything — a system-actor client connection is an ordinary, policy-bound principal. Only host code holding the server handle can make a system-authority call, and connections have no way to express one, so a client cannot mint privilege by construction. A client that thinks it needs system authority almost certainly wants an admin role grant (which passes the governance meta-gate *through* policy) or scopes instead. The maintenance surfaces — export, import, verify, rebuild — sit outside the per-operation gates and **require** system authority. ## Declarative access rules A schema (itself [a document](/concepts/schemas-as-documents/)) declares access rules in both directions: a **write** rule per operation — a separate rule for create, update, delete, rename, restore, and purge, plus a fallback for any write kind without its own (purge excepted — see [Defaults](#defaults)) — and a **read** rule governing who may see the type at all. This section covers the write rules; the read rule, whose default is deliberately different, lives under [Read filtering](#read-filtering). A rule is one of: anyone, any authenticated subject, nobody, a **role**, a named **grant**, or a named **subject** (exactly that principal — the condition per-document ACLs are built from) — singly or as a list of alternatives, and the same vocabulary serves read and write alike, at both the type grain and the [per-document grain](#per-document-entries). Roles form an ascending hierarchy — viewer, editor, admin — and a principal’s **effective role** is the highest of what the folder’s `sys:access` roles document assigns it (per subject, or via folder-wide defaults for authenticated and anonymous subjects) and any role carried in its connect-time grants. A document type with no write rules defaults to *any authenticated subject*; reads default the other way (see [Defaults](#defaults)). A [domain command](/concepts/domain-commands/) is a write kind of its own, `command:`, granted under the block’s `commands` map by name. Unlike the other write kinds it never falls back to the `write` rule: a command no rule names is nobody’s, so an app can keep raw updates to a type behind the admin role while editors change it only through the mutators it registered. The three roles are **fixed** — apps cannot define their own or reorder the hierarchy. App-specific privileges beyond the trio ride named grants (matched exactly, with no hierarchy). Two **meta-rules** protect the governance surface itself: writing `sys:access` or any schema document requires the admin role regardless of the access rules — an editor cannot rewrite a schema to grant themselves more. ### Allow schema editing while blocking changes to sys:access Schema editing requires the admin role, which also permits writing `sys:access`. To let an agent edit schemas while blocking changes to folder roles and per-document permissions, attach scopes like these to its **server-resolved principal**: ```ts const agentScopes = [ { kinds: ["read", "create", "update", "rename"], excludeDocIds: [ACCESS_DOCUMENT_ID], // "sys:access" }, ]; ``` Scopes restrict permissions; the principal must already hold the admin role. Multiple scopes are alternatives, so every scope that would otherwise permit writes to `sys:access` must exclude it. The document remains readable so the client can predict its own authorization. **This does not prevent all permission changes.** Schemas contain `access` rules, and an agent that can edit a schema can change those rules too — including opting a type into the purge that no rule permits by [default](#defaults). Per-document restrictions and principal scopes still apply. If permission changes require human approval, this configuration alone is insufficient. Two configurations go further: * **Freeze the schemas whose rules matter.** Add their document ids to the same exclusion (`sys:schema:`, which `schemaDocId` spells for you). The agent can still create *new* types — and a new type’s rules govern only the documents it is creating anyway — but it can no longer redefine an existing one, so today’s types stop being an editable permission surface. The cost is schema evolution: those types then change only through a human admin or the host. * **Keep the agent below admin.** Without the admin role it cannot write schema documents at all, and the host creates them with a [system-authority](#system-authority) direct call — the same privilege bridge per-document entries use for a non-admin claiming a private document. The app decides which types come into existence; the agent never holds the authority itself. ## Per-document entries The **object grain** of the model: the folder’s `sys:access` document carries an optional `documents` map of per-document entries, keyed by document id. An entry holds the same eight rule fields as a schema’s access block — the same vocabulary, one grain finer — and resolution is **most-specific-wins per operation**: a rule in the entry (the operation’s own, else the entry’s `write` fallback) *replaces* the document type’s rules for that operation; only where the entry is silent does the type grain apply. Replacing rather than composing is the point — it lets an entry restrict *below* what a principal’s folder role would allow, which is what per-document privacy needs. A private document is simply an entry naming one subject: ```json { "documents": { "todo_x7fk2": { "read": { "subject": "alice" }, "write": { "subject": "alice" } } } } ``` There is no owner concept — ownership *is* an entry naming one subject, and sharing is an `anyOf` naming two. Three properties follow from the design: * **Entries are facts about the id, not the document.** An entry may be written before the document exists — `create` resolves through it, so an entry admitting one principal **reserves the id**. It survives delete, restore, and purge, and applies again on recreation: a purge-and-recreate cannot shed an id’s access rules. Entries die only by an explicit `sys:access` edit. (A `write: "nobody"` entry is a persistent **freeze** — the stored dual of the `excludeDocIds` scope axis.) * **Entries hide from admins too.** Under most-specific-wins a `read: { subject: "alice" }` entry hides the document from folder admins as well. An admin regains access by editing the entry — a visible governance write — never by silent omniscience. The governance meta-gate itself is untouchable: entries targeting `sys:access` or schema documents are inert, and the shaped system documents stay exempt from read rules. * **Redaction shapes each entry to its reader.** A non-admin’s synced `sys:access` carries whole the entries whose rules it satisfies — parties sharing a document see each other, like any real ACL. An entry that does not admit the reader but leaves reads open (a lock on a visible document) arrives as an identity-free deny-stub, so the lock predicts without naming who holds it. An entry that hides its document is omitted entirely: a reader never learns that an id it may not see is governed at all (for a reservation, that would leak intent to create). **Who writes entries?** `sys:access` is admin-meta-gated, so an admin edits `documents` over the ordinary wire like any governance content — sharing and revocation are just synced edits, and revocation fires the [revocation sweep](#revocation). For a non-admin claiming a private document, the **host** writes the entry with a [system-authority](#system-authority) direct call — a privilege bridge at the same trust boundary as connect-time grant resolution: the app decides *who may mint*, the entry itself is ordinary declarative data, and the client then creates the document over the normal wire. ## Read filtering Visibility is governed the same way, enforced everywhere the server answers a read — on direct reads and per **recipient** on every broadcast. A subscribe or a history request for a document the principal may not read comes back as *not found* (a hidden deleted document included — never revealed as deleted), the `sys:index` and `sys:trash` listings omit it, live updates about it are withheld, and any listing built on the index never shows it. **Hidden existence** is deliberate: nothing a client can observe distinguishes “hidden from you” from “does not exist”. Two grains, mirroring writes: * A [per-document entry’s](#per-document-entries) **read rule** hides that one document — where per-document privacy and ownership live. * The schema’s **read rule** hides a whole document type from anyone below the required role. Its default is the **inverse** of the write kinds: with no read rule at either grain, *anyone admitted to the folder* reads — not merely any authenticated subject — and the write fallback does not apply. Read filtering is opt-in per document (an entry) or per type. Hiding is not locking — pair a read rule with the matching write rules (an entry naming a subject usually carries both). The read rule’s bar is a **role**, resolved from `sys:access` and connect-time grants exactly as the write rules resolve theirs — so `sys:access` is not a third read knob but the shared roles dial both axes read. Demote a subject there and read access falls together with write access in a single edit, which fires the [revocation sweep](#revocation). The system documents are **shaped, never blocked** — blocking one would strand the client. `sys:index` and `sys:trash` arrive filtered to the reader’s visible entries; schema documents stay readable to everyone admitted (validation and prediction depend on them); `sys:access` arrives redacted (below). Presence is *not* in this set — it is governed like any document (its presence schema’s `access` rules), so a channel can be read-blocked; the denial is an explicit `unauthorized` error. Writing without reading is otherwise legal: a “drop-box” type is readable only by admins but writable by any authenticated subject (every write kind except **update**). It works because the authoring connection always receives its own writes’ events, so an awaited write still resolves for its author. An **update requires read**. A patch runs against the document: its `test` operations pass or fail on values the writer can’t see, and a path that isn’t there is refused. So the server answers an update from a principal who can’t read the document as “not found”, the same answer as for a document that doesn’t exist, whatever the patch asks. `can` and `getCapabilities` predict it: `update` is false wherever `read` is. A create needs no read, since a new document starts empty. ## Defaults Every rule is optional, and the defaults are deliberately **asymmetric**: an untouched document type is easy to write and easy to read. Writes lean on “any authenticated subject”, reads lean on “anyone already in the folder” — so both write restriction and read filtering are **opt-in per type**. | Axis | Default when the type declares no rule of its own | | ------------ | -------------------------------------------------------------------------------------------------- | | **Write** | **any authenticated subject** | | **Read** | **anyone admitted to the folder** | | **Purge** | **nobody** — explicit opt-in per type | | **Commands** | **nobody** — each [domain command](/concepts/domain-commands/) is granted by name under `commands` | Writes resolve through the grains before that default applies: the per-document entry’s own rule for the kind, else the entry’s `write` fallback, else the block’s rule for the kind, else the block’s `write` rule; only when all four are absent does *any authenticated subject* take over. Reads have no per-kind rules and no `write` fallback at either grain — a document with no `read` rule on its entry or its type is simply visible to everyone admitted to the folder. **Purge inverts the write default entirely.** Hard deletion is irreversible, so a document with no `purge` rule (at either grain) cannot be purged over the wire by anyone — admins included — and the `write` fallback does not apply. The asymmetry is deliberate: in a folder with no `sys:access` document every authenticated subject acts as admin (see [Bootstrapping access](#bootstrapping-access)), so any role-based default would hand irreversible destruction to anyone authenticated in an ungoverned folder. A schema opts in explicitly, e.g. `access: { purge: { role: "admin" } }`. (The host’s own retention sweep runs under system authority and is unaffected.) Two things sit outside the table and override it: **scopes** on the principal (a hard cap checked first, for every actor) and the **governance meta-gate** (writing `sys:access` or a schema always needs admin, whatever the block says). And when the whole `sys:access` document is absent — not just a rule — a folder-wide default takes over instead, which is the next section. ## Bootstrapping access A folder **without** a `sys:access` document treats authenticated subjects as admin: governance is opt-in, and somebody has to be able to create the document that grants the right to create it. Apps that hand out roles should seed the roles document when the folder is created; an app whose roles ride connect-time grants alone should seed an **empty** one, or the bootstrap default quietly makes every authenticated subject an admin. On hosts that support it, this seeding runs once at first boot as a [system-authority](#system-authority) direct call and is never clobbered by later admin edits. ## Revocation When the facts move — a `sys:access` edit demotes a subject, a schema tightens a read rule — the server **sweeps** every live connection: newly-hidden documents are evicted (the client drops its copy; the document leaves the UI) and each connection’s filtered index, trash, and redacted `sys:access` view refresh. Per-document entry changes are `sys:access` edits, so they sweep the same way. Facts the engine can’t observe itself — grants in connection metadata — trigger the sweep on demand. Eviction only drops the subscription, so a re-grant heals on the next resubscribe: reveal, don’t push. **Mid-connection principal changes** are first-class: a host can rewrite a live connection’s principal — a demotion, a grant widening — and the engine re-sweeps reads and refreshes that connection’s served identity, which the client adopts on the spot (see [Capability-aware clients](#capability-aware-clients)). Enforcement, prediction, and capability UI all flip without a reconnect. Forcibly closing a connection remains for actual offboarding — prefer demotion for revocation, disconnection for removal. ## Managing access in practice A folder is one sync scope; an organization spans many. The split that keeps membership changes cheap: **org-wide facts live in your app’s own store and ride the connection; `sys:access` holds only folder-local exceptions.** * **Membership and admission** stay in the app’s database. Whether a user may connect to a folder at all *is* the read boundary, and their baseline role arrives as a grant minted at connect time. Onboarding a member is one row in your store — no folder is touched; each folder learns the identity lazily, at connect. Never fan a role change out across folders. * **Folder-local roles** go in `sys:access`: per-subject exceptions and folder-wide defaults. The server reads it fresh on every write and clients sync it live, so granting or revoking there takes effect on the very next write, no reconnect. * **Put ceilings in the document, floors in grants.** Effective role is the *max* of the two, so a grant only ever adds role — a role you routinely lower belongs in `sys:access`, where an edit takes effect on the next write. Grants can move mid-connection too: a host can rewrite a live connection’s grants and reauthorize in one motion, so even grant-carried role is revocable without closing the socket. Revoking admission itself (offboarding) blocks future connects at the app layer and closes live sockets. * **`sys:access` is redacted per reader.** A non-admin member’s synced copy carries only its *own* role entry, the folder defaults, and the [per-document entries](#per-document-entries) that admit it — never other subjects’ assignments or entries; admins read the full document (as do [system-authority](#system-authority) direct reads), and its history is admin-only. Redaction preserves every role-derived self-verdict by construction, so a member’s own capability prediction keeps working. But it also means a client can only predict for **its own principal**. Gate any member-management UI on the admin role. ## Capability-aware clients The same rule evaluation runs on the client — one shared, pure evaluation, so the two sides cannot drift: * **Identity is server-fed.** Every client subscribes a virtual, per-connection identity document carrying the principal the server actually enforces, and **adopts** it. Prediction and capability discovery therefore need no client-supplied principal (a pre-connect seed remains, but the first served answer corrects it), the “worker mints it, browser mirrors it” drift class is gone by construction, and a host-driven mid-connection change arrives as a fresh adopted identity on the live socket. * **Prediction.** A write the server would deny throws locally, before anything is sent, in the same window as optimistic validation. Prediction never falsely blocks: with no principal or unsynced facts it passes through, and the server stays authoritative — a stale “allowed” is corrected by the normal rollback. * **Discovery.** A client can ask whether the current principal may perform a given kind on a document type — answered as yes, no, or *unknown* (“don’t gate the UI yet”). The kinds span the five write kinds plus read; a read “no” predicts the document would be hidden. Delete, rename, and restore, held only by id, resolve their target’s type first, so they are discoverable the same way — not just create and update. There is deliberately no “is this document hidden” query — absence is the signal (hidden existence, again: a hidden document is indistinguishable from a nonexistent one). Omitted entries make “yes” optimistic Prediction runs the same declarative evaluation the server enforces, over the client’s synced facts. Redaction keeps that prediction exact wherever the document itself is visible: a [per-document entry](#per-document-entries) that does not admit the reader but leaves reads open — a **lock** — still reaches the reader as an identity-free deny-stub (subject conditions stripped, an emptied rule becoming `nobody`), so a locked document predicts read-only instead of offering an editor whose every write bounces. Only an entry that *hides* its document is omitted outright, and for those documents prediction resolves the type grain instead: * A predicted **“no” is definitive** — omission only ever removes DENYING rules the reader would have failed anyway, so nothing the client cannot see could re-allow it. * A predicted **“yes”** can be optimistic for a document the server hides: the write comes back denied through the normal rollback, and a read simply answers not-found — exactly the hidden-existence behavior everywhere else. Gate destructive UI on a definitive “no”, but treat “yes” as “not blocked yet”, not “guaranteed”. ## Staging and authorization [Staged sessions](/staged-sessions/overview/) compose with authorization rather than bypass it — **commit has no privilege of its own**. A commit drains through ordinary client writes, authorized exactly as the committer. That makes “the agent proposes, the human disposes” a *configuration*: scope an agent so it can stage freely but never drain a target, and let the human commit the same persisted changeset. * A **preflight** reports the predicted verdict per stage before committing, and the session’s own projection carries the same per-stage flag. * A denied stage fails as unauthorized and stays staged for retry. * Every stored event records the writing principal (subject and actor) — so “which edits did the AI make?” is answerable from the committed data, not from logs. ### Session-level attenuation A session can carry its own **scopes**, drawn from the same vocabulary as a principal’s scopes — for example, a worker session that may only ever update kanban cards, never create or delete and never touch another type, regardless of what its principal could otherwise do. Reads through that session are filtered too: an out-of-scope document reads as absent, and discovery documents such as `sys:index`, `sys:trash`, `sys:session`, and `sys:stage:` are shaped to the documents still readable through the session. Session scopes **intersect** the client’s principal scopes — both gates run — so a session can only ever *narrow* what its client may read or write through the session facade, never widen it, and never adopt a different identity: a session always writes as its client’s principal. An out-of-scope write is rejected **locally**, before it leaves the client, in the same unauthorized shape the principal-scope prediction raises. Session scopes are **advisory** — a cooperative guardrail for agent tooling and UI, not a security boundary. The underlying connection still holds whatever its principal may read, and real write enforcement stays server-side against that principal, so a caller that bypasses the session API is not stopped by session scopes; use principal scopes, resolved at connect time, for a hard cap. ## What this is not The declarative model spans three grains — folder roles, document-type rules, and [per-document entries](#per-document-entries) — and stops there. There are no **per-field** rules: read filtering is whole-document — a principal sees a document entirely or not at all. There are no **content-based** rules either (“readable once published”): verdicts are computed from the principal and the governance facts alone, never from document data — the pure-facts contract that lets the client predict every verdict. An app encodes a state transition by *writing the entry* when the state changes, keeping the fact in governance where it is auditable. And there is no group or org model in the engine: relationship-shaped access (“members of team X”) rides **grants**, resolved at the app’s trust boundary and matched by grant conditions at either grain. # System documents > The engine's own state, exposed as documents you read and subscribe to like any other. datadata exposes its own state as documents under the reserved `sys:` prefix. They ride the same subscribe/read path as your data — [documents all the way down](/introduction/documents-all-the-way-down/) — so a devtools panel, a sidebar, or an agent introspects the engine with the calls it already uses. Most `sys:` documents are **synthesized read-models**: the engine derives them and you never write them directly. The exceptions are **real, editable documents**, written through the same `createDocument`/`updateDocument` calls: schema documents ([schemas as documents](/concepts/schemas-as-documents/)) and the `sys:access` roles document ([authorization](/concepts/authorization/)). No `sys:` document can be deleted, editable or not. System documents are also commonly **shaped by authorization**: unlike a normal document, which every subscriber sees identically, a system document is shaped per reader. Listings hide the documents you can’t access, `sys:access` is redacted to your own entry, and `sys:principal` is your connection’s identity alone. Two clients subscribed to the same `sys:` document can hold different state. See [Authorization](/concepts/authorization/). ## Folder state * **`sys:index`** — every document in the folder the reader can access, with its type and [name](/concepts/documents-and-folders/). Renaming a document patches its entry here, so a subscriber sees the new name without reloading the document. The data source for a listing or sidebar. * **`sys:trash`** — every soft-deleted document the reader can access (id, type, name, [generation](/concepts/documents-and-folders/), when); the data source for a trash/restore UI. It updates live as documents are deleted, restored, and purged (a purge removes the entry — the signal that restore is no longer possible). A restore is aimed at the generation the client saw deleted — an entry’s, when the client holds no tombstone of its own — so a restore that sits queued never brings back a different document later created and deleted under the same id. ## Schemas * **`sys:schema:`** — the schema for a document type. A real document, written through the same `createDocument`/`updateDocument` calls and carrying its own [migration history](/concepts/schema-evolution/). * **`sys:schema:presence:`** — a named [presence](/concepts/presence/) channel’s cell schema. Presence is schema-gated too: no declared schema, no channel. ## Access * **`sys:access`** — the folder’s roles document: per-subject role assignments and folder-wide defaults. A real, editable document (writable only by admins, with admin-only history), but served **redacted per reader** — a non-admin syncs only its own entry plus the defaults, while admins and trusted server code read the full document. A folder without one treats authenticated subjects as admin; governance is opt-in. See [Authorization](/concepts/authorization/). ## Staged work * **`sys:session`** — a summary of all currently staged work in a [session](/staged-sessions/overview/). Its per-target entries are shaped by target read visibility; hidden stages are counted in aggregate without revealing their target ids. * **`sys:stage:`** — the full staged detail for one document. See [`sys:session` and `sys:stage`](/staged-sessions/sys-session-and-sys-stage/). ## Identity and client state * **`sys:principal`** — the connection’s own identity: the principal the server actually enforces (roles, grants, scopes). Synthesized per connection from its connect-time state and never stored. A client subscribes to it and **adopts** it, so prediction and capability discovery run against the real principal; a host-driven mid-connection change is re-served on the live socket. See [Authorization](/concepts/authorization/). * **`sys:presence::`** — ephemeral per-participant state (who’s here, cursors, what’s streaming) on a channel. Never stored durably: the server keeps an in-memory aggregate and each client owns its own cells. See [Presence](/concepts/presence/). * **`sys:presence-view::`** — a [session’s](/staged-sessions/overview/) own lens over that aggregate, shaped `{ self, peers }`: `self` is this session’s cell, `peers` is every other cell annotated with staleness from the liveness overlay. Synthesized on read and read+write — writing `self` publishes this session’s cell through to the aggregate — but never persisted or synced; it exists only through a session, so the live client reads it as null. * **`sys:client-docs-status`** — the client’s own view of what it’s subscribed to and where each document sits in the sync lifecycle. A client-side read-model, not folder state. # Staged sessions > A staging area on top of live sync — accumulate changes, preview conflicts, commit atomically. The live client sends every change immediately. That’s right for direct manipulation, but wrong for work that should be **reviewed before it lands** — an AI agent drafting edits across several documents, a form with a save button, a batch refactoring of a diagram. The **staged session** adds a staging layer for exactly that. ## Stage → review → commit A session wraps the live client. Through it you edit documents as usual, but changes are **staged**, not sent: * Reads through the session show the **staged view**: the live document — its **head** — kept current by ordinary sync, with your staged changes overlaid on top. You are never looking at a snapshot: as collaborators change the document underneath you, the overlaid read moves with them. (A `skipStaged` read bypasses the overlay when you want head exactly as it stands.) * Each staged document records the **base sequence** it was edited from, and every staged change carries guards asserting the values it was authored against. * Because head keeps moving, the session can tell you at any moment whether it has drifted somewhere that [conflicts with your staged work](/staged-sessions/conflict-preview-and-resolution/). When ready, **`commit()`** writes all staged changes atomically — all stages land or none do, and a commit is refused while any stage is blocked by a conflict. Or walk away: discard the session and nothing ever hits the live documents. Rich text participates fully. The session **forks** the target Y.Doc into a [bindable staged copy](/staged-sessions/staged-rich-text/) and hands you that, so a real editor (Tiptap, ProseMirror) drafts against the copy while the live document is untouched. At commit, the copy merges back into the target. ## The changeset can live in a document Here’s the datadata twist: give the session a **host document** and its state — the [changeset](/staged-sessions/changesets-and-lanes/) — is persisted inside that document, in a dedicated field. Because the changeset is then just document data: * **It survives reload.** Reopen the session over the same host document and the staged work resumes. * **It syncs.** Two collaborators (or a human and an [agent](/ai-agents/sessions-as-agent-workflow/)) each open a session over the same host document, and see and contribute to the same staged changeset, converging like any other concurrent document edit. Staged rich-text copies get live carets between participants for the same reason. * **It’s inspectable** through the synthesized [`sys:session` and `sys:stage` documents](/staged-sessions/sys-session-and-sys-stage/), like everything else in datadata. A natural host is the document that *motivates* the changes — for example, an AI conversation document hosting the changeset of edits the conversation is producing. The host’s schema declares the field, and it defaults to an empty changeset, so a host document needs nothing more when it is created: ```ts // In the host type's schema fields: { title: { type: "string" }, session: SESSION_CHANGESET_FIELD }, client.live.createDocument({ docId: hostDocId, type: "chatConversation", data: { title: "Refactor the intro" } }); const session = client.openStagingSession({ hostType: "chatConversation", hostDocId }); ``` The changeset is there from the moment the host exists — and a host stored before the field was added gets it the first time the server reads it — so every staged edit lands inside it, and edits from several authors merge instead of replacing one another. ## Keeping the editor away from the host Persisting the changeset requires the session’s principal to hold **write access to the host document** — the staged work rides ordinary host-document writes. Often that’s fine: the human who owns the conversation may edit the conversation. But when the session’s editor shouldn’t control the host — an AI agent staging edits into the conversation that drives it must not rewrite that conversation, or read its own staging machinery back — the app narrows the session, not the client: ```ts const session = client.openStagingSession({ hostType: "chatConversation", hostDocId, scopes: [ { kinds: ["read", "create", "update", "rename"], excludeDocTypes: ["chatConversation"] }, ], }); ``` `excludeDocTypes` is the scope’s **deny axis** — “everything this scope covers, except these types” — the shape an allow-list can’t express when the rest of the world is open-ended. With the hosting type excluded, the session treats every document of that type like any other out-of-scope document: it reads as absent, drops out of the `sys:index` / `sys:trash` listings, and every session-routed write targeting it (the staged Y.Doc lanes included) is rejected. The changeset persistence is untouched — it rides the client’s own capabilities underneath. Two things to keep straight: scopes are OR-ed, so the exclusion must appear in *every* scope that would otherwise cover the type; and this is per-session app policy, not a property of the host — a reviewer’s session over the same host can keep the conversation readable by simply not excluding it. The listings drop excluded documents rather than redacting them — ordinary [hidden existence](/concepts/authorization/#read-filtering) — so a type resolved through the session can’t distinguish hidden from nonexistent. An app gate that needs that distinction resolves through the client instead, whose listings this session-level attenuation doesn’t touch. ### Freezing the type’s schema too Excluding a docType does not fence off that type’s **schema document**. A schema document’s own type is `sys:schema`, not the type it defines, so an editor hidden from `chatConversation` documents can still redefine what a conversation *is* — and excluding `sys:schema` wholesale would take away schema authoring for every other type along with it. Name the document instead: ```ts const session = client.openStagingSession({ hostType: "chatConversation", hostDocId, scopes: [ { kinds: ["read", "create", "update", "rename"], excludeDocTypes: ["chatConversation"], // schemaDocId("chatConversation") spells this id for you excludeDocIds: ["sys:schema:chatConversation"], }, ], }); ``` `excludeDocIds` is the same deny axis one grain finer — freeze these documents, whatever their type. An app hiding a host type usually wants its schema frozen too, so reach for both together. Two properties make it a dependable freeze: docIds are stable (`rename` changes a document’s display name, never its id, and schema documents refuse rename outright), and `create` is gated on the same id, so freezing an id that nothing has claimed yet blocks the document from being created at all. Freezing is not blinding, though. The system documents — `sys:index`, `sys:trash`, `sys:access`, `sys:principal`, `sys:schema:*` — are exempt from read attenuation and stay readable however they’re scoped, because a client that couldn’t read the schema couldn’t validate or predict against the type. A frozen `sys:schema:chatConversation` means *no writes*, not invisibility. ## Sessions without a host document The host document is optional. Open a session without one — no host schema, no field to declare — and the changeset is held in memory instead. An in-memory session stages, previews conflicts and commits exactly like a hosted one — same reads, same guards, same atomic `commit()`. What it gives up is everything that followed from the changeset being document data: the staged work does not survive reload, does not reach other devices, and cannot be collaborated on. It is single-client by construction, so staged rich text has no carets to show. Dispose the session and the work goes with it. That suits staged work that lives and dies inside **one process** — a server script that stages a batch of edits and commits them in one pass, a single agent turn handled server-side, a test. Reach for a host document as soon as the staged work must outlive the client that made it, or be seen by anyone else: a form with a save button wants one, otherwise a reloaded tab drops everything the user typed. Either way `sys:session` and `sys:stage` work, since those are derived from the changeset rather than from the host document. Open question A changeset occupies one field of its host document, and every session opened over that field shares it — that is how participants collaborate on staged work. What’s missing is *named* changesets: addressing several by id, with an API to create and drop them. Parallel agent proposals and draft-vs-review lanes want that, and it stays deliberately deferred. See [Open questions](/known-issues/open-questions/). # Changesets & lanes > How staged work is represented — stages, two lanes of changes, and commit intents. A session’s persisted state is a **changeset**: a structured value stored in a field of the [host document](/staged-sessions/overview/). Its shape mirrors the [hybrid change model](/concepts/rich-text-with-yjs/) of datadata itself. ## Stages For every document touched by the session, the changeset holds a **stage**: * the document’s **base sequence** — the version the edits were computed against, * the document **type**, whether the stage **creates** or **deletes** the document, and a pending **rename** if there is one. That’s all — no copy of the document’s data travels with the stage. Each staged patch embeds its own RFC 6902 `test` guards (the values its author saw), which is what [conflict detection](/staged-sessions/conflict-preview-and-resolution/) runs on: the patches plus the live head, no stored base. Storing no base keeps the changeset proportional to the staged **edits**, not to the size of the documents being staged — a session touching large documents is still a small host document. When a resolution surface needs the base *value* (the three-way preview), it is resolved on demand — from the live head, an in-memory pin, or a replay of the document’s [event log](/architecture/storage-and-event-log/). ## Two lanes, two shapes Staged edits accumulate in two parallel lanes — and, deliberately, the lanes have different shapes: * **`changes`** — [JSON Patch](/concepts/changes-as-json-patch/) operations against structured data: a flat record set ordered by fractional index, the same [ordered-record pattern](/concepts/documents-and-folders/) used for collections everywhere in datadata. Each record is addressable — a change can be **amended** or **dropped** by id, not just wholesale — and carries the [generation](/concepts/documents-and-folders/) of the document it was authored against, and the [schema sequence](/staged-sessions/conflict-preview-and-resolution/#when-the-schema-moves-under-you) its patch was authored at. * **`yjsCopies`** — staged [Yjs](/concepts/rich-text-with-yjs/) content: **one live-seeded host sub-document *copy* per target Y.Doc**, recorded as `{ docId, yjsId, copyRef, stage }`. The record stores no update bytes — the copy holds the content as a real, synced sub-document, not a stored delta. The asymmetry is the point. Structured edits *are* addressable, so the JSON lane is a record set a review can amend or drop from by id. Collaborative text isn’t — it merges as a whole document — so the Yjs lane is a **forked copy** instead. [A prepared document’s `getYDoc`](/staged-sessions/staged-rich-text/) hands back a copy seeded from the live target, but stages nothing: the first *edit* promotes it to a `yjsCopies` entry anchored by a deterministic `copyRef`, so merely opening an editor leaves the changeset untouched. Editors edit the copy directly, and live sharing, cursors, offline editing and state-vector catch-up ride the host document’s ordinary Yjs sync for free. The copy also [tracks the live target](/staged-sessions/staged-rich-text/) during the session, so upstream edits surface to the editor as they happen rather than at commit. Both lanes converge under concurrent authorship, by different means. Patch records from two authors interleave by fractional index. The Yjs copy needs nothing extra: `copyRef` is **deterministic** — a stable id derived from the target’s `docId`, `yjsId` and [generation](/concepts/documents-and-folders/) — so concurrent sessions on the same target resolve to the **same single copy**. They collaborate on it natively, with no last-writer race over which copy is canonical. The generation in that id is what keeps two *incarnations* of a target apart. A copy forked before the target was purged and recreated under its id belongs to the document that no longer exists, so it resolves to a **different** sub-document than a copy forked after — the vanished document’s content can never mix into the one the recreated document’s authors are editing. That generation, and the one on each `changes` record, are also what make the `targetReplaced` block survive concurrent authorship. The stage record carries a generation too, but it is **one path** in the changeset: an author whose view of it predates a sibling’s staging writes it whole, marker included, so the sibling’s stale work would go unblocked and commit into the document that reused the id. Staged work is judged by the generation it carries itself — per change, and per copy — which no other author’s write can replace. For the same reason, every `changes` and `yjsCopies` record also notes the **stage it was written under**: the document type and base sequence. Whether a stage still exists is decided by whoever last touched it. When one author commits or drops a stage, it’s removed once the work *they* can see is gone. By then another author may have added work under that stage that hasn’t synced yet, for example an edit made offline. When that work arrives, it rebuilds its stage from the record instead of being dropped, so it shows up as staged work for everyone to review. If it was written against a create that has since been committed, it names no generation of the new document, so its stage shows up blocked as `targetReplaced`. `copyRef` is a `yjsRef`-typed value, so it’s [reference-counted](/concepts/references-and-integrity/) like any other Yjs reference. Its presence in the changeset anchors the copy sub-document against the host document’s orphan GC; clearing it reaps the sub-document. There’s no session-specific Yjs cleanup — committing or discarding just drops the entry, and the general reference GC reclaims the copy. ## Commit intents One more record set rides the changeset: **`commitIntents`**, written by `commit()` itself just before it drains. Each intent records — per stage — which staged change ids are draining under which pre-minted write event ids. It’s the changeset’s crash-safety ledger: with [offline persistence](/concepts/offline-persistence/) enabled, a commit that dies with its page leaves journaled writes that replay on the next boot, and the intent is how a resumed session recognizes them as *its own commit already in flight* — it waits for their outcome and finishes the commit’s bookkeeping instead of re-applying the staged work. You never write this lane yourself; it appears during a commit and clears when the commit’s writes settle. Stages leave the changeset one by one, each as its own write is acknowledged — not all together when the commit ends. In between, a stage’s changes are in the live document *and* still listed as staged, so the intent also tells a reader which changes to leave out: a session never folds a change its commit has already put into the live document it holds. The committing client knows from its own write — pending, or remembered as applied the moment its acknowledgement arrives, so a listener notified by that acknowledgement, or by a refusal’s rollback, reads the stage right. Every other client subscribed to the target knows from the target’s patch: it carries the write’s event id, the same one the intent records, so an observer reads the stage as landed from that patch on, even when the host’s patch that clears it arrives late. For a schema document every client can tell from the document alone — the intent records the keys its staged migrations add to the log, and finding them in the live log is the landed write. An observer whose copy of an ordinary target never saw the patch (it subscribed, or reloaded, between the target’s write and the host’s) has no such evidence, and may read a non-idempotent staged op (an append into an array) twice until the host’s patch arrives. ## Authorship is not staged A staged change carries only its id, order and patch — no per-change author or originating-turn metadata. Staging is transient: the changeset is cleared once every stage drains at commit (or the session is discarded), so anything stamped on a staged change is visible only for the window the change sits staged — not where a durable record would matter. Durable authorship lives on the [event log](/architecture/storage-and-event-log/) instead: every committed event records the writing principal (subject and actor). An app that wants to group staged changes by conversation turn keeps its own map keyed by staged change id — turns and tool calls are the app’s vocabulary, not the library’s. # Staged rich text > Bind a real editor to staged Yjs content — live cursors included — before anything commits. Staging [Yjs](/concepts/rich-text-with-yjs/) content through update callbacks works for programmatic edits — an agent appending to a draft. An interactive editor is different: Tiptap or ProseMirror holds a `Y.Doc` and renders keystrokes from it. For that, the session hands the editor a **real, long-lived `Y.Doc`**. ## The staged Y.Doc `session.prepareDocument(docId).getYDoc(yjsId)` hands back a **copy** of the live Y.Doc, seeded from live state. Bind y-prosemirror to it and the editor is editing the copy directly, while the live document stays untouched. Reading a copy stages nothing — the first local **edit** promotes it into the [changeset](/staged-sessions/changesets-and-lanes/), which is where the copy’s identity and lifecycle live. Repeat calls return the same copy, and so do two *separate* sessions on the same target. If the target is gone when that first edit lands — deleted, purged, or no longer visible to the session — the edit still applies to the copy, but nothing is staged. Editing never throws: the refusal is reported once on the client’s `onWriteError` channel, and the copy stays unstaged until an edit made after the document comes back promotes it, with everything typed in between. Because the copy is just another synced Y.Doc of the host document, everything Yjs already does, it does here for free: edits sync as ordinary Yjs updates, two sessions’ editors converge on each other’s staged content **before either commits**, and offline editing and exact state-vector catch-up ride the host document’s sync — with no delta capture or debounced fold in between. The copy also **tracks the live target during the session**: concurrent upstream edits to the live document are mirrored into it as they arrive, so an editor sees what changed upstream instead of only discovering it at commit. That relayed upstream is a **local overlay** — never persisted into the copy or synced to other participants, and re-derived on each acquire, so a server-side session survives Durable Object hibernation without holding a live subscription across it. At commit, the copy’s full state merges back into the live target — cleanly, because the copy was seeded from live, so re-applying it no-ops the seed and adds only the staged edits. `session.discard()` has nothing to un-apply: the copy was forked *from* the live document, never applied to it. Either way the changeset entry clears and the copy is reclaimed; re-acquire after a discard and you get a fresh copy, seeded again from current live state. ## Cursors ride presence `session.getYjsAwareness({ docId, yjsId })` is the cursor half: one shared y-protocols `Awareness` per staged doc, bridging the copy. Editors in the same session instance get caret visibility with no transport at all. For a persisted session, the cursors ride **this session’s own [presence](/concepts/presence/) cell** — the `sys:awareness` field of its `self` in the target’s presence document — so they travel between participants’ instances too. Two properties are deliberate: * **Ephemeral by construction.** Presence never touches storage: cursors vanish on disconnect and don’t survive reload. The staged *content* does. * **Resolvable only by participants.** The cursors are positions into the copy’s CRDT structs, and only session participants hold those structs. A non-participant who sees the cell cannot resolve them — they point into content it doesn’t have. ## Why this matters for review This is the missing half of [sessions as agent workflow](/ai-agents/sessions-as-agent-workflow/): an agent stages a draft, a human opens a real editor **on the staged text** and refines it — with live cursors if a second reviewer is in the session — and nothing touches the live document until commit. Review stops being read-only. # Conflict preview & resolution > Derived three-way conflicts, and the verbs for resolving them — amend, rebase, drop. While a session holds staged changes, the live documents keep moving — the **head**, the live document as it stands right now, advances underneath the [staged view](/staged-sessions/overview/). The session’s job is to make that visible **before** commit, not to surprise you at commit time. ## Conflicts are derived, not stored A stage stores only its [patches and base sequence](/staged-sessions/changesets-and-lanes/) — no snapshot of the document. Each staged `replace` or `remove` carries an RFC 6902 `test` [guard](/concepts/changes-as-json-patch/) asserting the prior value at the path it changes — an `add` has no prior value to assert and so [carries no guard](/known-issues/json-patch-rfc/#add-has-no-precondition). Detection is a fold: the staged patches are applied in order, guards included, onto the live head. A clean fold **is** the auto-merge — disjoint upstream changes need no attention. A failed fold — a guard that no longer matches, or an operation that no longer applies — marks the stage **blocked**. The fold needs nothing but the head and the patches, so it gives the same answer on every client, and it’s recomputed rather than persisted: a head change that turns out to be disjoint un-blocks the stage by itself. One deliberate strictness: a staged path that was edited and then reverted still carries its guards, so an upstream edit to it conflicts — the session expressed intent about that path even though the net change is nothing. ## The three-way preview A blocked stage is available as data — for a diff UI, or for an [agent](/ai-agents/sessions-as-agent-workflow/) to reason about: * **base** — the document as it was when staging began, * **ours** — base with the session’s changes applied, * **theirs** — the live document now. `getConflictPreview(docId)` returns a promise, because the base *value* is resolved rather than stored: it’s instant in the common cases (the live head still sits at the stage’s base sequence, or this client has watched the stage since it opened), and on a cold client it’s reconstructed by replaying the document’s [event log](/architecture/storage-and-event-log/) up to the base sequence. The replay is online-only — a cold, offline client’s preview waits for reconnect. Detection and commit gating never wait; they need no base. The session’s read view degrades the same way, but only where it must: on a real overlap, where no coherent head-plus-stack view exists, it serves the last coherent one (base plus the staged changes), and a cold client sees the live head un-overlaid until the replayed base lands. ## The resolution verbs * **Amend** a staged change — rewrite a pending change in place (fix the agent’s typo before committing). Async for the same reason as the preview: it re-derives the change against the state it was authored on. * **Drop** a staged change (`dropStagedChange`), or a whole stage (`dropStage`) — its staged edits, rich-text copies, rename, delete or create — discarding part of the staged work and leaving the session’s other stages alone. * **Rebase** a stage — adopt the current head as the new base and restage on top of it. The embedded guards are rewritten against the new base; a staged change the upstream edit already made identical is dropped, and a stage left with no surviving work is removed. * **Discard** the session — walk away; live documents were never touched. `commit()` is the gate: it refuses while any stage is blocked, and writes all stages atomically once none are. ## When the document moves under you A changeset’s `onUpstreamAdvance` policy decides where the line between *merge silently* and *surface it* falls. A hosted changeset carries it, so every session over that host applies the same policy: `autoMerge` unless the host document was created with `session: emptySessionChangeset({ onUpstreamAdvance: "block" })`. An in-memory session takes it as an `openStagingSession` option: * **`autoMerge`** — the default, and everything above. Upstream changes that still fold cleanly merge without a word; only a failed fold blocks the stage. * **`block`** — blocks the stage whenever the document advanced upstream at all, even when the staged patches would still apply. The reason the second policy exists is that a clean fold proves *structural* independence, not *semantic* independence. A staged change guards the paths it **wrote**, not the paths it **read**. An agent that reads a whole document and stages a one-field edit to `summary` produces a patch that folds cleanly over an upstream edit to `body` — and a summary that now describes text nobody can find. Disjoint on paper, stale in fact. So reach for `block` when a staged change was derived from more of the document than it touched: agent-authored edits, or a review gate where “somebody else moved this while you were staging” is itself something the reviewer should see. Because it keys off the document’s sequence rather than the guards, it also catches what the guards structurally cannot — two authors adding the same new key. You pay for all this in false positives: on a busy document, every upstream touch blocks and wants a rebase. That’s why `autoMerge` is the default. A stage blocked this way reports `reason: "upstreamAdvance"` rather than `physicalConflict` — nothing clashed, the patches still apply. It keeps reading head plus your staged changes, which is exactly what commit would write; only a real overlap falls back to base plus stack. And when a stage both overlaps *and* advanced, the overlap wins the reason: `physicalConflict` is the more actionable of the two. Both policies are live: staged work follows the head as it moves, and both ignore Yjs-only upstream movement (rich-text deltas merge onto any later state, so they conflict with nothing — an “advance” here means the JSON lane’s sequence, which they never touch). Neither ever auto-resolves an overlap — that stays a decision for you or the agent. ## When the schema moves under you A [migration](/concepts/schema-evolution/) is not a conflict. When a type’s schema gains a rename, remove or remap while a stage on one of its documents is open, the server brings the document forward — which advances its sequence and moves its shape out from under patches authored against the old one. Read naively, that is a stack that no longer folds. So staged work records the shape it was authored in: each change carries the `schemaSequence` of its type’s `sys:schema:` document at staging time, and the stage carries the one its base is at — the [`conformedSchemaSequence`](/concepts/schema-evolution/#migration-runs-on-read--and-writes-back) the live document itself names, which is also how the session knows a head is still a frame behind a schema change that just landed. Whenever the schema has moved since, the session reads the stack **brought forward** — every staged patch, and the base or head it folds onto, replayed through the migrations stamped since with the same engine the server uses for [pending writes](/concepts/schema-evolution/#pending-edits-survive-a-migration). The conflict check, the overlaid read, the three-way preview, `amend`, `rebase` and commit all see that one form: an edit to `title` staged before a `title → name` rename reads, previews and commits as an edit to `name`. What still fails to fold after that is a real overlap, and a stack that folds but fails the new shape (a field it now requires) is `schemaInvalid` as ever. The stored changeset is **not** rewritten. A stage is shared by every author on the host document, so rewriting it from each client would be a burst of identical writes and a window of mixed forms; the brought-forward form is a pure function of the stored stack and the schema’s append-only migration log, so every client derives the same one — and the same verdict. `stagedChanges` on a stage lists the patches as commit would fold them. A change whose every operation a `remove` migration deleted lists an empty patch, folds as nothing, and drains with its stage. `rebaseStage` and `amendStagedChange`, which rewrite patches anyway, write them at the current shape and re-stamp them. The same holds for a migration staged **in the session**. While a session stages a schema edit with migrations, it reads every document of that type — staged or not — as the migrations will leave it, so content is authored and validated in the shape it commits into, and changes staged before the schema edit are carried through it. Commit drains the schema document first, then waits for each target’s migration to arrive before folding its content. Content staged under such a migration is in a shape no schema sequence names yet: the live schema plus migrations that have not committed. The server stamps a migration with the sequence of the write that commits it, and any other schema write landing first takes the number a guess would have used. So the change records what is known — the live `schemaSequence` it sat on, and `stagedMigrations`, the keys of the staged migrations it was authored after. A reader takes those keys out of what the change is owed, wherever the migrations ended up: still staged, committed behind someone else’s schema write, or committed before an offline author’s change under them reached the changeset. A change staged between two staged migrations is carried through the second only. A document the session shows carried through staged migrations names no `conformedSchemaSequence` for the same reason. A schema’s migration log is a record keyed by each migration’s [key](/concepts/schema-evolution/), so a staged migration is an entry of its own: when another author commits a migration first, the two sit side by side in the log and replay in key order. What does not merge is reported rather than merged silently: a staged key that another author already used for a different migration, or one that sorts *before* a key committed meanwhile (the server only accepts new keys past the committed range). The patches still fold and the merged schema still validates, so neither is an overlap — but the live log already says the server would refuse the write, and the schema stage reports `reason: "migrationKeyRefused"`, with the key in `migrationKey` and the server’s refusal in `message`. It keeps reading head plus your staged changes, like an `upstreamAdvance`, and outranks one. Resolve it with `rekeyStagedMigration(schemaDocId, from, to)`, which gives the staged migration a later key. Do not amend the schema stage by hand instead: every change staged under the migration records its key, and a change naming a key that no longer exists is read as owed the migration again under the new one — a value a chained `remap` already moved is moved a second time. The session operation rewrites the schema stage and those records in one write; when the key was taken by another author, their entry stays and yours moves. A client that has not heard of the committed key yet can still commit into the refusal; the schema stage then fails and stays staged, to be rekeyed once the conflict shows. If the schema stage fails at commit, content staged under its migrations is held back with it and stays staged: written early, it would be migrated a second time once the schema did commit. One case is not detected: two authors migrating the **same field**. Content staged under your migration is carried through theirs afterwards; when the two do not commute, review the merged schema before committing (the `block` policy flags the schema stage for exactly that). A type validated against a bundled [static schema](/concepts/schemas-as-documents/) has no migration log, so its staged work records no sequence and is read as authored. ## When the document is replaced A stage records the [generation](/concepts/documents-and-folders/) of the document it was opened against. If that document is purged and a new one is created under the same id, the new one is a different generation — even when its sequence happens to match — and every stage on it reports `reason: "targetReplaced"`: staged edits, deletes, renames and staged rich-text copies alike — including a copy opened before the replacement and first edited after it. Nothing staged belongs to the new document, so the session shows it as it is: its reads and index ignore a staged delete or rename, `getConflictPreview` returns `null` (there is no base to merge), rebase refuses, and commit refuses while the stage remains. The only resolution is dropping the stage with `dropStage(docId)`. Gaps & open questions The resolution *vocabulary* (amend/rebase/drop) is settling, but higher-level patterns — field-level pick-and-choose merges, partial commits — are open questions being worked out in the apps. See [Open questions](/known-issues/open-questions/). # sys:session & sys:stage > Staged work is observable as documents — because of course it is. In keeping with [documents all the way down](/introduction/documents-all-the-way-down/), a session’s staged work is exposed as **synthesized, read-only documents** that any client can subscribe to with the ordinary document API. ## `sys:session` — the summary One document summarizing all staged work in the session: one entry per staged document, with enough metadata to render a “pending changes” list — what’s staged, what’s blocked, what would commit. The entry list is visibility-shaped. If the reader cannot read a target document — because of the connection principal or the session’s own scopes — that target is omitted from `stages`. The top-level summary still describes the whole changeset, so a UI can tell that commit-affecting work is hidden without learning the hidden targets: * `isDirty` and `isBlocked` include visible and hidden stages. * `hiddenStageCount` counts omitted stage entries. * `hasHiddenDirtyStages` and `hasHiddenBlockedStages` say whether hidden work is dirty or blocked. Visibility is judged against the schemas the **session** sees, not just the committed ones. A brand-new type whose `sys:schema:` document is staged in the same session counts as known, and that staged schema’s own access rule decides — so an agent that defines a type and creates an instance of it in one turn can read that instance back before either has committed. ## `sys:stage:` — the detail Per staged document, the full picture: the two [lanes](/staged-sessions/changesets-and-lanes/) — the ordered JSON Patch records and the staged Yjs copies — plus base metadata and conflict state. This is what a review UI renders when the user clicks into one pending document. For a target the reader cannot read, `sys:stage:` reads as absent. The detail document carries patches and staged copy references, so it follows the same whole-document visibility rule as the target it describes. ## Why synthesized documents? These documents are **projections** — computed from the changeset, never written directly. Exposing them as documents means: * A review sidebar is just another document subscription — same hooks, same reactivity as the rest of the app. * Devtools and agents introspect staged work with the API they already have. * No second “session inspection API” to design, version, or learn. It’s the same trick as `sys:index` and `sys:client-docs-status`: when the engine has interesting state, the answer is a document. # Why agents like datadata > A small referential API, schemas an agent can write, and changes an agent can explain. datadata is being honed against applications where AI agents and humans edit the same data. Several properties that are merely *nice* for humans turn out to be essential for agents. ## The API fits in a prompt The whole surface is a handful of referential verbs — subscribe, read, create, update, delete — over one kind of thing: the document. An agent’s tool definitions stay small, and there are no special cases to teach: the same calls work for user documents, [schemas](/concepts/schemas-as-documents/), and [system state](/staged-sessions/sys-session-and-sys-stage/). ## Agents can create their own document types Because schemas are documents at `sys:schema:`, an agent that needs a new shape of data can **write the schema and start using it** in the same session — designing a data model is just more document editing. The schema it writes is validated against the meta-schema, so a malformed design is rejected like any other bad write. A valid schema can still be a poor fit, so `lintSchema` (from the `schema` entry point) gives advisory warnings about a schema’s design. It flags an array of objects, where a record keyed by the entity’s id would let concurrent edits to different entries merge, with ordering in a `fractionalIndex` field. It also flags an `id` field that repeats the record key. A tool that lets an agent define schemas can return these warnings to the agent, and an eval can count how often its schemas come out clean. Runtime schema editing requires admin authorization, including authority to change the schema’s access rules. See [scoping schema editing](/concepts/authorization/#allow-schema-editing-while-blocking-changes-to-sysaccess) for how to block changes to folder roles and per-document permissions, and what that restriction leaves editable. ## Changes are legible Structured edits are [RFC 6902 patches](/concepts/changes-as-json-patch/) — compact, diff-shaped, and readable by the agent itself, by a reviewing human, and by the model judging a conflict. When a change commits, the [event log](/architecture/storage-and-event-log/) records which principal wrote it — human or agent — so “which edits did the AI make?” is answerable from the data. ## Staging is the natural agent workflow The [staged session](/staged-sessions/overview/) gives agents what they actually need: a place to accumulate multi-document work, a [three-way conflict preview](/staged-sessions/conflict-preview-and-resolution/) they can read as data, and an atomic, refusable commit — detailed in [Sessions as agent workflow](/ai-agents/sessions-as-agent-workflow/). ## The server can host the agent With [in-process clients](/architecture/in-process-clients/), an agent runs server-side as a first-class client — same API, no WebSocket. Its edits stream to the browser as it works, and the user can keep editing the same documents by hand at the same time. Both sides converge on the same live state. ## The docs are readable by agents too This site publishes an [`llms.txt`](/llms.txt) index. The full documentation is one Markdown file at [`/llms-full.txt`](/llms-full.txt), and [`/llms-small.txt`](/llms-small.txt) is the same without the asides and collapsed sections. Hand an agent whichever fits its context window. # Sessions as agent workflow > Agent stages, human reviews, commit is atomic — and the whole thing survives interruption. The [staged session](/staged-sessions/overview/) maps almost one-to-one onto how agent edits *should* land in shared data: proposed first, reviewed, then committed — never silently. ## The loop 1. **The agent stages.** Tool calls edit documents through a session; nothing touches live data. Multi-document work accumulates as a [changeset](/staged-sessions/changesets-and-lanes/). 2. **The human reviews.** The UI subscribes to [`sys:session` / `sys:stage`](/staged-sessions/sys-session-and-sys-stage/) and renders pending changes live, as the agent works. Diffs are JSON patches against a known base — reviewable, attributable, droppable one by one. 3. **Conflicts are data.** If live documents moved underneath, the [three-way preview](/staged-sessions/conflict-preview-and-resolution/) (base / head / staged) can be fed back to the agent to resolve — amend, rebase, or drop — or surfaced to the human. 4. **Commit is atomic and refusable.** All staged changes land together, or the commit is refused while blocked. A discarded session leaves no trace. ## Interruption-safe by construction Because the changeset persists **inside a host document** — say, the conversation document driving the work — the agent process can crash, the tab can close, the session can resume tomorrow: the staged work is just document data, already synced. A natural pattern is one conversation = one host document = one changeset, so the chat and its proposed edits travel as a unit. Hosting on the conversation doesn’t have to hand the agent the conversation: open the agent’s session with scopes [excluding the hosting docType](/staged-sessions/overview/#keeping-the-editor-away-from-the-host) (`excludeDocTypes`) and the agent can neither read conversations back nor edit, rename, or delete them — staging into one is all the host-write capability buys. Add the hidden type’s schema document to `excludeDocIds` (`sys:schema:chatConversation`) and it can’t redefine the type either — the one write path the docType exclusion doesn’t cover, since a schema document is typed `sys:schema` rather than the type it defines. The attenuation lives on the session, so every tool call routed through it is covered without wrapping the tool surface in an app-level gate. ## Both sides of the same session Changesets sync like any document data, so a human can amend the agent’s staged change — or stage alongside it — before either of them commits. Review isn’t a modal gate; it’s two clients converging on the same staged state — the agent, for instance, running [in-process on the server](/architecture/in-process-clients/) while the human reviews from a browser. The session treats them identically. That extends to rich text. The agent drafts into a staged Y.Doc; the human opens a real editor [bound to the same staged content](/staged-sessions/staged-rich-text/) and refines the draft in place — with live carets if another reviewer is in the session — and the live document still hasn’t moved. Review stops being read-only. # System overview > Session, client, server, the extension points between them, and the optional persistence behind the client. datadata is a library with a clean client/server split. An application talks to it through a **session**, and two deliberate extension points — the **event bus** and the **storage adapter** — decide where and how it runs. ``` flowchart TB accTitle: datadata's client/server split accDescr { A session sits on top of a client. The client writes a journal and cache to an optional persistence adapter, and exchanges events with an event bus. The bus exchanges events with the server, which persists through a storage adapter. Blob bytes bypass the bus: the client uploads and downloads them over HTTP through the host, which persists them in an optional object storage adapter. } session["session
live / staging
reads, writes, presence"] client["client
optimistic overlay"] persistence["persistence adapter
IndexedDB / memory
optional, shared by tabs
"] bus["event bus
WebSocket / in-process"] server["server
authoritative"] storage["storage adapter
SQLite / Postgres / memory"] objects["object storage adapter
R2 / S3 / fs / memory
optional
"] session --> client client -->|journal + cache| persistence client <-->|events| bus bus <--> server server --> storage client <-->|"blob bytes, HTTP
via the host"| objects ``` ## The parts * **Session** — what the application actually holds. Every read and write goes through one, in one of two kinds: a **live** session writes straight through, a **staging** session accumulates a [changeset to commit](/staged-sessions/overview/). Both expose the same document API — prepare a document, read it, write it, publish presence — so an app switches lanes without rewriting its call sites. Framework bindings like the React integration hand you a session too. * **Client** — one per connection, shared by the sessions layered on it. Owns subscriptions and the [optimistic overlay](/concepts/sync-and-optimistic-updates/), and speaks the protocol; a session’s writes land here before they reach the wire. * **Server** — the authority for one [folder](/concepts/documents-and-folders/): validates against [schemas](/concepts/schemas-as-documents/), assigns sequence numbers, persists, broadcasts. Enforces limits (document size, rates) and exposes a change callback for automation. * **Event bus** — routes [protocol events](/architecture/wire-protocol/) between clients and server. Implementations: WebSocket (one each for the Durable Object and Node hosts), [in-process](/architecture/running-in-memory/) (server-side agents, tests, demos), and a composite that combines both on one server. * **Storage adapter** — persistence behind the server: documents, the [event log](/architecture/storage-and-event-log/), and Yjs state. Two persistent implementations pass the same conformance suites: SQLite in a Durable Object (production) and Postgres behind the async server (not yet in production); tests use memory. * **Object storage adapter** — **optional**, and the one place bytes live outside the storage adapter: an object store (R2, S3-compatible, filesystem, memory) holding [blobs](/concepts/blobs/) — files referenced from document JSON by immutable handle. The server keeps the catalog and the authorization; the host streams the bytes over HTTP. Without one the server refuses schemas that declare blob fields and nothing else changes. * **Persistence adapter** — the storage adapter’s client-side counterpart, and **optional**: the pending-write journal and the document cache that let a reload replay its writes and render offline. IndexedDB in browsers, memory in tests. Without one the client is memory-only. See [offline persistence](/concepts/offline-persistence/). Tabs on the same `(folder, principal)` namespace share a persistence adapter, so two offline tabs converge with no server round-trip — the only path in this architecture where clients exchange data without the server. ## Where it runs today The production shape is [Cloudflare Workers + Durable Objects](/architecture/cloudflare-deployment/): one Durable Object per folder, embedding the server, its storage, and its WebSocket connections. The extension points exist precisely so that this is a deployment choice, and a second backend now proves it: [one Node process over Postgres](/architecture/node-deployment/), serving many folders from one database through the async server. It passes the same conformance suites as the Durable Object; it has not yet run in production. # Cloudflare deployment > The current production shape — one Durable Object per folder, SQLite storage, WebSockets. The production backend for datadata today is **Cloudflare Workers with Durable Objects**. It’s a natural fit for the model: a [folder](/concepts/documents-and-folders/) needs exactly one authority that orders its events, and a Durable Object is exactly that — a single-threaded instance with its own storage, addressable by id. ## The shape * **One Durable Object per folder.** The DO embeds the datadata server, so event ordering is free: there’s only one thread that could be ordering them. * **SQLite storage in the DO** via the storage adapter — document snapshots, the append-only [event log](/architecture/storage-and-event-log/), and Yjs state, colocated with the compute. * **WebSockets to clients.** Browsers connect to the DO directly; the WebSocket event bus tracks per-connection subscriptions and broadcasts to exactly the subscribers of each document. * **Auth at the door.** Clients present a token (user, folder, expiry) minted by the host application; the DO validates it before accepting the connection and attaches a principal to it. What that principal may do *inside* the folder is then engine-enforced — see [Authorization](/concepts/authorization/). * **Agents in the DO.** Server-side agents attach as [in-process clients](/architecture/in-process-clients/) inside the same DO, sharing the folder with WebSocket users via the composite event bus. A change callback is synchronous, so an agent’s turn — minutes of model calls — is started from it, not awaited in it. The DO base class’s `runInBackground(work, { label })` is what keeps the object in memory for such work: a promise on its own is invisible to the runtime, which would hibernate the DO mid-turn, and `ctx.waitUntil` has no effect in a Durable Object. While work is pending, the DO keeps a timer scheduled, which rules hibernation out, and its alarm fires every 30 seconds, which rules eviction out. Each firing logs what is pending: the count, the oldest item’s age and the labels, as a warning once the oldest has been pending for five minutes. Each call names its work with a label, which its failure log carries too, and can pass an `AbortSignal`: once it aborts, the object stops staying awake for the work. Hooks that live outside the DO class receive `this.runInBackground` as an argument. It is a bound function, and the package exports its type as `RunInBackground`. * **Blobs in R2, streamed by the Worker.** The Worker mounts the engine’s [blob](/concepts/blobs/) routes next to the WebSocket upgrade and holds the R2 binding: it asks the DO to authorize an upload or a download (metadata only), then streams the bytes itself, so no file body passes through the DO’s memory. One bucket serves every folder through per-folder key prefixes, and the DO’s alarm reclaims abandoned and unreferenced blobs — armed by activity, so an idle folder never wakes for it. ``` flowchart TB accTitle: One Durable Object per folder accDescr { Browsers connect over WebSockets, presenting a token minted by the host application which the Durable Object validates before attaching a principal to the connection. Inside one Durable Object per folder sit the composite event bus, the datadata server, and SQLite holding snapshots, the event log and Yjs state. Server-side agents attach to the same bus as in-process clients. Because one folder is one object, and one object is one thread, event ordering needs no coordination. Blob bytes take a separate path: browsers upload and download over HTTP through the Worker, which asks the Durable Object only to authorize and record each transfer and streams the bytes to and from an R2 bucket itself. } browsers["browsers
WebSocket"] worker["Worker
blob routes, HTTP"] r2[("R2
blob bytes")] subgraph do["one Durable Object per folder"] direction TB bus["composite event bus
per-connection subscriptions"] engine["datadata server"] sqlite[("SQLite
snapshots · event log · Yjs state")] agents["in-process agents"] bus --> engine engine --> sqlite agents --> bus end browsers -->|"token validated
at the door"| bus browsers <-->|"upload / download"| worker worker -->|"authorize + record
metadata RPC"| engine worker <-->|"stream bytes"| r2 ``` ## Properties that fall out * **Region-local consistency** — a folder lives where Cloudflare places its DO; all writes serialize there. * **Scale-out by folder** — thousands of folders mean thousands of small, independent DOs, not one big server. The flip side: a single folder’s throughput is bounded by its single DO. * **Hibernation-friendly costs** — idle folders cost nothing, and a folder with background work pending counts as busy until it settles. That keeps the instance resident without making the work durable — a deploy still ends a turn midway — so a host that marks work as in flight in a document closes out stale marks when a fresh instance starts. Each connection’s identity and subscription list ride its WebSocket’s hibernation attachment, which Cloudflare caps at 16,384 bytes — room for roughly 400 UUID-sized document ids. A connection subscribed to more than fit stays fully live, but the DO closes it on the next wake so the client reconnects and re-sends its subscriptions. # Node + Postgres deployment > The second backend — one Node process serving many folders out of one Postgres database. Conformance-tested against the same suites as Cloudflare; not yet in production. The second backend runs the same server in a **Node process over Postgres**. Supporting it meant rewriting the engine so one core serves both a synchronous storage adapter (SQLite in a Durable Object) and an asynchronous one (Postgres): every operation is written once as a storage-agnostic generator, and each backend supplies the interpreter that runs it. The two backends are then held to the same behavior by the same conformance suites. Not in production The Cloudflare backend is what our apps run on. The Node + Postgres backend passes every conformance suite but has not carried production traffic, and its Postgres driver today is [PGlite](https://pglite.dev/) — Postgres compiled to WebAssembly, embedded in the process, in memory or on a data directory. A `pg` / `postgres.js` driver implements the same small driver interface without touching the adapter, but nobody has written one yet. Part of that interface is saying which of the driver’s errors are transient, so a dropped connection, a timeout or a serialization failure is a transient storage failure, which an edit retries, rather than an error that rolls the edit back. ## The shape * **One process, many folders, one database.** Every table carries a folder id, and a folder registry owns exactly one async server and storage adapter per folder, created on first use. A folder’s operations run one at a time through the async server’s FIFO; different folders’ statements interleave freely. On shutdown the registry drains every folder’s queue before the host closes the shared database. * **Postgres storage** via the async flavor of the storage adapter — the same snapshots, [event logs](/architecture/storage-and-event-log/), Yjs state, and blob catalog as the Durable Object’s SQLite, statement for statement. * **WebSockets to clients** through a Node event bus with the same per-connection subscription bookkeeping as the Durable Object’s, and the same frame cap. * **Identity is the host’s.** A `resolvePrincipal` option turns each incoming request — WebSocket upgrade and blob routes alike — into a principal, the role the Worker plays in front of the Cloudflare backend. What that principal may do inside the folder is engine-enforced, as everywhere: see [Authorization](/concepts/authorization/). * **Blobs in a filesystem or S3-compatible store.** The [blob](/concepts/blobs/) handlers mount on the same HTTP server as the WebSocket upgrade, under the same principal resolver. The byte store is a directory beside the database by default, or any S3-compatible bucket (AWS S3, R2’s S3 API, MinIO) from environment variables. A timer replaces the Durable Object’s alarm for garbage collection: each tick asks the blob catalog which folders in the database have something to reclaim — every folder, whether or not the process has served it — and sweeps those, reclaiming abandoned uploads and unreferenced blobs once they are past the retention window. Folders with clean catalogs cost nothing, and a restart loses no garbage. ``` flowchart TB accTitle: One Node process, many folders accDescr { Browsers connect over WebSockets to one Node process. The host's principal resolver turns each request into a principal. Inside the process a folder registry holds one async datadata server per folder, each with its own Postgres storage adapter, all reading and writing one shared Postgres database whose tables are scoped by folder id. Blob bytes go over HTTP to the same process, which streams them to a filesystem directory or an S3-compatible bucket. } browsers["browsers
WebSocket · HTTP"] store[("filesystem / S3
blob bytes")] pg[("Postgres
PGlite today · tables scoped by folder")] subgraph node["one Node process"] direction TB resolver["principal resolver
host-owned"] registry["folder registry"] a["async server
folder A"] b["async server
folder B"] resolver --> registry registry --> a registry --> b end browsers --> resolver a --> pg b --> pg node <-->|"stream bytes"| store ``` ## Same engine, same limits The Postgres adapter defaults to the Durable Object’s size limits — 1.9 MB per document, per event, and per embedded Y.Doc — even though Postgres has no such row limit, and the limits are overridable per adapter. The default is an application contract: a document valid in one deployment is valid in every other that keeps it, so an [exported stream](/architecture/portable-event-streams/) imports across backends. The Yjs [write-behind](/architecture/storage-and-event-log/#write-behind-for-streamed-text) takes the same approach on both adapters, retry budget included: a burst whose flush fails is retried with backoff before it is dropped, and a dropped burst is reported to the host, since it is data subscribers already saw. ## One behavior on both backends A shared conformance suite — create, delete and restore, rebuild, export and import, purge, blobs — runs against every storage adapter. A document behaves the same whichever backend holds it, and an exported stream moves between them. ## Trade-offs against Cloudflare * **Placement and scaling are yours.** A Durable Object is placed and woken by the platform; a Node process runs where you run it, and a folder’s throughput is bounded by the process it lives in rather than by a per-folder object. * **No scale to zero.** The process runs whether or not any of its folders is active, where an idle Durable Object costs nothing. Presence and subscriptions live in process memory, so a restart is a reconnect for every client — the same [reconnect and replay](/architecture/reconnect-and-replay/) path as a Durable Object eviction. * **One database to operate.** Backups, retention, and the missing [compaction](/known-issues/limitations/) are one Postgres problem instead of thousands of SQLite files. # Wire protocol > The event vocabulary between client and server — small on purpose. The protocol between client and server is a small set of typed events. Application code never touches them directly. All events travel over an [event bus](/architecture/system-overview/) — WebSocket frames from a [Durable Object](/architecture/cloudflare-deployment/) or a [Node process](/architecture/node-deployment/), function calls [in-process](/architecture/running-in-memory/). The vocabulary is identical either way. ## Client → server | Event | Meaning | | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `doc:subscribe` | Start receiving changes for a document — carrying the last seen sequence and [generation](/concepts/documents-and-folders/) (if any) and per-Y.Doc **state vectors**, so each [lane](/concepts/two-sync-lanes/) answers with only what’s missing. That cursor only counts for the generation it names. Server replies with `doc:init` or `doc:resume` (or `doc:notfound`). | | `doc:unsubscribe` | Stop receiving changes. | | `doc:create` | Create a document — id, type, initial data, an optional name, the generation the client minted for its optimistic copy (which the server adopts; a create naming a purged generation is rejected `notFound`), the schema sequence the data was authored against, optionally with embedded Yjs documents. | | `doc:update` | Change a document — a [JSON Patch](/concepts/changes-as-json-patch/) and/or Yjs updates, an optional guard (sequence or patch), the generation and schema sequence it was written against, and the client’s event id for optimistic confirmation. A write naming a generation the document no longer has — its id was purged and reused — is rejected `notFound`; one naming an earlier schema sequence is [brought forward](/concepts/schema-evolution/) through the migrations stamped since. | | `doc:command` | Run a [domain command](/concepts/domain-commands/) — the command’s name and arguments instead of a patch; the server runs the registered mutator over the document it holds and broadcasts the resulting diff as a `doc:patch`. Carries the base sequence, the generation and, for an `exact` contract, a `"sequence"` guard. A refusal comes back as a `doc:error` of category `commandRefused` with the mutator’s code. | | `doc:delete` | Soft-delete a document — it leaves the index and the read path, but its history is retained for restore. Takes an optional sequence guard (“delete only if unchanged since I looked”) and, like `doc:update`, names the generation it was written against. | | `doc:restore` | Undo a soft delete — the document re-enters the index at the exact sequence it left, and subscribers receive a fresh `doc:init`. Like `doc:delete`, it names the generation it was aimed at — the one the client saw deleted, from its `doc:deleted` or the document’s `sys:trash` entry — and a restore naming a generation the document no longer has is rejected `notFound`. | | `doc:purge` | PERMANENTLY destroy a batch of soft-deleted documents (up to 100) and release their ids — a purged id afterwards answers `doc:notfound` like one that never existed, and may be created again. All-or-nothing — one bad or unauthorized target rejects the whole batch and destroys nothing. Requires the docType’s explicit `access.purge` rule (the default is nobody, admins included); success is one `sys:trash` removal patch at one sequence. | | `doc:rename` | Set a document’s name (the name lives in `sys:index`, not the document body). Patches the document’s `sys:index` entry; never touches its `data`, `sequence` or event log. Last-writer-wins, but — like `doc:update` — it names the generation it was written against when the client holds a copy, and a rename naming a generation the document no longer has is rejected `notFound`. | | `doc:get-events` | Fetch one page of a document’s [event logs](/architecture/storage-and-event-log/) — both lanes: the JSON patch events and the Yjs update events, after a per-lane cursor (`afterSequence`, `afterYjsSequence`), at most `limit` rows (500 at most). The answer’s `hasMore` says whether rows remain, and its `generation` which incarnation of the document the page came from; `getAllDocumentEvents` walks every page, starting over if the document is purged and re-created mid-walk. | | `doc:get-processed-writes` | Ask which of the client’s own writes on a document the server has processed — the answer, from the write de-duplication store, is what lets a client re-send safely behind its own later writes (see [reconnect and replay](/architecture/reconnect-and-replay/)). | | `clock:ping` | Ask for the server’s clock — sent only by a client that started its [connection clock](/concepts/shared-clock/). Carries nothing; the client times the round trip itself. | ## Server → client | Event | Meaning | | ---------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `doc:init` | Full JSON snapshot at subscribe time — sequence, generation, type, data; a client holding a copy of another generation of the id replaces it, embedded Y.Docs included. The Yjs lane rides along as exact per-Y.Doc diffs against the subscribe’s state vectors (full states when none were sent), plus tombstone notices and the server’s own vectors. A create/restore commit’s init also carries the document’s listing entry, so it is never readable-but-unlisted while its `sys:index` patch is in flight. | | `doc:patch` | An incremental change — patch and/or Yjs updates, the new sequence, and the originating client event id (so the originator can retire its [optimistic entry](/concepts/sync-and-optimistic-updates/)). May also carry `yjsDeleted` — ids of embedded Y.Docs the orphan GC reclaimed on this write, so subscribers drop those sub-docs immediately. | | `doc:error` | A rejected event — categorized (validation failure, malformed request, undeclared type, failed guard, limit, a refused command with its code), tied back to the client event id. Guard failures and command refusals are expected signals, not faults. | | `doc:notfound` | The subscribed document doesn’t exist. | | `doc:deleted` | The document was soft-deleted — pushed to live subscribers when a delete commits, and the answer to subscribing to an already-deleted document (deliberately distinct from `doc:notfound`). Names the deleted document’s type and generation, which the client keeps with the tombstone: they type a later restore or rename and aim it at the incarnation that was deleted. | | `doc:resume` | The client’s JSON state is current and the document is valid — no data re-transferred. The Yjs lane rides along independently: diffs and tombstones for whatever the client’s vectors were missing, plus the server’s own vectors. | | `doc:processed-writes` | The answer to `doc:get-processed-writes`: the asked-about event ids the server holds as processed, each with the sequence its write was answered at. | | `doc:resync` | Presence-only cold-start nudge — carries just the docId, asking a presence document’s subscribers to republish their cells. See [below](#presence-rides-the-same-events). | | `clock:pong` | The answer to `clock:ping`, to the asking connection only: the ping’s id and the server’s time (Unix epoch milliseconds) as it handled it. The one pair of events that is about the connection rather than a document. | A subscribe is the one exchange where the answer branches, and where the two lanes visibly answer independently: ``` sequenceDiagram accTitle: What a doc:subscribe earns back accDescr { A client subscribes carrying the last sequence and generation it saw and a state vector per embedded Y.Doc. If the document does not exist the server answers notfound, and if it is soft-deleted it answers deleted. Otherwise the JSON lane answers resume when the client's copy is the document's generation at its current sequence, transferring no data, or init with a full snapshot when the client is behind, sent no sequence, or holds another generation — a copy of a purged document whose id was reused, whose state vectors are then ignored. Either answer carries the Yjs lane alongside it as exact diffs against the client's vectors, plus the server's own vectors so the client can push back anything the server lacks. } participant c as Client participant s as Server c->>s: doc:subscribe — last seen sequence and generation
+ a state vector per embedded Y.Doc alt no such document s-->>c: doc:notfound else soft-deleted s-->>c: doc:deleted else client's copy is current — same generation, same sequence s-->>c: doc:resume — no JSON transferred else client is behind, sent no sequence, or holds another generation s-->>c: doc:init — full JSON snapshot end Note over c,s: either answer carries the Yjs lane alongside it:
exact diffs against the client's vectors, tombstones,
and the server's own vectors to push back against ``` ## Presence rides the same events [Presence](/concepts/presence/) adds just **one event of its own** — the `doc:resync` nudge. Otherwise a `sys:presence::` document is subscribed, initialized, patched, and updated through exactly the `doc:*` vocabulary above — a cell write is a `doc:update`, peers learn of it via `doc:patch`, and the per-session `sys:presence-view::` is a client-side projection the wire never sees. What differs is **server-side lifetime, not the protocol**: presence state lives only in memory, so it never appends to an event log (`doc:get-events` answers an explicit error — no history exists, ever) and is dropped when the publishing connection goes. That ephemerality surfaces on the wire in four ways: * **Liveness is a patch, not an event.** When a connection closes, the server doesn’t invent a disconnect event: it removes that connection’s cells, and peers receive the removal as an ordinary `doc:patch`. A reconnecting client simply republishes its cells. The `stale` flag in the view is client-side, with no wire vocabulary of its own: a client marks its peers stale while its own view is unverified (after it reconnects, or after a `doc:resync`) and hides any that aren’t confirmed within the [grace window](/concepts/presence/#liveness-and-the-grace-window). * **Ephemeral Y.Docs re-send in full.** A cell’s [streaming Y.Doc](/concepts/presence/#ephemeral-ydocs--streaming-transient-text) rides the normal Yjs lane inside `doc:update`/`doc:patch`, but because it is never logged it is the one place catch-up isn’t a state-vector diff: on the owner’s reconnect the server can’t reconstruct it, so the owner re-sends the Y.Doc’s full state. * **Cold starts nudge, they don’t reset.** When a transport host wakes from hibernation or restarts, its in-memory cell registry is gone, so it sends each still-subscribed presence document a `doc:resync`, asking its subscribers to republish. It is deliberately *not* an empty `doc:init`, which would clear the peers a client is still showing; instead the client arms its [staleness ledger](/concepts/presence/#liveness-and-the-grace-window) and republishes its own cell, so the roster refreshes without flickering empty. * **Sequences count in epochs.** A presence document’s sequence lives in memory too, and the server lets it go when the last cell leaves, so the next publish starts again at 1. Every presence `doc:init` and `doc:patch` therefore names the epoch its sequence counts in. A client that sees a frame that doesn’t follow its copy (a skipped sequence, or a new epoch arriving while it still holds cells confirmed by the server since its last wake) knows a frame went missing, applies what arrived, and resubscribes for a fresh `doc:init`. Cells retained after `doc:resync` are excluded from that check until a server patch reconfirms or removes them. The new epoch can therefore start at 1 without cutting short the grace period for peers still republishing. A missing frame within the new epoch still triggers a fresh snapshot. ## Design notes * **Binary on the wire.** Every event is one frame: a short header, a JSON envelope, and a table of byte payloads that the Yjs lanes index into, so a Yjs update is read in place from the frame the transport already holds. Both ends speak one frame version and deploy together. * **Every collection is bounded.** A frame is capped at 16 MiB, and the server checks the size of each collection in a client event — patch operations, Yjs lanes, update bytes, purge ids — before parsing it, so an oversized payload is refused up front with `sizeLimitExceeded`. A small patch can still ask for a lot (a `copy` that doubles a value, forty times over), so applying one is bounded too, and refused with the same category. * **Deltas dominate.** After `doc:init`, a subscriber receives patches. The exceptions are existence transitions, not content: a `doc:restore` pushes a fresh `doc:init` to live subscribers (as does the presence re-init that rides along with it), because the document’s re-entry can’t be expressed as a patch against a tombstone. * **One update, two payloads — one sequence.** A single `doc:update` carries structured patches and Yjs binary updates together, so a mixed edit (retitle + type in the body) is one event. But the [lanes version independently](/concepts/two-sync-lanes/): only the JSON half advances `sequence`; a Yjs-only update repeats it unchanged on its `doc:patch` broadcast. * **Errors are addressed, not broadcast.** Rejections return to the sender with its event id; other subscribers never see them. * **Orphaned Y.Docs are reclaimed live.** When a write orphans an embedded Y.Doc — a removed `yjsRef`, or a discarded session [copy](/staged-sessions/changesets-and-lanes/) — the reclaimed ids ride the same `doc:patch` as `yjsDeleted`, so live subscribers tear those sub-docs down at once rather than waiting for their next reconnect’s init/resume. * **Deletion is lifecycle, not content.** A delete never appends to the document’s own event log — it’s recorded in the folder’s membership log, so a delete → restore round-trip returns the document at the exact sequence it left, history intact. Deleted documents move from `sys:index` into the synthesized `sys:trash` document; subscribe to it and a trash UI updates live, with no dedicated listing request. The move is broadcast add-before-remove, so a row switches lists without ever vanishing in between. * **Sequences make hydration cheap.** Document state can be fetched outside the socket — say, server-rendered over plain HTTP — and handed to the client as preloaded state. When the client later subscribes, it sends the sequence and generation it already holds: the server answers `doc:resume` (no data) if nothing changed, or a fresh `doc:init` if the client is behind — or holds another generation of the id, so a document purged and recreated under the same id is never mistaken for the old copy. A document the server’s read [flags invalid](/concepts/schema-evolution/) is also answered `doc:init`, whatever the sequence: the flag travels on `doc:init` alone, so a `doc:resume` always means current *and* valid. JSON catch-up is deliberately a full snapshot, not incremental patches — simple over clever, at the cost of re-sending a document that moved one event. The same mechanism makes reconnects cheap — resubscribing with the cached sequence transfers nothing when nothing moved. * **Yjs catch-up is exact, either way.** The CRDT lane never falls back to snapshot-resending: whether the JSON lane earns an init or a resume, the Yjs lane answers the subscribe’s state vectors with precisely the missing diffs. So a reconnect after pure typing is a `doc:resume` plus a small diff — not a re-send of the document. A subscribe without vectors gets full Yjs states on `doc:init` (and, when the document holds Y.Docs, a sequence-matched subscribe without vectors falls back to a full init, since Yjs currency is unknowable without them). * **Catch-up is bidirectional.** `doc:init` and `doc:resume` carry the server’s own state vectors; a client holding Yjs content the server lacks — offline edits, or a write-behind crash window — pushes exactly the missing diffs back as an ordinary `doc:update`. See [Two sync lanes](/concepts/two-sync-lanes/). Gaps & open questions Writes buffered across a disconnect replay on reconnect in both lanes, but the buffer is in-memory by default — a page reload drops it unless an [offline persistence](/concepts/offline-persistence/) adapter is configured, and even then durability is best-effort; see [Limitations](/known-issues/limitations/). # Reconnect & replay > How buffered writes replay after a disconnect — de-duplication, the replay horizon, and the sweep for unanswered writes. When a client loses its connection, optimistic writes don’t stop — they queue in an in-memory buffer and replay when the socket returns. This page is the exact contract for that replay: how re-sends are de-duplicated, how long the guarantee holds, which writes are exempt, and how a write the server never answered is recovered while the connection stays up. By default the buffer is in-memory; with the opt-in [persisted write queue](/concepts/offline-persistence/) it survives a page reload too, and the reloaded client replays it **under exactly the rules on this page** — a reload is a reconnect, with no separate cold-start semantics. On reconnect the two [lanes](/concepts/two-sync-lanes/) reconcile differently: structured (JSON Patch) writes replay last-writer-wins, while the Yjs lane self-heals by CRDT merge — deltas merge cleanly no matter how late they arrive. ## What the reconnect rebuilds instead of replaying Only writes replay. Nothing else the client does while offline is held for later: the client sends no frames at all while the socket is down, and the reconnect rebuilds what the server needs from the client’s *current* state. * **Subscriptions** are re-sent with their current cursors. A document subscribed while offline is subscribed once, on reconnect. A document released while offline needs nothing, since the new connection never held it. * **Presence** cells are republished with their latest state, so peers never see an older state arrive after it. A cell cleared while offline stays gone. * **Event history reads** (`getDocumentEvents`) reject straight away while offline, the same answer a read in flight gets when the socket drops. Retry once the client is back online. A client starts offline, so the first connect works the same way. What a page subscribes to and writes before its socket first opens goes out on that connect, rebuilt from current state like any reconnect. The socket itself queues nothing: a frame sent while it is down is dropped, and the next connect sends what is still wanted. A host wiring its own transport calls `setConnected(connection)` every time the transport opens, the first time included. The `connection` it passes names that connection: the socket object, or anything else the host has one of per connection. The host passes it again with every event it hands over, `handleEvent(event, connection)`, as the connection the event arrived on. The client drops an event of any other connection, so a frame a replaced socket still delivers changes nothing. Without that, a late snapshot from the old socket would replace a newer copy of its document. Keeping one connection’s events in the order they arrived is still the host’s job. ## Re-sends are de-duplicated, within a bounded window Replay is safe to repeat: the server records each write’s client event id and re-acks a duplicate instead of applying it twice, so a write that committed but whose acknowledgement was lost can be re-sent. This covers every replayable write kind — updates, creates, deletes, restores and renames — so a replay also never re-applies over what happened in between: a delete replayed after someone restored the document re-acks with its live state instead of re-deleting it, a replayed restore doesn’t resurrect a document deleted since, and a replayed rename never clobbers a newer name. That dedup record is retained only for a bounded window. The record is of writes the server accepted. A refused update leaves none, so if its error is lost and the client replays it, the server judges it again against the document as it is by then. A guarded update whose guard failed the first time can pass the second time and be applied. The result is the one the write would have had if it had been delayed in transit, and nothing is applied twice. The client relies on the same rule: an update refused while an earlier one of its own is still unconfirmed is held and sent again behind it, because the refusal may only mean the earlier write had not arrived. A write that was *sent* and then sat unconfirmed longer than a client-side horizon (30 minutes, half the server’s retention) is dropped rather than replayed — but not silently. The optimistic entry rolls back to server truth and the outcome is reported as unconfirmed: an awaited write rejects, a fire-and-forget one hits `onWriteError`, and the caller refetches. That beats the two silent alternatives — a write that vanishes without a trace, or a replay that risks a double-apply once the dedup record has aged out. Writes made purely *offline* never reached the server, so they carry no such risk and always replay, however long you were away. When a write keeps failing with a transient storage error the client retries it a few times; during a systemic outage a circuit breaker stops piling on retries once many writes are failing at once. ## Sequence-guarded writes are exempt [Sequence-guarded](/concepts/changes-as-json-patch/#guards) writes replay at any age: their compare-and-set base means a late replay of a write that already committed is rejected benignly rather than applied twice, so they replay instead of being surfaced. The exception is when later unguarded work was stacked on top of one — an unguarded edit, or a rename, delete or restore of the same document — since all of it was authored assuming the CAS applied, so the whole chain is surfaced instead. ## Unanswered writes get swept too De-duplicated re-sending isn’t only for reconnects. A WebSocket delivers every frame in order for as long as it is open, but a server can still leave a write **unanswered on a live connection**: it fails while handling the write, or commits it and fails before the acknowledgement leaves. Heartbeats still flow, so nothing reconnects and nothing replays. The client therefore sweeps its own buffer: any write still unconfirmed after a bound (10 seconds by default) — an update, a create, a delete, a rename or a restore — is re-sent under its original event id, and keeps being re-sent each period until the server settles it. If the original write actually committed and only its acknowledgement went missing, the re-send draws a dedup re-ack instead of a second apply, and the client resyncs the document to fetch the data that acknowledgement would have carried. The write’s own patch keeps rendering in the local view until that resync lands, so the writer’s committed value never blinks out of its view, and an edit made meanwhile builds on it rather than on the stale base. The same holds when it is a create’s acknowledgement that went missing and an update pipelined behind it is answered first: the update’s effect stays in view over the staged create until the re-sent create’s answer lands. The same 30-minute horizon bounds this path too for updates; a create is never given up on, since a late one reconciles through the benign already-exists answer. A create the server never handled has a second consequence: every update sent behind it is refused as not found. The client treats that refusal as transient when it holds its own unconfirmed create for the document, keeps the update pending, and lets the sweep re-send the create and then the update, so the edit lands rather than being rolled back. The same reasoning covers two edits in a row. Because the server answers a document’s writes in the order they were sent, an answer for the second edit arriving while the first is still unconfirmed means the first was lost. The client then rebases the first past the second rather than re-sending it whole: whatever the second edit already decided is dropped from it, the rest is re-sent, and an edit with nothing left is settled as committed. Where the first edit wrote a whole value and the second changed something inside it, the first keeps that value with the second edit’s change folded in, so its re-send lands what the author saw rather than undoing the second edit. A rejection of the second edit in that situation is held rather than acted on, since it may only say that the patch met a document the lost edit never reached; the sweep re-sends both in order and the server answers again. When no answer for the second edit has arrived either, the client cannot tell whether that edit was lost too or was applied with only its acknowledgement lost, and re-sending both blind would be wrong in the second case: the duplicate of the second edit is de-duplicated, the first lands after it, and the earlier value wins on the server. So before re-sending an edit behind a later one it has already sent, the client asks (`doc:get-processed-writes`) which of the document’s unconfirmed edits the server processed. The answer comes from the same de-duplication store and names each processed edit with the sequence it committed at, so the client settles those exactly as their re-acknowledgements, which rebases the edits staged before them, and then sends whatever is left, in order. Edits made while the question is out wait behind them, and a question that goes unanswered for a period is asked again. The reconnect replay asks the same question for a document whose buffered edits include one sent before the socket dropped, instead of replaying them blind. Lifecycle operations are protected differently: a document’s create, deletes, renames and restores go on the wire **one at a time**. A delete staged on an unacknowledged create, or a restore staged on an unacknowledged delete, is recorded locally at once (the document reads deleted, the awaited promise is pending) but is sent only when the operation ahead of it has been answered. An unanswered operation can therefore neither strand the operations behind it, which were never sent, nor be overtaken by them when the sweep re-sends it. The cost is one round trip per operation in such a chain, which is rare; edits stay pipelined. A delete goes one step further: it also waits for the document’s own unconfirmed edits. Staging it keeps them (pending Y.Doc changes are flushed ahead of it), so an edit left unanswered is re-sent by the sweep and commits before the tombstone, which means a delete-then-restore undo brings it back. And a delete the server refuses, whether a guard miss or an authorization the client could not predict, leaves the document live with nothing lost. The document reads as deleted locally the moment the delete is staged either way. A delete from **another** client gets no such preparation — it can land while this client is still holding unconfirmed work — so the tombstone keeps that work rather than dropping it. Pending JSON edits stay for the server’s own answer to their re-send. The document’s Yjs content stays too: the live `Y.Doc`, its unflushed changes and its “changes pending” claim all survive the tombstone, held unsent while it stands (an update on a deleted document is refused, and a `Y.Doc` change cannot be rolled back once it is local) and released by the restore that undoes the delete — which re-sends the held changes and pushes back anything else the restored document is missing. Only evidence that the content can belong nowhere discards it: the document purged, or hidden from this client, or an answer naming a different [generation](/concepts/documents-and-folders/) of the id — a document that reused it. # Offline across tabs > How several tabs share one persisted store — journaling, adoption of dead tabs, the Yjs cache bus, and the mirror-never-merge rule for pending JSON writes. Tabs on the same folder and user share one persisted store. This page is the contract for how they coexist: how each tab’s writes stay durable, how a dead tab’s writes are picked up, and how live edits reach a sibling with no server in the loop. It assumes [offline persistence](/concepts/offline-persistence/) — the journal and the document cache are what all of it runs on. ## Every tab journals; the dead ones are adopted Every tab is **durable**: each session journals under its own owner id and holds a Web Lock named for it, so live tabs never collide. The lock doubles as a liveness signal — when a tab goes away its lock frees and a surviving session **adopts** its records, on the next page load or through a periodic adoption sweep (every 15 seconds by default), so a dead sibling’s unsent writes replay without waiting for anyone to reload. Adoption is crash-safe in the store itself, and a frozen tab misjudged as dead only produces a double-replay the protocol absorbs: replayed writes re-acknowledge idempotently under their original event ids. ## Offline tabs share Yjs edits through the cache For the Yjs lane — and only that lane, since only it is a CRDT — the shared cache doubles as a **local bus**. A record landing in it is broadcast to the namespace’s other sessions, and each sibling merges its Yjs snapshots into the live documents it already knows. Since every flush re-encodes those live documents, unsent edits included, **two offline tabs editing the same rich-text document converge with no server round-trip**: type in one, watch it render in the other, both still offline. Merging is what makes duplicate or racing deliveries benign, and merged content is as durable as native content — the receiving tab’s snapshot includes it, so **either** tab’s reconnect pushes the edits to the server. An edit survives its author’s tab dying offline, without waiting for adoption. A receiving tab that is online sends them itself. ## Pending JSON writes mirror between tabs The write journal broadcasts the same way, but pending JSON writes take the opposite discipline: **mirror, never merge**. They aren’t CRDTs — they are optimistic patches against a sequence, order-sensitive and rejectable, applied by the server in whichever order the tabs reconnect. So a sibling’s record is *rendered* — base, own pending, sibling pending — but never replayed, restamped, or merged into the receiving tab’s queue. Ownership moves exactly once, through adoption, when the author dies. When the author instead reconnects and its write acks, the record clears and the confirmed content arrives over the cache bus, which only advances the render forward. ``` flowchart TB accTitle: Two tabs over one shared store accDescr { Each tab journals its own writes into a store shared by every tab on the same folder and principal, holding a Web Lock as a liveness signal. Records landing in the store are broadcast to sibling tabs, where the two lanes are treated differently: Yjs records are merged into the sibling's live documents, so two offline tabs converge with no server round-trip, while pending JSON writes are only rendered and never merged into the sibling's queue. If a tab dies its lock frees and a surviving tab adopts its unsent records. } a["Tab A
own journal · own Web Lock"] store[("shared store
write journal + document cache
one per folder + principal")] b["Tab B
own journal · own Web Lock"] a -->|"journals every write"| store store -->|"Yjs records:
merged"| b store -->|"pending JSON writes:
rendered, never merged"| b a -.->|"A dies — B adopts
its unsent records"| b ``` The picture is symmetric — B journals and broadcasts to A the same way — and the server sits outside it entirely: whichever tab is online replays to it, and merged content is as durable as native content, so either tab’s reconnect carries the other’s edits. All of it stays invisible to the app: the pending-write count and the status document answer for the whole *namespace*, so a “changes pending” indicator is truthful across tabs without the app learning whose queue holds a write. Live edits, offline creates (discoverable by name, since they surface in `sys:index` reads), deletes, and staged-session work are all tab-coherent offline. Two caveats are inherent. Mirrored writes are another session’s **unconfirmed** work, and can bounce on that session’s reconnect exactly like your own optimistic writes. And **identical renders across tabs are not guaranteed while offline**: each tab composes its own pending writes first and the mirror on top, so disjoint edits render the same union everywhere, but two tabs editing the *same field* each keep showing their own value until a reconnect lets the server settle the order. The divergence is render-only — it changes nothing either tab journals or replays — and it is architectural rather than a bug: each tab is a full client with its own optimistic queue, and N queues over one base have no offline arbiter. Whether one should exist — a leader tab or SharedWorker owning a single queue — is tracked in [open questions](/known-issues/open-questions/). # Storage & event log > Snapshots for reads, an append-only event log for history — and no compaction yet. The server persists through a **storage adapter**; the production implementation is SQLite inside each folder’s Durable Object, and a [Postgres adapter](/architecture/node-deployment/) behind the async server is the second. The layout is simple and worth knowing because it explains several behaviors. Adapters come in two flavors, one engine. A **synchronous adapter** (in-memory, Durable Object SQLite) backs `DatadataServer`, whose operations complete on the caller’s stack. An **asynchronous adapter** — one whose methods return promises, like the Postgres adapter — backs `AsyncDatadataServer`, the Promise-returning twin: the same engine code executes every operation (the internals are written once as storage-agnostic generators), with the async server adding an internal FIFO so operations still run one at a time, in call order, however long each storage round-trip takes. A synchronous adapter satisfies the async interface as-is. The FIFO’s scope is one server, and one server serves one folder. A host that serves **many folders from one process** — many folders over one Postgres database — runs one async server per folder, owned by a `FolderRegistry`: it guarantees exactly one live server (and storage adapter) per folder, keeps different folders’ operations from serializing against each other, and at shutdown drains every folder’s queue before the host closes the shared database. ## What’s stored * **Document snapshots** — current JSON, type, and sequence per document. Reads and `doc:init` are served from here; nothing is replayed. * **The JSON event log** — every accepted structured change, appended as its [JSON Patch](/concepts/changes-as-json-patch/) with its sequence number and the attribution of whoever wrote it. This is the document’s history: an audit trail, the source for `doc:get-events`, and the raw material for future history features (diffs over time, undo, blame). * **The Yjs lane** — current Yjs binary state per embedded text document, plus its own append-only update log under its own per-document version, `yjs_sequence`. It mirrors the JSON lane exactly: the document’s `yjs_sequence` is its latest Yjs event’s, just as `sequence` is its latest JSON event’s. Each stored update carries its own attribution; a Y.Doc deletion is a tombstone row. The [two lanes are independent](/concepts/two-sync-lanes/): a Yjs update never touches `sequence`, the JSON log, or the snapshot. * **The blob catalog** — one row per [blob](/concepts/blobs/) (size, content type, SHA-256, lifecycle state, attribution) plus the document → blob reference edges, in the same adapter as the documents so an edge commits in the transaction of the write that created it. The bytes themselves live in a separate **object storage adapter** (R2, S3, filesystem, memory); nothing about them enters either event log. ## Attribution Every stored event, on both lanes, records the principal that wrote it: * **`subject`** — the principal’s subject, the user or service on whose behalf the write happened. Null for an anonymous write. * **`actor`** — what kind of code was acting: `user` (a frontend acting for the user), `agent:` (AI-generated), or `sys:` (server-side system code, e.g. `sys:schema-seed`). The two are independent, and neither confers privilege — [identity is never authority](/concepts/authorization/). Together they make “which edits did the AI make?” a question you answer from the data rather than from application logs. Attribution is written once and never rewritten. An [imported event](/architecture/portable-event-streams/) keeps its original author’s attribution rather than the importer’s, which is what lets a replayed stream still name the author of each event. It’s also why a write-behind burst stays single-author: a different author’s delta closes the open burst and starts a new one. Events stored before the authorization layer existed carry a null actor. Attribution is durable and belongs to accepted history. Staged changes carry no authorship of their own — staging is transient (the changeset clears at commit), so authorship is recorded here, on the committed event, not on the staged change; see [Changesets & lanes](/staged-sessions/changesets-and-lanes/). ## Consequences * **Reads are O(document), not O(history).** Snapshots mean a subscriber’s init cost doesn’t grow with a document’s age. * **History is first-class.** Because changes are stored as patches (not opaque blobs), history is inspectable — by devtools, by audit, by agents — and replayable elsewhere: see [Portable event streams](/architecture/portable-event-streams/). * **Writes are serialized per folder.** One Durable Object, one writer — the storage layer never sees concurrent writes to a folder. A multi-folder host gets the same property from one async server per folder; only different folders’ statements interleave, and every statement is folder-scoped. * **The logs are the recovery authority.** A document’s stored data is, by invariant, equal to the replay of its own logs — the JSON data from the patch log, each Y.Doc from the fold of its Yjs log. Snapshots are served for speed but are reproducible, and verify/rebuild/export/import tooling exists against that guarantee: a document export ships both lanes, and replay folds each lane in its own order. (Even [schema migrations](/concepts/schema-evolution/) applied on read are frozen into the log as ordinary events, so replay never re-derives them.) ## Write-behind for streamed text The Yjs lane is also where the one deliberate durability relaxation lives. Streamed CRDT deltas — an AI typing into a Y.Doc, a fast human — would mean a storage write per keystroke, so both persistent adapters **buffer a burst in memory and flush it as one merged transaction** into the Yjs log. A burst is single-author, so the merged rows keep correct attribution. The fold **shortens the log rather than perforating it**: the merged-away updates are never written, so the surviving rows take the next positions and stay contiguous, and the document’s `yjs_sequence` follows the log down with them. Coalescing is therefore invisible from outside storage — nobody can count how many keystrokes went into a merged row, and nobody needs to. The JSON lane never does this: data changes keep strict persist-before-broadcast ordering. The cost is a bounded window where a burst has been broadcast but not yet stored. A crash inside it loses the burst atomically (disk reverts to the pre-burst state; the document stays self-consistent) — but no longer permanently: any client that saw the broadcasts [pushes the missing content back](/concepts/two-sync-lanes/) on its next reconnect. Gaps & open questions The event logs — JSON Patch and Yjs alike — are **append-only with no compaction or pruning**. Long-lived, busy documents grow without bound; `doc:get-events` pages through them, but nothing ever shrinks them. This is the most concrete known scaling gap; see [Limitations](/known-issues/limitations/) and [Open questions](/known-issues/open-questions/) for the directions under consideration. # Portable event streams > A document's history is an independently replayable stream — tagged with schema versions, interpretable anywhere. The [event log](/architecture/storage-and-event-log/) is more than audit. It’s designed so that a folder’s history can be **consumed somewhere else** — a replica, a pipeline, another system — with guarantees that are deliberately narrow and therefore easy to honor. ## The guarantees **Per-document, per-lane total order.** A document’s history is two append-only streams, one per [sync lane](/concepts/two-sync-lanes/): the JSON patch events ordered by sequence number, and the Yjs update events in their own log order. Replaying the [JSON patches](/concepts/changes-as-json-patch/) in order reconstructs every state the structured data has ever had; folding the Yjs events reconstructs each embedded Y.Doc. A document export ships both streams, and an import replays each in order. The document’s [blobs](/concepts/blobs/) ride along outside the streams — metadata plus bytes, re-uploaded on import — since a blob has no history to replay, only content to carry. An exported Yjs event carries **no number at all**. Order is the only thing the Yjs log’s numbering ever meant — replay folds by order, and only across a delete-then-recreate of the same Y.Doc does even that matter — so the array carries it, and the importing store assigns its own. Nothing in the stream is a number the importer must honour, which is what makes it portable. **No cross-document ordering — on purpose.** datadata does not guarantee any event ordering *between* documents. That sounds like a weakness; it’s the property that makes the streams portable. Each document’s stream is self-contained: you can fetch it, ship it, and apply it without knowing anything about any other document in the folder, and replicas can consume documents independently, at different paces, without coordination. **Every event is tagged with its schema version.** Each event records the sequence of the governing `sys:schema:` document in effect when it was written. And since [schemas are documents](/concepts/schemas-as-documents/), the schema’s own history is *also* an event stream. Sync a document’s stream plus its schema document’s stream, and a remote location can interpret every historic version of the document with the exact schema it conformed to — a complete, semantically meaningful picture of the data over time, reconstructed with no access to the origin. **Migration events are ordinary events.** When [schema evolution](/concepts/schema-evolution/) rewrites a document on read, the rewrite is appended to the stream as a normal patch event. A consumer replays it like any other change — it doesn’t need to implement the migration engine to follow along. ## The folder boundary The schema version recorded on an event is the schema document’s sequence *within its folder* — a folder-local coordinate, not a global schema identity. There is no global schema timeline that every folder re-coordinatizes: a schema may evolve through versions A → B → C in source code, but a folder whose first boot [upserted](/concepts/schemas-as-documents/) C directly has no record of A or B at all — each folder holds only the slice of the evolution it actually lived through, and two folders can hold different, incomplete slices. Within a folder this is fully consistent: every event’s stamp resolves against that folder’s own schema history, so a folder’s streams are always coherent with each other. Portability is designed around that boundary. A **whole folder** (document streams plus schema-document streams) replays faithfully, and so does **part of one** — a subset of documents together with the schema documents they depend on, imported schema-first into a target that holds no independent history for those types. Moving a document into a folder whose schema evolved on its own is what datadata is **not** designed for, and import fails safely at that boundary. It validates *content*, at the right coordinate: the target folder’s schema document is replayed to the imported document’s conformed sequence, and the document’s reconstructed snapshot must validate against that historic schema. A faithful same-lineage export always passes — a stored snapshot conforms to the sequence it’s stamped with, even a document that has become invalid against the *current* schema — so a rejection means the events were produced under a schema the target folder never held, and the import is refused even when the two folders’ sequence spaces happen to overlap. An explicit relaxed mode accepts the document anyway (the replay is still faithful); its first read then flags it invalid like any other non-conforming document. Relaxed even covers a document stamped *beyond* the target’s schema history — it is checked against the newest schema the target holds instead. Only a type with no resolvable schema at all is always refused: that document could never be read. Making cross-folder transfer actually *work* remains a remap problem (match reconstructed schema content, not sequence numbers), and no remap tool exists yet. Purge’s guarantee stops at the folder boundary: it destroys *this folder’s* copy and releases the id — it cannot recall exports already taken. An export from before the purge still replays faithfully, and importing it afterwards is not a resurrection to refuse: it is an ordinary create of a new document under a freed id. ## What this enables * **Read replicas and mirrors** — feed a folder’s documents into another system, region, or store by tailing per-document streams. * **Analytics and pipelines** — events are inspectable JSON patches with schema context, not opaque blobs. * **Point-in-time reconstruction** — any historic version of any document, with the schema that governed it. ## What this is not This is **replication, not offline sync**: a one-way replay of server-confirmed history. Authority stays with the folder’s [server](/architecture/system-overview/) — a replica can’t write events of its own and reconcile later — a replica is read-only. That’s a different mechanism from a disconnected *client* catching up: offline client writes replay from the client’s own queue [on reconnect](/architecture/reconnect-and-replay/) — the JSON lane last-writer-wins, the Yjs lane merged by CRDT — and that queue can [outlive a reload](/concepts/offline-persistence/). This page describes how confirmed history travels. Gaps & open questions Consumers must genuinely not assume cross-document order — e.g. a reference may point at a document whose creation you haven’t applied yet, so apply [integrity rules](/concepts/references-and-integrity/) after convergence, not during replay. The Yjs stream is binary CRDT updates — replaying text requires Yjs, unlike the self-describing JSON patch lane. An export is **portable, not canonical**. A store whose [write-behind](/architecture/storage-and-event-log/) coalesces a burst logs one merged Yjs event where a store that doesn’t logs several, so the same document exported from two stores can differ in event count and in bytes. Both replay to the same Y.Docs, and the JSON lane is unaffected. So don’t diff two exports to compare documents — import them and compare the documents. And event retrieval is unpaged today, with [no log compaction](/known-issues/limitations/) — long histories arrive whole. # Running in memory > The whole engine — client, server, storage — in one process. Tests, demos, and server-side agents run this way. Because the [two extension points](/architecture/system-overview/) — event bus and storage adapter — both have in-memory implementations, a **complete datadata stack runs in a single JavaScript process**: real server, real clients, real optimistic updates, with microtask latency and no infrastructure. ## What it’s for * **Server-side agents.** The production use: [in-process clients](/architecture/in-process-clients/) attach agents directly to the server inside the Durable Object, alongside WebSocket users. * **Tests.** datadata’s test suite spins up real client/server pairs per test — multi-client convergence, conflict, and session scenarios run without mocks or network. * **Demos in the browser.** The [live demos](/demos/) on this site run the full engine in the page: several “clients”, one “server”, a latency slider standing in for the network. Nothing is simulated except the wire. The one thing the in-memory stack leaves out is the [blob](/concepts/blobs/) lane: a blob’s download URL is a plain string meant for ``, which no in-process transport can serve, so the backendless client has no blob support and its server refuses schemas that declare blob fields. ## The local mode isn’t a mock It’s easy to make a sync engine that *only* works distributed (everything needs the real backend) or one whose local mode is a mock that lies to you. Keeping one event vocabulary and one server core for both is what makes the in-memory mode faithful — what you observe in a test or a demo is the production code path, minus TCP. # In-process clients > Run a real client in the same process as the server — for server-side agents, bridges to outside data, tests, and demos. datadata’s client and server are connected by an **event bus** abstraction, not by a socket. One implementation routes events over WebSockets; another — the **in-process event bus** — dispatches them as plain function calls within a single JavaScript process. ## A real client, no network `connectInProcessClient` wires a full client to a full server in the same process. It is not a mock: the same optimistic-update machinery, the same validation, the same event ordering — just with microtask latency instead of network latency. ## Server-side agents The motivating use: an AI agent running **next to the server** (in the same [Durable Object](/architecture/cloudflare-deployment/) or [Node process](/architecture/node-deployment/) as the folder’s server) as a first-class client. A composite event bus lets in-process clients and WebSocket clients coexist on one server — so the agent stages edits through an in-process [session](/ai-agents/sessions-as-agent-workflow/) while humans watch the same live documents from their browsers, each seeing the other’s changes in real time. The server also exposes change callbacks — `onDocumentChange`, invoked on every create, update and delete, `onDocumentRestored`, invoked on every restore, and `onDocumentsPurged`, invoked once per purge — which is how agent-triggering, auditing, and downstream automation attach without polling. The callbacks are synchronous: long work is started from one, not awaited in it, and on Cloudflare it goes through the Durable Object’s `runInBackground` so the object outlives the event that started it (see [Cloudflare deployment](/architecture/cloudflare-deployment/)). They report exactly what reached storage. A write succeeds once it is stored: if a later step fails — telling subscribers, say — the server logs it and recovers the way it would from a dropped message, and the call still succeeds and the callback still fires. A call that rejects stored nothing, so it is always safe to retry, and a mirror or audit log fed from the callbacks never misses a stored write. ## Syncing with data outside the folder A [folder](/concepts/documents-and-folders/) is the unit of authority — one server owns it and orders its events. Applications routinely need it reconciled with data that lives elsewhere: your main application database, an external API, a system of record that predates the folder. Getting history *out* needs no client — that is what [portable event streams](/architecture/portable-event-streams/) are for, a one-way replay into another store. Getting changes back *in* is what needs one. Writes from outside have to enter through a client, so they are validated, ordered and broadcast like anyone else’s. Where that bridge runs decides what it can see. Outside the folder it connects over a WebSocket, the same way a browser does — and like a browser it [subscribes per document](/concepts/documents-and-folders/), so it only learns about documents it already knew to watch. Inside the folder’s own server process it can register `server.onDocumentChange` instead: every document write in the folder, subscribed or not, with the document in hand. Restores come through `server.onDocumentRestored`, which names the document but doesn’t carry it, so a mirror reads the restored document itself. A [schema migration](/concepts/schema-evolution/) isn’t one of those writes: a document is brought forward when it is next read (or, if a client subscribes to it, right after the schema write), and only the schema change fires the hook, so a mirror storing document data re-reads that type’s documents when its schema changes. The folder listings — `sys:index` and `sys:trash`, where renames show up — don’t fire it; a mirror that tracks them subscribes an in-process client to them, which keeps them current from patches. That, rather than the absent socket, is why a mirror wants to run in-process. The mapping between your documents and your tables stays application-specific. datadata supplies the client and the callback, not the bridge. ## Tests and demos The in-process bus is also why datadata is easy to exercise: the test suite runs real client/server pairs with no infrastructure, and the [live demos](/demos/) run the same way — a complete sync engine, client and server, inside your browser tab. In [Contingency](/demos/contingency/), the game’s simulation is an in-process client of its own: it moves the enemies, pays out gold, and writes the waves and lives. # Where datadata sits > datadata described along the dimensions the local-first community uses to compare sync engines. The local-first community has converged on a fairly standard set of dimensions for describing sync engines (as popularized by the [Local-First Landscape](https://www.localfirst.fm/landscape)). Here is datadata, answered along those lines. ## At a glance | Dimension | datadata | | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Maturity | Early-stage / pre-alpha. Working, actively reshaped, not open source yet. | | Distribution | TypeScript library (client + server), self-hosted. No hosted service. | | Data model | Documents — typed JSON with embedded Yjs text. [Schemas are documents too](/concepts/schemas-as-documents/). | | Schema management | Schema documents (optionally authored in TypeScript — typed, no codegen — and upserted at startup) with append-only migrations; documents migrate on next read with write-back; invalid documents flagged, not dropped; dynamic runtime schemas. | | Authority | Centralized — one server per [folder](/concepts/documents-and-folders/) orders all events. | | What syncs up | [JSON Patches](/concepts/changes-as-json-patch/) (optionally guarded) + Yjs updates. | | What syncs down | Initial snapshot, then patches + Yjs updates; reconnects catch the CRDT lane up with exact [state-vector diffs](/concepts/two-sync-lanes/). | | Replication granularity | Per-document subscriptions within a folder; no partial replication inside a document, no cross-folder sync. | | Data size | Documents (and single events) cap at 1.9 MB by default on both persistent backends — the Durable Object’s row limit, kept by the Postgres adapter for portability; model larger data as more documents. | | Conflict handling | [Layered](/concepts/conflicts/): server ordering + opt-in guards for structure; Yjs CRDT for rich text; three-way staged previews in sessions. | | Optimistic updates | Yes — core mechanism, with clean rejection. | | Offline reads | Whatever the client holds in memory; with the opt-in [persistence adapter](/concepts/offline-persistence/), a persisted document cache renders cached documents on an offline boot (`offline-cached` sync status) with pending writes re-projected on top. Schemas and the index are pinned, so validation works offline too. | | Offline writes | Writes made while disconnected replay on reconnect (structured writes last-writer-wins; locally held Yjs content [merges back](/concepts/two-sync-lanes/)); an opt-in [IndexedDB write queue](/concepts/offline-persistence/) makes the buffer durable across a page reload — a reload becomes a reconnect, under the same de-dup/horizon rules. | | Local query | Document reads + subscriptions — no query language, no cross-document select/project. An experimental projection layer reshapes a (fully synced) document into ergonomic views; writable lenses are alpha. For queries, the supported pattern is a derived store: project changes into another document or database and query that. | | Server persistence | Pluggable adapter: SQLite in a Durable Object (production) or Postgres behind an async server ([Node host](/architecture/node-deployment/); conformance-tested, not yet in production); snapshots + append-only event logs ([no compaction](/known-issues/limitations/)). | | Transport | WebSockets; [in-process bus](/architecture/running-in-memory/) for server-side agents, tests, demos. | | Auth | Token at the door (minted by host app), then engine-enforced [authorization](/concepts/authorization/) inside the folder: declarative role rules per docType, per-document access entries in `sys:access`, whole-document read filtering, scope caps for agents — and capability-aware clients that predict verdicts locally. | | Encryption | Transport-level only. No E2E encryption — incompatible with server-side validation as designed. | | Rich text | Embedded Yjs documents; ephemeral presence documents (`sys:presence::`, schema-validated) with a y-protocols awareness bridge over a cell’s `sys:awareness` field for editor cursors — in live documents and in [staged sessions](/staged-sessions/staged-rich-text/). | | Client platforms | Browser + server-side JS. React (Jotai) bindings; devtools panel. | | AI agents | A design focus: [sessions, attribution, in-process clients, agent-writable schemas](/ai-agents/why-agents-like-datadata/). | ## The one-sentence differentiator Everything — user data, schemas, the engine’s own state, staged work — is a document behind one small referential API, with an explicit staging layer (sessions) designed for humans and AI agents proposing changes to shared data. If those dimensions read like a fit, continue to [Design decisions](/comparison/design-decisions/); if some of the “No”s above are deal-breakers, [When to use it](/introduction/when-to-use/) names better-suited engines without hard feelings. # Design decisions > Why JSON Patch and not CRDTs, why schemas are documents, why server authority, why sessions and not branches. For sync-engine developers: the decisions that define datadata, what they buy, and what they cost. Disagreement is the point of publishing this — [argue with us](/contact/). ## Server-authoritative, not convergent **Decision.** One server per folder orders every change. No distributed merge for structured data. **Why.** Authority makes the hard things simple: validation has a place to stand (a write is checked against *the* current state, not a possible one), sequence numbers make history linear, and “what does the document say” has one answer. The apps datadata is honed against — collaborative tools with agents proposing changes — need *reviewable* conflicts more than they need serverless merging. **Cost.** No P2P, no E2E encryption, no CRDT-style merge of structured data (offline edits replay last-writer-wins, not a three-way merge). These aren’t deferred features; they’re the price of the model, paid knowingly. ## JSON Patch for structure, CRDT only for rich text **Decision.** Structured changes are RFC 6902 patches with opt-in [guards](/concepts/changes-as-json-patch/); only rich-text fields use CRDTs ([Yjs](/concepts/rich-text-with-yjs/)). **Why.** CRDTs guarantee convergence, not *intent*: two structurally valid merges can still be semantically wrong, and nobody gets asked. Patches are legible — to humans reviewing, to agents reasoning, to the [event log](/architecture/storage-and-event-log/) as audit — and guards turn “someone else changed this” into an explicit, handleable signal. Rich text is the exception because per-keystroke intent really is captured by CRDT merge semantics. **Cost.** Concurrent structured edits can reject and need retry or resolution. RFC 6902 itself has sharp edges we’ve had to design around — written up in [JSON Patch RFC issues](/known-issues/json-patch-rfc/). ## Changes are data, not domain events **Decision.** One format — the JSON Patch — is both the wire format for changes and the stored [event format](/architecture/storage-and-event-log/). There is no application-defined event vocabulary, and no reducer code that interprets events into state. **Why.** The [one-timeline principle](/introduction/one-timeline/), applied to changes. Classic event sourcing splits a change’s meaning between the event data and the source code that applies it — so the repository’s history becomes a second timeline again: replaying last year’s events through this year’s reducers is a synchronization problem, and every context that consumes the log must carry a compatible implementation. A patch needs no interpreter: the change *is* the data, and applying it is defined by a public RFC, not by the application. That makes history **bug-compatible** — if a mutation wrote something the wrong way, the log records exactly what happened and replays it identically forever, rather than being silently reinterpreted by newer code. And because one standard covers all synchronization, any context — a [replica](/architecture/portable-event-streams/), a pipeline, a debugger, an agent — can read and apply updates without knowing the schema or the application that produced them. **Cost.** Patches record *what* changed, not *why*. A domain event named `taskCompleted` carries reasoning; a `replace` at `/status` doesn’t, so intent that matters has to be carried by the app alongside the change. Reducer-style derived state and richer replay semantics (the strengths of event-sourcing engines) are off the table. ## Schemas are documents **Decision.** A document type’s schema lives at `sys:schema:`, edited through the same API as data. **Why.** It collapses an entire category of machinery — migration tools, admin APIs, deploy-coupled schema changes — into the machinery that already exists. Schemas sync live, version like data, and are [writable by agents](/ai-agents/why-agents-like-datadata/). The meta-schema keeps it from being anarchy. And it doesn’t cost TypeScript: schemas [authored in code](/concepts/schemas-as-documents/) are upserted into the schema documents at startup, with compile-time types inferred from the schema itself — no codegen. The deeper reason is **[history in one timeline](/introduction/one-timeline/)**. Code-only schemas split a document’s meaning across two histories — the repository’s and the data’s — so “what did this document mean six months ago?” requires correlating git commits with deploy times with document versions, and a field removed from a code schema is recorded nowhere except a commit. With schemas as documents, the schema’s history lives in the same event log as the data it governs: point-in-time introspection has one place to look, and a folder’s [exported history](/architecture/portable-event-streams/) carries its own interpretation with it instead of depending on a repository somewhere else. **Cost.** The engine must handle data written under old schema versions indefinitely ([documents migrate on their next read](/concepts/schema-evolution/), so cold documents keep old shapes), and runtime schema changes are a correctness surface that static registries never have. ## Sessions, not branches **Decision.** Staged work is a [changeset in a host document](/staged-sessions/overview/) — flat lanes of guarded changes plus per-stage metadata — not a fork of the folder’s history. **Why.** Branching a document store invites the full weight of merge semantics for a use case — “propose, review, commit” — that needs much less. A changeset *is itself document data*: it syncs, survives reload, supports multiple authors, and is observable through [sys:session / sys:stage](/staged-sessions/sys-session-and-sys-stage/) with zero new machinery. Conflicts are [derived](/staged-sessions/conflict-preview-and-resolution/) by folding the staged patches — each carrying its own `test` guards — onto the live head: no rebase graph, and no document snapshot in the changeset (the three-way preview’s base is resolved on demand, replayed from the event log when a cold client needs it). **Cost.** One level of staging, not arbitrary history surgery: no branches of branches, no cherry-picking across folders. So far the apps haven’t missed it. ## One API for everything **Decision.** Subscribe / read / create / update, over documents — including system state. No admin API, no separate metadata service. **Why.** Every concept added to a sync engine’s surface gets paid for by every application — and every [agent prompt](/ai-agents/why-agents-like-datadata/) — forever. Synthesizing engine state *as documents* means introspection, devtools, and review UIs are just subscribers. **Cost.** Read-models must be synthesized carefully (and documented as read-only), and “everything is a document” is a discipline that has to be defended in design review against every convenient exception. # Known issues & open questions > The gaps, in the open — aggregated here, flagged inline where you'd hit them. datadata is early-stage software. Rather than hiding the gaps behind a roadmap page, this site does two things: * **Flags issues in place.** Every page that touches a known gap carries a callout at the point where you’d hit it. * **Aggregates them here**, with the full write-ups. ## The sections * **[JSON Patch RFC issues](/known-issues/json-patch-rfc/)** — concrete problems we’ve hit applying RFC 6902 in a collaborative engine, and the workarounds baked into datadata’s design. * **[Limitations](/known-issues/limitations/)** — what’s missing, each classified as *by design*, *not yet*, or *open question*. * **[Open questions](/known-issues/open-questions/)** — the things we’re genuinely unsure about. If you’ve solved one of these in your engine, [we’d like to hear how](/contact/). ## Why publish this Two reasons. Application developers evaluating a sync engine deserve the failure modes up front — discovering them in production is how trust in this whole category dies. And sync-engine developers are the people most likely to have useful opinions about exactly these gaps; consider every page in this section an invitation. # JSON Patch RFC issues > Where RFC 6902 fights a collaborative sync engine, and what datadata does about it. datadata uses [JSON Patch (RFC 6902)](https://datatracker.ietf.org/doc/html/rfc6902) as its [structured change format](/concepts/changes-as-json-patch/). The RFC was designed for HTTP PATCH — one writer, one round-trip — and applying it in a concurrent, multi-writer engine surfaces real problems. This page collects the ones we’ve hit. ## Array operations address positions, not items **The issue.** RFC 6902 paths address array elements by index (`/items/3`). An index is only meaningful against the exact array the patch was computed from: any concurrent insert or remove earlier in the array silently retargets every following operation. The patch still *applies* — to the wrong elements. The `-` append token has the mirror problem: two concurrent appends both “succeed” with no way to express ordering intent between them. And it is worse than a retargeting risk, because a differ working from positions cannot see an insert at all. Inserting `{id: "z"}` at the head of `[{id: "x"}, {id: "y"}]` diffs to: ```json [ { "op": "replace", "path": "/items/1/id", "value": "x" }, { "op": "replace", "path": "/items/0/id", "value": "z" }, { "op": "add", "path": "/items/2", "value": { "id": "y" } } ] ``` One logical insert became two in-place rewrites of existing elements’ ids plus an append. Nothing here says “insert” — so nothing downstream, from conflict detection to the event log, can recover what the user meant. **datadata’s workaround.** Collections that matter aren’t arrays. They’re [flat records keyed by stable ids with fractional-index ordering](/concepts/documents-and-folders/) — so “insert between A and B” and “update item X” become single-key operations at disjoint leaf paths, which concurrent edits can’t retarget. The identity and the ordering have to live somewhere, and JSON Patch has nowhere to put either. So they go in the data model. Fractional indexes work well for us; the price is that anything ordered must be modelled as a record set rather than an array. That’s a standing constraint on schema design, and we lean on it enough that every mutable collection the engine itself maintains (the index, schema migration logs, session changesets) is keyed by id. Nested objects are spared, but only by an implementation detail: the differ recurses into them and emits ops at leaf paths, so two clients editing sibling keys of the same object touch `/cfg/a` and `/cfg/b` rather than both replacing `/cfg`. The format itself has no deep merge — a hand-authored `replace` of a container clobbers the whole thing. ## `test` is the only concurrency tool, and it’s blunt **The issue.** The RFC’s only precondition mechanism is the `test` op: assert a value, fail the whole patch otherwise. There is no way to express intent (“increment”, “insert after X”, “set if unset”), so any concurrent touch of a tested value rejects the entire patch — even when the writes were trivially compatible. **datadata’s workaround.** [Guard modes](/concepts/changes-as-json-patch/): auto-generated `test` ops scoped to the values actually changed (tolerating disjoint edits), or whole-document sequence guards when strictness is wanted. Rejection is treated as a benign, expected signal with a retry path — and the [staged session](/staged-sessions/conflict-preview-and-resolution/) turns it into reviewable three-way conflicts. But the underlying expressiveness gap is the RFC’s. ## `add` has no precondition **The issue.** Those auto-generated guards cover `replace` and `remove` — operations that have a prior value to assert. An `add` of a previously-absent member has nothing to test: the RFC offers no way to say “this key must not exist yet”, so the differ emits a bare `add` with no guard in front of it. Two clients concurrently adding the same new key both pass, and the later write silently overwrites the earlier one. Guarded writes are a complete same-path conflict detector for changes and deletions, but not for additions — and every layer built on guards (including [session conflict detection](/staged-sessions/conflict-preview-and-resolution/)) inherits that hole. **datadata’s position.** Open. Candidate fixes: extend the patch profile with an explicit absence test, hash-based value preconditions, or telling strict callers to use the coarse sequence guard (which does catch it). For id-keyed records with generated ids the collision is improbable by construction, which is why this is a sharp edge rather than a daily wound. ## Diff-generated patches lose intent **The issue.** The array insert above is one instance of a general problem: patches produced by diffing two states (how datadata’s client produces every one of them) are not unique. A `move` is indistinguishable from a `remove`+`add`; a small edit inside a string is a whole-value `replace`. Whatever the differ guesses becomes the recorded “change”, which degrades both conflict detection (coarser `test` ops than the actual edit warranted) and the [event log’s](/architecture/storage-and-event-log/) value as history. The whole-value `replace` also has a size cost: a one-character edit to a long text field ships and stores the entire new string on every change, so the wire and the event log bloat fast for frequently edited prose. **datadata’s position.** Live where it’s tolerable (structured fields are small), escape where it isn’t ([Yjs](/concepts/rich-text-with-yjs/) for text, where intent-per-keystroke is exactly what the CRDT captures). *** *If you’ve fought this RFC in your own engine, [compare notes with us](/contact/).* # Limitations > What datadata doesn't do — classified as by design, not yet, or open question. Each limitation is tagged: * **By design** — a consequence of the model. * **Not yet** — wanted, understood, unbuilt. * **Open question** — we don’t know the right answer yet; see [Open questions](/known-issues/open-questions/). | Limitation | Status | Notes | | --------------------------------------------------- | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | No per-field authorization | **By design** | The engine authorizes every write and filters every read ([Authorization](/concepts/authorization/)) across three declarative grains — folder roles, docType rules, and per-document `sys:access` entries — but visibility is whole-document: there are no per-field rules, and no content-based rules (verdicts never read document data). Whether per-field rules should exist is [open](/known-issues/open-questions/). | | Documents cap at 1.9 MB | **By design** | Both persistent backends cap a document, a single event, and each embedded Y.Doc state at 1.9 MB by default — just under the Durable Object SQLite value limit, and the [Postgres adapter](/architecture/node-deployment/) keeps the same default on purpose so a document valid in one deployment is valid in every other. Whole-document sync makes huge documents the wrong shape anyway: model large data as more documents. | | No queries across documents | **By design** | There is no query language — no select/project over multiple documents. The supported pattern is a derived store: listen to document changes and project them into another document (or another database), then query that. | | No partial replication within a document | **By design** | Subscriptions are per-document; a document syncs whole — including all its embedded Y.Docs (full states on first subscribe, [state-vector diffs](/concepts/two-sync-lanes/) on reconnects). There is no per-Y.Doc lazy loading. Folder = authority unit, document = sync unit. | | Folders are sized for thousands of documents | **By design** | A folder’s [listings](/concepts/system-documents/) — `sys:index` and `sys:trash` — are documents like any other, so a subscriber receives every entry it may read, built from one scan of the folder when it subscribes. That cost grows linearly with the folder, and the design target is up to a few thousand documents per folder. Past that, split the data across more folders. | | No cross-folder sync or federation | **By design** | A folder is one server’s world; nothing spans folders. | | No P2P, no decentralized authority | **By design** | The server-authoritative model is the point — see [Design decisions](/comparison/design-decisions/). | | No E2E encryption | **By design** | Server-side validation requires the server to read the data. | | Offline durability is best-effort | **By design** | The [persisted write queue and document cache](/concepts/offline-persistence/) live in browser storage, which browsers evict under pressure, prune for stale origins, and discard in private browsing. The degraded mode is memory-only behavior, never corruption — but “survives a reload” is a practical guarantee, not a durable one. | | Long-offline unguarded writes land last-writer-wins | **By design** | A queued unguarded JSON write replays silently on reconnect, however old. Protocol-legal, occasionally surprising. Offline work that deserves review belongs in a [staged session](/staged-sessions/overview/), whose reconnect produces a changeset instead. | | Cached reads outlive an access revocation | **By design** | A revoked principal can still read cached documents offline — the client already saw the data — until the next connection replaces them with `notfound`. The persistence namespace must be `(folder, principal)`-scoped and `clear()` belongs in logout. | | Blobs default to 16 MiB | **By design** | A [blob](/concepts/blobs/) is accepted at upload only up to `maxBlobSizeBytes`, which defaults to the ceiling a document export can carry inline — so every blob the platform holds stays exportable. A host may raise it consciously, trading away export of the larger blobs until a streaming export exists. | | No blobs without a served host | **By design** | The backendless [in-memory client](/architecture/running-in-memory/) has no blob lane: a blob’s download URL is a plain string for ``, which no in-process transport can serve. Demos and tests that need files run against a real host. | | Blob uploads stream through the host | **Not yet** | Bytes go client → host → object store, capped by the host’s request body limit. Presigned direct-to-store uploads and multipart are future work behind the same client contract, as are signed download URLs and serving user content from a cookie-less origin; today [the response headers](/concepts/blobs/#serving-uploads-safely) make same-origin serving safe. | | No event-log compaction | **Not yet** | Document and Yjs logs are append-only and unbounded (`doc:get-events` pages through them). The most concrete scaling gap. Yjs’s own item-level garbage collection is also unmanaged: deleted-item tombstones are not reclaimed from stored `Y.Doc` state, since native GC is only safe once every peer has seen the delete and datadata does not coordinate that condition across its stored log. See [Storage & event log](/architecture/storage-and-event-log/). | | No named changesets | **Open question** | A changeset occupies one field of its host document, shared by every session opened over that field. There is no way to address several changesets by id, nor an API to create and drop them — deliberately deferred until the apps demand it. | | Projections are per-document, read-mostly | **Open question** | A projection layer reshapes a document into more ergonomic views — but it operates on the full synced document (it is not partial sync), and writable lenses are very alpha. Whether projections should be writable at all is open. | | JS/TS only | **Open question** | datadata is a TypeScript library on both sides of the wire. Whether other-language clients are wanted is untested — nobody has asked. Supporting them would also turn the [wire protocol](/architecture/wire-protocol/) into a public contract, versioned and held stable for implementers outside this repo; today it is documented, not promised. | If one of these is the thing you came to check — that’s the page working as intended. If you think one of the classifications is wrong, that’s a conversation [worth having](/contact/). # Open questions > The things we're genuinely unsure about. Opinions wanted. These aren’t missing features — they’re design problems where we don’t yet trust any answer, including our own. If you’ve faced one of these in your engine, [we’d like to compare notes](/contact/). ## Compaction without losing history’s value The [event log](/architecture/storage-and-event-log/) is an asset — audit, provenance, future undo/diff features — until it’s a liability. What’s the right compaction contract? Snapshot + truncate loses blame; rolling windows lose old provenance; never compacting loses the disk. Yjs has the same question with different math. And staged sessions have quietly come to depend on the log: a cold client reconstructs a stage’s base for the [three-way preview](/staged-sessions/conflict-preview-and-resolution/) by replaying it, so truncation would degrade those previews (conflict detection and commit never need the log). ## What does offline actually mean here? For the CRDT half, the answer turned out to be structural: the [Yjs lane self-heals](/concepts/two-sync-lanes/) — a reconnecting client pushes back exactly the content the server lacks, because CRDT deltas merge cleanly however late they arrive. The JSON half is answered too: writes buffered while disconnected replay on reconnect, last-writer-wins (a guarded write is rejected if the server moved under it), re-sends de-duplicated by client event id and bounded by the replay horizon ([full replay rules](/architecture/reconnect-and-replay/)) — and with the opt-in [persisted write queue](/concepts/offline-persistence/) that buffer now survives a page reload, closing the durable-queue half of this question. What stays open is the *semantics* of long offline periods: a persisted unguarded queue replayed days later lands silently last-writer-wins (protocol-legal, humanly surprising), and a guarded one mostly means rejections, which pushes everything into [session-style three-way resolution](/staged-sessions/conflict-preview-and-resolution/). Is offline-then-review (a reconnect producing a *changeset to review* rather than silent writes) the right model for longer offline periods on server-authoritative engines? We suspect yes — staged sessions’ changesets are already server-durable, they [compose with the persisted queue](/concepts/offline-persistence/) (work staged offline survives a reload and resumes; the flow is integration-tested), and commit is deliberately an online act — one that is now also crash-safe: a durable commit intent lets a commit that died with its page finish its bookkeeping on resume instead of double-applying on retry. What hasn’t been proven is the *experience*: whether a changeset review on reconnect is what people actually want after days offline, or an interruption they’ll resent. A second sub-question: **who arbitrates while offline?** Tabs in one browser profile share their offline work over a cache bus, but each tab is still a full client with its own optimistic queue, so two tabs editing the *same field* keep showing their own value until a reconnect lets the server settle the order ([offline and multi-tab](/architecture/offline-multi-tab/)). The end-state we know of is a single arbiter per browser profile — a leader tab or SharedWorker owning one queue for all tabs — which would make renders converge offline, at the cost of a single point of failure and a large architecture change for a divergence that is render-only: nothing either tab journals or replays differs. We haven’t taken that trade. What would move it is evidence that people hit same-field divergence across their own tabs in practice, rather than the disjoint edits that already render the same everywhere. ## How far should declarative authorization stretch? Authorization is fully declarative: folder-role + docType rules in schema documents, per-document `sys:access` entries (ownership, sharing, id reservation), read filtering, scope caps — see [Authorization](/concepts/authorization/). What stays open is the boundary. Per-field rules and content-based rules (“readable once published”) remain outside the model: verdicts never read document data, which is what keeps the client’s capability prediction sound. Each further step re-runs the original trade-off — declarative-in-governance-documents is elegant and a big correctness surface; pushing the rest to app code keeps the engine small but pushes complexity onto every app. ## Conflict resolution UX patterns The [three-way preview](/staged-sessions/conflict-preview-and-resolution/) provides the data, but what do users actually want — field-level pick-and-choose? Restage-from-head and re-apply? Let the agent propose the merge and review *that*? The apps are the laboratory; no pattern has won yet. A concrete sub-question: **partial accept**. Per-document accept already falls out of the model (each stage commits independently), but accepting *some staged operations within one document* while keeping the rest staged needs a selective drain plus a rebase of the remainder. Per-item triage of AI suggestions is proven UX elsewhere — is it worth the machinery, or is whole-stage commit the right simplicity? ## Cross-folder portability A document’s [event stream replays anywhere](/architecture/portable-event-streams/), but its schema version stamps are folder-local coordinates — so import only works into a fresh or cloned folder that carries the schema document’s history. Importing into a folder whose schema evolved independently is a remap problem: translate sequence stamps by matching reconstructed schema *content*. Nothing extra needs to be stored to make that possible — but the remap tool doesn’t exist, and whether it’s ever needed (versus “clone whole folders” being the only real use case) is open. ## Should projections be writable? The projection layer reshapes a fully synced document into a more ergonomic view — picks, renames, sorts, groupings. A subset of those operations inverts cleanly, which makes **writable lenses** possible: edit the view, and the edit translates back into path-disjoint physical operations that merge cleanly. The mechanics exist at a very alpha level — but we genuinely don’t know whether they’re worth the effort, what the right shape is, or whether projections should simply stay read-only and writes always go through the physical document. If you’ve built (or abandoned) a lens layer over a sync engine, this is the question we’d most like to compare notes on. ## Multiplexed changesets A changeset occupies one [field of its host document](/staged-sessions/overview/), and every session opened over that field shares it. What’s missing is *named* changesets: addressing several by id, with an API to create and drop them. Parallel agent proposals and draft-vs-review lanes want exactly that, and it is repeatedly tempting and repeatedly deferred — is the added model complexity worth it, or is “one changeset, host documents are cheap” the right discipline? ## Should anything speak the protocol but TypeScript? datadata is TypeScript on both sides of the wire, and nobody has yet asked for otherwise — which is a fact about our sample, not evidence of anything. The design problem hiding behind the missing feature is what a second language would cost: the [wire protocol](/architecture/wire-protocol/) is documented today but not *promised*, and it evolves whenever the two lanes need it to. A client we don’t compile is a client we can’t migrate in the same commit, so supporting one means versioning the protocol as a public contract and holding it stable for implementers we can’t see. That is the trade — not “write a Rust client” but “freeze the wire protocol.” Whether the reach is worth the loss of velocity is the part we don’t know. ## How far does “schemas as documents” stretch? [Migration happens per document on its next read](/concepts/schema-evolution/), with the result written back — so cold documents accumulate pending migrations until someone loads them. “Is the migration done?” now has a clean answer: a budgeted validation sweep forces discovery over cold documents, and when it finishes, the invalid-document enumeration is complete for the current schema. The half before has a first answer: a dry run previews what a candidate schema would strand, heal or leave broken, document by document, without writing anything. It is advisory, so a schema write still never refuses on it. The half after has a first answer too, in the playground: list the flagged documents, open one, and fix every located violation in a single write that the server accepts only when the whole document is valid again. An agent gets the same through the playground’s MCP server: list the invalid documents with their located issues, preview a candidate schema, and repair with an ordinary update. What remains open is whether that holds up at scale: thousands of flagged documents, fixes an operator wants to apply in bulk, and repairs that need a person who knows the data rather than a picker. Nobody has hit these walls yet, which is not the same as the walls not existing. # Glossary > The vocabulary of datadata, in one place. **Attribution** — who wrote a stored event: the **principal**’s subject and actor, recorded on every event in both lanes and never rewritten. See [Storage & event log](/architecture/storage-and-event-log/). **Blob** — an immutable file (image, PDF, attachment) held in an **object storage adapter** and referenced from document JSON by a server-minted opaque handle in a `blobRef` field. The handle syncs like any field; the bytes move over HTTP and never enter a sync lane or an event log. Reference-counted at the write boundary, swept after a grace period, read-gated like the documents that reference it. See [Blobs](/concepts/blobs/). **Changeset** — the persisted state of a [staging session](/staged-sessions/overview/): stages plus two [lanes](/staged-sessions/changesets-and-lanes/) of staged changes, stored in a field of the host document. **Client** — the connection to one folder, owning reads, writes, subscriptions and the optimistic overlay. Live and staging **sessions** are opened on it. Connects over a socket or, [in-process](/architecture/in-process-clients/), by function calls. **Document** — the unit of data: `docId`, `type`, server-assigned `sequence`, JSON `data`, optionally with embedded Yjs documents. See [Documents & folders](/concepts/documents-and-folders/). **Event log** — the append-only history of a document’s accepted changes: one log per [sync lane](/concepts/two-sync-lanes/) — JSON Patch events under `sequence`, Yjs update events under `yjs_sequence`. See [Storage & event log](/architecture/storage-and-event-log/). **Folder** — the unit of synchronization and authority; one server instance owns one folder. In production that is a Durable Object per folder; on the Node backend, one process and one Postgres database hold many folders, each with its own server instance inside. **Fractional index** — a string key that sorts between its neighbors, used to order records in collections without array indices. **Generation** — which incarnation of a `docId` a document is: an opaque token minted when the document is created or imported, unchanged by edits, delete and restore. Purge releases an id, so a later create under it is a new generation — which is how a reconnecting client tells its copy of the purged document from the new one. See [Documents & folders](/concepts/documents-and-folders/). **Guard** — a precondition on an update: either *sequence* (document unchanged since base) or *patch* (RFC 6902 `test` ops on the touched values). See [Changes as JSON Patch](/concepts/changes-as-json-patch/). **Host document** — the document whose field stores a staging session’s changeset — typically the document that motivates the changes (e.g. a conversation). **In-process client** — a full client connected to the server by function calls instead of a socket. See [In-process clients](/architecture/in-process-clients/). **JSON lane** — the half of document sync carrying structured changes as [JSON Patch](/concepts/changes-as-json-patch/), versioned by `sequence`. See [Two sync lanes](/concepts/two-sync-lanes/). **Lane** — in sync: the JSON lane or the Yjs lane. In a changeset: `changes` (an ordered record set of JSON Patch records) or `yjsCopies` (one [staged Yjs copy](/staged-sessions/changesets-and-lanes/) per target Y.Doc). **Object storage adapter** — the optional byte store behind the server — R2, an S3-compatible store, a filesystem, or memory — holding [blobs](/concepts/blobs/) by opaque key. Bytes only; the catalog, the lifecycle and the authorization stay in the engine. **Optimistic update** — a change applied to the local view immediately, retired when the server confirms it (or discarded when rejected). See [Sync & optimistic updates](/concepts/sync-and-optimistic-updates/). **Presence** — ephemeral participant state (cursor, selection, display name), carried in named channels: a `sys:presence::` document riding the ordinary document wire — no bespoke protocol — read and written **per-session** through a synthetic `sys:presence-view::` view. Never stored. See [Presence](/concepts/presence/). **Principal** — who a client acts as: a **subject** (the user or service the write is on behalf of; null when anonymous) and an **actor** (what kind of code is acting — a user’s frontend, a named AI agent, server-side system code). The two are independent, and neither confers privilege: identity is never authority. Stamped onto every stored event as **attribution**. See [Authorization](/concepts/authorization/). **Sequence** — the per-document, server-assigned, monotonically increasing version number of the **JSON snapshot**. The Yjs lane has its own version, `yjs_sequence`. See [Two sync lanes](/concepts/two-sync-lanes/). **Session** — a document-editing surface on a **client**. Two kinds share it: a **live session**, whose writes land on the document immediately, and a **staging session**, which accumulates the same calls into a [changeset](/staged-sessions/changesets-and-lanes/) to preview and commit atomically. Commit, discard, stages and conflicts exist only on the latter. See [Staged sessions](/staged-sessions/overview/). **Soft delete** — removing a document from the index and read path while retaining its history, so it can be restored at the exact sequence it left. Deleted documents are listed in `sys:trash`. See [Documents & folders](/concepts/documents-and-folders/). **Stage** — a staging session’s record for one touched document: base sequence, type, create/delete kind, pending rename. No snapshot — the staged patches themselves live in the changeset’s lanes. **Staged Yjs copy** — a [staging session’s](/staged-sessions/overview/) staged Yjs content: the live target Y.Doc forked into a host sub-document that editors bind to and collaborate on natively. Returned by the `getYDoc` of the handle `session.prepareDocument` returns. Tracks the live target while the session runs, and is merged back into it at commit. See [Changesets & lanes](/staged-sessions/changesets-and-lanes/). **State vector** — Yjs’s compact summary of the content a peer already holds. Subscribes carry one per embedded Y.Doc so the server can answer with exact diffs; the server’s answers carry its own so the client can push back what it lacks. See [Two sync lanes](/concepts/two-sync-lanes/). **System document** — engine state exposed as a document under the reserved `sys:` prefix, read and subscribed to like any other. See [System documents](/concepts/system-documents/). **Three-way preview** — a derived conflict view comparing *base* (staged from), *head* (live now), and *staged* (base + staged changes). The base is resolved on demand — live head, an in-memory pin, or an event-log replay on a cold client. See [Conflict preview & resolution](/staged-sessions/conflict-preview-and-resolution/). **Write-behind burst** — a run of streamed Yjs deltas buffered in memory and flushed as one merged transaction into the Yjs lane’s log. Single-author, so attribution stays correct. The fold shortens the log rather than perforating it: the surviving rows keep consecutive positions, and the document’s `yjs_sequence` follows them down, so coalescing is invisible from outside storage. See [Storage & event log](/architecture/storage-and-event-log/). **Yjs lane** — the CRDT half of document sync: Yjs updates under their own per-document version and stored log, never advancing `sequence`. See [Two sync lanes](/concepts/two-sync-lanes/). **`yjs_sequence`** — the Yjs lane’s per-document version, and each Yjs log row’s position: the document’s value is its latest Yjs event’s, exactly as `sequence` is its latest JSON event’s. 0 for a document with no Yjs history, fully independent of `sequence`, and **server-side only** — no wire event carries it, because a CRDT lane gives clients nothing to do with a version. # References > The standards and building blocks datadata is built on. ## Standards & building blocks * [RFC 6902 — JavaScript Object Notation (JSON) Patch](https://datatracker.ietf.org/doc/html/rfc6902) — datadata’s structured change format, sharp edges and all. See [our issues with it](/known-issues/json-patch-rfc/). * [RFC 6901 — JSON Pointer](https://datatracker.ietf.org/doc/html/rfc6901) — the path syntax inside JSON Patch. * [Yjs](https://yjs.dev/) — the CRDT implementation embedded for [collaborative text](/concepts/rich-text-with-yjs/). * [Fractional indexing](https://www.figma.com/blog/realtime-editing-of-ordered-sequences/) (Figma’s write-up) — the ordering technique behind datadata’s [ordered records](/concepts/documents-and-folders/). * [Cloudflare Durable Objects](https://developers.cloudflare.com/durable-objects/) — the per-folder authority in the [production deployment](/architecture/cloudflare-deployment/). * [PGlite](https://pglite.dev/) — Postgres in WebAssembly, the driver behind the [Node + Postgres backend](/architecture/node-deployment/) today. ## Suggest something If there’s a paper or write-up that bears on one of our [open questions](/known-issues/open-questions/), [send it over](/contact/). # Contact > Not open source yet — but very much open to conversation. datadata is built by **Jonas Bengtsson**. I’m interested in talking to people who are evaluating, building, or stretching sync engines. datadata isn’t open source yet, so this site is the public interface for now. * **You’re evaluating sync engines** and want to know whether the model fits your app — or you’ve spotted a gap that isn’t documented in [Known issues](/known-issues/). Ask; I’ll give you a straight answer. * **You build sync engines** and want to push back on a [design decision](/comparison/design-decisions/) or compare notes on an [open question](/known-issues/open-questions/). Please do. * **You want to try it.** Early access ahead of open-sourcing is a possibility — get in touch. ## Reach me * Best: * X/Twitter: [@jonasb](https://x.com/jonasb) * GitHub: [jonasb](https://github.com/jonasb) * LinkedIn: [Jonas Bengtsson](https://www.linkedin.com/in/jonasbengtsson/) * In person: catch me at [Local-First Conf](https://www.localfirstconf.com/).