Skip to content

Schema evolution

Schemas are documents, so schema evolution is document editing — but evolution has rules of its own. datadata’s stance is evolution-first: schemas change freely while the app runs, migrations are explicit, and documents that no longer fit are flagged, never dropped.

A schema’s version is its document’s sequence number — there is no separate version field. Every accepted edit to sys:schema:<type> advances it, and every document tracks which schema sequence its data conforms to. Every event in a document’s history also records the schema sequence in effect when it was written — which is what keeps historic versions interpretable anywhere.

Because that number counts this folder’s accepted edits to the schema document, it is a folder-local coordinate: the same type can sit at different sequence numbers in two folders depending on when each folder’s schema was last edited. A version number only means something inside its own folder — there is no global schema version to compare across folders.

Data-shape changes are declared as migrations in the schema document. Three operations exist, deliberately minimal:

  • rename — move a field to a new name.
  • remove — delete a field. Removal is never inferred: dropping a field from the schema without a remove migration makes documents still holding it invalid, rather than silently discarding data.
  • remap — rewrite scalar values (old → new pairs). Collapsing several old values into one is allowed; one-to-many is not. Remaps are pure — the new value depends only on the old one.

migrations is a record — each migration sits under a key its author picks:

// As stored: `sequence` is the server's stamp, whatever the author wrote.
migrations: {
"001-rename-title": { op: "rename", sequence: 2, from: "title", to: "heading" },
"002-remove-body": { op: "remove", sequence: 4, path: "body" },
}

The log is append-only, enforced by the server: a schema write may add keys but never modify or drop a committed one, and the server stamps each new migration with the schema sequence it took effect at — authors don’t control the stamps.

A rename never replaces a field unasked. Where the target name may already hold a value, the rename has to say what happens then, with onConflict:

  • keep — nothing: both keys stay, and validation reports the one the schema doesn’t declare.
  • replace — the renamed value takes the name; the old occupant is dropped.
  • discard — the occupant keeps the name; the renamed field is dropped.
// Merge `nickname` into `displayName`, keeping displayName where both exist.
"003-merge-nickname": {
op: "rename", sequence: 5, from: "nickname", to: "displayName", onConflict: "discard",
},

The server requires onConflict exactly where a collision is possible and refuses it everywhere else, so it always means something. It is required when the target is declared in that container (in the schema being replaced, in any variant of a union, or because an earlier migration of the same write renamed something there), and when the container’s keys are data: a record’s keys, an object that keeps undeclared members, a json value. A rename of a field onto itself is always refused. Instead of a policy you can remove the target under an earlier key; to swap two fields, go through a temporary name (a → tmp, b → a, tmp → b). Merging two fields into one is the second rename with discard: a → c, then b → c leaves c holding whichever of the two was present, a first.

Whether a rename with a policy actually happened depends on each document, so an offline edit written against the old shape that touches either of its keys can’t be carried forward faithfully: it is rejected as a precondition failure to refetch and rebase. Edits elsewhere in the document carry over as usual.

Migrations reach nested data through dotted paths, with * standing for every member at that level (a rename’s at: "items.*" renames a field in every element of items). An array is entered through * only: the server refuses a migration whose path steps into an array by index (items.0). A position means a different element as soon as something is inserted or removed before it, so a migration aimed at one element could not be carried faithfully across the edits written against the old shape. A record is keyed by name, not position, so labels.en stays a valid path.

What authors do control is the key, and the keys’ lexicographic order is the replay order (English collation, case- and accent-sensitive — not numeric, so pad numbers: 010 sorts after 009, 10 sorts before 9). A new migration’s key must sort after every committed one — the server rejects a key that would land inside the committed range, since that would silently reorder replay for older documents, and rejects an empty key. Keys can be written by hand or generated by tooling; a convention like 002-rename-title sorts correctly and stays readable. Two authors adding migrations at once merge as long as their keys differ and the later write’s key sorts last; otherwise the later write is refused and needs a new key. Where several people author migrations concurrently, a prefix that sorts by authoring time — a date (2026-09-21-rename-title) or a ULID — makes the out-of-order case rare and a shared key all but impossible. A staged session reports a staged key the server would refuse before commit, and rekeys it in one operation.

Tooling that writes schema documents on an author’s behalf can reach the same verdict itself: @repo/datadata/schema exports compareMigrationKeys (the key order), orderedMigrations (a log’s entries in replay order) and migrationLogRefusal (the append-only verdict on a submitted log against the committed one, with the key it is about), so a host mints a key that sorts last and checks its log with the rule the server enforces, rather than re-implementing the collation. migrationShapeRefusal is the server’s other verdict, on the migrations a write adds, judged against the schema it replaces.

When the author can’t see the committed log at all — an LLM evolving a type through a tool, or a form that collects “the migrations for this change” — appendMigrations mints the keys:

import { appendMigrations } from "@repo/datadata/schema";
const migrations = appendMigrations(committedSchema.migrations, [
{ label: "rename-title", migration: { op: "rename", from: "title", to: "heading" } },
{ migration: { op: "remove", path: "body" } },
]);
// { …committed, "0003-rename-title": …, "0004": … } after a log ending in "0002-…"

It appends in the order given, for any committed log: it continues a counter the last key ends in (0002-… → 0003-…) and otherwise extends that key (zebra → zebra~0001-…), checking each key against the collation — under which simply appending characters to a key doesn’t always sort it later. The migrations it takes have no sequence, since the server stamps every new one. A label can’t start with a digit or contain ~, since the key would then read as something else.

parseMigrationKey reads a minted key back into its stem, counter and label. When a staged session reports that the committed log moved on under its migrations, re-mint them after the new log and keep their labels:

import { appendMigrations, parseMigrationKey } from "@repo/datadata/schema";
// staged: the [key, migration] entries the new log doesn't have yet
const reminted = appendMigrations(
movedSchema.migrations,
staged.map(([key, { sequence, ...migration }]) => ({
label: parseMigrationKey(key)?.label,
migration,
})),
);

Migration runs on read — and writes back

Section titled “Migration runs on read — and writes back”

When a document is read whose conformed sequence is behind the schema, the server brings it forward: it replays the migrations stamped after that sequence — in key order — validates the result (backfilling declared defaults for added fields), and — if anything changed — persists the migrated data back as a normal change with a new sequence number. Each document pays the migration cost once, on its next read, not on every read. The server’s onDocumentChange doesn’t fire for the write-back — the schema write that caused it is the change that gets announced.

That change’s patch says what the migrations did: a rename is a JSON Patch move from the old name to the new one — not a remove and an add — a remove is a remove, and a remap a replace. Backfilled defaults follow as ordinary adds.

The write-back doesn’t wait for the result to be valid. If the migrations rewrote the document but it no longer fits the schema, the migrated shape is still what’s persisted — it’s what every reader is handed, so it’s what the event log has to hold too — and the document is flagged alongside it. Only the flag says whether it conforms; the conformed sequence says how far the stored data has been carried.

Subscribed documents don’t wait for a read. A client that already holds a document never reads it again, so the schema write itself brings forward every document of that type with a live subscriber, right after it commits. Each subscriber receives the migration as an ordinary patch or, for a document the new schema flags invalid, a fresh doc:init carrying the flag. Either way, the migrated state reaches open clients as soon as the schema change lands. Beyond subscriptions there is no proactive bulk sweep: a document nobody reads keeps its old shape, and its pending migrations simply accumulate until it’s next loaded.

A copy says how far it has been carried. The schema change and a document’s migration reach a subscriber as two frames, the schema first, so for a moment a client holds the new schema and a copy still in the old shape. Every frame that carries a document’s data — doc:init, the doc:patch of each write, and doc:resume — names the document’s conformed sequence, and the client keeps it on its copy as conformedSchemaSequence (cached with the copy, so a reload in that moment still knows). A copy whose sequence trails its schema’s is one the server still owes a migration, and everything on the client that has to read it in the current shape replays from exactly there instead of guessing. It is “carried to”, not “conforms at”: it moves for a document the new schema flags, too. Documents of a static-schema type carry none — there is no migration log for them to be behind.

A client that was offline is told when it reconnects. Its resubscribe is a read like any other, and a document that read flags is always answered with a full doc:init carrying the flag — even when the schema change reshaped nothing and the client’s copy is current byte for byte. Validity isn’t a function of the sequence, so a matching one doesn’t earn the usual no-data doc:resume; the document is resent on each reconnect until it’s repaired. A doc:resume, in turn, tells the client the document is valid and clears a flag it held.

flowchart TB
  accTitle: What happens to a document on read
  accDescr {
    When a document is read, the server compares the sequence its data conforms
    to against the schema's current sequence. If it is current, the document is
    served as stored. If it is behind, the migrations stamped after that
    sequence are replayed in order and the result validated, backfilling
    declared defaults. A result that changed is persisted back as a normal change
    at a new sequence, whether or not it validates. A result that no longer fits
    the schema is delivered anyway, flagged with its violations.
  }
  read["document read"] --> behind{"conformed sequence<br/>behind the schema?"}
  behind -->|"no"| serve["deliver to the reader"]
  behind -->|"yes"| replay["replay the migrations stamped after<br/>that sequence, in key order"]
  replay --> changed{"did anything<br/>change?"}
  changed -->|"yes"| persist["persist back as a normal change<br/>at a new sequence"]
  changed -->|"no"| validate
  persist --> validate{"valid under<br/>the current schema?"}
  validate -->|"yes"| serve
  validate -->|"no"| flag["flag with located violations<br/><small>data intact · migrated shape persisted</small>"]
  flag --> serve

A client’s pending write — an edit still unsent, or one made offline and replayed on reconnect — is authored against the shape the client held. If a migration moves that shape before the write lands, the patch as written no longer fits: replace /title finds no title after title was renamed to name. Rejecting that as ordinary drift would silently discard offline work the user could not have prevented.

So every write names the schema sequence it was authored against (the schemaSequence on doc:create and doc:update — the sequence of the sys:schema:<type> document the client validated the write with), and the server brings the write forward through the migrations stamped since, exactly as it brings a stored document forward on read:

  • A rename moves the paths the patch targets (and the keys inside any value it writes). One with an onConflict policy may not have happened in a given document, so an edit touching either of its keys is rejected as a precondition failure to refetch and rebase; a value that contains its container is still carried, with the policy applied to it.
  • A remove drops the operations that edit the removed field — the schema author’s stated intent, not a loss — and strips the field from any value the patch writes. A patch emptied this way is acknowledged as a no-op, never rejected. A remove whose path lands on an array slot removes nothing (arrays are never left with holes), so an edit there is kept.
  • A remap rewrites the values the patch sets at the remapped path.

Guard test ops are rewritten with the rest, so a patch-guarded edit tests the migrated document for the migrated value. A create’s data is carried the same way. A raw move or copy is carried too, including one whose transferred value a migration rewrites inside of, as long as the migration rewrites it alike at both ends (reordering the elements of an array whose every element gets a field renamed, say). What cannot be carried, besides an edit to a policy rename’s keys, is a transfer the migration rewrites at one end but not the other, or one that a remove deletes at either end: the author moved the pre-migration value, which no longer exists, so such a write is rejected as a precondition failure to refetch and rebase, rather than applied as something the author did not write. datadata’s own clients never emit move or copy. Only a client validating against a bundled static schema sends no sequence — it has no migration log to name — and its writes apply as authored. A dynamic-schema client never falls into that case by accident: it refuses a write until the type’s schema document has synced, so every write it sends names a sequence.

The client carries its own pending copy of the write the same way. When its synced sys:schema:<type> document moves past the sequence a pending write names, it rewrites the pending patch (or a pending create’s data) with the same migrations, re-stamps it with the new sequence, and re-journals it — so the edit keeps rendering over the migrated copy while it waits for its acknowledgement, including a whole offline queue replayed after a migration, and a reload or another tab sees the rewritten write. The client never migrates its confirmed copy itself: that arrives from the server as an ordinary patch, and a pending update is rewritten once its copy’s conformedSchemaSequence says the copy is at the new shape (or the migrations it is still owed don’t touch it). An edit made in the moment between the schema frame and the migration frame is authored in the shape the copy is still in: the client checks it as the server will land it — carried through the owed migrations, then validated — stamps it with the copy’s sequence rather than the schema’s, and rewrites it with the rest when the migration arrives. Re-sending a rewritten write that was already on the wire is safe, because writes are deduplicated by event id. The server’s bring-forward remains the authority: any pending write the client cannot rewrite — it has no local copy of the document, say — goes out as authored and is brought forward on arrival.

Work held in a staged session is carried the same way, without being rewritten: each staged change records the sequence it was authored at, and the session reads the stack brought forward, so a migration never surfaces as a conflict.

The one exception is a presence schema (sys:schema:presence:<presenceType>). Presence data is ephemeral — never stored, so never brought forward — which means a migration could never fire. So a presence schema carries no migration log at all: a write that declares migrations is rejected. The field shape still evolves in any way you like (add, remove, retype, restructure) purely by editing the schema, since there is no stored presence for the change to strand; live cells simply re-validate against the new shape on their next write.

Invalid documents are flagged, not dropped

Section titled “Invalid documents are flagged, not dropped”

Schema edits are not checked against existing documents — you can tighten a type or add a required field freely, and documents that no longer fit become invalid. What happens then is asymmetric on purpose:

  • Writes are strict. A change that would leave a document invalid under the current schema is rejected (a schemaValidation error on the wire).
  • Reads are relaxed. An invalid document is still delivered — flagged with the violation and its data intact, with the flag recorded so the document can be enumerated and repaired later. The app (or an agent) decides how to repair it.

Invalidity is discovered on read — at schema-write time only the documents with a live subscriber are checked, and their subscribers see the flag appear, or disappear when a later edit loosens the schema again — so a document nobody has loaded since the tightening isn’t known to be broken yet. The discovered invalid documents are enumerable, so “what broke when we tightened the schema?” is a query over what reads have surfaced so far. When you need the full audit, a budgeted validation sweep forces that discovery: it reads every document whose conformed sequence trails the current schema (in host-sized batches, off the hot path), migrating the ones it can and flagging the rest — after which the enumeration is complete, and “is the migration done?” is answerable. Each flagged document carries its violations as located paths, attributed to the schema shape or to a broken cross-document reference.

Object types and discriminated unions choose how to treat keys the schema doesn’t declare:

  • reject — the default. An unknown field is a validation issue.
  • strip — accept the value but drop unknown fields from the validated output, for open-by-design shapes.
  • keep — accept and retain unknown fields, as long as the retained values are still JSON-compatible.