Schema evolution
Schemas are documents, so schema evolution is document editing — but evolution has rules of its own. datadata’s stance is evolution-first: schemas change freely while the app runs, migrations are explicit, and documents that no longer fit are flagged, never dropped.
Versioning
Section titled “Versioning”A schema’s version is its document’s sequence number — there is no
separate version field. Every accepted edit to sys:schema:<type> advances
it, and every document tracks which schema sequence its data conforms to.
Every event in a document’s history also records the schema sequence in
effect when it was written — which is what keeps historic versions
interpretable anywhere.
Because that number counts this folder’s accepted edits to the schema document, it is a folder-local coordinate: the same type can sit at different sequence numbers in two folders depending on when each folder’s schema was last edited. A version number only means something inside its own folder — there is no global schema version to compare across folders.
Migrations are explicit and append-only
Section titled “Migrations are explicit and append-only”Data-shape changes are declared as migrations in the schema document. Three operations exist, deliberately minimal:
rename— move a field to a new name.remove— delete a field. Removal is never inferred: dropping a field from the schema without aremovemigration makes documents still holding it invalid, rather than silently discarding data.remap— rewrite scalar values (old → new pairs). Collapsing several old values into one is allowed; one-to-many is not. Remaps are pure — the new value depends only on the old one.
migrations is a record — each migration sits under a key its author
picks:
// As stored: `sequence` is the server's stamp, whatever the author wrote.migrations: { "001-rename-title": { op: "rename", sequence: 2, from: "title", to: "heading" }, "002-remove-body": { op: "remove", sequence: 4, path: "body" },}The log is append-only, enforced by the server: a schema write may add keys but never modify or drop a committed one, and the server stamps each new migration with the schema sequence it took effect at — authors don’t control the stamps.
A rename never replaces a field unasked. Where the target name may
already hold a value, the rename has to say what happens then, with
onConflict:
keep— nothing: both keys stay, and validation reports the one the schema doesn’t declare.replace— the renamed value takes the name; the old occupant is dropped.discard— the occupant keeps the name; the renamed field is dropped.
// Merge `nickname` into `displayName`, keeping displayName where both exist."003-merge-nickname": { op: "rename", sequence: 5, from: "nickname", to: "displayName", onConflict: "discard",},The server requires onConflict exactly where a collision is possible and
refuses it everywhere else, so it always means something. It is required when
the target is declared in that container (in the schema being replaced, in any
variant of a union, or because an earlier migration of the same write renamed
something there), and when the container’s keys are data: a record’s keys, an
object that keeps undeclared members, a json value. A rename of a field onto
itself is always refused. Instead of a policy you can remove the target under
an earlier key; to swap two fields, go through a temporary name (a → tmp,
b → a, tmp → b). Merging two fields into one is the second rename with
discard: a → c, then b → c leaves c holding whichever of the two was
present, a first.
Whether a rename with a policy actually happened depends on each document, so an offline edit written against the old shape that touches either of its keys can’t be carried forward faithfully: it is rejected as a precondition failure to refetch and rebase. Edits elsewhere in the document carry over as usual.
Migrations reach nested data through dotted paths, with * standing for every
member at that level (a rename’s at: "items.*" renames a field in every
element of items). An array is entered through * only: the server
refuses a migration whose path steps into an array by index (items.0). A
position means a different element as soon as something is inserted or removed
before it, so a migration aimed at one element could not be carried faithfully
across the edits written against the old shape. A record is keyed by name, not
position, so labels.en stays a valid path.
What authors do control is the key, and the keys’ lexicographic order is the
replay order (English collation, case- and accent-sensitive — not numeric, so
pad numbers: 010 sorts after 009, 10 sorts before 9). A new
migration’s key must sort after every committed one — the server rejects a key
that would land inside the committed range, since that would silently reorder
replay for older documents, and rejects an empty key. Keys can be written by
hand or generated by tooling; a convention like 002-rename-title sorts
correctly and stays readable. Two authors adding migrations at once merge as
long as their keys differ and the later write’s key sorts last; otherwise the
later write is refused and needs a new key. Where several people author
migrations concurrently, a prefix that sorts by authoring time — a date
(2026-09-21-rename-title) or a ULID — makes the out-of-order case rare and a
shared key all but impossible. A
staged session
reports a staged key the server would refuse before commit, and rekeys it in
one operation.
Tooling that writes schema documents on an author’s behalf can reach the same
verdict itself: @repo/datadata/schema exports compareMigrationKeys (the
key order), orderedMigrations (a log’s entries in replay order) and
migrationLogRefusal (the append-only verdict on a submitted log against the
committed one, with the key it is about), so a host mints a key that sorts
last and checks its log with the rule the server enforces, rather than
re-implementing the collation. migrationShapeRefusal is the server’s other
verdict, on the migrations a write adds, judged against the schema it replaces.
When the author can’t see the committed log at all — an LLM evolving a type
through a tool, or a form that collects “the migrations for this change” —
appendMigrations mints the keys:
import { appendMigrations } from "@repo/datadata/schema";
const migrations = appendMigrations(committedSchema.migrations, [ { label: "rename-title", migration: { op: "rename", from: "title", to: "heading" } }, { migration: { op: "remove", path: "body" } },]);// { …committed, "0003-rename-title": …, "0004": … } after a log ending in "0002-…"It appends in the order given, for any committed log: it continues a counter
the last key ends in (0002-… → 0003-…) and otherwise extends that key
(zebra → zebra~0001-…), checking each key against the collation — under
which simply appending characters to a key doesn’t always sort it later. The
migrations it takes have no sequence, since the server stamps every new one.
A label can’t start with a digit or contain ~, since the key would then read
as something else.
parseMigrationKey reads a minted key back into its stem, counter and label.
When a staged session reports that the committed log moved on under its
migrations, re-mint them after the new log and keep their labels:
import { appendMigrations, parseMigrationKey } from "@repo/datadata/schema";
// staged: the [key, migration] entries the new log doesn't have yetconst reminted = appendMigrations( movedSchema.migrations, staged.map(([key, { sequence, ...migration }]) => ({ label: parseMigrationKey(key)?.label, migration, })),);Migration runs on read — and writes back
Section titled “Migration runs on read — and writes back”When a document is read whose conformed sequence is behind the schema, the
server brings it forward: it replays the migrations stamped after that
sequence — in key order — validates the result (backfilling
declared defaults for added fields), and — if anything changed —
persists the migrated data back
as a normal change with a new sequence number. Each document pays the
migration cost once, on its next read, not on every read. The server’s
onDocumentChange doesn’t fire for the write-back — the schema write that
caused it is the change that gets announced.
That change’s patch says what the migrations did: a rename is a JSON Patch
move from the old name to the new one — not a remove and an add — a
remove is a remove, and a remap a replace. Backfilled defaults follow
as ordinary adds.
The write-back doesn’t wait for the result to be valid. If the migrations rewrote the document but it no longer fits the schema, the migrated shape is still what’s persisted — it’s what every reader is handed, so it’s what the event log has to hold too — and the document is flagged alongside it. Only the flag says whether it conforms; the conformed sequence says how far the stored data has been carried.
Subscribed documents don’t wait for a read. A client that already holds a
document never reads it again, so the schema write itself brings forward every
document of that type with a live subscriber, right after it commits. Each
subscriber receives the migration as an ordinary patch or, for a document the
new schema flags invalid, a fresh doc:init carrying the flag. Either way, the migrated state reaches open
clients as soon as the schema change lands. Beyond subscriptions there is no proactive bulk sweep:
a document nobody reads keeps its old shape, and its pending migrations simply
accumulate until it’s next loaded.
A copy says how far it has been carried. The schema change and a document’s
migration reach a subscriber as two frames, the schema first, so for a moment a
client holds the new schema and a copy still in the old shape. Every frame that
carries a document’s data — doc:init, the doc:patch of each write, and
doc:resume — names the document’s conformed sequence, and the client keeps it
on its copy as conformedSchemaSequence (cached with the copy, so a reload in
that moment still knows). A copy whose sequence trails its schema’s is one the
server still owes a migration, and everything on the client that has to read it
in the current shape replays from exactly there instead of guessing. It is
“carried to”, not “conforms at”: it moves for a document the new schema flags,
too. Documents of a static-schema type carry
none — there is no migration log for them to be behind.
A client that was offline is told when it reconnects. Its resubscribe is a
read like any other, and a document that read flags is always answered with a
full doc:init carrying the flag — even when the schema change reshaped nothing
and the client’s copy is current byte for byte. Validity isn’t a function of the
sequence, so a matching one doesn’t earn the usual no-data doc:resume; the
document is resent on each reconnect until it’s repaired. A doc:resume, in
turn, tells the client the document is valid and clears a flag it held.
flowchart TB
accTitle: What happens to a document on read
accDescr {
When a document is read, the server compares the sequence its data conforms
to against the schema's current sequence. If it is current, the document is
served as stored. If it is behind, the migrations stamped after that
sequence are replayed in order and the result validated, backfilling
declared defaults. A result that changed is persisted back as a normal change
at a new sequence, whether or not it validates. A result that no longer fits
the schema is delivered anyway, flagged with its violations.
}
read["document read"] --> behind{"conformed sequence<br/>behind the schema?"}
behind -->|"no"| serve["deliver to the reader"]
behind -->|"yes"| replay["replay the migrations stamped after<br/>that sequence, in key order"]
replay --> changed{"did anything<br/>change?"}
changed -->|"yes"| persist["persist back as a normal change<br/>at a new sequence"]
changed -->|"no"| validate
persist --> validate{"valid under<br/>the current schema?"}
validate -->|"yes"| serve
validate -->|"no"| flag["flag with located violations<br/><small>data intact · migrated shape persisted</small>"]
flag --> serve
Pending edits survive a migration
Section titled “Pending edits survive a migration”A client’s pending write — an edit still unsent, or one made offline and
replayed on reconnect — is authored against the shape the client held. If a
migration moves that shape before the write lands, the patch as written no
longer fits: replace /title finds no title after title was renamed to
name. Rejecting that as ordinary drift would silently discard offline work
the user could not have prevented.
So every write names the schema sequence it was authored against (the
schemaSequence on doc:create and doc:update — the sequence of the
sys:schema:<type> document the client validated the write with), and the
server brings the write forward through the migrations stamped since,
exactly as it brings a stored document forward on read:
- A
renamemoves the paths the patch targets (and the keys inside any value it writes). One with anonConflictpolicy may not have happened in a given document, so an edit touching either of its keys is rejected as a precondition failure to refetch and rebase; a value that contains its container is still carried, with the policy applied to it. - A
removedrops the operations that edit the removed field — the schema author’s stated intent, not a loss — and strips the field from any value the patch writes. A patch emptied this way is acknowledged as a no-op, never rejected. A remove whose path lands on an array slot removes nothing (arrays are never left with holes), so an edit there is kept. - A
remaprewrites the values the patch sets at the remapped path.
Guard test ops are rewritten with the rest, so a patch-guarded edit tests
the migrated document for the migrated value. A create’s data is carried the
same way. A raw move or copy is carried too, including one whose transferred
value a migration rewrites inside of, as long as the migration rewrites it alike
at both ends (reordering the elements of an array whose every element gets a
field renamed, say). What cannot be carried, besides an edit to a policy
rename’s keys, is a transfer the migration rewrites at one end but not the
other, or one that a remove deletes at either end: the author moved the
pre-migration value, which no longer exists, so such a write is rejected as a
precondition failure to refetch and rebase, rather than applied as something the
author did not write. datadata’s own clients never emit move or copy. Only a
client validating against a bundled
static schema sends no sequence — it has no
migration log to name — and its writes apply as authored. A dynamic-schema client never
falls into that case by accident: it refuses a write until the type’s schema
document has synced, so every write it sends names a sequence.
The client carries its own pending copy of the write the same way. When its
synced sys:schema:<type> document moves past the sequence a pending write
names, it rewrites the pending patch (or a pending create’s data) with the same
migrations, re-stamps it with the new sequence, and re-journals it — so the edit
keeps rendering over the migrated copy while it waits for its acknowledgement,
including a whole offline queue replayed after a migration, and a reload or
another tab sees the rewritten write. The client never migrates its confirmed
copy itself: that arrives from the server as an ordinary patch, and a pending
update is rewritten once its copy’s conformedSchemaSequence says the copy is
at the new shape (or the migrations it is still owed don’t touch it). An edit
made in the moment between the schema frame and the migration frame is authored
in the shape the copy is still in: the client checks it as the server will land
it — carried through the owed migrations, then validated — stamps it with the
copy’s sequence rather than the schema’s, and rewrites it with the rest when the
migration arrives. Re-sending a rewritten
write that was already on the wire is safe, because writes are deduplicated by
event id. The server’s bring-forward remains the authority: any pending write
the client cannot rewrite — it has no local copy of the document, say — goes out
as authored and is brought forward on arrival.
Work held in a staged session is carried the same way, without being rewritten: each staged change records the sequence it was authored at, and the session reads the stack brought forward, so a migration never surfaces as a conflict.
Presence schemas don’t migrate
Section titled “Presence schemas don’t migrate”The one exception is a presence schema
(sys:schema:presence:<presenceType>). Presence data is ephemeral — never stored, so
never brought forward — which means a migration could never fire. So a presence
schema carries no migration log at all: a write that declares migrations is
rejected. The field shape still evolves in any way you like (add, remove,
retype, restructure) purely by editing the schema, since there is no stored
presence for the change to strand; live cells simply re-validate against the new
shape on their next write.
Invalid documents are flagged, not dropped
Section titled “Invalid documents are flagged, not dropped”Schema edits are not checked against existing documents — you can tighten a type or add a required field freely, and documents that no longer fit become invalid. What happens then is asymmetric on purpose:
- Writes are strict. A change that would leave a document invalid under
the current schema is rejected (a
schemaValidationerror on the wire). - Reads are relaxed. An invalid document is still delivered — flagged with the violation and its data intact, with the flag recorded so the document can be enumerated and repaired later. The app (or an agent) decides how to repair it.
Invalidity is discovered on read — at schema-write time only the documents with a live subscriber are checked, and their subscribers see the flag appear, or disappear when a later edit loosens the schema again — so a document nobody has loaded since the tightening isn’t known to be broken yet. The discovered invalid documents are enumerable, so “what broke when we tightened the schema?” is a query over what reads have surfaced so far. When you need the full audit, a budgeted validation sweep forces that discovery: it reads every document whose conformed sequence trails the current schema (in host-sized batches, off the hot path), migrating the ones it can and flagging the rest — after which the enumeration is complete, and “is the migration done?” is answerable. Each flagged document carries its violations as located paths, attributed to the schema shape or to a broken cross-document reference.
Unknown fields
Section titled “Unknown fields”Object types and discriminated unions choose how to treat keys the schema doesn’t declare:
reject— the default. An unknown field is a validation issue.strip— accept the value but drop unknown fields from the validated output, for open-by-design shapes.keep— accept and retain unknown fields, as long as the retained values are still JSON-compatible.