Skip to content

Reconnect & replay

When a client loses its connection, optimistic writes don’t stop — they queue in an in-memory buffer and replay when the socket returns. This page is the exact contract for that replay: how re-sends are de-duplicated, how long the guarantee holds, which writes are exempt, and how a write the server never answered is recovered while the connection stays up.

By default the buffer is in-memory; with the opt-in persisted write queue it survives a page reload too, and the reloaded client replays it under exactly the rules on this page — a reload is a reconnect, with no separate cold-start semantics.

On reconnect the two lanes reconcile differently: structured (JSON Patch) writes replay last-writer-wins, while the Yjs lane self-heals by CRDT merge — deltas merge cleanly no matter how late they arrive.

What the reconnect rebuilds instead of replaying

Section titled “What the reconnect rebuilds instead of replaying”

Only writes replay. Nothing else the client does while offline is held for later: the client sends no frames at all while the socket is down, and the reconnect rebuilds what the server needs from the client’s current state.

  • Subscriptions are re-sent with their current cursors. A document subscribed while offline is subscribed once, on reconnect. A document released while offline needs nothing, since the new connection never held it.
  • Presence cells are republished with their latest state, so peers never see an older state arrive after it. A cell cleared while offline stays gone.
  • Event history reads (getDocumentEvents) reject straight away while offline, the same answer a read in flight gets when the socket drops. Retry once the client is back online.

A client starts offline, so the first connect works the same way. What a page subscribes to and writes before its socket first opens goes out on that connect, rebuilt from current state like any reconnect. The socket itself queues nothing: a frame sent while it is down is dropped, and the next connect sends what is still wanted. A host wiring its own transport calls setConnected(connection) every time the transport opens, the first time included.

The connection it passes names that connection: the socket object, or anything else the host has one of per connection. The host passes it again with every event it hands over, handleEvent(event, connection), as the connection the event arrived on. The client drops an event of any other connection, so a frame a replaced socket still delivers changes nothing. Without that, a late snapshot from the old socket would replace a newer copy of its document. Keeping one connection’s events in the order they arrived is still the host’s job.

Re-sends are de-duplicated, within a bounded window

Section titled “Re-sends are de-duplicated, within a bounded window”

Replay is safe to repeat: the server records each write’s client event id and re-acks a duplicate instead of applying it twice, so a write that committed but whose acknowledgement was lost can be re-sent. This covers every replayable write kind — updates, creates, deletes, restores and renames — so a replay also never re-applies over what happened in between: a delete replayed after someone restored the document re-acks with its live state instead of re-deleting it, a replayed restore doesn’t resurrect a document deleted since, and a replayed rename never clobbers a newer name. That dedup record is retained only for a bounded window.

The record is of writes the server accepted. A refused update leaves none, so if its error is lost and the client replays it, the server judges it again against the document as it is by then. A guarded update whose guard failed the first time can pass the second time and be applied. The result is the one the write would have had if it had been delayed in transit, and nothing is applied twice. The client relies on the same rule: an update refused while an earlier one of its own is still unconfirmed is held and sent again behind it, because the refusal may only mean the earlier write had not arrived.

A write that was sent and then sat unconfirmed longer than a client-side horizon (30 minutes, half the server’s retention) is dropped rather than replayed — but not silently. The optimistic entry rolls back to server truth and the outcome is reported as unconfirmed: an awaited write rejects, a fire-and-forget one hits onWriteError, and the caller refetches. That beats the two silent alternatives — a write that vanishes without a trace, or a replay that risks a double-apply once the dedup record has aged out.

Writes made purely offline never reached the server, so they carry no such risk and always replay, however long you were away.

When a write keeps failing with a transient storage error the client retries it a few times; during a systemic outage a circuit breaker stops piling on retries once many writes are failing at once.

Sequence-guarded writes replay at any age: their compare-and-set base means a late replay of a write that already committed is rejected benignly rather than applied twice, so they replay instead of being surfaced. The exception is when later unguarded work was stacked on top of one — an unguarded edit, or a rename, delete or restore of the same document — since all of it was authored assuming the CAS applied, so the whole chain is surfaced instead.

De-duplicated re-sending isn’t only for reconnects. A WebSocket delivers every frame in order for as long as it is open, but a server can still leave a write unanswered on a live connection: it fails while handling the write, or commits it and fails before the acknowledgement leaves. Heartbeats still flow, so nothing reconnects and nothing replays. The client therefore sweeps its own buffer: any write still unconfirmed after a bound (10 seconds by default) — an update, a create, a delete, a rename or a restore — is re-sent under its original event id, and keeps being re-sent each period until the server settles it. If the original write actually committed and only its acknowledgement went missing, the re-send draws a dedup re-ack instead of a second apply, and the client resyncs the document to fetch the data that acknowledgement would have carried. The write’s own patch keeps rendering in the local view until that resync lands, so the writer’s committed value never blinks out of its view, and an edit made meanwhile builds on it rather than on the stale base. The same holds when it is a create’s acknowledgement that went missing and an update pipelined behind it is answered first: the update’s effect stays in view over the staged create until the re-sent create’s answer lands. The same 30-minute horizon bounds this path too for updates; a create is never given up on, since a late one reconciles through the benign already-exists answer.

A create the server never handled has a second consequence: every update sent behind it is refused as not found. The client treats that refusal as transient when it holds its own unconfirmed create for the document, keeps the update pending, and lets the sweep re-send the create and then the update, so the edit lands rather than being rolled back.

The same reasoning covers two edits in a row. Because the server answers a document’s writes in the order they were sent, an answer for the second edit arriving while the first is still unconfirmed means the first was lost. The client then rebases the first past the second rather than re-sending it whole: whatever the second edit already decided is dropped from it, the rest is re-sent, and an edit with nothing left is settled as committed. Where the first edit wrote a whole value and the second changed something inside it, the first keeps that value with the second edit’s change folded in, so its re-send lands what the author saw rather than undoing the second edit. A rejection of the second edit in that situation is held rather than acted on, since it may only say that the patch met a document the lost edit never reached; the sweep re-sends both in order and the server answers again.

When no answer for the second edit has arrived either, the client cannot tell whether that edit was lost too or was applied with only its acknowledgement lost, and re-sending both blind would be wrong in the second case: the duplicate of the second edit is de-duplicated, the first lands after it, and the earlier value wins on the server. So before re-sending an edit behind a later one it has already sent, the client asks (doc:get-processed-writes) which of the document’s unconfirmed edits the server processed. The answer comes from the same de-duplication store and names each processed edit with the sequence it committed at, so the client settles those exactly as their re-acknowledgements, which rebases the edits staged before them, and then sends whatever is left, in order. Edits made while the question is out wait behind them, and a question that goes unanswered for a period is asked again. The reconnect replay asks the same question for a document whose buffered edits include one sent before the socket dropped, instead of replaying them blind.

Lifecycle operations are protected differently: a document’s create, deletes, renames and restores go on the wire one at a time. A delete staged on an unacknowledged create, or a restore staged on an unacknowledged delete, is recorded locally at once (the document reads deleted, the awaited promise is pending) but is sent only when the operation ahead of it has been answered. An unanswered operation can therefore neither strand the operations behind it, which were never sent, nor be overtaken by them when the sweep re-sends it. The cost is one round trip per operation in such a chain, which is rare; edits stay pipelined.

A delete goes one step further: it also waits for the document’s own unconfirmed edits. Staging it keeps them (pending Y.Doc changes are flushed ahead of it), so an edit left unanswered is re-sent by the sweep and commits before the tombstone, which means a delete-then-restore undo brings it back. And a delete the server refuses, whether a guard miss or an authorization the client could not predict, leaves the document live with nothing lost. The document reads as deleted locally the moment the delete is staged either way.

A delete from another client gets no such preparation — it can land while this client is still holding unconfirmed work — so the tombstone keeps that work rather than dropping it. Pending JSON edits stay for the server’s own answer to their re-send. The document’s Yjs content stays too: the live Y.Doc, its unflushed changes and its “changes pending” claim all survive the tombstone, held unsent while it stands (an update on a deleted document is refused, and a Y.Doc change cannot be rolled back once it is local) and released by the restore that undoes the delete — which re-sends the held changes and pushes back anything else the restored document is missing. Only evidence that the content can belong nowhere discards it: the document purged, or hidden from this client, or an answer naming a different generation of the id — a document that reused it.