Persistence and networking · expert
Offline-first sync
How an app keeps working with no network. The local database is the truth. An outbox holds the writes the server has not seen yet. Idempotency keys make retries safe. A conflict rule decides who wins when two devices edit the same thing.
By the end you will be able to
- Give the offline-first answer in order — local source of truth, then the read path, then the write path
- Build an outbox with idempotency keys, and explain step by step what happens after a lost acknowledgement
- Pick a conflict rule per kind of data, and say out loud what each rule loses
A senior system-design question: "design a notes app that works on the subway." Most candidates lose it in their first sentence. Which first sentence passes?
The architecture in one diagram
UI ──reads──▶ Local store (SwiftData/Core Data/SQLite) ◀──applies── Sync engine ◀──▶ Server
──writes─▶ │ the ONLY source of truth ▲
└────── enqueues ──▶ Outbox ──drains────┘
Two words in that picture need pinning down before anything else.
Source of truth means the copy the app believes. When the local database and the server disagree, the screen shows the local copy. Reconciling the two is sync's job, and it happens later, off the screen.
The sync engine is the background code that does that reconciling. It has two halves: a push half that sends the user's edits up, and a pull half that brings other people's edits down. It contains no user interface at all.
Three consequences follow immediately, and each one answers a standard interview follow-up.
"What does the user see with no connection?" Everything. A read never touches the network, so there is nothing to fail. The @Query-driven screens from the SwiftData lesson (D2-02) are already this shape: the view is a live window onto the local store.
"What happens when they edit offline?" The edit is written to the local database and the screen updates immediately. That instant update has a name: optimistic UI, an interface that shows the result of an action before the server has confirmed it. In an online-first app, optimistic UI is a deliberate trick that needs a rollback path for when the server says no. Here you get it for nothing, because the local write is the real write. At the same moment, one row is added to the outbox — a note to self that the server has not been told yet.
"How does server data arrive?" The sync engine pulls the changes and writes them into the local store. The screens update because they observe that store, not because sync told them to. There is no line of sync code inside any view.
The write path: the outbox
Outbox is the pattern name worth saying out loud in the interview. The name comes from an email outbox, and it works the same way. Press send with no signal and the message does not vanish, and it does not fail. It sits in a folder until the phone can deliver it.
In an app, the outbox is a table in the same local database as everything else. Every local edit writes two things, inside one transaction so that either both land or neither does:
- The change itself, to the table the screens read.
- One row in the outbox table, describing the change as an operation — a small record of what the user did ("set title to X"), not a copy of the finished object.
Both writes are on disk before the function returns. That is the rule that makes the whole design work: if the process is killed one millisecond later, the edit and the intention to sync it both survive.
Here is the shape of one outbox row:
struct Operation: Codable {
let idempotencyKey: UUID // generated once, when this row is written,
// and never regenerated however often it is sent
let kind: String // "create-note", "edit-title", …
let entityID: UUID // which record the operation is about
let payload: [String: String] // the new values, e.g. ["title": "Groceries"]
let enqueuedAt: Date // when the user did it — used for ordering
}
The code that empties the outbox is the drain loop. It takes the oldest row, sends it, waits for the server's answer, and deletes the row only once the server confirms.
When a send fails, the loop waits and tries again. It waits using the exponential backoff from the networking lesson (D2-01): each attempt waits roughly twice as long as the one before it. That way a struggling server is not hammered by every phone at once.
Three design points come up in every interview on this topic.
Idempotency keys
Idempotent describes an operation you can apply many times and still get the result of applying it once. An idempotency key is how you get that behaviour across a network. It is a unique string that the client generates once, at the moment it writes the outbox row, and attaches to every attempt to deliver that row.
Two more words, because the sequence below depends on them.
An acknowledgement — everybody says ack — is the server's reply that means "I received your operation and I applied it".
A lost ack is the case where the server did the work and its reply never reached the phone. The connection dropped, or the request timed out, or the train entered a tunnel between the request and the response. The work happened. The news of it did not arrive.
Here is how a lost ack becomes a real production bug. Follow the sequence:
- The user renames a note. The app writes the new title locally and appends operation #7 to the outbox.
- The drain loop sends #7. The server applies it and sends back an ack.
- Wi-Fi drops while the ack is in flight. The phone times out.
- The phone now cannot tell two situations apart: "the server never received it" and "the server applied it and the reply was lost". From the client, those look identical. There is no way to find out.
- Operation #7 is still in the outbox, so the drain loop sends it again. That second send is the retry.
- A server with no idempotency key treats the retry as a brand-new request and applies it a second time. The user's note is renamed twice, or — in a payments app — the money is sent twice.
Step 6 is the only step you can change. Steps 1 to 5 are just the network being a network, so the client has to keep retrying. That policy has a name: at-least-once delivery — every operation reaches the server one or more times, but never zero times. The server then removes the duplicates itself: it stores the keys it has already applied, and when a key it has seen arrives again it skips the work and replies "already done".
Put the two halves together and you get the sentence worth saying in the interview: exactly-once effect from at-least-once delivery. You cannot make the network deliver each message exactly once. You can make repeated delivery change the data only once.
You have already met this rule inside a single device. SwiftData's @Attribute(.unique) (D2-02) turns a re-imported row into an update rather than a duplicate. Notification identifiers do the same job for scheduled alerts: scheduling a request whose identifier is already pending replaces that request instead of stacking a second reminder. P3-02 covers those later.
The idempotency key is that same rule applied across the network boundary. It is the single most important word in the sync interview.
Operations, not snapshots
A snapshot is the whole record as it looked after an edit. An operation is the change itself.
Two operations from two devices can both be applied: one sets the title, the other flips the star, and nothing is lost. Two snapshots cannot. Whichever snapshot arrives second overwrites the whole record, including the fields it never meant to touch, because it carries old values for those fields. Fine-grained operations cut down the number of genuine conflicts before any conflict rule has to run at all.
Ordering and dependency
"Edit note 7" must never reach the server before "create note 7". The usual rule is FIFO per entity — first in, first out, with one queue per record — so operations on the same note keep their order while unrelated notes drain in parallel.
Creates need one extra step. The phone invents a temporary local id, so that the screen has something to work with straight away. The server assigns the real id when the create finally lands, and its reply says which temporary id became which real id. The client then rewrites the operations still waiting in the outbox to use the real id.
The demo below is that whole story in one screen of code. The server keeps the set of keys it has already applied. The client keeps a row until it is acked. The ack for k2 is thrown away the first time it is sent, exactly as a tunnel would throw it away.
Read the output as the k2 story, one printed line at a time:
send k1— the server has never seen this key, so it applies the operation and printsserver: applied create note A. The ack arrives, and the client removes the row from the outbox.send k2— the server applies the edit and printsserver: applied edit title of A. Then its ack is dropped, which is whatdropAcksimulates.client: no ack — will retry k2— the client has no idea the edit landed. It keeps the row.send k2again — the server recognises the key it has already applied. It printsserver: k2 already applied — ack onlyand does not apply the edit a second time. This ack gets through, so the row finally leaves the outbox.send k3— the star, applied once and acked normally.
The last two lines report the outcome: the server's log holds create note A, edit title of A and star A, and the check prints edits applied once each: true. Four sends, three effects.
Now delete the seenKeys check from receive and run the same trace in your head. The retry in step 4 applies the edit again, so the server log reads create note A, edit title of A, edit title of A, star A, and the final check prints false. The user sees a rename they made once appear twice. Those few lines are the entire difference between a sync engine and a duplicate-data generator.
The read path: pull, apply, converge
Pull is the other half of the sync engine. It asks the server "what has changed since last time?" and writes the answer into the local database. Four things matter here, and the last two are what interviewers listen for.
The cursor
A cursor — also called a sync token — is a string the server hands back at the end of every pull. It is opaque: the client never reads it or reasons about it, it only stores it and sends it back next time. It means "you now have everything up to here".
The tempting alternative is to send a time instead: "give me everything changed since 14:05". Here is why that loses data. Follow the sequence:
- The phone's clock is five minutes fast. It is 14:00 on the server and 14:05 on the phone.
- The phone pulls, and stores "last synced at 14:05" — a time it read from its own clock.
- At 14:02 by the server's clock, somebody edits a note on another device. The server records that change at 14:02.
- The phone pulls again and asks for everything changed since 14:05. The 14:02 edit is older than that, so the server does not send it.
- The next pull asks for changes since an even later time. The 14:02 edit is never sent again.
The edit is not delayed. It is skipped, silently and permanently, and no error is raised anywhere. Device clocks are set by users, drift, and jump on time-zone changes. A server-issued cursor has none of those problems, because only one machine ever writes it.
Applying what comes back
The delta is the list of changes the server returns for your cursor. Applying it needs nothing new: upsert each record by its stable id, exactly as @Attribute(.unique) does in D2-02.
That makes the pull idempotent too. Applying the same delta twice leaves the database in the same state as applying it once, which is what makes it safe to retry a pull that failed halfway through.
Tombstones
A tombstone is a record that says "id X was deleted". It exists because a delta cannot express a deletion by leaving something out. Follow the sequence:
- Device A deletes note 7. The server removes the row from its table.
- Device B has been offline for a day. It comes back and pulls the delta for its cursor.
- The delta lists the notes that changed. Note 7 no longer exists, so there is nothing to put in the list.
- Device B hears nothing about note 7 and keeps showing it. Forever, because every later delta also says nothing about it.
In a delta protocol, leaving something out already means "unchanged". So deletion has to arrive as its own explicit record. The server keeps tombstones for a retention window — long enough to cover a plausible offline gap — then deletes them for good. A client whose cursor is older than that window is told to throw its cursor away and re-sync from scratch. Apple's CloudKit has an error for exactly this case, CKError.Code.changeTokenExpired.
Echo suppression
Your own edit comes back to you on the next pull, because from the server's point of view it is simply a change like any other. Applying it again is harmless: the upsert writes values the local store already has.
The bug appears in apps that compare the incoming record with the local one and then react to the difference. They show "edited on another device", or bump an unread badge, or play a sound. Then every edit you make announces itself back to you a few seconds later.
Echo suppression is the name for the cure. Stamp every operation with the id of the device that created it. When a change comes back carrying your own device id, apply it as normal, but do not notify anybody about it.
Conflicts: choose a rule you can defend
A conflict is two edits to the same record, made from the same starting point, where neither device saw the other's edit. Two people renamed the same note while both were offline. Somebody has to decide what the note is called now. To merge is to produce one record out of both edits.
These are the four rules to choose from, in the order you should offer them:
| Rule | How it works | Right for | What it costs |
|---|---|---|---|
| Last-write-wins (LWW) | the server orders the writes by timestamp or version number, and the later one replaces the earlier one completely | most fields in most apps | one edit is lost, and nobody is told |
| Field-level merge | last-write-wins applied per field — the title from device A, the star from device B | records with separate, independent fields | more bookkeeping: you must keep the version each edit started from |
| Domain merge | rules written for the kind of data — text merges line by line, sets take the union, counters add up (CRDTs are the formal version of this) | collaborative text, counters, shared lists | genuine engineering effort |
| Ask the user | keep both copies and make the person choose | rare conflicts where the stakes are high, such as documents | it interrupts people, and they dislike it |
CRDT stands for conflict-free replicated data type: a data structure designed so that any two replicas that have seen the same set of edits end up identical, whatever order those edits arrived in. It is the right tool for shared text, and it is far more work than most apps need.
The strong interview answer picks a rule per kind of data, rather than one rule for the whole app. It sounds like this: "field-level merge for the note's metadata. For the note body, a domain merge if we are shipping collaboration, otherwise last-write-wins with version history as the safety net."
A single blanket answer sounds inexperienced in one of two directions. "CRDTs everywhere" is over-engineering. "Just last-write-wins" ignores the edits it throws away.
One mechanism sits under all four rules. The server keeps a version number on every record, and every edit a client sends carries the version it was based on — its base version. If the base version matches the server's current version, nobody else edited in the meantime: apply the change and increase the version. If the base version is behind, somebody did edit in the meantime, and that is the conflict path where one of the four rules runs.
The block below compares the first two rules on exactly the same pair of edits.
The same two offline edits, run through two different conflict rules. Which edits survive each one?
struct Note {
var title: String
var starred: Bool
var version: Int
}
// Base record both devices synced at version 1:
let base = Note(title: "Groceries", starred: false, version: 1)
// Offline, device A renames; device B stars. Both based on v1:
let fromA = Note(title: "Groceries + pharmacy", starred: false, version: 1)
let fromB = Note(title: "Groceries", starred: true, version: 1)
// Strategy 1 — whole-record LWW, B arrives last:
var lww = fromA
lww = fromB
lww.version = 3
// Strategy 2 — field-level: keep each field that CHANGED from base.
var merged = base
if fromA.title != base.title { merged.title = fromA.title }
if fromB.title != base.title { merged.title = fromB.title }
if fromA.starred != base.starred { merged.starred = fromA.starred }
if fromB.starred != base.starred { merged.starred = fromB.starred }
merged.version = 3
print("LWW: \(lww.title) | starred \(lww.starred)")
print("merged: \(merged.title) | starred \(merged.starred)")The follow-up the interviewer always asks: "the user edits a note, immediately sends the app to the background, and the subway kills the connection. What guarantees that the edit eventually reaches the server?"
Run the whole interview on yourself, out loud or on paper: "design offline-first sync for a shared grocery-list app, used by two family members who are both often offline." Cover five things. The diagram. What each kind of edit puts in the outbox. Where the idempotency key comes from. The conflict rule for ticking an item off, for renaming an item, and for reordering the list — they are not the same rule. And the one place where you would accept losing an edit silently, plus why that place is acceptable. Give it twenty minutes, the length of the real thing.
You can now:
- Open the system-design answer with "the local database is the source of truth", and defend it
- Build an outbox with idempotency keys, and narrate the lost-ack sequence step by step
- Run a three-way field merge, and say exactly where it stops helping
- Choose a conflict rule per kind of data, the way someone who has shipped sync does
Next up: Keychain and security — where credentials live, and the storage questions that decide interviews.