Persistence and networking · advanced
Networking in production: transport, pagination and evidence
The networking questions that only appear after launch. Which HTTP version you are actually on, why a feed shows the same row twice, and how to answer "the API is slow" with evidence instead of a guess.
By the end you will be able to
- Explain what HTTP/2 and HTTP/3 change for a phone, and stop hand-rolling request batching
- Configure timeouts and connectivity waiting, and use reachability as a hint rather than a gate
- Paginate without drift, and diagnose a slow request from its own timing breakdown
Before every request, an app checks the current network path and shows "You're offline" when it is unsatisfied. What is wrong with that design?
Which HTTP version are you actually on
Every request you have written so far behaves identically whether it runs over HTTP/1.1, HTTP/2 or HTTP/3. URLSession negotiates the highest version both sides support and gives you the same API regardless. That is genuinely convenient, and it means most engineers never learn what changed — which matters, because one of the changes retired a piece of advice that is still repeated.
HTTP/1.1 carries one request at a time per connection. The reply must arrive before the next request goes out on that connection, so one slow response blocks everything queued behind it. That is head-of-line blocking: the item at the front of the line holds up the rest. Clients worked around it by opening several connections to the same host at once — commonly around six — and every one of those connections pays for its own DNS lookup and TLS handshake.
HTTP/2 replaced that with multiplexing: many requests and responses travel over a single connection at the same time, split into interleaved frames. One handshake, no per-request queue, and headers are compressed rather than repeated in full on every request.
This is the part with a practical consequence. The old advice — "batch your API calls, requests are expensive" — was true when each one might cost a fresh connection. Under HTTP/2 the marginal cost of one more request on an open connection is small, so a custom batching endpoint often buys much less than it costs in complexity on both sides. Measure before you build one.
HTTP/2 still has one weakness, and it is the one phones feel. Its streams are independent to HTTP, but they all sit on a single TCP connection, and TCP delivers strictly in order. Lose one packet and every stream waits for the retransmission. Head-of-line blocking moved down a layer rather than disappearing.
HTTP/3 fixes that by not using TCP. It runs on QUIC, a transport built on UDP that implements streams itself, so a lost packet stalls only the stream it belonged to. Two consequences matter on a phone:
- Loss tolerance. Mobile links drop packets far more than office ethernet, which is exactly the condition where independent streams pay off.
- Connection migration. A QUIC connection is identified by a connection id rather than by the IP-address-and-port pair TCP uses. A user walking out of the house, off Wi-Fi and onto cellular, keeps the same connection instead of re-establishing it. On TCP that switch costs a new handshake and, usually, a visible stall.
Apple's networking stack has supported HTTP/2 for many releases and added HTTP/3 in iOS 15. You do not opt in per request; the server advertises what it speaks and URLSession chooses.
What all of this means for your code is short:
- Reuse one
URLSession. Creating a session per request throws away connection reuse, and hands you a fresh DNS lookup and TLS handshake every time — several hundred milliseconds you did not need to spend. Hold one session in your networking layer and pass it around. - Do not hand-roll batching without measuring first.
- Check what you actually negotiated before optimising anything, using the metrics below.
networkProtocolNamereportsh2orh3, and a lot of "we're on HTTP/2" turns out to be wrong.
Timeouts, waiting, and the network you actually have
Three configuration values do most of the work, and the first two are routinely confused:
let configuration = URLSessionConfiguration.default
configuration.timeoutIntervalForRequest = 30 // stalled with no new data
configuration.timeoutIntervalForResource = 300 // total budget for the whole thing
configuration.waitsForConnectivity = true // queue rather than fail when offline
let session = URLSession(configuration: configuration)
timeoutIntervalForRequest is not a deadline for the request. It is how long the task may sit with no data arriving before it gives up — an inactivity timer that resets every time bytes turn up. Its default is 60 seconds.
timeoutIntervalForResource is the real deadline: the maximum time for the whole resource, from start to last byte, no matter how much progress is being made. Its default is seven days, which is sensible for a background download and far too generous for a screen a user is watching. If you want "give up after five minutes", this is the one to set.
waitsForConnectivity changes what happens when there is no usable network at the moment you send. The default is to fail immediately with a "not connected" error. Set it to true and the task waits instead, starting as soon as connectivity appears — with timeoutIntervalForResource as the ceiling on that patience. It pairs naturally with the pretest's rule: rather than checking the network yourself, hand the waiting to the session.
Two more flags describe the kind of connection your request is allowed to use:
allowsConstrainedNetworkAccess— setfalseand the request is refused when the user has turned on Low Data Mode. Right for prefetching, background refresh and high-resolution images; wrong for anything the user is waiting on.allowsExpensiveNetworkAccess— setfalseand the request is refused on connections the system considers expensive, typically cellular and personal hotspots.
Respecting Low Data Mode is a small piece of work that users genuinely notice, and it is the kind of detail an interviewer takes as evidence you have shipped something.
Pagination that does not drift
A feed loads twenty rows, then twenty more as the user scrolls. The obvious way to ask for the next twenty is to say how many to skip:
GET /feed?offset=0&limit=20
GET /feed?offset=20&limit=20
It works perfectly on a list that never changes, and every feed worth paginating changes constantly.
The feed is sorted newest first. The user loads page one, reads it for a few seconds, and during those seconds somebody posts a new row. Then they scroll, and the app asks for page two.
struct Row {
let id: Int
let title: String
}
struct Feed {
var rows: [Row]
func page(offset: Int, limit: Int) -> [String] {
let sorted = rows.sorted { $0.id > $1.id }
guard offset < sorted.count else { return [] }
let end = min(offset + limit, sorted.count)
return sorted[offset..<end].map { $0.title }
}
}
var feed = Feed(rows: [
Row(id: 5, title: "eggs"),
Row(id: 4, title: "drinks"),
Row(id: 3, title: "coffee"),
Row(id: 2, title: "bread"),
Row(id: 1, title: "apples"),
])
print("page 1:", feed.page(offset: 0, limit: 2))
feed.rows.append(Row(id: 6, title: "figs"))
print("page 2:", feed.page(offset: 2, limit: 2))The repair is to stop counting positions and start naming a place. Cursor pagination asks "give me what comes after this row", and the answer does not depend on how many rows were added or removed above it:
GET /feed?limit=2
GET /feed?after=4&limit=2A real API hands you the cursor as an opaque string rather than a row id — "eyJpZCI6NH0" — and that opacity is deliberate: it lets the backend change what a cursor encodes without breaking installed apps. Treat it as a token to store and send back, never as something to parse. D2-04 uses exactly the same idea for sync checkpoints, and for exactly the same reason.
Cursors have an honest cost. You cannot jump to page 7 — there is no cursor for a page nobody has fetched — so "load more" works and numbered pages do not. For an infinite feed that is no loss at all. If a screen genuinely needs page numbers, offsets are the trade-off you accept, with the drift on the record.
Choosing a realtime transport
When the server needs to tell the app something without being asked, four options are on the table:
| Approach | How it works | Reach for it when |
|---|---|---|
| Polling | Ask every N seconds | Updates are rare and lateness is cheap. Simple, and wasteful at scale |
| Long polling | Request that the server holds open until it has news | You need push-like latency and cannot deploy anything better |
| Server-Sent Events | One long-lived HTTP response streaming text/event-stream | Server-to-client only: prices, statuses, progress. Plain HTTP, so proxies and CDNs handle it |
| WebSocket | A connection upgraded to full two-way messaging | Both sides send often — chat, collaborative editing, order entry |
URLSessionWebSocketTask is the built-in client for the last one. It is genuinely easy to open and genuinely easy to get wrong, because a socket that looks open is not the same as one that works:
- Heartbeats. Send a periodic ping. Mobile connections die silently — a phone that loses signal in a lift leaves a socket that reads as connected and delivers nothing. Without a heartbeat the app shows stale data with no indication anything is wrong.
- Reconnect with backoff and jitter. The same discipline as D2-01's retries. A server restart otherwise brings every client back at the same instant.
- Resume from a sequence number. On reconnect, tell the server the last message you saw so it can send what you missed. Without it every reconnect silently loses data — which is precisely the sequence-gap problem D1-07 builds a whole architecture around.
The judgment worth carrying into an interview: a WebSocket is the answer when the client sends often too. If traffic is one-directional, Server-Sent Events give you most of the benefit with far less to operate, because it is ordinary HTTP the whole way down. And for updates that are rare but must arrive even when the app is closed, the right transport is a push notification that wakes the app to fetch — P3-02 and P3-03 cover that path.
Answering "the API is slow" with evidence
A product manager says the transfers screen is slow. The backend team says their endpoint answers in 40 milliseconds. Both are telling the truth, and without measurements the conversation goes nowhere.
URLSession collects the breakdown for you. Implement one delegate method and every request reports where its time went:
func urlSession(_ session: URLSession,
task: URLSessionTask,
didFinishCollecting metrics: URLSessionTaskMetrics) {
for transaction in metrics.transactionMetrics {
// domainLookupStart/End — DNS
// connectStart/End — TCP or QUIC setup
// secureConnectionStart/End — TLS handshake
// requestStart → responseStart — the server's own thinking time
// responseStart → responseEnd — downloading the body
print(transaction.networkProtocolName ?? "?", // "h2" or "h3"
transaction.isReusedConnection) // was a handshake avoided?
}
}
Those timestamps split one number everybody argues about into four that each point at a different owner:
TTFB is time to first byte — the gap between your request going out and the first byte of the reply arriving. It is the closest thing to "how long the server took", and it is the only phase the backend team owns.
Read the two reports together and the argument resolves itself. The endpoint really does answer in about 60 milliseconds, both times. The cold request spent 420 of its 500 milliseconds on DNS and TLS — the cost of not having a connection yet, which is the app's problem and is what session reuse removes.
Two habits make this usable beyond your own laptop:
Send a request id. A UUID per request in a header both sides log, as D2-07 suggested. When a user reports a failure at 14:32, one id joins your logs to the backend's without anyone grepping by timestamp and hoping.
Log the phase breakdown, not just the total. A total tells you something is slow. A breakdown tells you whose it is — and lets you notice that a region's DNS is failing over, or that connection reuse quietly stopped working after a refactor.
Testing the layer
D2-01 and P1-01 built the seam approach: a Transport protocol, a fake that returns the status and body a test needs, and no network anywhere. It stays the right default. It is fast, it cannot flake, and it tests your logic.
There is a second technique worth knowing, because it tests what the seam cannot. A URLProtocol subclass registered on a URLSessionConfiguration intercepts requests inside URLSession itself:
final class StubProtocol: URLProtocol {
static var handler: ((URLRequest) -> (HTTPURLResponse, Data))?
override class func canInit(with request: URLRequest) -> Bool { true }
override class func canonicalRequest(for request: URLRequest) -> URLRequest { request }
override func startLoading() {
guard let handler = StubProtocol.handler else { return }
let (response, data) = handler(request)
client?.urlProtocol(self, didReceive: response, cacheStoragePolicy: .notAllowed)
client?.urlProtocol(self, didLoad: data)
client?.urlProtocolDidFinishLoading(self)
}
override func stopLoading() {}
}
let configuration = URLSessionConfiguration.ephemeral
configuration.protocolClasses = [StubProtocol.self]
let session = URLSession(configuration: configuration)
The difference is what each one exercises. A seam fake replaces your networking layer, so nothing tests whether that layer builds the right request. URLProtocol sits below it, so the real URLRequest — method, headers, encoded body and all — arrives at the stub and can be asserted on.
So: seam fakes for view models and everything above them; URLProtocol for the networking layer itself. Use the first to prove the screen shows an error when a fetch fails, and the second to prove the Idempotency-Key header is actually on the request. Being able to say which belongs where — and why testing both at the top would be redundant — is the answer that lands in an interview.
Your app creates a fresh URLSession inside each repository method, then discards it. Metrics show DNS and TLS time on nearly every request. What is the single best fix?
An interviewer says: "Users report our feed sometimes shows duplicate posts, and it feels slow on cellular." You have the codebase but no reproduction. Write the plan. Say what you would measure first and with which tool, what the duplicate symptom alone already tells you about the pagination scheme, which two configuration values you would check before writing any code, and what you would ask the backend team for. Then name the one change you would ship first if you only had a day, and why that one.
You can now:
- Say what HTTP/2 and HTTP/3 change, and why reusing one
URLSessionis usually the largest easy win available - Set request and resource timeouts for what they actually mean, and let the session wait for connectivity instead of pre-checking it
- Paginate with a cursor, and explain the silent skip that offsets produce
- Choose a realtime transport, and name the three things a WebSocket needs before it is production-ready
- Turn "the API is slow" into a phase breakdown that names an owner
- Split your tests: seam fakes above the networking layer,
URLProtocolstubs for the layer itself
Next up: the D track is complete. D1-06 and D1-07 are the capstones to reread before a system-design interview — this unit is what they are built on.