Persistence and networking · core

Authentication: tokens, JWT and the sign-in flow

What the string in the Authorization header actually is, why anyone holding it can read it, and how a device that cannot keep a secret still signs a user in safely.

15 min read14 min practice0/2 exercises5 recall cards

By the end you will be able to

  • Read a JWT's three segments and its claims, and explain why signed does not mean private
  • Refresh a token before it expires rather than after a 401, allowing for a phone clock that is wrong
  • Explain the authorization code flow with PKCE, and why a native app needs it
Guess firstAnswering before you read makes the explanation stick — even when you get it wrong.

Your backend puts the user's email address, account tier and internal customer id into the JWT it issues. The token is stored in the Keychain and only ever sent over HTTPS. Who can read those three values?

The two tokens, briefly

D2-05 introduced the pair, and D1-06 made them work under concurrency. A one-paragraph recap, because everything below builds on them.

An access token is what proves who is calling. It rides on every request in the Authorization header, and it is deliberately short-lived — typically minutes to an hour. A refresh token has exactly one job: exchange it for a new access token when the old one expires. It lives much longer, and it is the more valuable of the two, which is why D2-05 stores it with kSecAttrAccessibleAfterFirstUnlockThisDeviceOnly and never in UserDefaults.

Short-lived access tokens are a containment decision. If one leaks, the window in which it is useful is small. The refresh token, which would be catastrophic to leak, travels rarely — only to the one endpoint that renews sessions.

Inside a JWT

JWT stands for JSON Web Token, and it is pronounced "jot". It is a way of packing facts about a signed-in user into a single string that can travel in an HTTP header.

Take one apart and it is three pieces joined by full stops:

eyJhbGciOiJIUzI1NiJ9  .  eyJzdWIiOiI4ODIxIiwiZXhwIjoxNzU1MDkwMDAwfQ  .  9f2c7ab41d
    header                              payload                          signature

The header says which algorithm signed the token. The payload carries the claims. The signature is computed by the server over the first two segments using a key only the server holds.

A claim is one statement about the user or the token, written as a JSON key and value. Seven have standard names, and you will meet them constantly:

ClaimShort forMeans
ississuerwho issued this token
subsubjectwho it is about — normally the user id
audaudiencewhich service is meant to accept it
expexpiration timethe instant it stops being valid
iatissued atwhen it was created
nbfnot beforedo not accept it before this instant
jtiJWT ida unique id for this token, so it can be revoked individually

exp, iat and nbf are all seconds since 1 January 1970 UTC — plain integers, not formatted dates.

Reading a token needs no key and no permission, which the simulation below makes uncomfortably clear. The runtime here cannot do base64, so the payload segment is written as key=value pairs instead of encoded JSON; every other step is what a real decoder does:

Loading runnable Swift…

Nothing in that code is privileged. String splitting recovered the user id, the tier and the issuer.

That second rule has a practical edge worth stating plainly, because it sounds stricter than it is. Reading tier to decide whether to show the Premium tab is fine: the worst case is a wrongly drawn tab, and the server still refuses the request behind it. Reading tier and letting the app unlock a paid feature locally is not fine — that check runs on a device the user controls.

Refresh before the 401, not after

The simplest possible design refreshes on failure: send the request, get a 401, refresh, retry. It works, and it costs every user a doubled request every time a token expires. You already have the information to avoid that, because exp says exactly when the token dies.

Refresh proactively, with a margin:

Loading runnable Swift…

The margin exists because a token that is valid now may not be valid when it arrives. A request sent with four seconds left, over a slow connection, reaches the server after it has expired. Sixty seconds is a common choice.

Proactive refresh does not let you delete the 401 handling. The token can be revoked from the server side at any moment — the user changes their password, an administrator ends the session, the account is locked — and no amount of reading exp predicts that. So you need both: the margin removes the routine case, and the 401 path catches everything else.

Predict, then runA wrong prediction you have committed to is worth more than a right answer you read.

Ledger refreshes when the access token is expired according to the device's clock. This user's phone clock is four minutes fast. The token is genuinely valid for another 200 seconds by the server's clock.

let serverNow = 1000
let phoneNow = serverNow + 240
let expiresAt = 1200

func isExpired(expiresAt: Int, now: Int) -> Bool {
    return now >= expiresAt
}

print("server thinks expired:", isExpired(expiresAt: expiresAt, now: serverNow))
print("phone thinks expired:", isExpired(expiresAt: expiresAt, now: phoneNow))

Two 401s, one refresh

When several screens hit a 401 in the same instant, they must not each start their own refresh. D1-06 builds that guarantee in full — the actor-reentrancy trap, the stored Task, and why a server that rotates refresh tokens treats a replayed one as theft and revokes the whole session. Reread it after this lesson; the two fit together directly.

Rotation is the piece to carry forward here: many auth servers issue a new refresh token with every refresh and immediately retire the old one. That makes a stolen refresh token detectable — if the thief uses it, the real user's next refresh presents a retired token, and the server can end every session. It also means the refresh path has no room for duplicated calls.

Signing in: the flow, and why a phone needs a special one

Everything so far assumed the app already had a token. Getting the first one is where the interesting design lives.

Four pieces of vocabulary make the rest readable. The resource server is your API — it holds the data. The authorization server is the service that authenticates users and issues tokens; it may be your own, or Google, or a bank's. The client is the thing asking for a token, which here is your app. And clients come in two kinds: a confidential client can keep a secret, because it runs on a server you control; a public client cannot.

A mobile app is always a public client. Anyone can download your app from the App Store, unzip the IPA and read every constant in the binary. Any "client secret" you ship is public the day you ship it. That single fact shapes the entire flow.

The flow that solves it is the authorization code flow with PKCE. PKCE — "Proof Key for Code Exchange", said as "pixie" — is the part that makes it safe for a client that cannot keep a secret.

Here is the whole sequence:

  1. Your app invents a large random string called the code verifier, and keeps it in memory.
  2. It hashes the verifier with SHA-256 and base64url-encodes the result. That is the code challenge.
  3. The app opens the authorization server's login page in the system browser, sending the challenge — never the verifier — plus a random state value.
  4. The user signs in there, with their password, their second factor, their passkey. Your app sees none of it.
  5. The browser redirects back to your app's URL scheme carrying an authorization code and the same state you sent.
  6. Your app checks state matches what it sent in step 3, then sends the code back to the authorization server — this time with the verifier.
  7. The server hashes the verifier it just received and compares it to the challenge from step 3. They match only for the app that invented the verifier, so it issues the tokens.

Step 7 is the whole point. An attacker who intercepts the authorization code — through a malicious app that registered the same URL scheme, for instance — cannot exchange it, because they never saw the verifier and cannot work it out from the challenge.

Loading runnable Swift…

The state value from step 3 does a separate job. It is a random string your app generates per sign-in and checks on the way back. Without it, an attacker can send your app a redirect carrying their authorization code; your app exchanges it and quietly signs the user into the attacker's account, where everything the user then does is visible to them. Checking state costs three lines, and exercise 2 writes them.

The API that does this properly

On Apple platforms you do not build the browser part yourself. ASWebAuthenticationSession presents the login page in a Safari-backed view that your app cannot read, and hands you the redirect URL when it is done:

import AuthenticationServices

let session = ASWebAuthenticationSession(
    url: authorizationURL,                 // carries the challenge and the state
    callbackURLScheme: "ledger"            // ledger://auth?code=…&state=…
) { callbackURL, error in
    // Check state, then exchange the code for tokens.
}
session.presentationContextProvider = self
session.prefersEphemeralWebBrowserSession = false   // true = no shared cookies
session.start()

Two properties are worth knowing. The system shows a consent sheet naming the domain, because the session can share cookies with Safari — which is what lets a user already signed in on the web skip typing anything. Setting prefersEphemeralWebBrowserSession = true opts out of that sharing, which is what you want for an app where two people might sign in on one device.

The comparison that gets asked:

ASWebAuthenticationSessionWKWebView login
Can the app read the passwordNoYes — which is the objection
Shares the Safari sessionOptionallyNo
Passkeys and provider 2FAWorkOften break
RFC 8252 compliantYesNo
Provider termsAcceptedFrequently banned

Sign in with Apple

Apple's own identity provider, through AuthenticationServices. Two things make it worth its own mention. It gives users a private relay email, so they can sign up without handing over a real address — which means your backend must treat the relay address as a real, permanent identifier and must not attempt to email it from anywhere except a registered domain. And App Review guideline 4.8 requires it in apps that offer third-party sign-in with comparable features; P2-01 covers where that bites at submission time.

What signing out has to guarantee

Signing out is not "delete the token and show the login screen". Ship that and users stay signed in to a system that thinks they left. A complete sign-out does five things:

  1. Tell the server. Call the revocation endpoint so the refresh token is retired. Without this the session lives on until it expires naturally, and a copy of the refresh token still works.
  2. Delete both tokens from the Keychain. Note that Keychain items survive app deletion — D2-05's point — so an incomplete sign-out can outlive the app itself.
  3. Clear the caches holding personal data. The URLCache from D2-03 and any on-disk models from D2-02 or D2-06. A cached feed that renders after sign-out shows one user's balance to the next person holding the phone.
  4. Cancel work in flight. A refresh already running will otherwise write a fresh token into the Keychain seconds after you cleared it, silently signing the user back in.
  5. Reset the in-memory state. New view models, empty navigation stack, nothing carried across.

Point 4 is the one that gets shipped broken most often, and it is invisible in testing because it needs a sign-out during a refresh to reproduce.

Proving the request came from your app

One more layer, and it matters for fintech specifically. A token proves who the user is. Nothing so far proves the request came from your app rather than a script replaying your API with a stolen token.

App Attest, in the DeviceCheck framework, is Apple's answer. The app asks the Secure Enclave to generate a hardware-backed key, Apple's servers vouch that it belongs to a genuine, unmodified build of your app, and your server can then require a signed assertion on requests that matter. It raises the cost of scripted abuse substantially — credential stuffing, scraping, bot signups — without touching the honest user's experience.

The honest counterweight: it is not a jailbreak detector and it is not unbreakable. It authenticates the app instance, so it is worth the integration effort on high-value endpoints — signup, payments, transfers — and rarely worth it on a public content feed.

Check yourself

An engineer proposes putting "role": "admin" in the JWT and having the iOS app read it to decide whether to allow the admin actions in the app. What is the strongest objection?

Loading exercise…
Loading exercise…
Explain itSaved on this device. Never graded.

Explain to a backend engineer on your team why the iOS app cannot enforce the role claim, even though the token is signed and arrives over HTTPS. Then explain the opposite case: why the app reads that same claim to decide which tabs to show, and why that is not a contradiction. Finish with the one sentence you would put in the API documentation to stop this coming up again.

Checkpoint

You can now:

  • Take a JWT apart, name its claims, and explain why signed does not mean private
  • Draw the line between reading a claim to decide what to draw and using one to decide what is allowed
  • Refresh on a margin rather than on a 401, and say why the device clock cannot be trusted
  • Walk the authorization code flow with PKCE, and say what state and the verifier each defend against
  • List what a complete sign-out has to do, including the refresh already in flight

Next up: the networking questions that only appear after launch — which HTTP version you are actually on, pagination that drifts, and how to find out where a slow request spent its time.

How well do you know this now? Rating yourself honestly, then being tested on it, is how you find out where your intuition is wrong.