Quality and craft · advanced

Diagnostics: sanitizers, the memory graph and logs

The tools that find code that is wrong rather than code that is slow — what each sanitizer detects, how to read a memory graph, and how to log in a way that survives production without leaking your users' data.

13 min read12 min practice0/2 exercises6 recall cards

By the end you will be able to

  • Choose the right sanitizer for a symptom, and explain why a clean run is not a proof
  • Read a memory graph, and define a leak precisely enough to tell it apart from ordinary growth
  • Log with levels and privacy so production problems are diagnosable and users' data stays out of the log
Guess firstAnswering before you read makes the explanation stick — even when you get it wrong.

You enable Thread Sanitizer, run the full test suite, and it reports nothing. What have you proved?

Two kinds of problem, two kinds of tool

P1-03 covered the tools that answer "why is this slow?" — Time Profiler, Allocations, the Hangs instrument, signposts.

This lesson covers the other question: "why is this wrong?" Crashes with no obvious cause, memory that never comes back, a value that changes when nothing should have changed it, a screen that renders nothing. Different symptoms, different instruments, and a different way of working — because these tools mostly do not measure anything. They watch for a rule being broken and tell you the moment it happens.

The sanitizers

A sanitizer is a mode you switch on in the scheme editor — Product ▸ Scheme ▸ Edit Scheme ▸ Run ▸ Diagnostics. Xcode rebuilds your app with extra instrumentation, so the resulting binary is slower and larger and is never something you ship. You turn one on for an afternoon when you have a specific suspicion.

ToolCatchesCostReach for it when
Thread Sanitizer (TSan)Data races — two threads touching the same memory with no synchronisation, one of them writingHeavy: several times slower, several times more memoryValues change unpredictably; crashes that never reproduce twice; anything @unchecked Sendable
Address Sanitizer (ASan)Use-after-free, buffer overruns, double free, stack-use-after-returnRoughly twice as slow, plus large memory overheadEXC_BAD_ACCESS with a useless stack; corruption that moves when you add a print
Undefined Behavior Sanitizer (UBSan)Misaligned pointers, invalid enum and boolean values, C integer overflowLightYou bridge to C or Objective-C, or wrap a third-party framework
Main Thread CheckerUIKit and AppKit called off the main threadNegligible — on by default for debug runsAlways. Read its purple warnings rather than dismissing them
Zombie ObjectsMessages sent to a deallocated Objective-C objectMemory grows for the whole run, since nothing is truly freedOver-release crashes in Objective-C or older bridged code
Malloc Scribble / Guard MallocReads of freed memory; writes past an allocationHeavyASan pointed at heap corruption and you need a second opinion

Two facts about running them:

TSan and ASan cannot be enabled together. They both take over memory management in incompatible ways, so Xcode makes them mutually exclusive. Pick the one matching your symptom: threading, or memory validity.

Main Thread Checker is already on. It is the one in that table you are using right now without knowing it. Those purple runtime issues that appear in the Issue Navigator saying a UIKit method was called on a background thread are real bugs that happen to have worked so far.

The pretest's point, made concrete:

Loading runnable Swift…

retryDrain and exportCSV are exactly the kind of path a test suite skips — an error recovery route and an export nobody wrote a test for. Both are broken. The tool said nothing, and it was not wrong to say nothing.

This is why the technique that works is run your app under TSan, do not just run your tests. Use the features by hand for twenty minutes with the sanitizer on. You will cover paths no test reaches.

The Memory Graph Debugger

Run the app, exercise the screen, and press the Debug Memory Graph button in the debug bar. Execution pauses and Xcode shows every object alive right now, with arrows for the references between them.

Before reading one, get the definition exact, because it is the difference between two very different bugs:

A leak is memory that is unreachable from any root and still has not been freed.

"Root" means something the app inherently holds: the app delegate, a scene, a global, a singleton. ARC frees an object when its last strong reference goes — so an object nothing can reach should be gone. If it is still there, something is holding it in a circle that ARC cannot break.

Loading runnable Swift…

That is what the Memory Graph Debugger computes and draws. The pair with the purple ! badge is the cycle; clicking either one shows the arrows forming it, and the fix is the one S3-02 taught — make one of the two references weak or unowned, and it is the child's reference to its parent that gives way.

Two settings make the graph far more useful, both under Edit Scheme ▸ Run ▸ Diagnostics:

  • Malloc Stack Logging. Without it the graph tells you what is alive; with it, selecting an object shows the backtrace of where it was allocated. That turns "some URLSession delegate is leaking" into a file and a line.
  • Zombie Objects for the opposite failure — an object freed too early rather than too late.

The View Hierarchy Debugger

The visual counterpart, next to the memory graph button: Debug View Hierarchy. It pauses the app and shows the view tree exploded in 3D, with every view's frame, constraints and clipping.

It answers questions that are almost impossible to reason about from code. Why is nothing visible — is the view absent, or present with a zero-size frame, or sitting outside its parent's bounds, or at alpha 0, or behind something opaque? All four look identical on the device and are instantly distinguishable here. It also shows the constraints on a selected view, which is the fastest route to the ambiguous-layout problem from U6-02.

Logs that survive production

print() is fine while you are looking at Xcode's console and useless everywhere else. It has no levels, no structure, no privacy control, and it does not exist in the logs you can collect from a user's device.

The system logger does:

import os

private let logger = Logger(subsystem: "com.ledger.app", category: "transfers")

logger.debug("building request for \(endpoint)")
logger.notice("transfer submitted")
logger.error("transfer failed: \(status)")
logger.fault("keychain unreadable — session cannot continue")

Levels decide what is kept. debug is for development and is not persisted. info is kept only when logs are actively being collected. notice — the default — error and fault are persisted to disk and appear in a sysdiagnose, which means they are what you actually have to work with when a user reports a problem from a device you have never touched.

Privacy is enforced, and the defaults are not obvious. Values interpolated into a log message are redacted according to type: a dynamic string is replaced with <private> in the persisted log, while numbers are shown. You override it per value:

logger.error("transfer \(id, privacy: .public) failed for \(iban, privacy: .private)")
logger.notice("user \(userID, privacy: .private(mask: .hash)) upgraded")

The masked-hash form is the one worth knowing: it writes a stable hash instead of the value, so you can tell that the same user appears in twenty log lines without ever learning who they are.

Loading runnable Swift…

Reading them back: Console.app shows a connected device's live log with filtering by subsystem and category — which is why setting those two strings properly is worth the thirty seconds. On the command line, log stream --predicate 'subsystem == "com.ledger.app"' follows it live, and log collect pulls an archive off a device.

Check yourself

An app crashes with EXC_BAD_ACCESS about once every fifty launches. The stack trace points at a different, innocent-looking function every time. Which tool first, and why?

Check yourself

Memory climbs steadily while a user browses a photo feed and never comes back down. The Leaks instrument reports nothing. What is the most likely explanation?

Loading exercise…
Loading exercise…
Design itSaved on this device. Never graded.

An engineer on your team says "memory is going up while people scroll the feed, I think we have a leak." Write the diagnosis plan you would hand them. Say which tool you would open first and what result would confirm or eliminate "leak" as the explanation, what you would open second if it is eliminated, and what fix each of the two outcomes implies. Then write the one sentence you would add to the team's code review checklist so this class of bug is caught before it ships.

Checkpoint

You can now:

  • Pick the sanitizer that matches a symptom, and state what a clean run does and does not prove
  • Say what Swift 6 left for Thread Sanitizer to cover, and why
  • Define a leak by reachability, and tell it apart from unbounded growth — with the right tool and fix for each
  • Use the memory graph with Malloc Stack Logging, and the view hierarchy debugger for invisible views
  • Log with levels and privacy that keep a production problem diagnosable and users' data out of a sysdiagnose

Next up: the Instruments catalogue — every template worth knowing, and the four call-tree settings that turn an unreadable trace into an answer.

How well do you know this now? Rating yourself honestly, then being tested on it, is how you find out where your intuition is wrong.