Quality and craft · advanced
Diagnostics: sanitizers, the memory graph and logs
The tools that find code that is wrong rather than code that is slow — what each sanitizer detects, how to read a memory graph, and how to log in a way that survives production without leaking your users' data.
By the end you will be able to
- Choose the right sanitizer for a symptom, and explain why a clean run is not a proof
- Read a memory graph, and define a leak precisely enough to tell it apart from ordinary growth
- Log with levels and privacy so production problems are diagnosable and users' data stays out of the log
You enable Thread Sanitizer, run the full test suite, and it reports nothing. What have you proved?
Two kinds of problem, two kinds of tool
P1-03 covered the tools that answer "why is this slow?" — Time Profiler, Allocations, the Hangs instrument, signposts.
This lesson covers the other question: "why is this wrong?" Crashes with no obvious cause, memory that never comes back, a value that changes when nothing should have changed it, a screen that renders nothing. Different symptoms, different instruments, and a different way of working — because these tools mostly do not measure anything. They watch for a rule being broken and tell you the moment it happens.
The sanitizers
A sanitizer is a mode you switch on in the scheme editor — Product ▸ Scheme ▸ Edit Scheme ▸ Run ▸ Diagnostics. Xcode rebuilds your app with extra instrumentation, so the resulting binary is slower and larger and is never something you ship. You turn one on for an afternoon when you have a specific suspicion.
| Tool | Catches | Cost | Reach for it when |
|---|---|---|---|
| Thread Sanitizer (TSan) | Data races — two threads touching the same memory with no synchronisation, one of them writing | Heavy: several times slower, several times more memory | Values change unpredictably; crashes that never reproduce twice; anything @unchecked Sendable |
| Address Sanitizer (ASan) | Use-after-free, buffer overruns, double free, stack-use-after-return | Roughly twice as slow, plus large memory overhead | EXC_BAD_ACCESS with a useless stack; corruption that moves when you add a print |
| Undefined Behavior Sanitizer (UBSan) | Misaligned pointers, invalid enum and boolean values, C integer overflow | Light | You bridge to C or Objective-C, or wrap a third-party framework |
| Main Thread Checker | UIKit and AppKit called off the main thread | Negligible — on by default for debug runs | Always. Read its purple warnings rather than dismissing them |
| Zombie Objects | Messages sent to a deallocated Objective-C object | Memory grows for the whole run, since nothing is truly freed | Over-release crashes in Objective-C or older bridged code |
| Malloc Scribble / Guard Malloc | Reads of freed memory; writes past an allocation | Heavy | ASan pointed at heap corruption and you need a second opinion |
Two facts about running them:
TSan and ASan cannot be enabled together. They both take over memory management in incompatible ways, so Xcode makes them mutually exclusive. Pick the one matching your symptom: threading, or memory validity.
Main Thread Checker is already on. It is the one in that table you are using right now without knowing it. Those purple runtime issues that appear in the Issue Navigator saying a UIKit method was called on a background thread are real bugs that happen to have worked so far.
The pretest's point, made concrete:
retryDrain and exportCSV are exactly the kind of path a test suite skips — an error recovery route and an export nobody wrote a test for. Both are broken. The tool said nothing, and it was not wrong to say nothing.
This is why the technique that works is run your app under TSan, do not just run your tests. Use the features by hand for twenty minutes with the sanitizer on. You will cover paths no test reaches.
The Memory Graph Debugger
Run the app, exercise the screen, and press the Debug Memory Graph button in the debug bar. Execution pauses and Xcode shows every object alive right now, with arrows for the references between them.
Before reading one, get the definition exact, because it is the difference between two very different bugs:
A leak is memory that is unreachable from any root and still has not been freed.
"Root" means something the app inherently holds: the app delegate, a scene, a global, a singleton. ARC frees an object when its last strong reference goes — so an object nothing can reach should be gone. If it is still there, something is holding it in a circle that ARC cannot break.
That is what the Memory Graph Debugger computes and draws. The pair with the purple ! badge is the cycle; clicking either one shows the arrows forming it, and the fix is the one S3-02 taught — make one of the two references weak or unowned, and it is the child's reference to its parent that gives way.
Two settings make the graph far more useful, both under Edit Scheme ▸ Run ▸ Diagnostics:
- Malloc Stack Logging. Without it the graph tells you what is alive; with it, selecting an object shows the backtrace of where it was allocated. That turns "some
URLSessiondelegate is leaking" into a file and a line. - Zombie Objects for the opposite failure — an object freed too early rather than too late.
The View Hierarchy Debugger
The visual counterpart, next to the memory graph button: Debug View Hierarchy. It pauses the app and shows the view tree exploded in 3D, with every view's frame, constraints and clipping.
It answers questions that are almost impossible to reason about from code. Why is nothing visible — is the view absent, or present with a zero-size frame, or sitting outside its parent's bounds, or at alpha 0, or behind something opaque? All four look identical on the device and are instantly distinguishable here. It also shows the constraints on a selected view, which is the fastest route to the ambiguous-layout problem from U6-02.
Logs that survive production
print() is fine while you are looking at Xcode's console and useless everywhere else. It has no levels, no structure, no privacy control, and it does not exist in the logs you can collect from a user's device.
The system logger does:
import os
private let logger = Logger(subsystem: "com.ledger.app", category: "transfers")
logger.debug("building request for \(endpoint)")
logger.notice("transfer submitted")
logger.error("transfer failed: \(status)")
logger.fault("keychain unreadable — session cannot continue")
Levels decide what is kept. debug is for development and is not persisted. info is kept only when logs are actively being collected. notice — the default — error and fault are persisted to disk and appear in a sysdiagnose, which means they are what you actually have to work with when a user reports a problem from a device you have never touched.
Privacy is enforced, and the defaults are not obvious. Values interpolated into a log message are redacted according to type: a dynamic string is replaced with <private> in the persisted log, while numbers are shown. You override it per value:
logger.error("transfer \(id, privacy: .public) failed for \(iban, privacy: .private)")
logger.notice("user \(userID, privacy: .private(mask: .hash)) upgraded")
The masked-hash form is the one worth knowing: it writes a stable hash instead of the value, so you can tell that the same user appears in twenty log lines without ever learning who they are.
Reading them back: Console.app shows a connected device's live log with filtering by subsystem and category — which is why setting those two strings properly is worth the thirty seconds. On the command line, log stream --predicate 'subsystem == "com.ledger.app"' follows it live, and log collect pulls an archive off a device.
An app crashes with EXC_BAD_ACCESS about once every fifty launches. The stack trace points at a different, innocent-looking function every time. Which tool first, and why?
Memory climbs steadily while a user browses a photo feed and never comes back down. The Leaks instrument reports nothing. What is the most likely explanation?
An engineer on your team says "memory is going up while people scroll the feed, I think we have a leak." Write the diagnosis plan you would hand them. Say which tool you would open first and what result would confirm or eliminate "leak" as the explanation, what you would open second if it is eliminated, and what fix each of the two outcomes implies. Then write the one sentence you would add to the team's code review checklist so this class of bug is caught before it ships.
You can now:
- Pick the sanitizer that matches a symptom, and state what a clean run does and does not prove
- Say what Swift 6 left for Thread Sanitizer to cover, and why
- Define a leak by reachability, and tell it apart from unbounded growth — with the right tool and fix for each
- Use the memory graph with Malloc Stack Logging, and the view hierarchy debugger for invisible views
- Log with levels and privacy that keep a production problem diagnosable and users' data out of a sysdiagnose
Next up: the Instruments catalogue — every template worth knowing, and the four call-tree settings that turn an unreadable trace into an answer.