Quality and craft · advanced

The Instruments catalogue: which template, and how to read it

Every Instruments template worth knowing, matched to the question it answers — plus the four call-tree settings that turn an unreadable trace into a name, and the rules that stop you profiling a lie.

12 min read12 min practice0/2 exercises6 recall cards

By the end you will be able to

  • Choose the Instruments template that matches a symptom, across the full catalogue
  • Read a call tree — self against total time — and configure the four settings that make it legible
  • Profile honestly, and turn a finding into a regression gate rather than a one-off fix
Guess firstAnswering before you read makes the explanation stick — even when you get it wrong.

You profile the app in the Simulator with a Debug build, find that decodeTransactions dominates the Time Profiler, and spend two days optimising it. On device, nothing measurably improves. What went wrong?

The catalogue

P1-03 introduced six templates through the symptoms that send you to them. Here is the full set worth knowing, with the ones you already met marked so you can skip them:

TemplateThe question it answers
Time ProfilerWhere is CPU time going?P1-03
HangsWhat blocked the main thread, and for how long?P1-03
AllocationsWhat is being allocated, and does it come back?P1-03
LeaksIs anything unreachable and unfreed?P1-05
App LaunchWhat happens before the first frame?P1-03
Animation HitchesWhy does scrolling stutter, and by how much?new
System TraceWhat are my threads actually doing — running, blocked, waiting on a lock?new
NetworkWhich requests, how big, how long, how many connections?new
File ActivityWhat is touching the disk, and on which thread?new
Energy LogWhat is draining the battery?new
SwiftUIWhich views re-evaluate their body, and how often?P1-04
Core DataWhich fetches run, how long, and are faults firing per row?new
Metal / GPUIs the GPU the bottleneck rather than the CPU?new
os_signpost / Points of InterestHow long did my own operation take?P1-03

Four of the new ones deserve their reason spelled out.

Animation Hitches is the one most directly tied to what users complain about. A hitch is a frame that arrives later than it should have; the template reports hitch time ratio — milliseconds of hitch per second of scrolling — which turns "it feels janky" into a number you can hold to a target and track over releases. Time Profiler tells you what the CPU was doing; Animation Hitches tells you whether it mattered, and separates a hitch caused by your commit phase from one caused by the render server.

System Trace is the one to reach for when CPU time looks fine and the app is still slow. It shows thread states — running, preempted, blocked — so it distinguishes work from waiting. A thread blocked on a lock or a disk read burns almost no CPU and shows up as nearly nothing in Time Profiler, which is exactly why a slow app can profile as idle.

File Activity catches synchronous disk I/O on the main thread — a Data(contentsOf:), a Core Data save, a UserDefaults synchronise. Disk is orders of magnitude slower than memory, and this is a common cause of hangs that Time Profiler under-reports for the same reason System Trace exists.

Energy Log matters because battery complaints have causes no other template names: location updates left running, a timer firing all night, network requests waking the radio repeatedly. D2-09's point about connection reuse is an energy question as much as a latency one — the cellular radio stays in a high-power state for seconds after each transmission, so many small scattered requests cost far more than the same bytes sent together.

Turning that catalogue into a decision is the skill:

Loading runnable Swift…

The default case is deliberate. When you genuinely do not know, Time Profiler is the right broad start — it is the cheapest way to find out whether you are even CPU-bound, which tells you which of the others to open next.

Reading a call tree

Open a Time Profiler trace and the first thing you meet is a tree of function names with times beside them, most of which belong to the system. Two numbers and four settings turn it into an answer.

Self time is time spent executing inside a function's own code. Total time is self time plus everything it called. The distinction is the whole skill: a function with huge total time and near-zero self time is not slow — it is a caller of something slow, and optimising it achieves nothing.

Loading runnable Swift…

Read the top block and main looks like the problem at 289ms — it is simply where everything happens. Read the bottom block and the two real culprits are named immediately: decodeJSON at 180ms and formatDate at 90ms. Same data, one setting apart.

The four settings that make a trace legible

In the Time Profiler's detail pane, under Call Tree options:

  • Invert Call Tree. Puts the deepest frames at the top, ranked by self time — the bottom block above. This is usually the first thing to turn on, because it answers "what is actually burning CPU" without any tree walking.
  • Hide System Libraries. Removes every frame that is not your code. What remains is the line you can actually change. Turning this on is the moment most traces become readable.
  • Separate by Thread. Splits the tree per thread, so main-thread work — the only work that can cause jank — is not averaged in with background work.
  • Flatten Recursion. Collapses recursive frames into one entry, so a recursive function does not spread its cost across forty nested rows.

The habit worth building: Invert + Hide System Libraries + Separate by Thread, then look only at the main thread. That combination answers most performance questions in about a minute.

One more, on the timeline rather than the detail pane: drag to select a time range and every number recalculates for just that window. Profiling a whole two-minute session tells you almost nothing; selecting the 800 milliseconds where the stutter happened tells you everything. Reproduce the symptom, then select exactly it.

Hitches, measured

Since Animation Hitches is the template most tied to user complaints, it is worth being able to interpret its headline number:

Loading runnable Swift…

Expressing it as a ratio rather than a count is what makes it comparable: a long scroll and a short one become the same measurement, so it works as a target and as a regression check. The thresholds above are a reasonable working guide rather than a specification — the number that matters is your own app's, tracked over releases.

Profiling honestly

Five rules, each of which exists because breaking it has cost somebody a week:

  1. Release build, real device. The pretest's lesson. Optimisation off and desktop hardware are two different lies.
  2. Realistic data. Ten rows profile nothing. Profile the account with four years of transactions, because complexity that is invisible at ten rows is the whole problem at ten thousand — the algorithmic point P1-03 makes with its complexity predict.
  3. Reproduce first, measure second. Know the exact gesture that shows the symptom before recording, then select that time range in the timeline.
  4. Change one thing. Two optimisations at once and you cannot attribute the improvement — or notice that one of them made things worse.
  5. Re-measure after the fix. The loop is not "profile, fix"; it is "profile, fix, profile again". Roughly a third of plausible optimisations do nothing, and you only find out by measuring.

Making the finding stick

A fix you cannot detect regressing will regress. Two mechanisms turn a profiling session into something permanent, and both belong in CI (P2-02):

  • XCTOSSignpostMetric measures the intervals you already marked with OSSignposter in P1-03, so "feed refresh" becomes a number the test suite asserts on.
  • XCTMemoryMetric and XCTClockMetric put a performance test around a specific operation with a baseline recorded from a known-good run.

There is an honest caveat worth stating in an interview: performance tests are noisy on shared CI hardware, so treat them as trend detection with generous thresholds rather than a hard pass/fail gate. The alternative — MetricKit, from P2-02 — gathers real percentiles from real devices in the field, which is slower feedback and far more truthful.

Check yourself

The Time Profiler shows UITableView.layoutSubviews with 2,100ms total time and 3ms self time. Where is the problem?

Check yourself

An app feels sluggish, but Time Profiler shows the main thread using barely any CPU during the slow period. Which template next, and what are you looking for?

Loading exercise…
Loading exercise…
Design itSaved on this device. Never graded.

Your team's app has one complaint from users — "it gets slow after I've been in it a while" — and no reproduction. Write the investigation plan. Say which template you open first and what result would send you to which template second, how you would build a reproduction that makes the symptom appear on demand, and what data volume you would use. Then describe the single measurement you would add to CI afterwards so that whatever you find can never quietly come back.

Checkpoint

You can now:

  • Match a symptom to a template across the full catalogue, and know why Time Profiler is the right broad start
  • Tell self time from total time, and configure the four call-tree settings that make a trace legible
  • Recognise the slow-but-idle case and reach for System Trace instead of guessing
  • Profile honestly — Release, on device, with realistic data, one change at a time, re-measured
  • Turn a finding into a regression gate, and say honestly how much to trust it

Next up: the P track's quality unit is complete. P1-01's tests prove behaviour, P1-05's tools find what is wrong, and these find what is slow — three different questions, three different instruments.

How well do you know this now? Rating yourself honestly, then being tested on it, is how you find out where your intuition is wrong.