Compare a failing build with the last one that passed
Most build failures are easier to explain than to diagnose from scratch. Start from the last build that passed for the same platform and profile, list what changed in five places, and treat what stayed the same as evidence too.
A release build fails during the iOS archive step with four hundred lines of log. The usual next move is to read the log from the bottom up and search for the first red line. Sometimes that works. Often the first error is a symptom three steps removed from the cause, and twenty minutes later you are reading forum threads about a linker warning that has been in every build for a year.
There is a cheaper first question: what is different from the last build that passed? A build that worked on Tuesday and fails on Thursday did not break for no reason. Something moved in between, and the list of things that can move is shorter than the log. This week Ubriot started answering that question on every failing build. Deciding what belongs in the answer turned into a checklist that works on any CI system.
Pick the right last good build
The comparison is only as good as the baseline. A green Android build says nothing about an iOS failure, because the two share almost no build steps. A debug build that passed says little about a release build, which signs differently, strips differently and often runs a different set of scripts. So the baseline should be the most recent successful build of the same app, on the same platform, with the same build profile. Anything looser gives you a list of differences that mostly do not matter.
If there is no such build, say so plainly. A first release build of a new profile has nothing to compare against, and pretending otherwise sends people hunting through changes that were never relevant. In that case the log really is the place to start.
Five places to look
Code. Did the branch change? Did the commit change, and if so, what is in the range between the two? This is the most common answer and the easiest to check, but only if both builds recorded the exact commit they built. A build that recorded only a branch name cannot tell you which commits it contained, and a tool should admit that rather than guess.
The machine. Was the failing build run on a different build machine from the last good one? Did the machine's operating system get updated in between? A macOS point release or a new Xcode command line tools install can break a build that has not changed at all, and it is the cause people check last because nobody remembers doing it.
Setup. Were signing credentials added, replaced, copied from another app or removed? Was a provisioning profile regenerated? Did someone change an app setting such as the bundle identifier or the folder the project lives in? These changes happen outside the repository, so they never show up in a commit range, yet they are behind a large share of signing and upload failures.
Timing. Did a phase that normally takes two minutes take nine? A dependency install that suddenly runs four times longer often points to a cache that was lost or a registry that is struggling, even when the failure itself happens later.
Memory. Did peak memory use jump? A build that used three gigabytes last week and five this week may be failing on a machine that simply ran out, and the error message rarely says that in so many words.
Set thresholds so noise stays quiet
Build timings wobble. A step that takes 90 seconds one day and 110 the next has not changed in any way that matters, and a comparison that lists every wobble buries the real change under false ones. We only flag a phase as slower when it took at least twice as long as in the last good build and the difference is more than a minute. We only flag memory when peak use rose by half and by more than a gigabyte. Those numbers are judgement calls, not laws, but having some threshold is what keeps the list short enough to read.
We also leave out the phase that failed. A failed step stops early, so its time is meaningless. What is useful about that step is the other side of the comparison: the last good build got through it, and how long it took.
What stayed the same is evidence too
A list of changes reads best next to a list of things that were checked and found unchanged. Same branch. Same build machine. No credential or setting changes. When those lines are there, the reader can stop worrying about whole categories and look harder at what is left.
The most interesting case is when nothing tracked changed at all. If the failing build used the very same commit on the same machine with the same setup, the cause is almost certainly something resolved at build time: a dependency declared with a version range that picked up a new release, a remote resource that changed or was briefly down, or a step that is simply flaky. That is a narrow and useful conclusion. It tells you to look at the lockfile and the network, not the diff.
Doing this without a tool
You can run the same check by hand on any CI system. Find the last successful run for the same platform and profile. Write down both commits and run a git log between them. Compare the runner names and their OS versions in the two logs. Ask whoever manages signing whether anything was replaced in between. Put the step durations side by side. It takes ten minutes, and it is usually ten minutes better spent than the first ten minutes of reading the log.
In Ubriot the comparison appears on the failing build's row on the builds page, with the commit range and a link to the GitHub compare view when both commits were recorded, the setup changes in date order, and the checks that came back unchanged.