A queued build should say why it has not started
Queued is one word for four different situations: waiting in line, a free machine that has not claimed the job, no machine online, and no machine ever set up. Each needs a different response. How to tell them apart and what to do about each.
It is late on a Friday and a fix has to reach TestFlight before the weekend. The build was triggered twenty minutes ago and the dashboard still says Queued. Is it fifth in line behind a batch of nightly builds? Is the Mac that runs iOS builds asleep after an update restart? Has the job been sitting in front of an idle machine that simply has not noticed it? From the outside these look identical, and the only honest advice anyone can give is to wait and see, which is the one thing the person with the deadline cannot afford.
We changed Ubriot's builds page this week so that every queued build carries a one line reason and a suggested next step. Working out what those reasons should be turned into a useful way to think about any build queue, including one you run yourself.
Four situations hiding behind one word
A build that has not started is in one of four states, and they can be told apart by asking three questions in order. Has any machine that can run this kind of build ever connected? Is one online now? Is one free?
Waiting in line. Machines are online and all of them are busy, and there are other builds ahead of yours. This is the healthy case. The right response is to do nothing except perhaps check how long it is likely to be.
Free but not picked up. A machine is online and idle, nothing is ahead of your build, and yet a few minutes have passed and it has not started. This is the case that most deserves attention, because it means something is wrong, and waiting longer will not fix it. Usually the machine is stuck finishing or cleaning up its previous job. Sometimes the build is asking for something the free machine cannot give it, such as source code from a path that machine was never told it may read.
No machine online. Machines that can run this build exist, but none has been heard from recently. The build will start when one comes back, and the useful facts are when one was last seen and whether anyone is going to bring it back.
No machine ever. Nothing capable of this build has ever connected. That is a setup problem, not a delay, and the build will wait forever until it is solved.
Count per platform, not per queue
A queue position is only meaningful among builds that compete for the same machines. If you have four Android builds and one iOS build waiting, the iOS build is not fifth in line. It is first, in a line of its own, because Android builds cannot use the macOS machine and iOS builds cannot use anything else. A single global position number makes iOS builds look stuck behind work that will never touch the machine they need. Ubriot counts position separately for iOS, Android and web for this reason.
Estimates should be rough and say so
Once a build is genuinely waiting in line, the next question is how long. A reasonable estimate needs only three numbers: how many builds are ahead, how many capable machines are online, and how long a recent build on this platform usually takes.
For the last of those we use the median of recent finished builds rather than the average. Build times have a long tail. One cold build that had to download every dependency from scratch can take three times as long as a normal one, and a single run like that drags an average upward for days. The median ignores it.
Even so, the result is a rough figure and the page says so in those words. It cannot know that the build ahead of yours is a large app or that the next one will hit a slow network. If there are no recent builds on that platform, we show no estimate at all, because a guess dressed up as a number is worse than an honest blank.
Match the advice to the state
The point of naming the state is to change what the person does next. Waiting in line needs nothing more than a way to watch the build. Free but not picked up needs a short wait and then a cancel and rebuild, plus a pointer to the likely cause. No machine online needs the last seen time, so you can tell an overnight sleep from a machine that has been gone for a week, and once a build has waited half an hour we say plainly that cancelling it may be the better route. No machine ever needs setup instructions, not patience.
Getting this wrong in either direction is costly. Telling someone to wait when a machine is stuck wastes their afternoon. Telling them to cancel when they are third in line throws away their place.
If you run your own runners
None of this depends on our product. If your team runs self-hosted CI runners, the same three questions work as a first check whenever a job sits in a queue: is any runner registered with the labels this job asks for, has one checked in recently, and is one idle? If a runner is idle and the job still is not moving, look at the runner before you look at the queue. And if your dashboard only ever says Queued, it is worth adding the last time each runner checked in somewhere people can see it. That one timestamp settles most of the questions that otherwise end up in a team chat.
On Ubriot the reason appears under each queued build on the builds page, with the position, the number of machines online and idle, when one was last seen, and a rough start time when there is enough history to give one.