Skip to content

Advertisement

DevOps Society
DevOpsAnalysis

Anthropic’s CI Jobs Grew 25 Times in Six Months. Faster Runners Are the Wrong Answer.

Anthropic published the numbers in September. Their build volume grew 25 times in six months because a model now writes most of the code their engineers merge. The instinct when queues lengthen is to buy more capacity, which is the expensive answer to the wrong question.

Published your local timeupdated

Anthropic’s CI Jobs Grew 25 Times in Six Months. Faster Runners Are the Wrong Answer.

In September, Anthropic's engineering team published a number worth putting in front of anyone funding a coding assistant rollout. Their continuous integration job volume grew 25 times in six months. Not 25 percent. Twenty five times.

The cause is not mysterious. Claude now writes about 80 percent of the code their engineers merge, and those engineers ship roughly eight times as much code per quarter as they did between 2021 and 2025. More changes means more pull requests, more pull requests means more test runs, and the thing that absorbs all of it is the build system.

Almost every engineering organisation currently rolling out assistants is measuring the upside in developer hours saved. Very few have modelled what happens to the system that sits downstream of all those extra changes. That system has a bill attached to it, and the bill is usage based.

Why the obvious fix is the wrong one

When builds start queuing, the instinct is to buy more capacity. Bigger runners, more of them, a higher tier. It works, it is quick, and it is the most expensive possible answer.

Anthropic went the other way. Rather than running everything faster, they built a service that works out which tests a given change could actually affect and runs only those. A listener records results from every run, and a selector decides what is worth running next time based on what the change touched and how those tests have behaved historically. Their first version ran as a single process and could not keep up, so they rebuilt it to spread the work across several machines.

The distinction matters to anyone holding a budget. Buying capacity scales your spend with your change volume, forever. Running fewer tests breaks that link. One is an operating cost that grows with adoption. The other is a one off piece of engineering.

What this does to a normal company's bill

Work the arithmetic on your own numbers rather than taking anyone's word for it. On GitHub's hosted runners, a standard Linux machine costs $0.006 a minute, and an Enterprise Cloud plan includes 50,000 minutes a month, which is a little over 800 hours.

Take a team using 40,000 minutes a month today, comfortably inside the included allowance and therefore invisible on anyone's report. Multiply the change volume by ten, which is well short of what Anthropic saw, and that becomes 400,000 minutes. After the included allowance, that is around $2,100 a month, appearing from nothing. Multiply by twenty five and it is roughly $5,700 a month.

Those figures are illustrative rather than a forecast, and your own mix of runner sizes will move them a lot. The point is the shape. A line item that was zero because you never came close to the allowance becomes a real number, and it arrives gradually enough that nobody treats it as an event.

The second cost is less visible and usually larger. When the queue lengthens, every engineer waits longer for every change, including the ones the assistant had nothing to do with. You can end up spending more on the build system and losing the time you bought.

The honest caveat

Anthropic is an extreme case and it would be dishonest to present their numbers as a prediction. A company where a model writes four fifths of the merged code is not where most organisations are, and may not be where most are headed. Their 25 times is the far end of the curve, not the middle of it.

What travels is the direction, not the multiple. If your engineers are producing meaningfully more changes than they were a year ago, your build system is absorbing that increase whether or not anyone has looked. The useful exercise is not to guess at your multiple. It is to go and measure the one you already have.

What to do about it

  • Pull your build minutes for the last twelve months and plot them. If the line has bent upwards in the last two quarters, you have your answer and you did not need anyone's benchmark to get it.

  • Find out what your median wait time is before a change starts building, not the average. The average hides the bad mornings, and the bad mornings are what people remember.

  • Ask whether every test needs to run on every change. For most codebases the honest answer is no, and nobody has had a reason to revisit it until now.

  • Before approving more capacity, ask what proportion of the tests in a typical run could possibly have been affected by the change. If nobody knows, that is the project, not the runners.

  • If you are planning an assistant rollout, put a build capacity line in the business case. It is the one predictable cost of the thing and it is routinely left out.

The part worth sitting with

There is a quieter point in Anthropic's write up. The rebuild of their test selection service took three weeks, where they estimate it would previously have taken closer to a quarter, because the assistants helped build it.

So the tooling creates the problem and then helps solve it, and both sides compound. That is a reasonable description of most infrastructure work over the next two years. The teams that come out ahead will not be the ones that adopted earliest. They will be the ones who noticed what the adoption did downstream and fixed it before it became a line on a budget review.

Our September hiring report found AI mentioned in nine out of ten infrastructure job postings, which is a fair indication of how widely this is being adopted, and how little of it has been thought through.

Sources: Anthropic's engineering post on scaling test impact analysis, published 14 September 2026, and GitHub's published Actions billing rates. Checked on 5 October 2026.

Advertisement

Follow DevOps Society on LinkedIn

Practical infrastructure engineering in your feed.

Follow

Written by

DevOpsSociety Editorial Team

Editorial Team

The DevOpsSociety Editorial Team covers DevOps, cloud infrastructure, Kubernetes, AI infrastructure, platform engineering, cybersecurity, FinOps, and modern engineering practices. We publish practical insights, technical guides, architecture analysis, and research for engineers and technology leaders.

More from DevOpsSociety →
The Infrastructure Briefing

Get the infrastructure briefing.

Practical DevOps, cloud, AI infrastructure and engineering insights, delivered weekly. Read by engineers and engineering leaders.

No spam. Unsubscribe anytime.

Anthropic’s CI Jobs Grew 25 Times in Six Months. Faster Runners Are the Wrong Answer., DevOps Society