# What our agents cost in CI time

Field notes · October 1, 2026 · Henry Kobutra with Codex

Looking back at September 21, 2026.

![Dozens of ivory and blue rails converge from the left into a single narrow orange shared route.](/log/images/agents-made-our-ci-bill-a-product-decision/cover.webp)

From September 16 to 21, Hydrant was starting 55 to 80 check runs a day. A full run used about 22 runner-minutes across quality, Node, Worker and two browser jobs.

That was the cost of an agent workflow that could capture, refine, build and ship many small issues. It was also a product decision hiding inside a workflow file: which evidence did we need for each kind of change?

## Every push was priced like a release

Draft pull requests ran the same broad suite as final candidates. A small documentation edit could wait on two browser jobs. Backend-only work waited too, even though the browser fixtures did not import the Worker.

Failures multiplied the effect. In a snapshot of 300 runs, 102 failed. The component-browser job was red in 72 and was the only red job in 34. Some branches ran the full suite 9 to 16 times as agents pushed review fixes.

Those numbers describe runner time and delay. They are not a dollar claim. We did not have a verified Blacksmith spend figure for that snapshot, and the organisation's GitHub-billed Actions usage netted to zero.

## We made uncertainty expensive

The new rules give drafts the quality job. Marking the final head ready starts the applicable suites. Documentation-only changes keep quality; backend-only changes add Node and Worker; frontend, shared code, browser tests, scripts, workflows, packages, configuration and mixed or unknown changes run all five.

The allowlists matter. If the classifier cannot prove that a change is narrow, it chooses full coverage.

After merge, main may reuse the ready pull request's heavy checks only when the two Git trees are identical and the earlier run is a successful, eligible, non-draft run at that exact head. Missing or ambiguous evidence sends main through every suite. The successful main run still exists as release evidence and cites the run whose coverage it reused.

The first clean sequence measured 1.48 runner-minutes for the draft, 20.45 for the ready head and 1.48 for main: 23.41 in total. That is one observed sequence, not a promised rate.

We kept focused checks before the first push and full hosted evidence for the final candidate. The saving came from deciding when the candidate was real, what a file could affect and how identical code could carry evidence forward. More agents made those decisions impossible to leave implicit.

Contact: bots@hydrant.dev
