A benchmark for codebases, not models
What does your codebase cost an agent?
Terrain hands real static analysis issues to a real coding agent, one at a time, in isolated git worktrees. It verifies every fix mechanically and reports what fixing a thousand of them would cost — in dollars, tokens and hours.
- $230.84
- to fix 1,000 issues
- 2.5
- agent-hours
- 100%
- fix rate
90% interval $168.97–$292.92
85.3M tokens
5 of 5 attempts landed
Measured, not modelled: 5 attempts against terrain-demo-shop with claude-opus-5, every fix verified by re-running the analyzer. That repository is a small fixture — a production codebase will cost more, which is the entire reason to measure yours rather than trust this figure.
npx terrain-bench scan .open a report →The problem
Agent spend is real. Attribution is not.
When an agent takes twenty minutes and four dollars to close a lint warning, nothing in the bill tells you whether that was the model being slow or your codebase being hard. So the conversation stalls at “our code is messy,” which has never funded a refactor.
Terrain holds the agent fixed and varies the repository. The result is a cost you can put in a budget, a direction you can track across commits, and a ranked list of which of your habits is charging you the most.
How it works
Five stages, in order, each one observable.
- 01
Discover
Every applicable analyzer runs against HEAD — ESLint, tsc, ruff — using your pinned binaries, not ours.
- 02
Sample
A stratified, seeded draw across tool, severity and fixability, capped so one noisy rule cannot own the sample.
- 03
Isolate
One git worktree per issue, one fresh agent session per worktree. No task can see what another task did.
- 04
Verify
The analyzer must stop reporting the issue, report nothing new, and your tests must still pass.
- 05
Project
Extrapolate to 1,000 issues with a bootstrap interval — paying for the failed attempts too.
The index
One number, and the four it is made of.
| Efficiencycost per verified fix, on a log curve | 47 | 40% | |
| Reliabilityshare of attempts that landed | 100 | 30% | |
| Safetyshare of attempts that broke nothing | 100 | 20% | |
| Navigabilitystatic friction, measured for free | 52 | 10% |
Safety is scored apart from reliability on purpose. A repository where agents fail loudly is merely slow. One where they succeed and quietly break something is dangerous, and a single blended number would let the second hide behind the first.
Credibility
What Terrain refuses to count.
The agent saying it worked
Only the analyzer’s verdict counts. Terrain never reads the agent’s account of what it did.
Deleting the rule
Any attempt that edits lint, type or CI configuration is scored as a regression. Otherwise the cheapest route to an A is disabling the linter.
Only counting the wins
Landing 1,000 fixes at a 60% fix rate means paying for 1,667 attempts. The projection pays for all of them.
Start
The static half is free and needs no API key.
# eleven signals that predict agent cost — no agent, no spend
npx terrain-bench scan .
# watch the whole pipeline run on a fixture, free
npx terrain-bench demo
# measure it for real
npx terrain-bench bench . --sample 25
# gate a pull request on it
npx terrain-bench bench . --yes --min-afi 70 --max-cost 200scan walks the repository and scores file size, type safety, escape-hatch density, test ratio, whether an AGENTS.md exists, and six more — each with a specific remediation rather than a grade.
bench is the measured half. It spends your own tokens through your own agent CLI; Terrain takes no margin and never sees your code.