TERRAIN

A benchmark for codebases, not models

What does your codebase cost an agent?

Terrain hands real static analysis issues to a real coding agent, one at a time, in isolated git worktrees. It verifies every fix mechanically and reports what fixing a thousand of them would cost — in dollars, tokens and hours.

$230.84
to fix 1,000 issues

90% interval $168.97–$292.92

2.5
agent-hours

85.3M tokens

100%
fix rate

5 of 5 attempts landed

Measured, not modelled: 5 attempts against terrain-demo-shop with claude-opus-5, every fix verified by re-running the analyzer. That repository is a small fixture — a production codebase will cost more, which is the entire reason to measure yours rather than trust this figure.

npx terrain-bench scan .open a report →
$0.168 → $0.324 per attemptx: file · y: rule · elevation: costpeak: no-else-return · format.js

The problem

Agent spend is real. Attribution is not.

When an agent takes twenty minutes and four dollars to close a lint warning, nothing in the bill tells you whether that was the model being slow or your codebase being hard. So the conversation stalls at “our code is messy,” which has never funded a refactor.

Terrain holds the agent fixed and varies the repository. The result is a cost you can put in a budget, a direction you can track across commits, and a ranked list of which of your habits is charging you the most.

How it works

Five stages, in order, each one observable.

  1. 01

    Discover

    Every applicable analyzer runs against HEAD — ESLint, tsc, ruff — using your pinned binaries, not ours.

  2. 02

    Sample

    A stratified, seeded draw across tool, severity and fixability, capped so one noisy rule cannot own the sample.

  3. 03

    Isolate

    One git worktree per issue, one fresh agent session per worktree. No task can see what another task did.

  4. 04

    Verify

    The analyzer must stop reporting the issue, report nothing new, and your tests must still pass.

  5. 05

    Project

    Extrapolate to 1,000 issues with a bootstrap interval — paying for the failed attempts too.

The index

One number, and the four it is made of.

Efficiencycost per verified fix, on a log curve
4740%
Reliabilityshare of attempts that landed
10030%
Safetyshare of attempts that broke nothing
10020%
Navigabilitystatic friction, measured for free
5210%

Safety is scored apart from reliability on purpose. A repository where agents fail loudly is merely slow. One where they succeed and quietly break something is dangerous, and a single blended number would let the second hide behind the first.

Credibility

What Terrain refuses to count.

The agent saying it worked

Only the analyzer’s verdict counts. Terrain never reads the agent’s account of what it did.

Deleting the rule

Any attempt that edits lint, type or CI configuration is scored as a regression. Otherwise the cheapest route to an A is disabling the linter.

Only counting the wins

Landing 1,000 fixes at a 60% fix rate means paying for 1,667 attempts. The projection pays for all of them.

Start

The static half is free and needs no API key.

# eleven signals that predict agent cost — no agent, no spend
npx terrain-bench scan .

# watch the whole pipeline run on a fixture, free
npx terrain-bench demo

# measure it for real
npx terrain-bench bench . --sample 25

# gate a pull request on it
npx terrain-bench bench . --yes --min-afi 70 --max-cost 200

scan walks the repository and scores file size, type safety, escape-hatch density, test ratio, whether an AGENTS.md exists, and six more — each with a specific remediation rather than a grade.

bench is the measured half. It spends your own tokens through your own agent CLI; Terrain takes no margin and never sees your code.

explore a real report →