Did the agent follow the decision?

This is dogfood. Our own benchmark, on our own repositories. 30 decisions our team made between June and August 2026. A coding agent was asked about each one 5 times in each of four setups, with and without Align. It is not a customer's data and it is not an independent audit.

Run on 2026-09-21. Every run used claude-sonnet-5. 600 runs, 53.4 dollars of agent spend. Every answer was then graded three times, by 3 models from two labs - the scores below say which judge produced them.

The result

+53.3 points on the decisions that were never written down

Not against an agent with nothing. Against one already reading the codebase with grep, file reads and globs. Same 30 tasks, same model, same tools. The only difference was the decision graph.

That figure is one of the four splits, and it is deliberate. Pooled across all 30 tasks the graph was ahead by 10.7 points, and that pooled number does not clear our own bar: its interval is -1.3 to 23.3, which includes zero. A previous run of this page led with the pooled figure because it did clear it then. This run it does not, so the page leads with the comparison that does.

Search alone followed the governing decision in 71% of runs (106 of 150). With the graph as well, 81% (122 of 150). For scale, the same agent with no tools at all managed 36% (54 of 150) - that is the floor, not the comparison.

Those three percentages are claude-haiku-4-5's scoring. Every percentage on this page belongs to a judge, and two other judges scored the same answers differently - see three judges, two labs. Task by task the graph was ahead on 10, behind on 5, level on 15. Fifteen ties out of thirty is most of the reason the pooled interval is as wide as it is.

Then we had a model from a different lab re-mark all 600 answers. Nothing was re-run - same transcripts, different judge. Here the result is awkward for us and we are going to say it plainly: both OpenAI judges put the pooled gap outside zero (12 and 9.3 points), and the Anthropic judge whose numbers this page actually prints does not (10.7 points, -1.3 to 23.3). "A judge from another lab agrees with us" would be a true sentence and a misleading headline, so it is not the headline. The split above is.

Where it matters most: on the tasks whose governing decision was never written in the repository - settled in a meeting, a ticket or a thread - search alone got 33.3% and the graph got 86.7%. No amount of grepping finds a decision that is not in the files.

Everything below is how we measured it, what we got wrong, and every number behind those two. The method starts here. We publish the run that disagrees with us as well as the ones that do not.

We ran it three times, and got three answers

This page used to publish a different number. It said Align followed the governing decision in 91.3% of runs, from a run we call v4. We then ran the same frozen tasks against the same build twice more. Here is all three, because the spread between them is more useful than any one of them.

RunDateBareRules fileRepo searchAlign
v42026-09-2038.7%51.3%72%91.3%
v52026-09-2129.3%47.3%70%80.7%
v62026-09-2136%43.3%70.7%81.3%

The two later runs agree with each other to within 0.7 of a point and disagree with the first by about ten. So v4 is the odd one out. Nothing got worse.

Here is the part that should stop you trusting any of it too far. Look at the bare column. That agent has no tools, no repository and no graph - it sees the question and nothing else. It moved 9.3 points down between the first two runs and 6.7 points back up, and both of those moves pass the same significance test we would use to report a result. Nothing we shipped can reach that arm. Whatever is moving it is moving every other arm too.

This is a real limit on what the benchmark can tell you. Thirty tasks cannot separate a ten-point move from run-to-run variation, so we cannot yet say that something we changed moved this number. Fixing it needs a frozen snapshot of the graph and a design that toggles one thing at a time. That work is queued. It is not done.

Three judges, two labs

A model grades every answer, so the obvious objection is that an Anthropic model marking an Anthropic model's homework will be generous. We had the same 600 transcripts graded three times to find out. Nothing was re-run; only the judge changed.

The same 600 transcripts, graded three times.
JudgeLabSearch aloneSearch + graphGapUnclear
claude-haiku-4-5Anthropic70.7%81.3%+10.7 (-1.3 to 23.3), spans zero9%
gpt-5OpenAI66%78%+12 (1.3 to 23.3)17%
gpt-4o-miniOpenAI75.3%84.7%+9.3 (0.7 to 20)18%

The score depends on who marks it. 81.3% under one judge and 78% under another, on identical answers. gpt-5 returns "unclear" on 17% of runs against claude-haiku-4-5's 9%, and unclear counts for neither side, so every arm falls together when it does. Treat any single percentage on this page as belonging to its judge.

The gap moves less than the score, and this run it still moves enough to matter. It is 10.7 points under Anthropic's judge, 12 under OpenAI's and 9.3 under their second, so the three are within three points of each other where the scores span six. But 2 of the 3 intervals clear zero and 1 does not, so the three judges do not agree on whether there is a gap at all.

The one that does not clear zero is the judge whose scores this page prints. claude-haiku-4-5 puts it at 10.7 points, -1.3 to 23.3. The two OpenAI judges both clear it. So "a judge from another lab agrees with us" is true and is the wrong thing to lead with, and it is why the headline on this page is a split rather than the pooled gap. An earlier run had this the other way round, with the two frontier judges clear of zero and the small one spanning it. We are not switching to whichever judge currently suits.

The third row is a small model judging a frontier model's output, and it stays in the table for the reason it always did: dropping a judge because it disagrees is choosing the judges that suit us. Note it is no longer the outlier. It returns "unclear" most often of the three, at 18%, and its gap is one of the 2 that clear zero.

Why the gap is a paired comparison

Both arms ran the same 30 tasks, so the two results are not independent samples - they are the same tasks, twice. Some tasks are hard for everything, which widens each arm's own interval without saying anything about which arm is better. Comparing them task by task removes that shared difficulty, because it is the same task on both sides of the subtraction.

So the interval on the gap - -1.3 to 23.3 - is a bootstrap of the per-task difference over 30 tasks. This run it includes zero, which is why the pooled gap is not the headline. The absent split, tested the same way, does not.

We got this wrong first, and the correction still stands even though the answer has changed. An earlier version of this page compared each arm's interval, saw them overlap, and concluded the difference was not there. That reasoning is wrong whatever the data says: overlapping per-arm intervals do not tell you about a paired difference, and on the run this page carried at the time the paired test did clear zero, so the mistake cost us a real result.

On this run the two readings happen to agree, and that is a coincidence rather than a vindication. The per-arm intervals overlap and the paired difference includes zero. If we had kept the wrong method we would be reporting the right answer here by luck, which is the least useful way to be right. The test is still the paired one.

What "followed the decision" means

Each task is a question an engineer might ask an agent, where one decision our team made governs the right answer. The judge reads the agent's written answer next to the text of that decision and returns one of three verdicts: complies, violates, or unclear. Unclear is never counted for either side. A run that hit the turn cap of 20 is kept in the data and graded unclear rather than re-run.

Nothing was executed. No arm could write files or run commands, so this measures whether the written answer matched the decision. Nothing checked that code worked. The times in the raw data were recorded at a concurrency of 6 and are not what a user would wait.

The four arms

Same model, same 30 questions, same judge, same turn cap. Only what the agent had changed.

ArmWhat the agent had
Barethe question only. No tools, no repository, no graph. Every run was a single turn.
Rules filethe question plus the repository's agent rules file, still no tools. Every run was a single turn.
Repo searchread-only tools over the repository (grep, read, glob) and nothing else.
Alignthe same repository tools plus Align's decision graph over MCP.

The four kinds of task

The 30 tasks are split by where the governing decision lives, because that is what a repository search can and cannot see.

SplitTasksWhat it tests
Present11the governing decision is stated somewhere in the repository
Absent6the decision was made elsewhere and the repository never states it
Reversed6the repository's own text still says the earlier, superseded thing
Scoped7a decision that looks relevant applies to a different instance or scope

Every cell

Runs per cell, the three verdicts, the compliance rate and its exact 95% interval. Errors are runs that hit the turn cap. They are inside the unclear column.
SplitArmRunsCompliesViolatesUnclearRate95% intervalErrors
PresentBare553219458.2%44.1 to 71.30
PresentRules file552822550.9%37.1 to 64.60
PresentRepo search55487087.3%75.5 to 94.70
PresentAlign55496089.1%77.8 to 95.90
AbsentBare3022086.7%0.8 to 22.10
AbsentRules file30591616.7%5.6 to 34.70
AbsentRepo search301017333.3%17.3 to 52.82
AbsentAlign30264086.7%69.3 to 96.20
ReversedBare30114153.3%0.1 to 17.20
ReversedRules file30619520%7.7 to 38.60
ReversedRepo search301812060%40.6 to 77.30
ReversedAlign301812060%40.6 to 77.30
ScopedBare351916054.3%36.6 to 71.20
ScopedRules file35269074.3%56.7 to 87.50
ScopedRepo search35305085.7%69.7 to 95.20
ScopedAlign35296082.9%66.4 to 93.40
All splitsBare15054692736%28.3 to 44.20
All splitsRules file15065592643.3%35.3 to 51.70
All splitsRepo search15010641370.7%62.7 to 77.82
All splitsAlign15012228081.3%74.2 to 87.20

Where a baseline won

Repo search beat Align on one of the four kinds of task this run. It was the scoped split, where a decision that looks relevant turns out to apply somewhere else: 85.7% to 82.9%. On reversed the two arms finished on an exact tie, 60% each. So all four splits, in order: clearly ahead on absent (86.7% against 33.3%), narrowly ahead on present (89.1% against 87.3%), level on reversed, and behind on scoped.

The previous run had no baseline wins at all, and this page said so at the time. Saying it this way round costs us something. That is the only reason saying it the other way round was worth anything.

The 2 runs that hit the turn cap

All 2 are in the two arms that had tools (2 in repo search, 0 in Align). They stay in the data as unclear, which lowers those arms' rates without recording a wrong answer.

With them excluded, which is not the analysis we registered, repo search is at 71.6% (106 of 148) and Align at 81.3% (122 of 150). Note which way that moves: excluding them helps repo search and leaves Align where it was, because all 2 are in the repo search arm. It closes about a point of the pooled difference and leaves the rest standing.

What it cost, and where that is a claim

Every run records what it spent. Nobody had looked until now, and it goes the way you would hope: giving the agent the decisions is cheaper than making it search, because it stops searching sooner.

Arm$ per run$ per correct answerMean wall-clock
Bare$0.0163$0.045416.0s
Rules file$0.0334$0.077121.3s
Repo search$0.1687$0.238744.8s
Align$0.1376$0.169232.7s

One cost comparison here is a claim and the other is not. On the absent split, where the decision was never written in the repository, Align costs significantly less per run than repo search in all three runs: v4 -0.25, v5 -0.22, v6 -0.18 dollars, and not one of those intervals crosses zero. Repo search spends roughly nine times more per correct answer there, hunting for something that is not in the files.

Pooled across all four splits, it is not a claim. The paired difference is -0.0311 dollars with an interval of -0.0708 to 0.0075, which includes zero, and it only excluded zero in the first of the three runs. So we will not tell you Align is cheaper without saying which kind of task. The per-run column above is real and it is one run.

Wall-clock sits between the two. Align finished faster in every run, and the paired test clears zero in two of the three, so the seconds are in the table and there is no speed result attached to them.

There is no token figure on this page, and that is a gap rather than a choice. The runs record what they spent and not the tokens behind it, so we can show you dollars and cannot show you tokens. Dollars move when a price list moves and tokens do not, so the token version is the more useful one and it is not measured yet. Also note this is agent spend only: what the judge cost to grade all 600 answers is not recorded anywhere.

What could be wrong with this, and what we got wrong first

  • Our first run was invalid, and we published the correction. In v1 the harness loaded our own MCP configuration into every arm, so the arm without Align had Align. The bare arm took more than one turn in 146 of 150 runs, which a zero-tool arm can only do by calling a tool. v1.1 is the same frozen protocol with that leak closed and per-run tool counts recorded. So it can be checked: bare and the rules file made zero tool calls in 300 of 300 runs.
  • The judge is a model. A haiku-class model at temperature 0 decides complies or violates. We read the judge's reasoning on every one of Align's non-complying runs, checked the checkable claims against the live repository, and found two where the answer looks right and the verdict looks wrong. Both are in the human review below, reported separately.
  • The tasks are frozen, the repository is live. The governing decisions were frozen on 2026-09-13+54e821fd97dff321. An agent with the live repository can answer with something newer than the frozen text and be marked as contradicting it. It happened on one task in this run, the design-token parity item.
  • The tasks are ours. Decisions from our own repositories, made by our own team, some of them public. The model's published training cutoff is January 2026. The decisions are from June to August 2026. A public corpus is still not a customer's.
  • One model. Everything here is one Claude model. We ran two other models on the same 600-run grid. One leg's intervals overlapped at this size. The other was dominated by infrastructure errors. Neither is a second result and neither is quoted.

Human review of the judge

A person read 34 verdicts, sampled across arms and verdicts, with the answer, the decision and the judge's reasoning side by side. Not blind: the reviewer saw the verdict before choosing. They agreed with the judge on 32 of the 34 they could call. 34 verdicts were read by hand and 32 agreed with the judge. Two did not. The review was not blind: the reviewer saw each verdict and its reasoning first, so the two judgements are not independent and this is a sanity check, not validation of the judge. The direction of the two disagreements was not recorded, so this cannot say whether the judge is too generous or too harsh, and the figures on this page are not adjusted for them. The aggregate here is still the judge's scoring. Reviewed on 2026-09-22.

Task by task

Compliant runs out of 5 per arm. A number in brackets is how many of the five hit the turn cap.
TaskSplitBareRules fileRepo searchAlign
present-brain-dockerfile-layerPresent0/50/55/55/5
present-gh-token-empty-form-gatePresent5/55/54/55/5
present-history-depth-fetch-preconditionPresent4/55/55/55/5
present-hnsw-ef-search-floorPresent3/52/55/54/5
present-image-build-provenance-referrer-capPresent5/52/55/55/5
present-import-connector-check-parityPresent0/50/55/54/5
present-jira-confluence-webhook-fail-closedPresent5/55/55/53/5
present-preview-deploy-align-bot-prodPresent0/50/53/55/5
present-route-auth-allowlist-tenant-scopePresent5/53/55/55/5
present-schema-guard-required-versionsPresent0/51/51/53/5
present-security-definer-tenant-lookupPresent5/55/55/55/5
absent-align-cli-decided-at-schema-v4Absent1/50/50/5(2)3/5
absent-align-cli-preview-check-disabledAbsent0/50/50/53/5
absent-consultancy-channel-motion-spikeAbsent0/54/55/55/5
absent-contractor-insource-icp-parkedAbsent1/51/55/55/5
absent-infra-ec2-spot-vcpu-quotaAbsent0/50/50/55/5
absent-preview-check-disabledAbsent0/50/50/55/5
reversed-both-checks-claimReversed0/51/54/51/5
reversed-cosmic-ray-not-evaluated-claimReversed0/50/50/51/5
reversed-design-token-parity-not-gated-claimReversed0/51/50/51/5
reversed-teams-webhook-handler-dead-route-claimReversed0/53/54/55/5
reversed-ui-lint-not-covered-claimReversed0/51/55/55/5
reversed-zoom-webhook-fail-open-claimReversed1/50/55/55/5
scoped-ali947-title-vs-body-misreadScoped4/55/55/55/5
scoped-align-check-pin-bump-action-onlyScoped0/50/55/54/5
scoped-align-gate-app-vs-workflow-pairScoped5/55/55/55/5
scoped-altitude-backfill-vs-producer-decisionScoped5/55/55/55/5
scoped-conflict-severity-tier-retirement-vs-ali245Scoped0/55/50/50/5
scoped-design-tokens-type-scale-exceptionScoped0/51/55/55/5
scoped-telemetry-anonymous-vs-ali443Scoped5/55/55/55/5

Reproducing it

The harness, the 30 tasks, the 600 raw runs and the 600 verdicts live in our private monorepo, checksummed. The tables on this page are the aggregates from that run, copied into this site and pinned by a test that re-adds every cell; the aggregates themselves re-derive upstream from the raw rows with no API key and no spend. The protocol was written down and committed before the first run, and the v1.1 ceilings before the re-run. If you want to see the rows, ask us and we will go through them with you on a call.

Align icon align.tech