PR review is not a prompt
Note: this is part one of two, and I have kept the wrong turns in, because they were more useful than the destination. Feedback very welcome.
Last Light has had a PR review workflow since almost the beginning. It is one of the first things you build once you have a harness that can watch a repository: a webhook fires on pull_request, you hand a coding agent the diff and a skill that tells it what a good review looks like, it writes some findings to a JSON file, and a deterministic phase posts them back to GitHub as a formal review.

Here it is on a real pull request. It is confident, it is specific, it uses the right vocabulary, and it approved with no inline comments at all. What you cannot tell by looking at it is whether it read the code or wrote a confident summary of the description.
This did not feel like a hard problem. Anthropic ship a code review skill. Matt Pocock has written about wiring one up. Half the agentic engineering community has a code-review/SKILL.md in a repo somewhere, and they all look roughly the same: a rubric, some severity levels, an instruction to look for the usual suspects, and a schema for the output. Some get more sophisticated and ask for specific sub-agents to look across security, maintainability etc; and mine was no different. It ran, it produced sensible looking reviews, and for a few months I was quietly pleased with it.
The reason I am writing this post is that it was not working, I did not know it was not working, and the thing that told me was not a bug report. It was real user feedback comparing with a human reviewer and then measurement.
The uncomfortable bit: it looks fine until you repeat and measure it
Here is the shape of the problem with reviewing anything with an LLM. A bad review and a good review look identical from the outside. Both are well written. Both are confident. Both use the right vocabulary. If the bot says “this looks good to me, nice separation of concerns” then you have no idea whether it read the code properly or whether it pattern-matched on the diff and produced a plausible noise.
And Last Light was saying that a lot. Over one representative period: 94 pr-review runs across 43 pull requests, 71% of them APPROVE, and of 59 approvals, 58 carried zero inline findings.
You can read that two ways. Either the code going through was genuinely clean, or the reviewer was not finding anything. I had been reading it the first way for months.
Building a golden dataset
The only way to settle it is to compare the bot against a human on the same code, and to do that you need “gold” answers and evals: real review comments, written by a named human, on a known commit SHA. If you want to know more about evals, check out the great blog post by my nearform colleagues Alfonso Graziano: From AI prototype to production: how to build evals for reliable agents .
Lastlight already ships with a rich evals harness; that allows you to test any of your workflows against real data to assess if the prompts, skills or workflows actually do what you expect them to do reliably.
Now I had real feedback from a maintainer of skillspro, an internal Nearform repository with a healthy review culture. They had given some clear feedback that Lastlight had completely missed a critical set of review items. This meant that I had a set of PRs where a human reviewer had reviewed a specific sha that was still retrievable and their exact review comments and line references.
That left me with eight cases containing 25 gold findings. This collapses to four pull requests, since review rounds on the same PR are correlated, with a three case blind split held out so I could tell overfitting from improvement. It is a small dataset. It is nowhere near enough to make a general claim, and I will come back to that. But it is enough to answer the only question I actually had, which was “is this thing working at all”.
How any of this gets scored
Every number in this post comes out of the same shape, and it is worth thirty seconds because the rest of the post is numbers and lots of definitions that likely make no sense unless you are already deep in this world. The shape is borrowed from Martian’s Code Review Bench, which is the public benchmark this field is currently ranked on: take a real merged pull request, treat the review comments a human actually left on it as the right answers, run the tool on the same commit, and have a judge model decide which of the tool’s findings correspond to which of the human’s.
That gives three numbers.
- Recall - of the defects the human found, what share did the bot also find? This is the miss rate, and it is the one I care about. A reviewer that misses most things is not a reviewer.
- Precision - of everything the bot posted, what share was a real defect? This is the noise rate. A reviewer nobody reads because it cries wolf is also not a reviewer.
- F1 - the two combined into one number, punishing you for being bad at either. It is what leaderboards rank on, and it is the number I do worst at.
Two details that matter later. I pool rather than average: micro-recall is total matched divided by total gold across all the pull requests, so a pull request carrying five defects counts five times as much as one carrying a single defect. And the match is decided by a model. A judge reads what the bot posted and what the human wrote and rules on which pairs up. That is standard for this kind of benchmark and it is the softest part of the whole setup. Mine spent weeks crediting the bot for reporting that something was correctly handled when the human’s comment said it was broken, because the judge was matching on the topic of the sentence rather than on which way its claim pointed.
One more label you will see in the tables. Three of the eight cases - twelve of the twenty five defects - are a blind split: held out, never read while I was changing anything. That is what tells a number going up on the other five apart from me quietly fitting the tool to defects I had memorised. Where a table has a blind row, it is recall on those three cases alone.
One structural flaw is worth knowing before any precision number below. The gold is not a complete list of the defects in the pull request - it is what one human happened to write down that day. A bot that finds a genuine problem the human never mentioned is scored as a false positive for it. That pushes every precision figure down, and it pushes hardest on reviewers that say a lot, which by the end of this post is mine.
One of twenty five
So I ran it, on Sonnet 4.6. The baseline arm cost $5.65 and produced this:
| Metric | Result |
|---|---|
| Gold findings | 25 |
| Findings posted | 2 |
| Findings matched | 1 |
| Recall | 0.040 |
| Precision | 0.500 |
| F1 | 0.074 |
| Recall on the blind split | 0.000 |
It posted nothing at all on five of the seven recall cases.
And that 0.040 turned out to be generous. Two days later I found that my own fixtures had been handing the agent its previous review of the same pull request, so part of what it “found” it was reading back off a note it had already written. With that repaired, and re-run on Haiku 4.5 - the model everything below uses - the original reviewer matched none of the twenty five. I ran it again and it matched none of them again. A third run, over seven of the eight cases, posted nothing at all.
Zero of twenty five is the number the rest of this measures against. One of twenty five is just the number that got me to look.
The part that bothered me most was not the number. It was what the transcripts showed (lastlight evals records and allows you to mine every interaction the agent takes as if you were directly steering it). This was not a lazy agent. It ran 54 to 68 turns per review. It opened files, followed imports, wrote a genuinely good cross-file trace of what the change touched, and then concluded “no findings” and threw all of it away.
It was not failing from lack of effort. It was failing at something else, and I did not yet know what.
Three theories, all wrong
What follows is about $38 of model spend and three rounds of work, each of which killed a theory I was confident about. I am writing them out because the falsifications turned out to be worth more than anything that worked.
Theory one: it is hallucinating, so it needs verification. This is the obvious one, and it is what most of the literature was pointing at. I built a version that generated findings, ran them through a machine-checked evidence gate where every claim had to be quote-backed, and then handed the survivors to a fresh-context adjudicator.
Mechanically it worked perfectly. Every disposition became quote-backed and machine checked. Posted findings went from 2 to 8. And recall went from 1 of 25 to 2 of 25, while F1 halved, the precision canary went from 1.00 to 0.00, and it cost 2.4 times as much. I deleted the machinery.
Theory two: it is a prompt problem. Three rounds of instruction work. A 22 item checklist got acknowledged in a single sentence and skipped entirely. A 17 row ledger got honestly discharged, row by row, and still produced zero findings. The model was not disobeying. It was doing exactly what I asked, at the level of abstraction I asked it, and that level was too high to find anything.
Theory three: it needs a bigger model. This one is worth dwelling on because it is the reflex answer for everything right now. It is wrong here, and there is published evidence: Haiku 4.5 beats Sonnet 4.6 on review recall, 41.2% against 22.1%. Martian’s leaderboard shows a roughly 28 point gap attributable to scaffolding at a fixed model class. The scaffolding is the variable. The model is not.
Three dead theories later, the actual diagnosis is embarrassingly simple:
The model’s question set does not contain the questions an experienced human reviewer would ask themselves when they are looking at the code: Discovery is the ceiling, not verification.
You cannot prompt your way to a question about a file you never knew was connected to the diff. All three of my theories were about improving the answers. None of them were about improving the questions.
So how good is the state of the art anyway?
At this point I went looking for what the ceiling actually is, half expecting to find that everyone else had solved this and I had missed a memo.
I had not.
- CR-Bench (arXiv:2603.11078) puts GPT-5.2 single-shot review agents at 27.0% recall and 3.6% precision.
- c-CRAB (arXiv:2603.23448) grades a review by synthesising a test from each human comment and checking whether it goes from failing to passing after the fix, which removes the judge problem entirely. Claude Code scores 32.1%, Devin Review 24.8%, PR-Agent 23.1%.
Two things fall out of that. The first is that my 0.040 is genuinely bad and I should not feel good about it. The second is more interesting: frontier precision on this task is 3 to 5%. Every serious reviewer in the field over-generates massively and eats the noise. My baseline posted 2 findings across 8 pull requests, which I had been reading as admirable restraint. It is the exact opposite of what the field does.
There is a lovely detail in the CR-Bench numbers that made me rethink the whole approach. Adding a second pass whose only instruction is “find what the first pass missed” moves GPT-5.2 from 27.0% to 32.8% recall. It also drops signal-to-noise from 5.11 to 1.95. You buy recall with noise. That is the trade, and it is not optional.
The products are good. They are also not prompts and skills.
Here is the thing that reframed it. The commercial products in this space are meaningfully better than a bare prompt, and when you read how they work, not one of them is a skill.
CodeRabbit clones the repo into an isolated microVM, compiles the project, runs twenty to fifty linters and analysers in parallel, builds a fresh call graph per review, and then has a separate judge model score every finding against that gathered context and drop what it cannot ground. Greptile maintains a semantic graph of the whole repository before a pull request even arrives, and its TREX work spins a sandbox to actually execute the behaviour in question and attaches the logs as evidence. Cursor Bugbot, the only one publishing before-and-after numbers, went from 0.2 to 0.5 resolved bugs per pull request by moving from a fixed pipeline to an agent pulling context at runtime, across 40 experiments in which, in their words, many intuitive changes surprisingly regressed the metrics.
Compiled projects, sandboxes, call graphs, parallel analysers, separate judges, executable oracles. Not one of them is a rubric in a markdown file.
I had been comparing my skill to other people’s skills, when the thing I should have been comparing it to was other people’s pipelines.
What I built instead
The ASCII sketch I had been carrying around in the plan / spec process with Claude Code is now nine phases with an actual specification, so here it is properly.
| Phase | Costs | What it actually does |
|---|---|---|
prepare | - | Installs the pull request’s dependencies into a throwaway copy of the code. Sounds like plumbing; it is not. A TypeScript project that inherits its config from an installed package cannot resolve without it, so the compiler quietly drops that project and two of the five question families go blind. It is gated on a flag that ships off, and it was skipped in every run behind every number here. |
facts | - | Runs static analysis over exactly the lines the diff touched and writes ~140 KB of machine-readable evidence: which symbols changed, everywhere that calls them, which exported signatures moved and who consumes them from outside the diff, which constants are referenced properly against where the same value is hard-coded, dependency changes, and scanner hits. |
seed | - | Turns that evidence into a numbered list of concrete questions, each naming both ends of a possible defect - “this value is set at A; what checks it at B?” A question with only one end is thrown away and counted: it sends the model hunting rather than looking, and the measurements below say that can cost more than it finds. |
survey ×5 | model | Five cheap model passes run at once, each handed exactly one family of questions and told to over-produce. Each writes to its own append-only file, so no pass can overwrite, argue with, or quietly drop another’s output. |
falsify | model | Takes the claims and tries to settle them by running something - writing a probe, executing it, keeping the transcript. A claim may only be deleted later if there is a transcript proving it failed to reproduce. This was designed, but on later tests proved not helpful - but remains behind a flag. |
review | model | An independent read of the diff that is not allowed to see anything the phases above produced, because a second opinion that has read the first one is not a second opinion. With the pipeline off this is the original reviewer byte-for-byte, so the off switch is a real off switch and not a different product. With the pipeline on it is a deliberately short pass. |
adjudicate | model | A fresh model that has seen none of the earlier reasoning reads every claim and decides which are real, how confident to be, and which deserve a comment. It must account for every single claim - agents shown the reasoning that produced a false report overwhelmingly fail to reject it, which is why this one is not allowed to see it. |
reconcile | - | A deliberately stupid safety net. Anything the adjudicator failed to account for gets filed privately, and any deletion whose transcript does not actually exist is put back. Nothing can be lost by a model simply running out of turns. |
post-review | - | Decides what a human actually sees: at most ten inline comments at the defect site, at most five more in the review body, everything else kept in a private record. Then posts exactly once, and writes down everything it chose not to say. |

One fact, from a real pull request on cal.com. The field doing the most work is inDiff. A symbol that changed and is only used inside the diff is usually fine; a symbol that changed and is used at four sites the author never opened is where the defects live, and no amount of reading the diff will tell you which one you have.
And here is one of the survey passes discharging its questions. Every stated obligation comes back either quoted to a line or explicitly marked absent, and anything it could not settle becomes a hypothesis for the phases downstream:

Each survey pass gets exactly one of these five families and this is where the bulk of the work and spend happens. A question only exists if the deterministic layer found something specific to hang it on, which is what stops the whole thing from being a checklist:
| Family | A question is raised when | And it asks |
|---|---|---|
| contract | an exported shape moved, and something outside the diff consumes it | does every consumer the author did not touch still satisfy it? |
| enforcement | a value is defined on one side of a boundary | who checks it on the other side? |
| security | a changed symbol sits in a file a scanner also flagged | does any path into it carry attacker-controlled input? |
| state | a changed symbol is used at sites the diff never touched | what about ordering, lifecycle and cache invalidation at those sites? |
| spec | the pull request body or a linked issue states a criterion | does the change actually do that? |
The first four are minted from static analysis, so they are only ever about code that really moved and really has consumers. spec is the odd one out: its questions come from what a human said they were going to do in the linked issue / PR, which is the axis a standards review structurally cannot check.
There are four ways to close one of these questions, and only one of them is a clean bill of health: quote the line that answers it, quote the line and name the gap it still leaves, say it can only be settled by running something, or say that every candidate was read and no such line exists.
That last one is the point of the whole exercise. It is not a failure to answer the question. It is the finding.
A sixth family, tests, was cut on the way to shipping. It needed a coverage report that only the probe phase can produce, the probe phase is off by default, and so it had produced zero artifacts in every measured run. It was paying a fifth of the fan-out to write “not measured”.
That family is where a deliberate boundary sits, and it is worth stating outright: the pipeline does not run the test suite, the linter or the type checker, and it is not going to. CI has already run all three on this exact commit, on a matrix I cannot reproduce on one machine, and a bot that re-derives a red check is spending a maintainer’s attention on something they saw before it arrived. So the review is handed CI’s result as evidence and forbidden from restating it. A failing check can be cited when it confirms a finding of the review’s own - this fails typecheck on line 42, which is the same issue as finding 2 - and a passing one buys silence on whether the thing builds, so the whole pass goes on judgement instead. The coverage extractor obeys the same rule: it reads a report if the repository already produced one and intersects it with the diff, and it never runs a suite to make one. The one execution the design does reserve is the targeted probe - running one specific thing to settle one specific question - which is different work from re-running the gate.
Two rules in that diagram I would have got wrong without the measurements.
The deterministic layer generates hypotheses. It does not filter them. This is theory one, inverted. BitsAI-CR (arXiv:2501.15134) independently reproduced the result I had paid for: a review filter raised precision from 54.5 to 67.1 and cut recall from 45.5 to 39.8.
A question that names only one end of a defect is a coin flip. Every defect mechanism has two ends: a place where something is introduced, and a place where something should have caught it. “Does this handler validate its input?” is one end. It names where the value enters and leaves the model to go hunting for whatever ought to check it - which is the open-ended search the original reviewer was already losing. “parseLimit now returns a string, and three call sites outside the diff still treat it as a number - do any of them bound it?” is both ends, and it is answered by reading two named places.
IRIS (arXiv:2405.17238) measured what that distinction is worth on a vulnerability benchmark, where the two ends are where untrusted data enters and where it does damage. A static analyser on its own found 27 of 120; paired with a model supplying both ends it found 55. Give away one end and the result depends entirely on which one: one half scored 36, the other 24, which is worse than the analyser working alone. Half a mechanism is not a weaker question, it is one that can cost you findings - which is why seed throws those away and counts them rather than asking them.
And one number sets the ceiling over all of it. Take a wider pool - 137 findings real reviewers left on 50 pull requests across cal.com, Grafana, Keycloak, Discourse and Sentry - keep the 99 that can be pinned to a changed line at all, pull out the identifiers each human mentioned, and ask whether the deterministic layer even names the thing the finding is about. On the TypeScript half that is 46.2%. Naming is necessary and nowhere near sufficient, but it is the difference between “the information is not there” and “the information is there and nothing is using it”.
That is a TypeScript number, and it should be: all of this rests on a compiler that can say which exported signature moved and who consumes it from outside the diff, and I have one of those for one language. On the Java, Ruby, Go and Python half of the same pool it is 2.7% - no symbols to walk, so the static families mint nothing. A small per-language parser wins most of that back and the extension point is already sitting there with one entry in it, but that is later work. For now a non-TypeScript review marks itself degraded and lists what it could not see, rather than going quiet and confident.
The seam between what you can test and what you cannot
The pipeline has a hard line through the middle of it, and it turned out to be the most useful property of the design, and actually one of the most useful parts of designing agentic workflows in general.
Everything up to and including seed is deterministic. Same commit in, same facts out, same questions out. I can say that with more confidence than I usually can about software, because when I ran the identical configuration three times, the generated questions came back byte-identical on all eight cases - same identifiers, same families, same counts. That half behaves like ordinary code. I can unit-test it, diff it, pin it with fixtures and reason about it.
Everything after it is not. Handed that byte-identical brief, the models produced 18 hypotheses on one run and 43 on the next for the same pull request. 10 against 23 on another. And across three identical runs of the whole pipeline, matched defects went 8, then 2, then 5.
Two things follow from that, and they are why I would build it this shape again.
Push work across the seam wherever you can. Not because deterministic code is cleverer - it plainly is not, it cannot review anything - but because it is the only half that holds firm under repetition. Every question the static layer mints is a question the model does not have to think of, and one I can regression-test forever. The model’s job shrinks to the part only it can do: looking at the code and forming a judgement.
On the other side of the seam, repetition is the only instrument you have. There is no unit test for “does this prompt find defects”. A single run of a non-deterministic system is an anecdote. I have a number I like and a number I do not - 0.320 and 0.080 - from the same code, on the same fixtures, on the same evening, and nothing but running it again distinguished them.
That is what an eval actually buys you, and it is not a score. It is the ability to tell a change from a draw. I did not appreciate the difference until I had a result I wanted to believe.
Where it lands
Eight pull requests, 25 real defects. The first measurement from the top of this post, then the original reviewer measured properly, then the pipeline, then the pipeline after one more change I will come to:
| first measurement | original reviewer | pipeline | + the seeding fix | |
|---|---|---|---|---|
| model | Sonnet 4.6 | Haiku 4.5 | Haiku 4.5 | Haiku 4.5 |
| fixtures | broken | repaired | repaired | repaired |
| defects matched | 1 of 25 | 0 and 0 of 25 | 9 and 7 | 9 and 11 |
| found at least once | 1 of 25 | 0 of 25 | 12 of 25 | 17 of 25 |
| generated, posted or not | - | - | 12 of 25 | 21 of 25 |
| findings posted | 2 | 1 and 1 | 41 and 39 | 56 and 60 |
| recall | 0.040 | 0 and 0 | 0.360 and 0.280 | 0.360 and 0.440 |
| precision | 0.500 | 0 and 0 | 0.220 and 0.179 | 0.161 and 0.183 |
| F1 | 0.074 | 0 and 0 | 0.273 and 0.219 | 0.222 and 0.259 |
| recall on the blind split | 0.000 | 0.000 | 0.500 and 0.333 | 0.333 and 0.333 |
| cost per pull request | $0.71 | $0.24 and $0.29 | $2.14 | $2.11 |
| cost for all eight | $5.65 | $1.93 and $2.28 | $17.11 | $16.90 |
The first column is the run from One of twenty five, sitting here so the two tables tie together rather than as something to compare against: a different model, and measured before I found the fixture bug. It is also a small lesson in reading precision on its own. 0.500 is the best precision anywhere in that table, and it is one correct finding out of two.
The second column is the one to sit with. Across eight pull requests containing 25 real defects, the original reviewer posted one finding and matched none of them - and it did that twice, on the same repaired fixtures, for a fifth of the pipeline’s price. Both runs posted their single finding on the same pull request. This is not one unlucky sample; it is what the thing does.
So the pipeline finds real defects the prompt never did, and it has never lost a finding the baseline had already caught. That is the thesis, and it holds.
The change that finally moved recall
What changed is not a model or a prompt. It came from looking at the defects the runs of lastlight had ever found and noticing something dull about them: every rule I had for generating a question required a reference from outside the diff. A value defined and used entirely within newly added code has no such reference, so no question was ever asked about it. Four of the never-found defects sat in files no question touched at all.
So seed learned two new ways to mint a question - one for symbols whose every use is inside the diff, one for route and hook registration order - and the result is the third column:
- defects generated: 12 of 25 to 21 of 25. Paired per defect, that is ten gained against one lost, at p = 0.006. It is the only change in this entire project that is significant on its own terms.
- defects actually posted: 12 to 17 of 25, nine gained against four lost, at p = 0.133. Real, not significant.
- mean recall 0.320 to 0.400, which is the first change to clear its own run-to-run noise band. The tool that scores these calls it KEEP, and it has said REVERT or INDISTINGUISHABLE to everything before it.
- and it costs the same money: $16.90 against $17.11.
And the bill for it
Precision fell from 0.220 to 0.161 and posted volume went up 40%, from 41 findings to 56. F1 went down, not up - 0.273 to 0.222 on the better of each pair - which is the honest way to read that column and the reason I do not lead with it.
That is not a surprise, it is the trade the field says it is. You buy recall with noise, and I bought some. If you rank on F1 this change was a regression; I am optimising for the misses, so I took it. It is also the exact opposite of the previous change, which bought precision and no recall - which tells you something uncomfortable about how much of this is one dial.
Two things are still true and neither is fixed by this.
Discovery moved; adjudicating did not. Twenty one defects generated against seventeen posted, and the significance sits on the first number, not the second. The pipeline is finding things and then failing to tell anyone about roughly a fifth of them. A fifth is tolerable and I nearly left it there. It does not stay at a fifth.
The per-run expectation barely moved. Across repeats, 17 of 25 were found at least once and only 3 of 25 every time - down from 4. A user does not get the union; they get one run, and which run they get is close to a draw.
And one earlier improvement is partly circular: the question catalogues were derived from defects the pipeline had historically missed, including some in the held-out split. Real improvement on those cases, but not evidence it generalises - which is why the next section exists.
Does any of it generalise?
Everything above is eight pull requests from one private repository, tuned against its own gold. That is not evidence of anything general, and I have been saying so throughout. So I ran the same pipeline, unchanged, on ten cal.com pull requests from Martian’s public review benchmark - 31 defects, open-source code, gold written by other people, none of it used to build or tune any part of this.
Paired defect by defect it gained seven findings against the original reviewer and lost none, at a one-sided p of 0.008. Ten pull requests is still ten pull requests. But it is the first result here that was not measured on the material it was built from, and it went the right way. Part two comes back to those same ten, with better numbers and the full set of cautions they deserve.
What I had at this point
Eight private pull requests on which the original reviewer matched none of twenty five defects and the pipeline found seventeen of them at least once. Ten public ones where it gained seven and lost nothing, measured against gold the pipeline had never seen before. A deterministic half I can regression-test forever and a model half I can only sample.
I had one diagnosis that had survived everything I threw at it: the model’s question set does not contain the human’s questions, so mint the questions. The obvious next move was more of the same. Find the defects nobody has ever found, work out what question would have caught them, mint that question too.
I did exactly that, twice. Both times the questions got arguably better, and both times the reviewer got worse.
Part two is where the constraint went when it stopped being discovery, the ten configurations I ran to find it, and the single prompt rule that turned out to be burying the answers.
Previous posts in this thread: The harness is the product, What is harness engineering, and Meet agentic-pi.