Every team that installs an AI reviewer has the same argument about a month later. One person says the bot caught a real bug last week. Another says it is mostly noise and they stopped reading it. Both are telling the truth about the PRs they happened to look at. Neither has a number.
You can get the number in an afternoon. GitHub already records what happened to every comment the bot left: whether the code under it changed, whether someone resolved the thread, whether anyone replied. This note gives you a script that reads those records for every merged pull request in a window and prints one rate. It also shows what that rate looked like on two busy public repositories this month, and what to do with yours.
The one number
Every inline thread the bot opens ends in one of three ways.
- Addressed. The code under the comment changed after the comment was posted, or the bot itself confirmed a fix in a later commit.
- Dismissed. The code did not change, but a person engaged: they resolved the thread or replied to it. Someone read the comment and said no.
- Ignored. The code did not change, nobody replied, and nobody resolved it. The PR merged anyway.
The acted-on rate is addressed threads over all threads. It is the automated version of the Day 5 count in add AI to your review pipeline: how many comments the author kept, across every PR instead of the last ten.
The other two buckets matter as much as the rate. Dismissed is healthy friction: the team read the comment and disagreed. Ignored is the warning sign. It means comments are landing on a PR that nobody owns.
The script
It needs the GitHub CLI, logged in, and jq. It works on any repository your token can read, public or private. It only reads.
#!/usr/bin/env bash
# review-ledger.sh OWNER/REPO BOT_LOGIN [DAYS]
# Sorts every inline thread the bot opened on merged PRs into
# addressed, dismissed, or ignored. Needs gh (logged in) and jq.
set -euo pipefail
repo=$1 bot=$2 days=${3:-30}
since=$(date -u -v-"$days"d +%F 2>/dev/null || date -u -d "$days days ago" +%F)
gh api graphql --paginate \
-f q="repo:$repo is:pr is:merged merged:>=$since" \
-f query='
query($q: String!, $endCursor: String) {
search(query: $q, type: ISSUE, first: 20, after: $endCursor) {
pageInfo { hasNextPage endCursor }
nodes { ... on PullRequest {
number
reviewThreads(first: 100) { nodes {
isResolved isOutdated
comments(first: 30) { nodes { author { login } body } }
} }
} }
}
}' --jq '.data.search.nodes[]' |
jq -s --arg bot "$bot" --arg since "$since" '
[ .[] | .number as $pr | .reviewThreads.nodes[]
| select(.comments.nodes[0].author.login == $bot)
| .comments.nodes[0].body as $body
| { pr: $pr,
severity: ($body | split("\n")[0]
| if test("critical|blocker"; "i") then "critical"
elif test("major"; "i") then "major"
elif test("minor"; "i") then "minor"
elif test("trivial|nitpick|nit:"; "i") then "trivial"
else "untagged" end),
outcome:
(if .isOutdated or ($body | test("addressed in commit"; "i"))
then "addressed"
elif .isResolved
or ([.comments.nodes[1:][] | select(.author.login != $bot)] | length > 0)
then "dismissed"
else "ignored" end) } ]
| def tally: {
threads: length,
addressed: ([.[] | select(.outcome == "addressed")] | length),
dismissed: ([.[] | select(.outcome == "dismissed")] | length),
ignored: ([.[] | select(.outcome == "ignored")] | length) }
| . + { acted_on: (if .threads == 0 then "n/a"
else "\(.addressed * 100 / .threads | floor)%" end) };
{ since: $since,
prs_with_threads: ([.[].pr] | unique | length),
all: tally,
by_severity: (group_by(.severity)
| map({ (.[0].severity): tally }) | add) }'
Save it as review-ledger.sh and pass the repository, the bot’s login, and how many days back to look:
chmod +x review-ledger.sh
./review-ledger.sh mastra-ai/mastra coderabbitai 14The bot login is the name GraphQL uses, without the [bot] suffix you see in the UI. CodeRabbit is coderabbitai. GitHub Copilot’s reviewer is copilot-pull-request-reviewer. If you are not sure, list the commenters on one PR the bot reviewed and drop the suffix:
gh api repos/OWNER/REPO/pulls/123/comments --jq '.[].user.login' | sort -uHow it decides
Addressed uses two signals. The first is GitHub’s own isOutdated flag, which flips when a later push changes the lines a comment points at. The second is a line some bots, CodeRabbit included, write into their own comment when they see a fix land: “Addressed in commit …”. Either one counts.
You need both, because each one misses fixes the other catches. On the Mastra sample below, the two signals agreed on 68% of threads. Of the rest, 299 threads were fixed somewhere other than the commented line (the bot confirmed the fix, but the line stayed the same). Another 220 had their line rewritten without the bot confirming anything. Outdated alone gave 55%. The bot’s own marker alone gave 60%. Together they give 73%.
Severity is read from the first line of the comment. CodeRabbit opens every inline comment with a tag such as “Functional Correctness | Major | Quick win”, so the script buckets on critical, major, minor, and trivial. Bots that do not tag severity land in untagged. That is fine. You still get the overall rate.
Only merged PRs count. An open PR has not finished deciding. A closed one may have been abandoned for reasons that have nothing to do with review.
A worked example: Mastra and VS Code
Here is the script’s output from 8 October 2026 on two public repositories that run an AI reviewer on almost every pull request. Mastra uses CodeRabbit. VS Code uses GitHub Copilot’s reviewer. The windows are short because both repos merge a lot: Mastra merged more than 800 pull requests in two weeks.
| Repository · bot · window | PRs | Threads | Addressed | Dismissed | Ignored |
|---|---|---|---|---|---|
| Mastra · CodeRabbit · 14 days | 530 | 1,643 | 1,203 (73%) | 226 (14%) | 214 (13%) |
| VS Code · Copilot · 7 days | 307 | 630 | 366 (58%) | 238 (38%) | 26 (4%) |
These are two different results, and neither is bad.
Mastra acts on most of what CodeRabbit says. Nearly three threads in four led to a code change. The leak is in the ignored column, and the severity split shows where it is:
| Mastra severity | Threads | Acted on | Ignored |
|---|---|---|---|
| Major | 535 | 77% | 6% |
| Minor | 960 | 71% | 16% |
| Trivial | 144 | 66% | 17% |
Four threads were tagged Critical, and all four were addressed. Major comments are rarely ignored. Minor and Trivial ones are ignored at almost three times that rate. Minor is also the biggest bucket, with 960 of the 1,643 threads, so most of the reading burden sits in the tier the team trusts least. If this were your team, that is the first knob to try: fold Minor into the summary comment, or stop emitting Trivial, and see whether the Major rate holds.
VS Code reads everything and argues back. Only 4% of Copilot’s threads were ignored, which suggests the team treats the comments as something to close out before merge. But 38% were resolved or answered with no code change. That is not noise in the sense of when AI review is noise, where nobody reads anything. It is a review seat with an owner that disagrees a lot. Those 238 threads are the tuning list: read twenty of them, and the rules to turn off will be obvious.
These numbers describe two repositories in early October 2026. They are not benchmarks for CodeRabbit or Copilot. A different team, with a different PR template, would get different numbers from the same bot. That is the point of measuring your own.
Reading your numbers
Start with ignored, not acted-on. A high ignored share means the comments land where nobody owns them, and no prompt change fixes that. Above roughly one in five, put the line from the noise guide into your PR template: the author resolves or dismisses automated comments before asking for review. Then measure again in two weeks.
Then look at acted-on, with the same thresholds the noise guide uses for a hand count:
- Half or more, with low ignored: keep the seat. Spend your tuning time on the dismissed threads.
- A quarter to a half: narrow it. Turn off the lowest severity tier, exclude generated paths and files the linter owns, and run the script again on the next window.
- Under a quarter after you narrowed: stop. The seat is costing reading time and not paying it back.
Last, compare the severity tiers. If Major is not acted on clearly more often than Trivial, the severity labels are decoration, and your reviewers are right to skim them.
What the number does not tell you
Outdated is not the same as fixed. A refactor that rewrites the same lines for another reason will mark the thread addressed. On Mastra, the two signals disagreed on about a third of threads, so treat any single rate as plus or minus ten points. Use it for trends: the same repo, the same script, month over month. Before you act on a surprising number, open ten addressed threads at random and check them by hand.
It does not count what the bot missed. A reviewer that comments rarely and is always right will score well even if it sleeps through the bug that pages you on Saturday. Pair this rate with the one escape metric you already track, such as reverts or incidents traced to a merged PR.
It only sees inline threads. Comments posted in a review summary, or as a plain PR comment, do not have an outdated flag, so the script skips them. Most reviewers put their real findings inline. If yours does not, the rate covers less than you think.
GitHub search stops at 1,000 results, and the query reads at most 100 threads per PR. On a busy repo, shorten the window rather than trusting a truncated month.
Run it this week
Pick the repo where the bot argument is loudest. Run the script for the last 30 days. Write the three numbers (acted on, dismissed, ignored) in the channel where the argument happens. Name one person who owns the review seat, as described in what is an AI code pipeline, and let them make one change: a tier turned off, a path excluded, or a line added to the PR template.
Run it again in two weeks. If the number moved the right way, you have a seat that earns its keep. If it did not, you have the evidence to turn it off without another argument.
