Review is the job now.
When agents write most of the first drafts, the scarce skill isn't typing code, it's judging it. Here's how I review code I didn't write: triage first, a full-diff read, Claude's first pass, a cold second opinion, and a loop that learns from every human comment.

When I added coding agents to my workflow this summer, the biggest change wasn't how code got written. It was how much of my day went to reading it.
In my setup post I described a cockpit that sends each ticket to its own agent session and gets back a draft pull request with evidence attached. That works. It also means the pile of work in front of me is no longer a list of tickets. It's a queue of diffs I didn't type. This post is about how I work through that queue without letting quality slip.
01 — The situationThe bottleneck moved
Review has always been part of how I work. At finsit I spent a lot of time making critical flows dependable with Cypress and Jasmine, and the people I worked with there credited me with the courage to give honest feedback. At YieldComputer I defined a testing strategy across unit, integration and end-to-end tests and tried to lift standards across teams without having a formal title to do it. At InSpace I lead the Experience Team, where a review is often the last moment to ask whether an interface is really right for the person using it.
What changed with agents is the ratio. When agents do the drafting, writing stops being the slow part. Judging is. A browser-tab routine (open the notification, scroll the diff, lose your place, switch tabs to check a caller) doesn't survive that. So I rebuilt review the way I'd treat any bottleneck: a flow with clear stages and little friction between them.
- Triage in gh-dash
- Claude's first pass
- Full-diff read
- My pending review
- Submit
- Loop feeds the author
- Checklist learns
02 — The queueTriage before reading
I start from a queue, not from notifications. My PR list lives in gh-dash, a terminal dashboard for GitHub, and it's split by the action each PR needs rather than by repository: CI failing, changes requested, ready to merge, needs my review, and involved for threads I'm part of but don't own.
A red build and a PR waiting on my approval need completely different kinds of attention, and one mixed inbox gives each of them half. Triage is quick, and it decides where the deep reading goes.
From any PR in that list, a handful of keys do the rest:
| i | Review in lazygitOpen the whole PR as one staged diff in a throwaway checkout. |
| I | Claude reviews firstRun an independent AI review in the background. Its findings land in my pending review, not on the PR. |
| f | SubmitApprove, request changes or comment, with an optional summary. |
| R | RebaseUpdate the branch on top of its base without leaving the queue. |
| T | Jump to the authorFocus the agent session that wrote this PR, or reopen it in its worktree if it's gone. |
03 — The readRead the whole change
The web diff shows you files. I want to see the change: every file at once, with the surrounding code one keystroke away. So i runs a small script called pr-review. It fetches the PR head into a detached, throwaway worktree, then soft-resets it to the merge-base with the base branch. The result is that lazygit shows the entire PR as staged changes, as if I'd written it myself and was about to commit. My own checkout is never touched, and the worktree is deleted when I quit.
# the PR head, detached, in its own throwaway worktree $ git worktree add --detach <tmp> <pr-head> # rewind to where the branch left base, keeping the changes staged $ git reset --soft $(git merge-base origin/<base> <pr-head>) $ lazygit # the whole PR is now one staged diff
Inside that session, three keys build a real GitHub review without opening a browser:
| c | Comment on a linePick an added, removed or unchanged line (or the whole file) and write the comment. It becomes a thread in my pending review. |
| C | Show pending commentsEverything I (and Claude, if it went first) have drafted so far. |
| S | SubmitComment, approve or request changes. Only now does anyone else see anything. |
Two details make this trustworthy. The line picker builds its list from git's Myers diff, so the hunks line up with the diff GitHub accepts comments on and every line number I pick is one GitHub will take. And because everything goes into a pending review, I can draft, reread and delete before a single word reaches the author. A review is a piece of writing, and I'd rather edit it before I send it.
04 — The first passClaude first, never last
Before reading a larger PR, I can hand it to Claude first with I. In the background, Claude checks out the PR on its own and reviews it against that repository's checklist. It works read-only: it can read files, grep and run git diff, git log and gh pr view, but editing and writing are switched off. When it's done, I get a macOS notification, and its findings are already sitting in my pending review.
The instructions shape what it writes. Every finding has to be anchored to a line inside the diff and written as the comment the author will read: the problem, the concrete input or state that breaks it, and what to change. Judgement calls are phrased as suggestions. If a finding can't be attached to a line GitHub accepts, it's kept as a file-level comment rather than silently dropped.
What this buys me is focus. There are too many examples to pick one, but they fall into a clear pattern:
- An error branch that swallows the failure, or reports success when part of the work was skipped.
- A caller, type or test that wasn't updated along with the code it depends on.
- Tests that only cover the happy path, or would still pass with the new logic deleted.
- Empty, loading and error states nobody designed for.
- Whether this is the right change for the people using it.
- Product intent the ticket never wrote down.
- Trade-offs that live outside the diff: other teams, other services, the roadmap.
- Whether the interface is consistent with the rest of the product.
Having the first list flagged before I open the diff leaves my attention for the second, which only a person with context can judge.
05 — Complex workA cold second opinion
For work the cockpit rates as complex or critical, there's one more reviewer before the PR is handed to me. A fresh agent gets the PR reference and the repo's checklist, and nothing else: no plan, no reasoning, no conversation history. It reviews what the diff is, not what the author meant. One line in its instructions has become a rule I use myself:
“If something only makes sense with context you weren't given, that's a finding. A human reviewer won't have that context either.”
It works through the same lenses I'd use, and it's told to try to break the change rather than confirm the happy path:
- Claim vs diffDoes the code deliver every Definition of Done item the PR ticks off?
- CorrectnessEdge cases, error paths, empty states, races, partial failure.
- Hidden couplingThe caller the diff forgot to update.
- ChecklistThis repository's known traps, one by one.
- Test adequacyWould these tests fail without the change?
- Security & dataAuthorisation on new paths, checks the server must repeat.
Before it reports anything, it has to try to refute its own findings and drop whatever it can't back with a concrete failure scenario. It only approves when it has zero confirmed findings, and it never approves anything on GitHub. My favourite line in its brief: a short honest report beats invented nitpicks. That's true for human reviewers too.
06 — The human partWhat I look for
Automation takes care of a lot of the mechanical work. What's left is judgement, and that's where my experience earns its keep. These are the questions I bring to every diff:
- Is it right for the person using it?"Does it work" is where I start, not where I stop. Who is using this, what are they trying to achieve, and what is the interface really asking of them? An interface has to tell one story: a badge that says "failed" next to text that says "due shortly" is a bug, even when every test is green.
- Would the tests catch a regression?A test that still passes when you delete the check it protects proves nothing. I look for failure paths, not just happy ones, and for the second interaction: open, act, close, then open again. Every bug fix should come with a test that failed before the fix.
- Does it fail closed?An error from a check that guards a write has to stop the write. A job that half-worked must not report success. And anything the interface checks, the server has to check again.
- Is the interface clear?Names that promise what they deliver (a
parsefunction can fail, atofunction can't), one responsibility per function, and shared components instead of one-off overrides. A comment explaining a block of code is often a sign that the block wants to be its own function. - What does it break for someone else?Removing or renaming anything other code depends on is a breaking change, and the PR should say so out loud.
None of these are new ideas. Most of them I learned the slow way, in teams where quality came from steady discipline rather than heroics. What's new is that I have time to ask them of every change.
07 — After the reviewClosing the loop
Once a review is submitted, a loop in the cockpit takes over. It checks every open PR on a timer: CI status, the review decision, and any new human comments since it last looked. Anything actionable goes straight back to the agent session that wrote the code, with a brief that stands on its own: the PR, exactly what fired (the failing checks with trimmed logs, or each comment quoted with its file and line), and a standing instruction. Fix it, re-run the gate, self-review only the new diff, push, then reply to each thread.
It works through red CI first, then requested changes, then merge conflicts, and it limits how many sessions it wakes at once so the laptop stays usable. It never merges, closes or marks a PR ready for review on its own. Those are my calls.
The checklist that learns
The last step is the one that compounds. When a human review comment applies beyond its PR, the loop adds it to that repository's review checklist. Every future task reads that checklist while it codes and again during self-review, and Claude's first pass reviews against it too. So a comment I make once becomes something the next PR is checked for automatically.
There's one hard rule: only human feedback goes into the checklists. Never bot comments, never the loop's own opinions. The system can't be allowed to mark its own homework.
08 — HonestlyThe trade-offs
- Review fatigue is real. More drafts arrive than I could ever write, and attention doesn't scale the way agents do. My defences are structural: the queue is triaged by action, how much process a task gets depends on its complexity so small changes stay small, and the mechanical checks are done before I start reading.
- Rubber-stamping is the risk I watch most. When an AI reviewer has already been through a PR, it's tempting to skim. So the AI never approves: its findings are drafts in my review, the cold reviewer can't approve on GitHub, I still read the full diff, and teammates still review. Agents open drafts; people decide what merges.
- AI reviewers can be noisy. That's why they have to defend every finding with a concrete failure case, and why I delete freely before submitting.
- It costs something. Every AI review has a price, and I can see it per review. I spend it where the risk justifies it, not by reflex.
- It's my own glue. A shell script, a lazygit config, a few keybindings and skills. When it breaks, I'm the one who fixes it.
The shift I've accepted is simple. Writing code is now often the cheap part. Deciding whether it's right, for the codebase and for the person on the other side of the screen, is the job. It deserves the same care and tooling I'd put into anything else I ship.
How do you review?
Questions, pushback or your own take? I read every email.


