●Jan Layola
Blog/04/AI agents · Process

Plans before code.

Every task my coding agents pick up starts read-only, with a written Definition of Done and a plan. Here's why, what goes into them, and how the ceremony scales so a copy fix takes minutes while a risky change gets a cold second review.

Illustration: a read-only plan for extracting a shared date-range picker, with a review comment on a misread business rule

When people first hand work to a coding agent, the instinct is to type the request and let it run. I understand the appeal. But everything I've learned about shipping software says the expensive mistakes happen before the first line of code: misreading the goal, missing an edge case, touching the wrong part of the system.

So in my setup, no agent writes code until it has told me, in writing, what "done" means and how it will get there. This post covers where that habit comes from, what it looks like in practice, and how I keep it from turning into bureaucracy. (I wrote about the whole setup separately.)

01 — Where it comes fromWriting down "correct"

At finsit, part of Wolters Kluwer, I grew from junior into someone the team relied on for critical processes. The lesson that stuck from that time is that you can only hand work to someone else when tests and clear writing define what "correct" means. A good ticket, a clear acceptance criterion and a test that fails when it should are what let a new colleague change code they didn't write without breaking it.

At YieldComputer I learned the other half: teams get fast when the safety lives in the process, not in someone's memory. Nobody should have to remember to check; the pipeline checks.

An agent is the most extreme version of a new colleague. It's capable, it's fast, and it has none of the context that lives in the team's heads. So the question I care about isn't "can it write this?" but "can I trust what comes back?" My answer is the same as it would be for a person: agree on what done looks like, agree on the approach, and then let them work.

The same idea, at work
NOVA, the platform my team builds at InSpace, promises its clients: "AI does the work. You make the call." Nothing ships without their yes. I hold my own agents to the same rule.

02 — The defaultRead-only by default

Every task session starts in its own git worktree, on its own branch, in plan mode: it can read the code, search and investigate, but it can't edit a file or run a command that changes anything. It stays that way until I approve its plan.

That sounds slow, but reading is cheap and wrong code is expensive. A plan is the cheapest place to be wrong, and there's a worked example of why below.

Once the plan is approved, a shared settings file pre-approves the tools the workflow needs (editing files, the package manager, git, the GitHub CLI, the test commands), so the agent doesn't stop to ask for every edit. Approval is one decision at the right moment, not fifty interruptions.

03 — The contractThe Definition of Done

The first thing a task session does is read the ticket and turn it into a Definition of Done: a checklist of observable outcomes. Not "implement the filter", but things you can check:

definition of done
# observable outcomes, not activities
- [ ] user can <do the thing>
- [ ] <component> renders when <condition>
- [ ] no regression in <the flow next door>

That list isn't paperwork. It travels through the whole task, and every later stage answers to it:

  1. Definition of Done
  2. Plan addresses it
  3. Code satisfies it
  4. Verification walks it
  5. PR reports each item
waits for me

In the pull request, each item is ticked only if it was actually verified, in a test or in the running app. Vague acceptance criteria become a clarifying question before any planning. And for a trivial change, the Definition of Done is one line.

04 — The approachWhat goes in a plan

Before planning, the agent investigates. The rule is that it should never ask what the code or the ticket already answers. Whatever is still genuinely open (scope, edge cases, design choices) goes into one batched round of questions. Then comes the plan, which is short and always has the same shape:

  • Files to touch and the approach, in a few lines.
  • Risks: what could break, and what the change sits next to.
  • The Definition of Done, restated, so I can see the plan actually covers it.
  • The test strategy: which new or updated tests will cover the change. New behaviour ships with tests that would fail without it. The only exemptions are changes with no testable logic (copy, styling values, config), and each one has to be stated with its reason in the PR. "This area has no tests yet" isn't an exemption; it's a reason to write the first one.

Two details make this work in practice. The first is batching every wait. Each time an agent stops to ask me something, it can sit there for hours while I'm busy elsewhere, so it gets one round of questions and one plan approval, never a trickle of stops.

The second is an honest check on effort. If the investigation shows the work is heavier than the ticket suggested, like a data-model change hiding behind a "small" ticket, the plan says so and suggests restarting at a higher tier. It's advice, not a blocker: I decide.

05 — A worked exampleA plan, reviewed

Here's what a plan for a refactor looks like. It's an illustration rather than a real ticket, but it's the kind of work my team does all the time at InSpace: several reporting screens in NOVA each grew their own copy of a date-range picker, and the task is to replace them with one shared component in the design system. A standard-tier task, so the plan waits for me:

plan · standard tier
task  Move the date-range picker into the design system

definition of done
- [ ] one DateRangePicker in design-system/date-range-picker/
- [ ] every reporting screen uses it; the local copies are deleted
- [ ] presets, keyboard use and focus order behave as before
- [ ] no visual change on any screen

files
  design-system/date-range-picker/
    new component, stories, tests
  features/reporting/*/filters.tsx
    swap to the shared picker
  features/reporting/lib/presets.ts
    delete, moved into the component

approach
  1. start from the most complete copy
  2. normalise the presets into one list
  3. compute ranges as [today − n days, today], in the browser's timezone
  4. swap one screen at a time, deleting each old copy after its swap

risks
  - screens pass dates in two formats (Date and ISO string)
  - while I'm here: align the picker with the new input styles

tests
  - unit: each preset's range, month boundaries, leap years
  - interaction: keyboard selection and focus return, in Storybook

open questions
  none

It looks thorough, and most of it is good. It also makes three decisions it had no business making on its own. They're highlighted above, and these are the comments I'd leave:

  • A business rule, read as codenormalise the presets into one list · [today − n days, today]The copies aren't duplicates by accident. In this example, the product defines "Last 30 days" as the 30 complete days before today, because today's numbers are still coming in, and one screen's "This month" means month-to-date while another shows the full calendar month. "Normalising" them would pass every test the agent wrote, since the tests would encode the same assumption, and quietly change numbers clients look at. Keep each definition exactly as it is. Where two copies disagree, that's a question for product, not a refactoring decision.
  • An assumption about something outside the repoin the browser's timezone · open questions: noneHow dates are interpreted, and which ranges the reporting API accepts, is decided in another service. The agent can't see it, so it shouldn't guess it. A refactor that touches an external contract and has no open questions is the tell. Send exactly the dates each screen sends today, and list the rest as questions.
  • Scope creepwhile I'm here: align the picker with the new input stylesNot in this task. A refactor with a visual change in it is two changes in one diff, and the Definition of Done already says "no visual change". Separate ticket.

After one round of comments, the parts that changed read like this:

plan · revised
- 2. normalise the presets into one list
+ 2. keep every preset's definition exactly; list copies that disagree
- 3. compute ranges as [today − n days, today], in the browser's timezone
+ 3. send exactly the dates each screen sends today
- while I'm here: align the picker with the new input styles
+ out of scope: restyling (separate ticket)
+ tests: each screen sends the same query before and after the swap
- open questions: none
+ open questions:
+   - which timezone does the reporting API expect? (lives outside this repo)
+   - two screens define "This month" differently: which one is right?

Those are the two mistakes I look for first in any plan, and they're the same ones a capable new colleague makes: treating a business rule as an implementation detail, and filling a gap outside the repo with a plausible guess. Both look perfectly reasonable in code, and CI passes because the tests share the assumption. Reading the plan took a minute. Catching the same thing in review means reading a diff across every reporting screen. Catching it after release means explaining to clients why their numbers moved.

06 — ProportionalityScaling the ceremony

There's a trap in all of this. Apply the full process to every task and a one-line copy fix takes an afternoon. My first version did exactly that: small fixes went through the same heavyweight process as real features. The fix was to make depth scale with a complexity tier that I pick when I dispatch the task:

TierQuestionsPlanSelf-reviewOn top
Trivialcopy, a rename, a constantNone; assumptions stated in the planThree lines, approved at a glanceReads its own diff against the checklistOne screenshot, if visual
Scopeda prop, a known bugAt most oneShort, approved at a glanceA medium-effort code reviewA quick check in the running app
Standarda new page, a real featureOne batched roundWaits for my approvalFull: code review, simplify, checklistWalks the whole Definition of Done
Complexcross-cutting, data model, authOne batched roundWaits for my approvalCode review and checklist, thinking harderIndependent cold second review
Criticalrare, severe if brokenOne batched roundWaits for my approvalSame, plus a small adversarial panelCapped parallel agents and a cold second review

The tier also picks the model and how hard it thinks, from a light, fast model for trivial work to the strongest one at maximum effort for the rare critical change. When I'm torn between two tiers, I pick the higher one: under-powering a hard task costs more than a bit of extra thinking.

One iteration I'm glad I made: tier up for thinking, not for fleet size. Complex work used to fan out into several parallel agents. Now only critical work does, and even that is capped. One strong session plus an independent review catches more per unit of machine than a fleet, and keeps the laptop usable.

07 — Complex & critical workA cold second opinion

For complex and critical tasks, the draft PR gets a second reviewer before it reaches me. It's a fresh agent that is deliberately given only the PR and the repo's review checklist: no plan, no reasoning, no conversation history. It reviews what the diff is, not what the author meant. Its instructions put it better than I can:

“If something only makes sense with context you weren't given, that's a finding. A human reviewer won't have that context either.”

It works through a fixed set of lenses and actively tries to break the change rather than confirm the happy path:

  • Claim versus diff. Does the code deliver every Definition of Done item the PR ticks off? An unmet ticked claim is the most serious kind of finding.
  • Correctness. Edge cases, error paths, empty states, races and partial failures.
  • Hidden coupling and the checklist. The caller the diff forgot, and this repo's known review traps, item by item.
  • Test adequacy. Would the new tests fail without the change? Are failure paths covered, or only the happy path?
  • Security and data integrity. Authorisation on new server paths, and checks that only exist in the client.

It has to try to refute its own findings, and anything it can't back with a concrete failure scenario gets dropped. The author fixes what's confirmed, one more cold review checks only the fixes, and if findings survive that, the task comes to me instead of looping. The reviewer is read-only: it never edits, pushes or approves, and its report is posted on the PR for the human reviewers.

08 — Trade-offsWhat it costs, honestly

What it buys
  • Misread rules and guessed contracts get caught in a paragraph, not in a rewrite.
  • PRs arrive with evidence, so reviewing them is checking, not detective work.
  • I stay in control of direction without having to stay in control of every line.
  • Trust grows: each agent's output becomes predictable in shape.
What it costs
  • Waiting. A plan that sits unread blocks a task, which is why waits are batched.
  • Attention. Rubber-stamped plans are theatre, which is why small plans stay small and the real ones get read.
  • Judgement. A wrong tier wastes time either way, which is why I err higher and the effort check exists.
  • It doesn't replace human review. It makes it faster.

The line from my task instructions that sums it up is this one:

“The deliverable is not a PR. It's a PR a human reviewer approves without rework.”

Everything before the code (the Definition of Done, the plan, the tier) exists to make that sentence true more often. It's not a new idea. It's what good teams have always done with a new colleague. The agents just made it impossible to skip.

Written by
Senior Engineer, Experience Team Lead at InSpace · Breda, NL

I build NOVA's client platform and the design system behind it, and write about product engineering, frontend craft and shipping software with coding agents.

How do you brief your agents?

Questions, pushback or your own take? I read every email.