<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Jan Layola — Blog</title>
    <link>https://birrejan.me/blog/</link>
    <atom:link href="https://birrejan.me/blog/feed.xml" rel="self" type="application/rss+xml" />
    <description>Jan Layola's blog on product engineering, frontend craft, release automation and shipping software with AI coding agents.</description>
    <language>en-gb</language>
    <lastBuildDate>Tue, 29 Sep 2026 09:00:00 +0000</lastBuildDate>
    <item>
      <title>Review is the job now</title>
      <link>https://birrejan.me/blog/review-is-the-job/</link>
      <guid isPermaLink="true">https://birrejan.me/blog/review-is-the-job/</guid>
      <pubDate>Tue, 29 Sep 2026 09:00:00 +0000</pubDate>
      <description>When agents write most of the first drafts, the scarce skill is judging code, not typing it. How I review code I didn't write: triage, a full-diff read, Claude's first pass, a cold second opinion, and a loop that learns from every human comment.</description>
      <content:encoded><![CDATA[<p><img src="https://birrejan.me/blog/img/review-is-the-job.jpg" alt="Illustration: a code diff with a pending first-pass review comment, next to a review queue" width="1600" height="900"></p>
<p class="lede">When I added coding agents to my workflow this summer, the biggest change wasn't how code got written. It was how much of my day went to reading it.</p>
        <p>In <a href="https://birrejan.me/blog/my-setup/">my setup post</a> I described a cockpit that sends each ticket to its own agent session and gets back a draft pull request with evidence attached. That works. It also means the pile of work in front of me is no longer a list of tickets. It's a queue of diffs I didn't type. This post is about how I work through that queue without letting quality slip.</p>

        <!-- ─────────────── 01 ─────────────── -->
        <section class="layer" id="moved">
          <h2><span class="n">01 — The situation</span>The bottleneck <em class="s">moved</em></h2>

          <p>Review has always been part of how I work. At finsit I spent a lot of time making critical flows dependable with Cypress and Jasmine, and the people I worked with there credited me with the courage to give honest feedback. At YieldComputer I defined a testing strategy across unit, integration and end-to-end tests and tried to lift standards across teams without having a formal title to do it. At InSpace I lead the Experience Team, where a review is often the last moment to ask whether an interface is really right for the person using it.</p>
          <p>What changed with agents is the ratio. When agents do the drafting, writing stops being the slow part. <strong>Judging is.</strong> A browser-tab routine (open the notification, scroll the diff, lose your place, switch tabs to check a caller) doesn't survive that. So I rebuilt review the way I'd treat any bottleneck: a flow with clear stages and little friction between them.</p>

          <ol class="pipeline" aria-label="My review flow">
            <li><span class="step">Triage in gh-dash</span></li>
            <li><span class="step tier">Claude's first pass</span></li>
            <li><span class="step wait">Full-diff read</span></li>
            <li><span class="step wait">My pending review</span></li>
            <li><span class="step wait">Submit</span></li>
            <li><span class="step">Loop feeds the author</span></li>
            <li><span class="step">Checklist learns</span></li>
          </ol>
          <div class="legend-row"><span class="wait">only I can do this</span><span class="tier">optional, runs in the background</span></div>
        </section>

        <!-- ─────────────── 02 ─────────────── -->
        <section class="layer" id="triage">
          <h2><span class="n">02 — The queue</span>Triage before <em class="s">reading</em></h2>

          <p>I start from a queue, not from notifications. My PR list lives in <strong>gh-dash</strong>, a terminal dashboard for GitHub, and it's split by the action each PR needs rather than by repository: <em>CI failing</em>, <em>changes requested</em>, <em>ready to merge</em>, <em>needs my review</em>, and <em>involved</em> for threads I'm part of but don't own.</p>
          <p>A red build and a PR waiting on my approval need completely different kinds of attention, and one mixed inbox gives each of them half. Triage is quick, and it decides where the deep reading goes.</p>
          <p>From any PR in that list, a handful of keys do the rest:</p>

          <div class="rij-group">In gh-dash</div>
          <table class="keys">
            <tr><td><kbd>i</kbd></td><td><strong>Review in lazygit</strong>Open the whole PR as one staged diff in a throwaway checkout.</td></tr>
            <tr><td><kbd>I</kbd></td><td><strong>Claude reviews first</strong>Run an independent AI review in the background. Its findings land in my pending review, not on the PR.</td></tr>
            <tr><td><kbd>f</kbd></td><td><strong>Submit</strong>Approve, request changes or comment, with an optional summary.</td></tr>
            <tr><td><kbd>R</kbd></td><td><strong>Rebase</strong>Update the branch on top of its base without leaving the queue.</td></tr>
            <tr><td><kbd>T</kbd></td><td><strong>Jump to the author</strong>Focus the agent session that wrote this PR, or reopen it in its worktree if it's gone.</td></tr>
          </table>
        </section>

        <!-- ─────────────── 03 ─────────────── -->
        <section class="layer" id="diff">
          <h2><span class="n">03 — The read</span>Read the whole <em class="s">change</em></h2>

          <p>The web diff shows you files. I want to see <strong>the change</strong>: every file at once, with the surrounding code one keystroke away. So <kbd>i</kbd> runs a small script called <code>pr-review</code>. It fetches the PR head into a <strong>detached, throwaway worktree</strong>, then soft-resets it to the merge-base with the base branch. The result is that lazygit shows the entire PR as staged changes, as if I'd written it myself and was about to commit. My own checkout is never touched, and the worktree is deleted when I quit.</p>

          <div class="term" role="group" aria-label="Terminal">
            <div class="term-bar"><div class="lights"><i></i><i></i><i></i></div><span>pr-review</span></div>
<pre><span class="c"># the PR head, detached, in its own throwaway worktree</span>
<span class="p">$</span> git worktree add --detach &lt;tmp&gt; &lt;pr-head&gt;
<span class="c"># rewind to where the branch left base, keeping the changes staged</span>
<span class="p">$</span> git reset --soft $(git merge-base origin/&lt;base&gt; &lt;pr-head&gt;)
<span class="p">$</span> lazygit   <span class="c"># the whole PR is now one staged diff</span></pre>
          </div>

          <p>Inside that session, three keys build a real GitHub review without opening a browser:</p>

          <div class="rij-group">Inside the lazygit review</div>
          <table class="keys">
            <tr><td><kbd>c</kbd></td><td><strong>Comment on a line</strong>Pick an added, removed or unchanged line (or the whole file) and write the comment. It becomes a thread in my <em>pending</em> review.</td></tr>
            <tr><td><kbd>C</kbd></td><td><strong>Show pending comments</strong>Everything I (and Claude, if it went first) have drafted so far.</td></tr>
            <tr><td><kbd>S</kbd></td><td><strong>Submit</strong>Comment, approve or request changes. Only now does anyone else see anything.</td></tr>
          </table>

          <p>Two details make this trustworthy. The line picker builds its list from git's Myers diff, so the hunks line up with the diff GitHub accepts comments on and every line number I pick is one GitHub will take. And because everything goes into a pending review, I can draft, reread and delete before a single word reaches the author. A review is a piece of writing, and I'd rather edit it before I send it.</p>
        </section>

        <!-- ─────────────── 04 ─────────────── -->
        <section class="layer" id="first-pass">
          <h2><span class="n">04 — The first pass</span>Claude first, <em class="s">never last</em></h2>

          <p>Before reading a larger PR, I can hand it to Claude first with <kbd>I</kbd>. In the background, Claude checks out the PR on its own and reviews it against that repository's checklist. It works <strong>read-only</strong>: it can read files, grep and run <code>git diff</code>, <code>git log</code> and <code>gh pr view</code>, but editing and writing are switched off. When it's done, I get a macOS notification, and its findings are already sitting in my pending review.</p>
          <p>The instructions shape what it writes. Every finding has to be anchored to a line inside the diff and written as the comment the author will read: <strong>the problem, the concrete input or state that breaks it, and what to change.</strong> Judgement calls are phrased as suggestions. If a finding can't be attached to a line GitHub accepts, it's kept as a file-level comment rather than silently dropped.</p>

          <div class="carry"><div class="kicker">Why pending, not posted</div>Claude's findings are a draft in <em>my</em> review, under my name, and nobody sees them until I submit. That keeps the accountability where it belongs. I read each one, keep what's right, rewrite what's unclear and delete what's noise. It's a first pass, never a verdict.</div>

          <p>What this buys me is focus. There are too many examples to pick one, but they fall into a clear pattern:</p>
          <div class="fit">
            <div class="yes"><div class="kicker">A first pass usually catches</div><ul>
              <li>An error branch that swallows the failure, or reports success when part of the work was skipped.</li>
              <li>A caller, type or test that wasn't updated along with the code it depends on.</li>
              <li>Tests that only cover the happy path, or would still pass with the new logic deleted.</li>
              <li>Empty, loading and error states nobody designed for.</li>
            </ul></div>
            <div class="no"><div class="kicker">Still needs me</div><ul>
              <li>Whether this is the right change for the people using it.</li>
              <li>Product intent the ticket never wrote down.</li>
              <li>Trade-offs that live outside the diff: other teams, other services, the roadmap.</li>
              <li>Whether the interface is consistent with the rest of the product.</li>
            </ul></div>
          </div>
          <p>Having the first list flagged before I open the diff leaves my attention for the second, which only a person with context can judge.</p>
        </section>

        <!-- ─────────────── 05 ─────────────── -->
        <section class="layer" id="cold">
          <h2><span class="n">05 — Complex work</span>A cold second <em class="s">opinion</em></h2>

          <p>For work the cockpit rates as complex or critical, there's one more reviewer before the PR is handed to me. A fresh agent gets the PR reference and the repo's checklist, <strong>and nothing else</strong>: no plan, no reasoning, no conversation history. It reviews what the diff <em>is</em>, not what the author <em>meant</em>. One line in its instructions has become a rule I use myself:</p>

          <p class="pull">“If something only makes sense with context you weren't given, that's a finding. A human reviewer won't have that context either.”</p>

          <p>It works through the same lenses I'd use, and it's told to try to break the change rather than confirm the happy path:</p>
          <ul class="rij-lenses">
            <li><strong>Claim vs diff</strong>Does the code deliver every Definition of Done item the PR ticks off?</li>
            <li><strong>Correctness</strong>Edge cases, error paths, empty states, races, partial failure.</li>
            <li><strong>Hidden coupling</strong>The caller the diff forgot to update.</li>
            <li><strong>Checklist</strong>This repository's known traps, one by one.</li>
            <li><strong>Test adequacy</strong>Would these tests fail without the change?</li>
            <li><strong>Security &amp; data</strong>Authorisation on new paths, checks the server must repeat.</li>
          </ul>
          <p>Before it reports anything, it has to try to refute its own findings and drop whatever it can't back with a concrete failure scenario. It only approves when it has zero confirmed findings, and it never approves anything on GitHub. My favourite line in its brief: <em>a short honest report beats invented nitpicks.</em> That's true for human reviewers too.</p>
        </section>

        <!-- ─────────────── 06 ─────────────── -->
        <section class="layer" id="look-for">
          <h2><span class="n">06 — The human part</span>What I <em class="s">look for</em></h2>

          <p>Automation takes care of a lot of the mechanical work. What's left is judgement, and that's where my experience earns its keep. These are the questions I bring to every diff:</p>
          <ol class="benefits">
            <li><div><strong>Is it right for the person using it?</strong>"Does it work" is where I start, not where I stop. Who is using this, what are they trying to achieve, and what is the interface really asking of them? An interface has to tell one story: a badge that says "failed" next to text that says "due shortly" is a bug, even when every test is green.</div></li>
            <li><div><strong>Would the tests catch a regression?</strong>A test that still passes when you delete the check it protects proves nothing. I look for failure paths, not just happy ones, and for the second interaction: open, act, close, then open again. Every bug fix should come with a test that failed before the fix.</div></li>
            <li><div><strong>Does it fail closed?</strong>An error from a check that guards a write has to stop the write. A job that half-worked must not report success. And anything the interface checks, the server has to check again.</div></li>
            <li><div><strong>Is the interface clear?</strong>Names that promise what they deliver (a <code>parse</code> function can fail, a <code>to</code> function can't), one responsibility per function, and shared components instead of one-off overrides. A comment explaining a block of code is often a sign that the block wants to be its own function.</div></li>
            <li><div><strong>What does it break for someone else?</strong>Removing or renaming anything other code depends on is a breaking change, and the PR should say so out loud.</div></li>
          </ol>
          <p>None of these are new ideas. Most of them I learned the slow way, in teams where quality came from steady discipline rather than heroics. What's new is that I have time to ask them of every change.</p>
        </section>

        <!-- ─────────────── 07 ─────────────── -->
        <section class="layer" id="loop">
          <h2><span class="n">07 — After the review</span>Closing the <em class="s">loop</em></h2>

          <p>Once a review is submitted, a loop in the cockpit takes over. It checks every open PR on a timer: CI status, the review decision, and any new human comments since it last looked. Anything actionable goes straight back to the <strong>agent session that wrote the code</strong>, with a brief that stands on its own: the PR, exactly what fired (the failing checks with trimmed logs, or each comment quoted with its file and line), and a standing instruction. Fix it, re-run the gate, self-review only the new diff, push, then reply to each thread.</p>
          <p>It works through red CI first, then requested changes, then merge conflicts, and it limits how many sessions it wakes at once so the laptop stays usable. It never merges, closes or marks a PR ready for review on its own. Those are my calls.</p>

          <h3>The checklist that learns</h3>
          <p>The last step is the one that compounds. When a human review comment applies beyond its PR, the loop adds it to that repository's review checklist. Every future task reads that checklist while it codes and again during self-review, and Claude's first pass reviews against it too. So a comment I make once becomes something the next PR is checked for automatically.</p>
          <p>There's one hard rule: <strong>only human feedback goes into the checklists.</strong> Never bot comments, never the loop's own opinions. The system can't be allowed to mark its own homework.</p>
        </section>

        <!-- ─────────────── 08 ─────────────── -->
        <section class="layer" id="tradeoffs">
          <h2><span class="n">08 — Honestly</span>The <em class="s">trade-offs</em></h2>
          <ul>
            <li><strong>Review fatigue is real.</strong> More drafts arrive than I could ever write, and attention doesn't scale the way agents do. My defences are structural: the queue is triaged by action, how much process a task gets depends on its complexity so small changes stay small, and the mechanical checks are done before I start reading.</li>
            <li><strong>Rubber-stamping is the risk I watch most.</strong> When an AI reviewer has already been through a PR, it's tempting to skim. So the AI never approves: its findings are drafts in my review, the cold reviewer can't approve on GitHub, I still read the full diff, and teammates still review. Agents open drafts; people decide what merges.</li>
            <li><strong>AI reviewers can be noisy.</strong> That's why they have to defend every finding with a concrete failure case, and why I delete freely before submitting.</li>
            <li><strong>It costs something.</strong> Every AI review has a price, and I can see it per review. I spend it where the risk justifies it, not by reflex.</li>
            <li><strong>It's my own glue.</strong> A shell script, a lazygit config, a few keybindings and skills. When it breaks, I'm the one who fixes it.</li>
          </ul>
          <p>The shift I've accepted is simple. Writing code is now often the cheap part. Deciding whether it's right, for the codebase and for the person on the other side of the screen, is the job. It deserves the same care and tooling I'd put into anything else I ship.</p>
        </section>]]></content:encoded>
      <category>AI agents</category><category>Code review</category>
    </item>
    <item>
      <title>A menu bar that knows about my meetings</title>
      <link>https://birrejan.me/blog/menu-bar/</link>
      <guid isPermaLink="true">https://birrejan.me/blog/menu-bar/</guid>
      <pubDate>Wed, 23 Sep 2026 09:00:00 +0000</pubDate>
      <description>A personal tool that touches my calendar deserves the same care as a product feature. How my menu bar keeps important meetings visible while I'm in focus mode, and keeps private titles off a shared screen.</description>
      <content:encoded><![CDATA[<p><img src="https://birrejan.me/blog/img/menu-bar.jpg" alt="Illustration: a menu bar showing a focus timer and a meeting pill, while a notification is silenced" width="1600" height="900"></p>
<!-- ─────────────── 01 ─────────────── -->
        <section class="layer" id="heads-down">
          <h2><span class="n">01 — The situation</span>Heads <em class="s">down</em></h2>

          <p>The best part of my day is the stretch where I'm deep in a diff or a design problem and nothing interrupts me. To protect it, I switch notifications off, either with macOS Focus or by starting a 25- or 50-minute focus timer in my own menu bar. Then I get on with the work.</p>
          <p>The catch is obvious once you say it out loud. A calendar reminder <em>is</em> a notification, so the one thing Focus mode reliably silences is the thing I can't afford to miss: the meeting where a client, a teammate or a product decision is waiting for me. I wanted deep focus <strong>and</strong> a guarantee that the important meeting still reaches me.</p>
          <p>The second requirement came straight after. Anything that can see my calendar and my screen is handling personal data, mine and the people I meet with, so it has to be built with the care I'd expect from any product feature that touches personal data. Privacy and security have been interests of mine for a long time, and a personal tool doesn't get a pass.</p>
          <p>Both answers ended up in the same place: the menu bar. My desktop runs on <strong>AeroSpace</strong>, a tiling window manager that keeps its own workspaces, so macOS's own bar can't show them anyway. I cover the rest of the setup in <a href="https://birrejan.me/blog/my-setup/">My setup, and how I got here</a>. The bar is <strong>SketchyBar</strong>, an empty strip where every item is a small script. In June it was SketchyBar's sample config. By September it had become a small, tested piece of software built around these two ideas.</p>
        </section>

        <!-- ─────────────── 02 ─────────────── -->
        <section class="layer" id="meetings">
          <h2><span class="n">02 — Focus</span>The meeting that <em class="s">matters</em></h2>

          <p>The trick is that the meeting pill isn't a notification at all. Focus mode filters notifications; it doesn't hide the menu bar. The <strong>MTG</strong> pill is drawn by SketchyBar from data a small calendar helper keeps up to date, so it appears whether Focus is on, off or halfway through a 50-minute session. The bar doesn't know or care what Focus is doing, and that's the point.</p>

          <ul>
            <li><strong>It shows up early.</strong> MTG appears 30 minutes before a meeting starts, with a countdown in minutes (<code>12m</code>) that becomes <code>NOW</code> at the start time.</li>
            <li><strong>It stays until it stops being useful.</strong> The pill goes away five minutes after the start or when the meeting ends, whichever comes first, so a late join is still one click away.</li>
            <li><strong>It sits next to the focus timer, not under it.</strong> Starting a focus session adds a <code>FOCUS</code> countdown to the bar and hides nothing. When a meeting comes into range, both pills are visible side by side. The timer ends with a short chime and <em>DONE</em>, and clicking <em>DONE</em> starts the break.</li>
            <li><strong>One meeting at a time.</strong> If two meetings overlap, it shows the one that starts first.</li>
          </ul>

          <table class="keys">
            <tr><td><kbd>click</kbd></td><td><strong>Join</strong>Opens the call link: Google Meet, Zoom, Teams, Webex, Whereby, Jitsi or a Slack huddle. If the event has no call link, it switches to my calendar workspace instead, so I still land on the details.</td></tr>
            <tr><td><kbd>right-click</kbd></td><td><strong>Details</strong>Shows the meeting's time, the same join action and a shortcut to record.</td></tr>
            <tr><td><kbd>mic</kbd></td><td><strong>Record with Granola</strong>A microphone next to MTG starts a new Granola note with transcription on. It's a launcher, not a recording indicator: Granola owns capture, pause and stop.</td></tr>
          </table>

          <p>The last rule matters as much as the pill itself: <strong>a wrong meeting is worse than no meeting.</strong> If the calendar helper hasn't refreshed in the last three minutes, or calendar access hasn't been granted, the pill simply isn't drawn. It won't offer a link from a cache that has gone stale. The test for it is named exactly that: <em>stale meeting never opens old link</em>.</p>
        </section>

        <!-- ─────────────── 03 ─────────────── -->
        <section class="layer" id="calendar">
          <h2><span class="n">03 — Privacy</span>A calendar is <em class="s">personal data</em></h2>

          <p>Behind the pill is a native helper, written in Swift, that reads the calendars macOS already syncs. It refreshes every minute and whenever a calendar changes. A calendar holds names, client meetings, private appointments and call links, so I held the helper to the same standard as a feature that handles someone else's data:</p>
          <ol class="benefits">
            <li><div><strong>Read-only.</strong>macOS calls the permission "Full Access", but the helper only reads events. It never creates, edits or deletes one, and the permission prompt says exactly that.</div></li>
            <li><div><strong>As little as possible, for as short as possible.</strong>It only looks at the next 24 hours and skips all-day, cancelled and declined events. For each meeting it keeps four things: start, end, title and a vetted call link. Notes and locations are scanned for a link, but never stored.</div></li>
            <li><div><strong>Owner-only storage.</strong>The cache is written atomically into an owner-only folder under <code>~/Library/Caches</code>, and no credentials live in the dotfiles.</div></li>
            <li><div><strong>Links are data, not commands.</strong>A join link must be HTTPS, on a known meeting domain, with no embedded username or password. It's checked once when the helper stores it and again before it's opened, then passed to <code>open</code> as a single argument. The controller's first line says it plainly: <em>"External names and URLs never become shell code."</em></div></li>
            <li><div><strong>Quiet logs.</strong>When something fails, the log records the type of error and nothing else: no titles, no URLs, no arguments.</div></li>
          </ol>
          <p>None of this is exotic. It's data minimisation, least privilege and input validation, the same rules I'd ask for in a code review. They just apply to a menu bar this time.</p>
        </section>

        <!-- ─────────────── 04 ─────────────── -->
        <section class="layer" id="share">
          <h2><span class="n">04 — Screen sharing</span>Safe to <em class="s">share</em></h2>

          <p>The other place a menu bar leaks is a screen share. Everyone on the call sees the app you're in, the song that's playing and the title of your next meeting. Three things keep that under control:</p>
          <ul>
            <li><strong>Presentation mode</strong> (<span class="combo"><kbd>⌥</kbd><kbd>⇧</kbd><kbd>P</kbd></span>) hides app, media and meeting titles with one shortcut. The meeting pill keeps its countdown but drops the title, and even its menu says "Upcoming meeting". An orange <strong>PRESENT</strong> pill stays visible so I can't forget it's on, clicking it turns it off, and the mode survives a reload of the bar.</li>
            <li><strong>No flash of private titles.</strong> Window-focus events are coalesced, so an event already in flight can't briefly show an app name after presentation mode has been switched on.</li>
            <li><strong>A mic/camera card</strong> appears only while one of them is actually live: amber for the microphone, red for the camera. It asks CoreAudio and CoreMediaIO the same question that drives macOS's own orange and green dots, rather than guessing from which apps are open, and it ignores output-only devices so music never lights it up.</li>
          </ul>
          <p>On the laptop screen a compact profile goes further and drops app, media and meeting titles entirely, because the space beside the camera notch is too small for them.</p>
        </section>

        <!-- ─────────────── 05 ─────────────── -->
        <section class="layer" id="quiet">
          <h2><span class="n">05 — Design rule</span>Quiet by <em class="s">default</em></h2>

          <p>A bar that shouts all day would undo the focus it's meant to protect. The rule that shapes everything is a comment at the top of the config:</p>
          <p class="pull">“Ordinary information sits directly on the bar. Only focus and alerts get a surface.”</p>
          <ul>
            <li><strong>Metric cards stay hidden until they matter.</strong> CPU only appears above 70%, in amber, and turns red at 90%. Memory also appears above 70%, but its colour follows the kernel's memory-pressure level, which is what actually predicts slowdowns. Clicking either opens <code>btop</code>. With coding agents building in the background, that's my early warning.</li>
            <li><strong>Battery only speaks when it has something to say.</strong> It hides while charging unless it's low, shows time left on battery, and shows the charge rate while charging.</li>
            <li><strong>One accent colour means "attention".</strong> Orange marks the focused workspace, the meeting pill and presentation mode, and it matches the outline JankyBorders draws around the focused window.</li>
          </ul>
          <p>Getting here took a round trip. In August I added network and disk cards; in September I took them out again, along with the list of apps in every workspace. If it doesn't help me decide something right now, it doesn't get a place.</p>
        </section>

        <!-- ─────────────── 06 ─────────────── -->
        <section class="layer" id="built">
          <h2><span class="n">06 — Under the hood</span>Built like a <em class="s">product</em></h2>

          <p>There's no framework and nothing exotic. The pieces are small, and each has one job:</p>
          <div class="term" role="group" aria-label="Project layout">
            <div class="term-bar"><div class="lights"><i></i><i></i><i></i></div><span>~/.config/sketchybar</span></div>
<pre>sketchybarrc            <span class="c"># the bar, its defaults, which items exist</span>
items/*.sh              <span class="c"># declare each item and when it updates</span>
plugins/*.sh            <span class="c"># fill items in from AeroSpace, pmset, vm_stat…</span>
plugins/toolbar.py      <span class="c"># one controller: menus, timers, meetings, privacy</span>
native/ToolbarBridge    <span class="c"># Swift: calendars, audio outputs, screens</span>
bin/capture_state       <span class="c"># Swift: is the mic or camera live?</span>
tests/test_toolbar.py   <span class="c"># privacy, meetings, launching, timers, profiles</span></pre>
          </div>
          <ul>
            <li><strong>Source, not binaries.</strong> The Swift helpers are compiled from source on first run and whenever that source changes, so the repo only carries code I can read.</li>
            <li><strong>One owner for shared state.</strong> The Python controller takes a file lock, writes its state atomically with owner-only permissions, and is the only thing that decides what's visible. A display change can't interrupt a timer or switch presentation mode off.</li>
            <li><strong>Cheap when idle.</strong> The controller only ticks every second while a timer is running; otherwise it checks in every few seconds.</li>
            <li><strong>Works on any desk.</strong> The bar measures the camera notch instead of hard-coding it, gives the cards that won't fit beside it a twin on the other side, and switches layout profiles when I dock or undock.</li>
            <li><strong>Tested like a feature.</strong> The test suite reads like a spec for the two promises in this post. A few of its names:</li>
          </ul>
          <ul class="mb-tests" aria-label="Examples of test names">
            <li>meeting window</li>
            <li>meeting excludes stale ended and old</li>
            <li>join link validation</li>
            <li>stale meeting never opens old link</li>
            <li>presentation survives refresh and hides titles</li>
            <li>granola visibility and render never start recording</li>
          </ul>
          <p>A menu bar with a test suite sounds like overkill until you remember what it touches: my calendar, my microphone and whatever is on my screen when I share it.</p>
        </section>

        <!-- ─────────────── 07 ─────────────── -->
        <section class="layer" id="changing">
          <h2><span class="n">07 — Iteration</span>Reproducible, and still <em class="s">changing</em></h2>

          <p>The bar lives in <a href="https://github.com/birrejan/mac-setup">mac-setup</a> next to the rest of my dotfiles, and GNU Stow symlinks it into <code>~/.config</code>. Editing the bar is editing the repo, and a new Mac gets it with the same single command as everything else. AeroSpace starts SketchyBar and JankyBorders at login.</p>

          <ol class="timeline">
            <li><span class="when">June</span><div><strong>The sample config.</strong> Workspaces with their app names, the date, volume, battery and CPU, always on.</div></li>
            <li><span class="when">August</span><div><strong>Awareness.</strong> Notch detection and twin cards, the mic/camera indicator, colour thresholds, and network and disk cards.</div></li>
            <li><span class="when">September</span><div><strong>Focus and privacy.</strong> The meeting pill and Granola shortcut, the focus timer, presentation mode, the read-only calendar helper, display profiles and the tests. Network, disk and the per-workspace app lists went back out.</div></li>
            <li class="now"><span class="when">Next</span><div><strong>Whatever gets in my way.</strong> That's been the only roadmap so far.</div></li>
          </ol>

          <h3>The honest trade-offs</h3>
          <ul>
            <li><strong>It's a lot of machinery for a menu bar.</strong> Shell, Python and Swift in one folder is powerful, but I'm the only maintainer.</li>
            <li><strong>It depends on the calendar being right.</strong> The pill is only as good as what's synced to macOS. A meeting that isn't on my calendar won't show up, and a stale sync hides the pill on purpose.</li>
            <li><strong>It leans on other people's behaviour.</strong> Granola's desktop link could change in a future version, and macOS permission prompts shift between releases.</li>
          </ul>

          <h3>Standing on shoulders</h3>
          <ul class="credits">
            <li><a href="https://github.com/FelixKratz/SketchyBar">SketchyBar</a><span>the bar itself</span></li>
            <li><a href="https://github.com/nikitabobko/AeroSpace">AeroSpace</a><span>tiling window manager</span></li>
            <li><a href="https://github.com/FelixKratz/JankyBorders">JankyBorders</a><span>focus outline</span></li>
            <li><a href="https://github.com/aristocratos/btop">btop</a><span>where the cards lead</span></li>
            <li><a href="https://www.granola.ai">Granola</a><span>meeting notes</span></li>
            <li><a href="https://developer.apple.com/sf-symbols/">SF Symbols</a><span>the icons</span></li>
          </ul>
          <p>If you want to borrow one idea, take this one: put the signal you can't miss somewhere that Focus mode doesn't reach, and treat everything it reads as private.</p>
        </section>]]></content:encoded>
      <category>Tooling</category><category>Privacy</category><category>macOS</category>
    </item>
    <item>
      <title>Release pipelines that make shipping boring</title>
      <link>https://birrejan.me/blog/release-pipelines/</link>
      <guid isPermaLink="true">https://birrejan.me/blog/release-pipelines/</guid>
      <pubDate>Tue, 15 Sep 2026 09:00:00 +0000</pubDate>
      <description>From building on a laptop and dragging files onto a server to a pipeline nobody has to think about: why automating releases changes how a team behaves, and how to introduce it without formal authority.</description>
      <content:encoded><![CDATA[<p><img src="https://birrejan.me/blog/img/release-pipelines.jpg" alt="Illustration: before, a folder dragged onto a server by hand; after, small changes flowing through build, test, deploy and observe" width="1600" height="900"></p>
<p class="lede">Nobody gets excited about a release pipeline. When one works well, nobody talks about it at all, and that's exactly the point.</p>
        <p>At <strong>YieldComputer</strong> I automated the entire release pipeline and defined the team's testing strategy across unit, integration and end-to-end tests. We deployed far more often, and production got noticeably calmer. The results are in the <a href="https://birrejan.me/#work">Work section of my homepage</a>. This post covers what doesn't fit in a bullet point: why automating releases changes how a team behaves, what I think a good pipeline needs, and how you introduce one when nobody has given you the authority to.</p>

        <!-- ─────────────── 01 ─────────────── -->
        <section class="layer" id="scary">
          <h2><span class="n">01 — The problem</span>Big releases are <em class="s">scary</em></h2>

          <p>My starting point will look familiar to a lot of small teams: <strong>build the app locally, then drag and drop the files onto the server.</strong> It worked, most of the time. But every release depended on whatever was on that laptop, on one person remembering the steps, and on a lot of care. Going back to the previous version meant doing it all again, by hand, under pressure.</p>
          <div class="carry"><div class="kicker">Before → after</div><strong>Before:</strong> build on a laptop, copy files to a server, hope nothing was missed. <strong>After:</strong> every change is built, tested and deployed the same way by the pipeline, and nobody touches the server by hand.</div>
          <p>Releases rarely get scary because of one bad decision. They get scary through a loop, and every step in it looks sensible from the inside:</p>

          <div class="rp-loops">
            <div class="rp-loop fear">
              <div class="kicker">The fear loop</div>
              <ol>
                <li>Releasing takes effort and attention.</li>
                <li>So the team releases less often.</li>
                <li>So changes pile up, and every release carries more of them.</li>
                <li>So each release is riskier and harder to debug.</li>
              </ol>
              <div class="again"><b>↺</b>…which makes releasing take even more effort.</div>
            </div>
            <div class="rp-loop calm">
              <div class="kicker">The confidence loop</div>
              <ol>
                <li>Releasing is automated and cheap.</li>
                <li>So changes ship small and often.</li>
                <li>So each release is easy to verify and easy to undo.</li>
                <li>So people trust releases.</li>
              </ol>
              <div class="again"><b>↺</b>…and ship smaller still.</div>
            </div>
          </div>

          <p>When a release is big and risky, extra checks, freeze windows and a senior engineer on standby are all reasonable responses. The trouble is that each of them makes releasing more expensive, so the next release gets bigger.</p>
          <p>You can't get out of that loop with courage. You get out by making releases <strong>cheap</strong>: cheap enough that a small change is worth shipping on its own, and safe enough that shipping it isn't a decision anyone has to agonise over.</p>

          <p class="pull">Most product quality comes from boring discipline, not heroics.</p>
        </section>

        <!-- ─────────────── 02 ─────────────── -->
        <section class="layer" id="behaviour">
          <h2><span class="n">02 — The effect</span>Behaviour follows the <em class="s">pipeline</em></h2>

          <p>What surprises people is that a good pipeline changes the team more than it changes the code. Three things shift once releasing stops being an event:</p>
          <ul>
            <li><strong>Changes get smaller.</strong> When a release costs almost nothing, there's no reason to bundle. A change small enough to review in one sitting is also small enough to reason about when something goes wrong.</li>
            <li><strong>Fear goes down.</strong> When rolling back is a routine command rather than an emergency meeting, "what if it breaks?" stops being a reason to wait. It becomes "how will we know if it breaks?", which is a much better question.</li>
            <li><strong>Feedback gets faster.</strong> Small releases reach users sooner, so you learn sooner whether a change did what you hoped it would.</li>
          </ul>
          <p>That last point is where the pipeline meets the product. At YieldComputer, alongside the pipeline and the testing strategy, I introduced product analytics and instrumentation, and helped move the team from pure development towards product-oriented thinking. Those things reinforce each other. It's hard to ask "did this actually help anyone?" when a change takes weeks to reach users, and easy when it reaches them the same day.</p>
          <p>Shipping often is only half the job. The other half is seeing what happened.</p>
        </section>

        <!-- ─────────────── 03 ─────────────── -->
        <section class="layer" id="anatomy">
          <h2><span class="n">03 — The anatomy</span>What a good pipeline <em class="s">contains</em></h2>

          <p>The tools change from team to team, but the shape doesn't. This is what I look for, and what I build when it isn't there:</p>

          <ol class="pipeline" aria-label="One path to production">
            <li><span class="step wait">Pull request</span></li>
            <li><span class="step">Checks</span></li>
            <li><span class="step">Merge</span></li>
            <li><span class="step">Build once</span></li>
            <li><span class="step">Deploy</span></li>
            <li><span class="step">Observe</span></li>
            <li><span class="step tier">Roll back</span></li>
          </ol>
          <div class="legend-row"><span class="wait">waits for a human review</span><span class="tier">only when needed, but always ready</span></div>

          <ol class="benefits">
            <li><div><strong>Tests at the right layers.</strong>Most checks should be fast unit tests on logic. Fewer integration tests should cover the contracts between pieces, and a small set of end-to-end tests the journeys that must never break. At finsit I learned what that last layer is worth, using Cypress and Jasmine to cut regressions in critical flows. Flip the pyramid, with slow end-to-end tests for everything, and the pipeline loses the team's trust.</div></li>
            <li><div><strong>One path to production.</strong>Every change takes the same route, urgent ones included. A hotfix is the normal pipeline with priority, not a side door. Side doors are where incidents hide.</div></li>
            <li><div><strong>Build once, promote what you tested.</strong>The thing you deploy should be the thing that passed the checks. Rebuilding separately for each environment is a quiet way to ship something nobody tested.</div></li>
            <li><div><strong>Rollback is a workflow, not a ritual.</strong>Undoing a release should be a documented, automated path that anyone on the team can run, ideally one that has been exercised before you need it. If rollback needs one specific person, you don't have rollback. You have a dependency.</div></li>
            <li><div><strong>Observe after you ship.</strong>A release isn't done when it's deployed. It's done when you've seen it behave: errors, performance, and the product events that tell you people can still do what they came to do.</div></li>
            <li><div><strong>Fast enough that nobody routes around it.</strong>A slow pipeline gets bypassed, and a bypassed pipeline protects nothing. Speed is a safety feature.</div></li>
          </ol>
        </section>

        <!-- ─────────────── 04 ─────────────── -->
        <section class="layer" id="adoption">
          <h2><span class="n">04 — The adoption</span>Leading without a <em class="s">title</em></h2>

          <p>At YieldComputer I was a technical leader without formal authority. That turns out to be a good fit for this kind of work, because a pipeline isn't something you can mandate anyway. People adopt it when it's obviously the easier way to do their job. The approach I use is simple to describe, even if it takes patience:</p>
          <ul>
            <li><strong>Find the real bottleneck.</strong> Before proposing anything, look at where the time goes and where the fear sits. It's rarely where the loudest complaint is.</li>
            <li><strong>Build the fix, don't pitch it.</strong> A working pipeline that ships a real change beats any slide about how one could work.</li>
            <li><strong>Make the new path the easy path.</strong> Use good defaults, write short docs, and pair on the first few releases. If the automated way takes more steps than the manual one, people will keep doing it by hand, and they'll be right to.</li>
            <li><strong>Keep it visible.</strong> Show status where the team already looks. Trust grows when people can see what the pipeline did and why.</li>
            <li><strong>Then find the next bottleneck.</strong> Once releasing is boring, something else becomes the slowest part. That's progress, not a failure.</li>
          </ul>
          <p>It's also how standards rise across teams without anyone announcing new standards. A colleague summed it up more kindly than I would:</p>
          <p class="pull">“He consistently improved team processes and efficiency through practical tooling and automation.”<span class="rp-cite">Shiva, Backend Developer &amp; Scrum Master at YieldComputer</span></p>
        </section>

        <!-- ─────────────── 05 ─────────────── -->
        <section class="layer" id="costs">
          <h2><span class="n">05 — The trade-offs</span>What it <em class="s">costs</em></h2>

          <p>Boring shipping isn't free, and pretending it is how pipelines get abandoned halfway:</p>
          <ul>
            <li><strong>It slows you down before it speeds you up.</strong> Building the pipeline and the test suite takes time away from features. That should be a conscious trade the team agrees to, not something done on the side.</li>
            <li><strong>Flaky tests are worse than missing ones.</strong> A test that fails at random teaches everyone to ignore red. Fix it or quarantine it the same week.</li>
            <li><strong>A pipeline is a product.</strong> It needs an owner, maintenance and the occasional refactor. When nobody owns it, it gets slow, and then it gets bypassed.</li>
            <li><strong>Automation doesn't replace judgement.</strong> A green pipeline means the checks passed, not that the change was a good idea. Review still matters.</li>
          </ul>
        </section>

        <!-- ─────────────── 06 ─────────────── -->
        <section class="layer" id="today">
          <h2><span class="n">06 — Today at InSpace</span>Same ideas, new <em class="s">team</em></h2>

          <p>At <strong>InSpace</strong> I lead the Experience Team working on NOVA, and much of my own work now runs through coding agents, each in its own isolated worktree. I wrote about <a href="https://birrejan.me/blog/my-setup/">the whole setup</a> separately. What matters here is that a room full of agents needs exactly what a team needs from a pipeline: a cheap way to check work while iterating, one check you can trust before anything goes up for review, and no duplicated effort.</p>
          <p>So the same rules apply, one level down:</p>

          <div class="rp-map" role="table" aria-label="Pipeline principles and how they apply to agents">
            <div class="row headrow" role="row"><span class="head" role="columnheader">Pipeline principle</span><span class="head" role="columnheader">How it shows up with agents</span></div>
            <div class="row" role="row"><span class="p" role="cell">Fast feedback</span><span role="cell">While iterating, a task runs a quick gate: a typecheck plus only the tests its change touches.</span></div>
            <div class="row" role="row"><span class="p" role="cell">One trustworthy check</span><span role="cell">The full gate runs once, before the pull request, not after every edit.</span></div>
            <div class="row" role="row"><span class="p" role="cell">Don't duplicate CI</span><span role="cell">Browser-based suites already run as their own CI jobs, so the local gate only re-adds them when a change touches code that can break them.</span></div>
            <div class="row" role="row"><span class="p" role="cell">Evidence over claims</span><span role="cell">Every draft PR carries its Definition of Done, test output and a self-review summary.</span></div>
            <div class="row" role="row"><span class="p" role="cell">Undo is routine</span><span role="cell">After the merge, deploying to production and rolling back are both workflows, not manual procedures.</span></div>
          </div>

          <div class="carry"><div class="kicker">What I carried forward</div>Make the right path the cheap path. It worked for a team releasing software, and it works for agents writing it.</div>

          <p>Boring is a feature. If releases are the least interesting part of your week, your pipeline is doing its job, and your team can spend its attention where it matters: on the people using what you ship.</p>
        </section>]]></content:encoded>
      <category>Delivery</category><category>Leadership</category>
    </item>
    <item>
      <title>Plans before code</title>
      <link>https://birrejan.me/blog/plans-before-code/</link>
      <guid isPermaLink="true">https://birrejan.me/blog/plans-before-code/</guid>
      <pubDate>Wed, 09 Sep 2026 09:00:00 +0000</pubDate>
      <description>Every task my coding agents pick up starts read-only, with a written Definition of Done and a plan. What goes into one, what reviewing it catches, and how the process scales from a copy fix to a risky change.</description>
      <content:encoded><![CDATA[<p><img src="https://birrejan.me/blog/img/plans-before-code.jpg" alt="Illustration: a read-only plan for extracting a shared date-range picker, with a review comment on a misread business rule" width="1600" height="900"></p>
<p class="lede">When people first hand work to a coding agent, the instinct is to type the request and let it run. I understand the appeal. But everything I've learned about shipping software says the expensive mistakes happen before the first line of code: misreading the goal, missing an edge case, touching the wrong part of the system.</p>
        <p>So in my setup, no agent writes code until it has told me, in writing, what "done" means and how it will get there. This post covers where that habit comes from, what it looks like in practice, and how I keep it from turning into bureaucracy. (I wrote about <a href="https://birrejan.me/blog/my-setup/">the whole setup</a> separately.)</p>

        <!-- ─────────────── 01 ─────────────── -->
        <section class="layer" id="correct">
          <h2><span class="n">01 — Where it comes from</span>Writing down <em class="s">"correct"</em></h2>

          <p>At finsit, part of Wolters Kluwer, I grew from junior into someone the team relied on for critical processes. The lesson that stuck from that time is that <strong>you can only hand work to someone else when tests and clear writing define what "correct" means.</strong> A good ticket, a clear acceptance criterion and a test that fails when it should are what let a new colleague change code they didn't write without breaking it.</p>
          <p>At YieldComputer I learned the other half: teams get fast when the safety lives in the process, not in someone's memory. Nobody should have to remember to check; the pipeline checks.</p>
          <p>An agent is the most extreme version of a new colleague. It's capable, it's fast, and it has none of the context that lives in the team's heads. So the question I care about isn't "can it write this?" but <strong>"can I trust what comes back?"</strong> My answer is the same as it would be for a person: agree on what done looks like, agree on the approach, and then let them work.</p>

          <div class="carry"><div class="kicker">The same idea, at work</div>NOVA, the platform my team builds at InSpace, promises its clients: "AI does the work. You make the call." Nothing ships without their yes. I hold my own agents to the same rule.</div>
        </section>

        <!-- ─────────────── 02 ─────────────── -->
        <section class="layer" id="read-only">
          <h2><span class="n">02 — The default</span>Read-only <em class="s">by default</em></h2>

          <p>Every task session starts in its own git worktree, on its own branch, in <strong>plan mode</strong>: it can read the code, search and investigate, but it can't edit a file or run a command that changes anything. It stays that way until I approve its plan.</p>
          <p>That sounds slow, but reading is cheap and wrong code is expensive. <strong>A plan is the cheapest place to be wrong</strong>, and there's a worked example of why below.</p>
          <p>Once the plan is approved, a shared settings file pre-approves the tools the workflow needs (editing files, the package manager, git, the GitHub CLI, the test commands), so the agent doesn't stop to ask for every edit. Approval is one decision at the right moment, not fifty interruptions.</p>
        </section>

        <!-- ─────────────── 03 ─────────────── -->
        <section class="layer" id="dod">
          <h2><span class="n">03 — The contract</span>The Definition <em class="s">of Done</em></h2>

          <p>The first thing a task session does is read the ticket and turn it into a <strong>Definition of Done</strong>: a checklist of observable outcomes. Not "implement the filter", but things you can check:</p>

          <div class="term" role="group" aria-label="Definition of Done example">
            <div class="term-bar"><div class="lights"><i></i><i></i><i></i></div><span>definition of done</span></div>
<pre><span class="c"># observable outcomes, not activities</span>
- [ ] user can &lt;do the thing&gt;
- [ ] &lt;component&gt; renders when &lt;condition&gt;
- [ ] no regression in &lt;the flow next door&gt;</pre>
          </div>

          <p>That list isn't paperwork. It travels through the whole task, and every later stage answers to it:</p>
          <ol class="pipeline" aria-label="How the Definition of Done travels">
            <li><span class="step">Definition of Done</span></li>
            <li><span class="step wait">Plan addresses it</span></li>
            <li><span class="step">Code satisfies it</span></li>
            <li><span class="step">Verification walks it</span></li>
            <li><span class="step">PR reports each item</span></li>
          </ol>
          <div class="legend-row"><span class="wait">waits for me</span></div>

          <p>In the pull request, each item is ticked <strong>only if it was actually verified</strong>, in a test or in the running app. Vague acceptance criteria become a clarifying question before any planning. And for a trivial change, the Definition of Done is one line.</p>
        </section>

        <!-- ─────────────── 04 ─────────────── -->
        <section class="layer" id="plan">
          <h2><span class="n">04 — The approach</span>What goes <em class="s">in a plan</em></h2>

          <p>Before planning, the agent investigates. The rule is that it should <strong>never ask what the code or the ticket already answers</strong>. Whatever is still genuinely open (scope, edge cases, design choices) goes into <strong>one batched round of questions</strong>. Then comes the plan, which is short and always has the same shape:</p>

          <ul>
            <li><strong>Files to touch</strong> and the approach, in a few lines.</li>
            <li><strong>Risks</strong>: what could break, and what the change sits next to.</li>
            <li><strong>The Definition of Done</strong>, restated, so I can see the plan actually covers it.</li>
            <li><strong>The test strategy</strong>: which new or updated tests will cover the change. New behaviour ships with tests that would fail without it. The only exemptions are changes with no testable logic (copy, styling values, config), and each one has to be stated with its reason in the PR. "This area has no tests yet" isn't an exemption; it's a reason to write the first one.</li>
          </ul>

          <p>Two details make this work in practice. The first is <strong>batching every wait</strong>. Each time an agent stops to ask me something, it can sit there for hours while I'm busy elsewhere, so it gets one round of questions and one plan approval, never a trickle of stops.</p>
          <p>The second is an <strong>honest check on effort</strong>. If the investigation shows the work is heavier than the ticket suggested, like a data-model change hiding behind a "small" ticket, the plan says so and suggests restarting at a higher tier. It's advice, not a blocker: I decide.</p>
        </section>

        <!-- ─────────────── 05 ─────────────── -->
        <section class="layer" id="example">
          <h2><span class="n">05 — A worked example</span>A plan, <em class="s">reviewed</em></h2>

          <p>Here's what a plan for a refactor looks like. It's an illustration rather than a real ticket, but it's the kind of work my team does all the time at InSpace: several reporting screens in NOVA each grew their own copy of a date-range picker, and the task is to replace them with one shared component in the design system. A standard-tier task, so the plan waits for me:</p>

          <div class="term pbc-plan" role="group" aria-label="Example plan for a refactor">
            <div class="term-bar"><div class="lights"><i></i><i></i><i></i></div><span>plan · standard tier</span></div>
<pre><span class="h">task</span>  Move the date-range picker into the design system

<span class="h">definition of done</span>
- [ ] one DateRangePicker in design-system/date-range-picker/
- [ ] every reporting screen uses it; the local copies are deleted
- [ ] presets, keyboard use and focus order behave as before
- [ ] no visual change on any screen

<span class="h">files</span>
  design-system/date-range-picker/
    <span class="c">new component, stories, tests</span>
  features/reporting/*/filters.tsx
    <span class="c">swap to the shared picker</span>
  features/reporting/lib/presets.ts
    <span class="c">delete, moved into the component</span>

<span class="h">approach</span>
  1. start from the most complete copy
  2. <span class="flag">normalise the presets into one list</span>
  3. <span class="flag">compute ranges as [today − n days, today], in the browser's timezone</span>
  4. swap one screen at a time, deleting each old copy after its swap

<span class="h">risks</span>
  - screens pass dates in two formats (Date and ISO string)
  - <span class="flag">while I'm here: align the picker with the new input styles</span>

<span class="h">tests</span>
  - unit: each preset's range, month boundaries, leap years
  - interaction: keyboard selection and focus return, in Storybook

<span class="h">open questions</span>
  <span class="flag">none</span></pre>
          </div>

          <p>It looks thorough, and most of it is good. It also makes three decisions it had no business making on its own. They're highlighted above, and these are the comments I'd leave:</p>

          <ul class="pbc-catch">
            <li><span class="kind">A business rule, read as code</span><span class="line">normalise the presets into one list · [today − n days, today]</span>The copies aren't duplicates by accident. In this example, the product defines <strong>"Last 30 days" as the 30 complete days before today</strong>, because today's numbers are still coming in, and one screen's "This month" means month-to-date while another shows the full calendar month. "Normalising" them would pass every test the agent wrote, since the tests would encode the same assumption, and quietly change numbers clients look at. Keep each definition exactly as it is. Where two copies disagree, that's a question for product, not a refactoring decision.</li>
            <li><span class="kind">An assumption about something outside the repo</span><span class="line">in the browser's timezone · open questions: none</span>How dates are interpreted, and which ranges the reporting API accepts, is decided in another service. <strong>The agent can't see it, so it shouldn't guess it.</strong> A refactor that touches an external contract and has no open questions is the tell. Send exactly the dates each screen sends today, and list the rest as questions.</li>
            <li><span class="kind">Scope creep</span><span class="line">while I'm here: align the picker with the new input styles</span>Not in this task. A refactor with a visual change in it is two changes in one diff, and the Definition of Done already says "no visual change". Separate ticket.</li>
          </ul>

          <p>After one round of comments, the parts that changed read like this:</p>

          <div class="term pbc-plan" role="group" aria-label="The revised parts of the plan">
            <div class="term-bar"><div class="lights"><i></i><i></i><i></i></div><span>plan · revised</span></div>
<pre><span class="del">- 2. normalise the presets into one list</span>
<span class="add">+ 2. keep every preset's definition exactly; list copies that disagree</span>
<span class="del">- 3. compute ranges as [today − n days, today], in the browser's timezone</span>
<span class="add">+ 3. send exactly the dates each screen sends today</span>
<span class="del">- while I'm here: align the picker with the new input styles</span>
<span class="add">+ out of scope: restyling (separate ticket)</span>
<span class="add">+ tests: each screen sends the same query before and after the swap</span>
<span class="del">- open questions: none</span>
<span class="add">+ open questions:</span>
<span class="add">+   - which timezone does the reporting API expect? (lives outside this repo)</span>
<span class="add">+   - two screens define "This month" differently: which one is right?</span></pre>
          </div>

          <p>Those are the two mistakes I look for first in any plan, and they're the same ones a capable new colleague makes: <strong>treating a business rule as an implementation detail</strong>, and <strong>filling a gap outside the repo with a plausible guess</strong>. Both look perfectly reasonable in code, and CI passes because the tests share the assumption. Reading the plan took a minute. Catching the same thing in review means reading a diff across every reporting screen. Catching it after release means explaining to clients why their numbers moved.</p>
        </section>

        <!-- ─────────────── 06 ─────────────── -->
        <section class="layer" id="tiers">
          <h2><span class="n">06 — Proportionality</span>Scaling the <em class="s">ceremony</em></h2>

          <p>There's a trap in all of this. Apply the full process to every task and a one-line copy fix takes an afternoon. My first version did exactly that: small fixes went through the same heavyweight process as real features. The fix was to make <strong>depth scale with a complexity tier</strong> that I pick when I dispatch the task:</p>

          <table class="pbc-tiers">
            <thead><tr><th>Tier</th><th>Questions</th><th>Plan</th><th>Self-review</th><th>On top</th></tr></thead>
            <tbody>
              <tr><td>Trivial<span>copy, a rename, a constant</span></td><td data-label="Questions">None; assumptions stated in the plan</td><td data-label="Plan">Three lines, approved at a glance</td><td data-label="Self-review">Reads its own diff against the checklist</td><td data-label="On top">One screenshot, if visual</td></tr>
              <tr><td>Scoped<span>a prop, a known bug</span></td><td data-label="Questions">At most one</td><td data-label="Plan">Short, approved at a glance</td><td data-label="Self-review">A medium-effort code review</td><td data-label="On top">A quick check in the running app</td></tr>
              <tr><td>Standard<span>a new page, a real feature</span></td><td data-label="Questions">One batched round</td><td data-label="Plan" class="wait">Waits for my approval</td><td data-label="Self-review">Full: code review, simplify, checklist</td><td data-label="On top">Walks the whole Definition of Done</td></tr>
              <tr><td>Complex<span>cross-cutting, data model, auth</span></td><td data-label="Questions">One batched round</td><td data-label="Plan" class="wait">Waits for my approval</td><td data-label="Self-review">Code review and checklist, thinking harder</td><td data-label="On top">Independent cold second review</td></tr>
              <tr><td>Critical<span>rare, severe if broken</span></td><td data-label="Questions">One batched round</td><td data-label="Plan" class="wait">Waits for my approval</td><td data-label="Self-review">Same, plus a small adversarial panel</td><td data-label="On top">Capped parallel agents and a cold second review</td></tr>
            </tbody>
          </table>

          <p>The tier also picks the model and how hard it thinks, from a light, fast model for trivial work to the strongest one at maximum effort for the rare critical change. When I'm torn between two tiers, I pick the higher one: under-powering a hard task costs more than a bit of extra thinking.</p>
          <p>One iteration I'm glad I made: <strong>tier up for thinking, not for fleet size</strong>. Complex work used to fan out into several parallel agents. Now only critical work does, and even that is capped. One strong session plus an independent review catches more per unit of machine than a fleet, and keeps the laptop usable.</p>
        </section>

        <!-- ─────────────── 07 ─────────────── -->
        <section class="layer" id="cold">
          <h2><span class="n">07 — Complex &amp; critical work</span>A cold <em class="s">second opinion</em></h2>

          <p>For complex and critical tasks, the draft PR gets a second reviewer before it reaches me. It's a fresh agent that is deliberately given <strong>only the PR and the repo's review checklist</strong>: no plan, no reasoning, no conversation history. It reviews what the diff <em>is</em>, not what the author <em>meant</em>. Its instructions put it better than I can:</p>
          <p class="pull">“If something only makes sense with context you weren't given, that's a finding. A human reviewer won't have that context either.”</p>
          <p>It works through a fixed set of lenses and actively tries to break the change rather than confirm the happy path:</p>
          <ul>
            <li><strong>Claim versus diff.</strong> Does the code deliver every Definition of Done item the PR ticks off? An unmet ticked claim is the most serious kind of finding.</li>
            <li><strong>Correctness.</strong> Edge cases, error paths, empty states, races and partial failures.</li>
            <li><strong>Hidden coupling and the checklist.</strong> The caller the diff forgot, and this repo's known review traps, item by item.</li>
            <li><strong>Test adequacy.</strong> Would the new tests fail without the change? Are failure paths covered, or only the happy path?</li>
            <li><strong>Security and data integrity.</strong> Authorisation on new server paths, and checks that only exist in the client.</li>
          </ul>
          <p>It has to <strong>try to refute its own findings</strong>, and anything it can't back with a concrete failure scenario gets dropped. The author fixes what's confirmed, one more cold review checks only the fixes, and if findings survive that, the task comes to me instead of looping. The reviewer is read-only: it never edits, pushes or approves, and its report is posted on the PR for the human reviewers.</p>
        </section>

        <!-- ─────────────── 08 ─────────────── -->
        <section class="layer" id="costs">
          <h2><span class="n">08 — Trade-offs</span>What it <em class="s">costs</em>, honestly</h2>

          <div class="fit">
            <div class="yes"><div class="kicker">What it buys</div><ul>
              <li>Misread rules and guessed contracts get caught in a paragraph, not in a rewrite.</li>
              <li>PRs arrive with evidence, so reviewing them is checking, not detective work.</li>
              <li>I stay in control of direction without having to stay in control of every line.</li>
              <li>Trust grows: each agent's output becomes predictable in shape.</li>
            </ul></div>
            <div class="no"><div class="kicker">What it costs</div><ul>
              <li>Waiting. A plan that sits unread blocks a task, which is why waits are batched.</li>
              <li>Attention. Rubber-stamped plans are theatre, which is why small plans stay small and the real ones get read.</li>
              <li>Judgement. A wrong tier wastes time either way, which is why I err higher and the effort check exists.</li>
              <li>It doesn't replace human review. It makes it faster.</li>
            </ul></div>
          </div>

          <p>The line from my task instructions that sums it up is this one:</p>
          <p class="pull">“The deliverable is not a PR. It's a PR a human reviewer approves without rework.”</p>
          <p>Everything before the code (the Definition of Done, the plan, the tier) exists to make that sentence true more often. It's not a new idea. It's what good teams have always done with a new colleague. The agents just made it impossible to skip.</p>
        </section>]]></content:encoded>
      <category>AI agents</category><category>Process</category>
    </item>
    <item>
      <title>From guesswork to evidence</title>
      <link>https://birrejan.me/blog/guesswork-to-evidence/</link>
      <guid isPermaLink="true">https://birrejan.me/blog/guesswork-to-evidence/</guid>
      <pubDate>Tue, 01 Sep 2026 09:00:00 +0000</pubDate>
      <description>Product analytics is a way to answer questions, not a way to collect data. How I take a team from opinions to evidence: start from decisions, keep the vocabulary small, close the loop, and design for privacy from day one.</description>
      <content:encoded><![CDATA[<p><img src="https://birrejan.me/blog/img/guesswork-to-evidence.jpg" alt="Illustration: scattered opinions converging into a single rising line, with a decision marked on it" width="1600" height="900"></p>
<p class="lede">Most product teams have strong opinions about their users and very few ways to check them. Product analytics closes that gap, but only if you treat it as a way to answer questions rather than a way to hoard data.</p>
        <p>I've introduced analytics to a team before, and the tool turned out to be the easy part. This post covers the rest: what to measure, how to name it, how to make sure someone actually looks, and how to do all of it without treating your users' privacy as an afterthought.</p>

        <!-- ─────────────── 01 ─────────────── -->
        <section class="layer" id="opinions">
          <h2><span class="n">01 — Where it started</span>Opinions are <em class="s">cheap</em></h2>

          <p>At <strong>YieldComputer</strong> I introduced product analytics and instrumentation to the team. Before that, we could see what we shipped but not how people used it, so prioritisation leaned on judgement, anecdotes and whoever made the strongest case in the room. That isn't a criticism of anyone. Without evidence, judgement is all you have.</p>
          <p>The change I cared about most wasn't a dashboard. It was the team moving from pure development to <strong>product-oriented thinking</strong>: asking who a feature is for, what it should change, and how we'll know. Analytics gave that shift something to stand on. "I think users want this" became "here's what users actually do", and prioritisation got a basis the whole team could see.</p>

          <div class="carry"><div class="kicker">What I took from it</div>Installing a tracker takes an afternoon. Getting a team to agree on what it wants to know, and to act on the answer, is the actual work.</div>
        </section>

        <!-- ─────────────── 02 ─────────────── -->
        <section class="layer" id="decisions">
          <h2><span class="n">02 — Decisions first</span>Start from the <em class="s">decision</em></h2>

          <p>The most common way to start is also the worst: install a tracker, switch on everything it can capture, and hope insight falls out. It doesn't. You end up with a large pile of events that nobody trusts and nobody reads.</p>
          <p>I work backwards instead:</p>

          <ol class="pipeline" aria-label="From decision to event">
            <li><span class="step">Decision</span></li>
            <li><span class="step">Question</span></li>
            <li><span class="step">Signal</span></li>
            <li><span class="step">Event</span></li>
            <li><span class="step">Review</span></li>
          </ol>

          <p>First, name a decision the team will actually have to make. Then the question behind it, the signal that would answer that question, and only then the event that produces the signal. The last step is agreeing when you'll look at it again. A few illustrative examples:</p>

          <table class="guesswork-ladder">
            <thead><tr><th>Decision</th><th>Question behind it</th><th>Signal that answers it</th></tr></thead>
            <tbody>
              <tr><td data-label="Decision">Invest more in onboarding?</td><td data-label="Question behind it">Where do new accounts stall before their first success?</td><td data-label="Signal that answers it">Drop-off between the steps of the first-run flow</td></tr>
              <tr><td data-label="Decision">Keep, fix or remove a feature?</td><td data-label="Question behind it">Do people who find it come back to it?</td><td data-label="Signal that answers it">Repeat use by accounts that tried it once</td></tr>
              <tr><td data-label="Decision">Is the new flow better than the old one?</td><td data-label="Question behind it">Do more people finish it, with fewer errors?</td><td data-label="Signal that answers it">Completion and error events, split by variant</td></tr>
            </tbody>
          </table>

          <p>My rule is simple: <strong>if I can't name the decision, I don't add the event.</strong> That alone keeps the tracking plan small enough to trust.</p>
        </section>

        <!-- ─────────────── 03 ─────────────── -->
        <section class="layer" id="vocabulary">
          <h2><span class="n">03 — A small vocabulary</span>Name things <em class="s">once</em></h2>

          <p>Events are an API between the product and everyone who reads the data, so I treat them like one: a small vocabulary, a consistent grammar, and changes that get reviewed like code.</p>

          <div class="term" role="group" aria-label="Example event schema">
            <div class="term-bar"><div class="lights"><i></i><i></i><i></i></div><span>tracking plan · illustrative</span></div>
<pre><span class="c"># object_action · past tense · snake_case</span>
report_exported            { format: "csv", surface: "dashboard" }
onboarding_step_completed  { step: "connect_site" }
filter_applied             { filter: "date_range", surface: "table" }

<span class="c"># properties describe the context, never the person:</span>
<span class="c"># no emails, names, free text or anything that identifies someone</span></pre>
          </div>

          <ul>
            <li><strong>Object, then action, in the past tense.</strong> Events record something that happened. <code>export_clicked</code> tells you about a button; <code>report_exported</code> tells you about an outcome.</li>
            <li><strong>Properties add context, not identity.</strong> Which surface, which format, which step. They never say who the person is.</li>
            <li><strong>One tracking plan, in the repo.</strong> It changes through pull requests, so a rename is a reviewed change rather than a silent break in every chart that depends on it.</li>
            <li><strong>Instrumentation is part of "done".</strong> If a feature exists to change behaviour, the pull request that ships it also ships the events that show whether it did.</li>
          </ul>
        </section>

        <!-- ─────────────── 04 ─────────────── -->
        <section class="layer" id="path">
          <h2><span class="n">04 — Focus</span>Instrument the path that <em class="s">matters</em></h2>

          <p>You don't need to measure the whole product on day one. You need to measure the path that decides whether someone gets value: from arriving, to the first moment it works for them, to coming back.</p>
          <p><strong>Activation</strong> is usually where the biggest and cheapest wins hide, because every new user walks through it. Everything that goes wrong there goes wrong for everyone.</p>
          <p>So I start with one path, instrumented end to end and checked against reality before anyone draws conclusions from it. <strong>One funnel you trust beats ten you argue about.</strong> Then I widen, one question at a time.</p>
        </section>

        <!-- ─────────────── 05 ─────────────── -->
        <section class="layer" id="loop">
          <h2><span class="n">05 — Rituals</span>A metric nobody looks at is <em class="s">noise</em></h2>

          <p>The failure after launch is quieter than the one before it: dashboards that exist and nobody opens. A number only matters if it's attached to a moment when someone will act on it, so I tie metrics to the rituals the team already has:</p>
          <ul>
            <li><strong>Planning.</strong> Every item meant to change behaviour says which signal it expects to move, and in which direction.</li>
            <li><strong>Retros and release reviews.</strong> We look at what moved, and just as importantly at what didn't.</li>
            <li><strong>An owner per metric.</strong> Someone who notices when it breaks, drifts or stops making sense.</li>
          </ul>

          <h3>Numbers need conversations</h3>
          <p>Quantitative data tells you <em>what</em> is happening and <em>where</em>. It rarely tells you <em>why</em>. For that I pair it with qualitative evidence: watching real sessions (with inputs masked), reading support conversations, and talking to the people who use the product.</p>
          <p class="pull">“The numbers tell me where to look. The conversations tell me what I'm looking at.”</p>
        </section>

        <!-- ─────────────── 06 ─────────────── -->
        <section class="layer" id="privacy">
          <h2><span class="n">06 — Privacy by design</span>Track behaviour, not <em class="s">people</em></h2>

          <p>Privacy and security are long-standing interests of mine. In 2024 I took Cisco's Cybersecurity Foundations, which covers privacy and data confidentiality, and my own accounts live on Proton. So when I introduce analytics, privacy isn't a legal checkbox at the end. It's part of the design:</p>

          <ol class="benefits">
            <li><div><strong>Consent first.</strong>No tracking before someone has said yes, and saying no is as easy as saying yes.</div></li>
            <li><div><strong>Collect the minimum.</strong>Every property has to earn its place against a question. If it can't, it goes.</div></li>
            <li><div><strong>No personal data in events.</strong>No emails, names or free text in properties, and identifiers that are pseudonymous rather than personal.</div></li>
            <li><div><strong>Mask by default.</strong>Session replays and heatmaps hide inputs and sensitive text unless someone deliberately decides otherwise.</div></li>
            <li><div><strong>Tooling that fits the jurisdiction.</strong>In Europe that means EU hosting, clear processing terms and sensible retention. PiwikPro, which we used at YieldComputer, is built for exactly that world.</div></li>
            <li><div><strong>Keep data only as long as the question needs it.</strong>Raw events aren't an archive. Retention should follow purpose.</div></li>
          </ol>

          <p>Done this way, privacy and good analytics pull in the same direction. Fewer, better events are easier to trust, and easier to defend.</p>
        </section>

        <!-- ─────────────── 07 ─────────────── -->
        <section class="layer" id="engineering">
          <h2><span class="n">07 — How engineering changes</span>Small, <em class="s">observable</em> steps</h2>

          <p>Once a team can see behaviour, the way it builds changes too. Features ship in smaller slices, because each slice can be checked. Risky changes go behind <strong>feature flags</strong>, so you can roll them out to a few accounts, watch, and then widen or roll back. When a question is genuinely open, you run an <strong>experiment</strong> instead of an argument.</p>
          <p>The most valuable change is the least visible one: <strong>deciding what not to build</strong>. Evidence makes it much easier to push back, gently, when a feature is solving the wrong problem, because the conversation is about what users do rather than whose opinion wins.</p>

          <div class="carry"><div class="kicker">What I do now</div>I stay close to design and product, prototype early, and ship in small, observable steps. Each step answers a question before the next one gets built.</div>
        </section>

        <!-- ─────────────── 08 ─────────────── -->
        <section class="layer" id="tradeoffs">
          <h2><span class="n">08 — The honest part</span>What can go <em class="s">wrong</em></h2>
          <ul>
            <li><strong>Vanity metrics.</strong> Page views and total sign-ups go up and to the right and tell you almost nothing. Prefer signals tied to a decision.</li>
            <li><strong>Instrumentation debt.</strong> Events rot like code: buttons get renamed, flows get removed, naming conventions get half-migrated. Without a tracking plan in the repo and a review step, your charts quietly start lying.</li>
            <li><strong>Over-reading small numbers.</strong> In B2B products especially, a handful of accounts can swing a chart. Treat small samples as a prompt for a conversation, not as a verdict.</li>
            <li><strong>Correlation dressed up as causation.</strong> People who adopt a feature are often simply your most engaged users. Experiments or careful comparisons beat eyeballing a trend.</li>
            <li><strong>Dashboards as theatre.</strong> A wall of charts can make a team feel data-driven without changing a single decision.</li>
          </ul>
        </section>

        <!-- ─────────────── 09 ─────────────── -->
        <section class="layer" id="today">
          <h2><span class="n">09 — Today</span>Evidence as a <em class="s">product</em></h2>

          <p>At <a href="https://inspace.io">InSpace</a>, the Experience Team builds <strong>NOVA</strong>, where evidence is the product itself. Clients open it to see how their organic traffic, their rankings and their citations in AI answers are moving. That makes these rules doubly relevant. Every number we put on a client's screen should answer a question they actually have, mean the same thing everywhere it appears, and lead somewhere they can act.</p>
          <p>For my own product work, these days I reach for <strong>PostHog</strong>. Having events, funnels, feature flags and session replays in one place makes the whole loop much easier to close.</p>

          <div class="recap">
            <div><div class="kicker">If you're starting tomorrow</div>Write down the next few decisions your team has to make and the question behind each one. Instrument only what answers them.</div>
            <div><div class="kicker">Then keep it honest</div>Put the tracking plan in the repo, give every metric an owner and a ritual, and keep personal data out from day one.</div>
          </div>
        </section>]]></content:encoded>
      <category>Product</category><category>Analytics</category>
    </item>
    <item>
      <title>Home Assistant, from user to contributor</title>
      <link>https://birrejan.me/blog/home-assistant/</link>
      <guid isPermaLink="true">https://birrejan.me/blog/home-assistant/</guid>
      <pubDate>Tue, 18 Aug 2026 09:00:00 +0000</pubDate>
      <description>I've run Home Assistant since 2018, improving my setup a little at a time. In 2025 I started contributing to its frontend, and reviewers I'd never met taught me about consistency, scope and evidence.</description>
      <content:encoded><![CDATA[<p><img src="https://birrejan.me/blog/img/home-assistant.jpg" alt="Illustration: three Home Assistant history charts with a single zoom selection spanning all of them" width="1600" height="900"></p>
<p class="lede">I've been running Home Assistant since 2018. For most of that time I was just a user, improving my setup a little at a time. In September 2025 I opened my first pull request against its frontend, and the review taught me more than writing the code did.</p>
        <p>This post is about that arc: from living with a product to helping build it, and what a huge open-source frontend teaches you when the reviewers are people you've never met.</p>

        <!-- ─────────────── 01 ─────────────── -->
        <section class="layer" id="user">
          <h2><span class="n">01 — 2018 · A user first</span>A product you <em class="s">live in</em></h2>

          <p>A smart home is an unusual kind of product. You're the engineer, the product manager and the main user all at once, and the feedback loop is your own daily routine. When something is off, you don't see it in a dashboard. You notice it at home.</p>
          <p>What keeps me on Home Assistant is simple: <strong>it runs locally, it's open source, and I can read how it works.</strong> Privacy and security are long-standing interests of mine, and a home that needs someone else's cloud to switch on a light has never sat right with me.</p>
          <p>I've also never tried to build the "finished" smart home in one go. Like <a href="https://birrejan.me/blog/my-setup/">my work setup</a>, it has changed in small steps over the years: make one change, live with it, and keep what earns its place.</p>

          <ol class="timeline">
            <li><span class="when">2018</span><div><strong>A user first.</strong> I started running Home Assistant, and I've been improving my setup a little at a time ever since.</div></li>
            <li><span class="when">Sep 2025</span><div><strong>First pull request.</strong> I forked the frontend and proposed synchronised zoom across the charts in the history panel.</div></li>
            <li><span class="when">Oct 2025</span><div><strong>Shipped.</strong> The feature went out in Home Assistant 2025.10 and made the release notes.</div></li>
            <li><span class="when">Autumn 2025</span><div><strong>Follow-ups and fixes.</strong> A fix for my own feature's rough edges, a bug left behind by a component migration, and a clearer energy card.</div></li>
            <li><span class="when">Dec – Jan</span><div><strong>Login and automations.</strong> Forgiving whitespace in 2FA codes, and fractions of a second in automation timings.</div></li>
            <li class="now"><span class="when">Today</span><div><strong>Still a user first.</strong> My own setup keeps changing, and the frontend is where I go when something could be better for everyone.</div></li>
          </ol>
        </section>

        <!-- ─────────────── 02 ─────────────── -->
        <section class="layer" id="first-pr">
          <h2><span class="n">02 — Sep 2025 · Getting merged</span>The first <em class="s">pull request</em></h2>

          <p>What turned me from user into contributor was the possibility itself. Home Assistant is one of the biggest open-source projects around, and its frontend is open to anyone willing to learn a corner of it. Understand one part well enough and you can improve it for a huge number of people who'll never know your name. For someone who builds interfaces for a living, that's a rare kind of leverage.</p>
          <p>My first real contribution was a feature for the history panel. When you have several charts there, <a href="https://github.com/home-assistant/frontend/pull/26898">zooming into one now zooms all of them</a> to the same time range, so you can compare what different entities were doing at the same moment.</p>
          <p>Writing it was the easy part. Getting it merged taught me five things:</p>

          <ol class="benefits">
            <li><div><strong>The process is part of the product.</strong>My first attempt didn't last a minute. The project's bot flagged commits that weren't linked to a GitHub account, so the contributor licence agreement check couldn't pass. I closed it and opened a clean one with properly linked commits. Reading the contribution guide first is part of the work.</div></li>
            <li><div><strong>Someone else might be building the same thing.</strong>Another contributor pointed out an earlier pull request that looked similar. Rather than compete, I offered to coordinate with its author. They ended up testing mine and caught a regression: the hover tooltips on the graphs had disappeared. I fixed it and posted a recording.</div></li>
            <li><div><strong>Use the system's components, at the right layer.</strong>A maintainer's review reshaped the code. Instead of dispatching zoom events from the parent, add a public zoom method on the base chart component and let each chart pass calls through to it. Use the project's own floating action button rather than a generic icon button. Respect the device's safe-area insets. Remove the props nothing reads. None of that was about whether zoom <em>worked</em>. It was about whether the code fitted a UI built by a very large community.</div></li>
            <li><div><strong>Performance on real devices is part of "done".</strong>Before approving, a maintainer asked how it behaved with lots of charts on a phone. I tested it on my phone with all of my own entities loaded. It was slow with that many charts, but usable, and another maintainer confirmed the slowness wasn't coming from the zoom sync.</div></li>
            <li><div><strong>Evidence beats argument.</strong>At one point a reviewer saw three bugs I couldn't reproduce. Instead of debating it, I posted a screen recording of everything working. Their build turned out to be the problem, and the next review was an approval.</div></li>
          </ol>

          <p>It was labelled <em>Noteworthy</em> and shipped in <a href="https://www.home-assistant.io/blog/2025/10/01/release-202510/">Home Assistant 2025.10</a>, where the release notes put it simply:</p>
          <p class="pull">“When you have multiple charts in the history panel, zooming in on one chart will now automatically zoom in on all other charts as well.”</p>
        </section>

        <!-- ─────────────── 03 ─────────────── -->
        <section class="layer" id="after">
          <h2><span class="n">03 — After the merge</span>Shipping is the <em class="s">start</em></h2>

          <p>After it merged, rough edges showed up. Graph lines were drawn over the new reset button, a second scrollbar could appear, and when several history graphs shared a panel, the wrong reset button responded. Within a week of the merge I opened <a href="https://github.com/home-assistant/frontend/pull/27133">a follow-up fix</a>, and it was merged the next day.</p>
          <p>The same thread was a lesson in scope. When I asked the contributor who'd reported the issues to test the fix, they asked for more: syncing charts on dashboards too, and moving the reset button into the toolbar. They were good ideas, but my pull request was a bug fix. Syncing dashboards would mean wiring the logic into every page, which deserved its own feature request rather than a PR too big to review and test. I checked the button on a small iPhone screen to answer the size concern. The double scrollbar turned out to be an existing issue in the virtualised list, not something I'd introduced, so I said so and left it for its own issue.</p>

          <div class="carry"><div class="kicker">What I took from it</div>Owning your regressions quickly and keeping the fix small are both ways of respecting reviewers' time. "Not in this PR" is a complete answer when it comes with a reason.</div>
        </section>

        <!-- ─────────────── 04 ─────────────── -->
        <section class="layer" id="small">
          <h2><span class="n">04 — The smaller ones</span>Small fixes, <em class="s">sharp lessons</em></h2>
          <p>After that I kept an eye out for small things. Not all of them merged, and the ones that didn't taught me just as much.</p>

          <ul class="home-assistant-prs">
            <li><span class="status merged">Merged</span><div>
              <span class="ttl"><a href="https://github.com/home-assistant/frontend/pull/27056">Sort installed add-ons by name on the logs page</a></span>
              <span class="body">Add-ons in the logs picker now appear alphabetically, after the built-in log sources.</span>
              <span class="lesson">A list in no particular order makes you read every item.</span>
            </div></li>
            <li><span class="status merged">Merged</span><div>
              <span class="ttl"><a href="https://github.com/home-assistant/frontend/pull/27134">Fix the target picker's buttons after the tooltip migration</a></span>
              <span class="body">A migration of the tooltip component had prefixed element IDs, so the target picker's remove and expand buttons received the wrong ID. I fixed it and asked a maintainer to check for the same pattern elsewhere.</span>
              <span class="lesson">In a big component library, a migration's bugs rarely come alone.</span>
            </div></li>
            <li><span class="status merged">Merged</span><div>
              <span class="ttl"><a href="https://github.com/home-assistant/frontend/pull/28086">Add total consumption to the energy usage graph card</a></span>
              <span class="body">The solar and gas cards already showed a total, but this one didn't. A maintainer explained why: this chart also shows grid return and battery charging, so a bare number would be ambiguous. We made it read as energy <em>used</em>, and kept the card's layout in line with its siblings.</span>
              <span class="lesson">A number without a label is a question.</span>
            </div></li>
            <li><span class="status merged">Merged</span><div>
              <span class="ttl"><a href="https://github.com/home-assistant/frontend/pull/28616">Trim whitespace from 2FA input before validation</a></span>
              <span class="body">A correct 2FA code with a stray space was being rejected. My first version trimmed input too broadly. Review pointed out it would affect far more than the 2FA field, so I moved it into the login functions.</span>
              <span class="lesson">Keep the fix as small as the bug.</span>
            </div></li>
            <li><span class="status merged">Merged</span><div>
              <span class="ttl"><a href="https://github.com/home-assistant/frontend/pull/29058">Accept decimal seconds in time inputs</a></span>
              <span class="body">Automation triggers and conditions can now use fractions of a second. The discussion was all about details: whether to show a separate milliseconds field, and how to pad a value like 0.2. I proposed keeping the existing formatting for consistency, and a maintainer decided to skip padding when there are decimals.</span>
              <span class="lesson">Bring options, and let the people who own the design make the call.</span>
            </div></li>
            <li><span class="status closed">Closed</span><div>
              <span class="ttl"><a href="https://github.com/home-assistant/frontend/pull/27207">A friendlier error for duplicate dashboard URLs</a></span>
              <span class="body">Another contributor had fixed the same thing in parallel while I was on holiday without my laptop, so I closed mine in favour of theirs. The review thread also found the real fix a layer down: the backend should return error codes the frontend can translate, not raw messages.</span>
              <span class="lesson">Sometimes the right fix lives one layer below yours.</span>
            </div></li>
            <li><span class="status closed">Closed</span><div>
              <span class="ttl"><a href="https://github.com/home-assistant/frontend/pull/28090">Back-button protection for unsaved changes</a></span>
              <span class="body">Review showed my approach left an extra entry in the browser history, so leaving a page would take two presses of the back button. The pull request was closed.</span>
              <span class="lesson">A closed PR can teach more than a merged one.</span>
            </div></li>
          </ul>
        </section>

        <!-- ─────────────── 05 ─────────────── -->
        <section class="layer" id="lessons">
          <h2><span class="n">05 — Lessons</span>What a big frontend <em class="s">teaches</em></h2>
          <p>Home Assistant's frontend is one of the largest open-source user interfaces I know of, maintained in public by people who don't know you. Contributing to it sharpened habits I now rely on every day:</p>
          <ol class="benefits">
            <li><div><strong>Write pull requests for strangers.</strong>Reviewers don't know you, your context or your intentions. Before-and-after screenshots, a screen recording and a clear "why" do the persuading.</div></li>
            <li><div><strong>Consistency is a thousand small agreements.</strong>Use the project's own components, keep sibling cards in sync, respect safe areas. No single rule makes a UI feel coherent; following all of them does. It's the same work I now do with the Experience Team at InSpace, building components and patterns that scale across NOVA.</div></li>
            <li><div><strong>Fix things at the right layer.</strong>A public method on the base component, error codes from the backend: the best review comments moved my fix to where it belonged.</div></li>
            <li><div><strong>Scope is a kindness.</strong>Small, reviewable changes get merged. Bigger ideas become their own issues.</div></li>
            <li><div><strong>Maintainers make the call.</strong>Bring evidence and options, then accept the decision. It's their product, and their time is the project's scarcest resource.</div></li>
            <li><div><strong>Closed isn't failed.</strong>Two of my pull requests were closed without merging, not counting the one the bot caught, and both left me knowing something I didn't before.</div></li>
          </ol>
        </section>

        <!-- ─────────────── 06 ─────────────── -->
        <section class="layer" id="local">
          <h2><span class="n">06 — Still a user</span>Local-first, <em class="s">slowly</em></h2>
          <p>Contributing hasn't changed why I use Home Assistant. It still runs locally, I still improve my setup in small steps, and I still care more about it being dependable than clever.</p>
          <p>It also connects to my day job more than I expected. At InSpace I lead the Experience Team, building the interface NOVA's clients use every day. Home Assistant shows what that looks like at the scale of a huge community: the details <em>are</em> the product, and the best interfaces are the ones nobody has to think about.</p>
          <p>If you want to see the work, <a href="https://github.com/home-assistant/frontend/pulls?q=is%3Apr+author%3Abirrejan">my pull requests are public</a>. And if you're tinkering with a smart home of your own, I'm always up for swapping notes.</p>
        </section>]]></content:encoded>
      <category>Smart home</category><category>Open source</category>
    </item>
    <item>
      <title>My setup, and how I got here</title>
      <link>https://birrejan.me/blog/my-setup/</link>
      <guid isPermaLink="true">https://birrejan.me/blog/my-setup/</guid>
      <pubDate>Tue, 04 Aug 2026 09:00:00 +0000</pubDate>
      <description>From agency work in Barcelona to leading the Experience Team at InSpace: how each job shaped the machine, desktop, terminal and coding-agent cockpit I work in today, and why it keeps changing.</description>
      <content:encoded><![CDATA[<p><img src="https://birrejan.me/blog/img/my-setup.jpg" alt="Illustration: four stacked layers labelled Machine, Desktop, Terminal and Agents, with agent sessions marked working, waiting and idle on the top layer" width="1600" height="900"></p>
<p class="lede">Every setup is a record of the problems its owner got tired of. Mine is no different. The window manager, the scripts and the agents are each there because, at some point in some job, I lost time, quality or focus to the thing they now handle.</p>
        <p>So instead of opening with a list of tools, I'll start with the story. The tour comes after, for anyone who wants to borrow something.</p>

        <!-- ─────────────── 01 ─────────────── -->
        <section class="layer" id="barcelona">
          <h2><span class="n">01 — 2019 · Barcelona</span>Own the whole <em class="s">path</em></h2>

          <p>My first job in software came right after a web development bootcamp at Ironhack. I joined <strong>Nakima</strong> in Barcelona as a fullstack developer, and it was agency work at its busiest: web apps in React and Angular, mobile apps in React Native, backends on Firebase and MongoDB, and one project that ran facial recognition on the device itself.</p>
          <p>What stayed with me had little to do with frameworks. I owned client requirements from start to finish (scoping, delivery, the next iteration), and I was often the person translating between designers, the backend team and the client. Two lessons came out of that. First, the job isn't writing code, it's <strong>shortening the distance between someone's problem and a working solution</strong>. Second, <strong>switching context is expensive</strong>. Every time you jump between projects you pay a re-entry cost, and nobody budgets for it.</p>

          <div class="carry"><div class="kicker">What I carried forward</div>Every task in my setup gets its own isolated space: its own branch, worktree, terminal tab and agent session. Switching between them costs a keystroke, not a morning.</div>
        </section>

        <!-- ─────────────── 02 ─────────────── -->
        <section class="layer" id="remote">
          <h2><span class="n">02 — 2021 · Remote</span>Trust comes from <em class="s">tests</em></h2>

          <p>In 2021 I joined <strong>finsit</strong>, part of Wolters Kluwer, working remotely. I built frontend modules on top of internal libraries, contributed to C# services, and spent a lot of my time making critical flows dependable with Cypress and Jasmine. I also modernised shared components and tooling, and that sped up the whole frontend team more than any single feature I shipped.</p>
          <p>That's where I grew from junior into someone the team relied on for critical processes. It's also where I learned that <strong>you can only hand work to someone else when tests and clear writing define what "correct" means</strong>. That someone might be a teammate, a new hire or, as it turned out, an agent. The other lesson was that shared tooling is some of the highest-leverage code you can write.</p>

          <div class="carry"><div class="kicker">What I carried forward</div>Nothing in my setup ships on an agent's word. Every task writes down its Definition of Done, adds tests, passes a gate and checks itself against a review checklist before I ever see it.</div>
        </section>

        <!-- ─────────────── 03 ─────────────── -->
        <section class="layer" id="eindhoven">
          <h2><span class="n">03 — 2024 · Eindhoven</span>Pipelines over <em class="s">heroics</em></h2>

          <p>At <strong>YieldComputer</strong>, from 2024 to mid-2026, I automated the release pipeline, defined a testing strategy across unit, integration and end-to-end tests, and introduced product analytics so we could prioritise from evidence instead of opinion. The results are on <a href="https://birrejan.me/#work">my homepage</a>, but the lesson is simpler: <strong>teams don't get fast by working harder</strong>. They get fast when shipping is safe and boring, and when that safety lives in the pipeline instead of in someone's memory.</p>
          <p>It's also where I learned to lead without formal authority. You find the bottleneck, build something that removes it, and make it easy for people to adopt. Then you look for the next bottleneck.</p>

          <div class="carry"><div class="kicker">What I carried forward</div>My agent workflow is a release pipeline for tickets: dispatch, plan, implement, gate, review, evidence, draft PR. I run it the way I learned to lead: find the bottleneck, remove it, repeat.</div>
        </section>

        <!-- ─────────────── 04 ─────────────── -->
        <section class="layer" id="inspace">
          <h2><span class="n">04 — 2026 · InSpace</span>Make it <em class="s">effortless</em></h2>

          <p>On 1 July I left YieldComputer and joined <a href="https://inspace.io">InSpace</a>. InSpace builds <strong>NOVA</strong>, an autonomous AI platform for SEO and GEO (generative engine optimisation). It researches, writes, optimises and publishes the content that gets brands found, both on Google and in the answers of ChatGPT, Gemini, Perplexity and Claude.</p>
          <p>As a senior engineer, I lead the <strong>Experience Team</strong>. We build the platform our clients see and use every day, from their dashboards to the moment they approve a page. Interfaces that feel effortless take precision and taste. Much of my work is building <strong>components and patterns that scale across NOVA</strong>, keeping the whole product consistent as it grows.</p>
          <p>I also try never to stop at "does it work?" I want to know who's using it, what they're trying to achieve, and what the interface is really asking of them. That's why I'm as comfortable in a product conversation as in a pull request.</p>
          <p>NOVA's promise to clients is <strong>"AI does the work. You make the call."</strong> Nothing ships without the client's yes, and every correction teaches NOVA, so it needs less checking next time. That turns out to be a pretty good description of how I build it, too.</p>

          <div class="carry"><div class="kicker">Where it all comes together</div>Owning the whole path, trusting tests and building pipelines all show up here: shared components with interaction tests, patterns that scale across a large product, and a workflow that ships a lot of interface without lowering the bar.</div>
        </section>

        <!-- ─────────────── 05 ─────────────── -->
        <section class="layer" id="agents">
          <h2><span class="n">05 — 2026 · The workflow</span>Enter the <em class="s">agents</em></h2>

          <p>Claude had been my daily co-pilot for a while, for pairing on architecture, drafting and refactoring. A new team, a new codebase and a product that is itself built on AI changed the question from <em>"can it help me write this?"</em> to <em>"can it take this ticket, and can I trust what comes back?"</em> Answering it took everything from the chapters above, and it didn't happen in one go:</p>

          <ol class="timeline">
            <li><span class="when">June</span><div><strong>Make the base reproducible.</strong> Before my first day at InSpace, I turned the machine itself into code: one command rebuilds my Mac.</div></li>
            <li><span class="when">July</span><div><strong>One task, one worktree.</strong> In my first weeks on the team: first a script that gives every task its own worktree and dev server. Then a cockpit that sends each task to its own tab and agent session, and a monitor that shows which ones are waiting on me.</div></li>
            <li><span class="when">August</span><div><strong>Past the pull request.</strong> The bottleneck moved to CI and review. A loop now watches every PR and sends failures and comments back to the session that wrote the code. Complex work gets a cold second review, and real review comments became checklists that every future task reads.</div></li>
            <li><span class="when">September</span><div><strong>Respect the machine.</strong> With many sessions on one laptop, builds started fighting over the cores, so expensive commands now queue for a slot. Small fixes were going through the heavyweight process, so how deep a task goes now depends on how complex it is. The menu bar learned about meetings.</div></li>
            <li class="now"><span class="when">This week</span><div><strong>Review gets its own tools.</strong> A PR opens as a staged diff in lazygit, Claude's first pass lands in my pending review, and one key jumps from a PR to the session that wrote it.</div></li>
          </ol>

          <p>None of this was planned up front. <strong>Each step dealt with the bottleneck the previous one exposed.</strong> That's the same loop I used to run with teams, only now some of the team is software.</p>
        </section>

        <!-- ─────────────── 05 ─────────────── -->
        <section class="layer" id="today">
          <h2><span class="n">06 — The tour</span>The setup <em class="s">today</em></h2>
          <p>Here's where it stands this month, from the bottom up.</p>

          <h3>The machine</h3>
          <p>A MacBook Pro, usually docked to an ultrawide monitor. Everything runs on that one laptop, including every agent session, and it's configured entirely from <a href="https://github.com/birrejan/mac-setup">mac-setup</a>, a public repo:</p>

          <div class="term" role="group" aria-label="Terminal">
            <div class="term-bar"><div class="lights"><i></i><i></i><i></i></div><span>fresh mac</span></div>
<pre><span class="p">$</span> bash &lt;(curl -fsSL https://raw.githubusercontent.com/birrejan/mac-setup/main/bootstrap.sh)
<span class="c"># preflight → homebrew → dotfiles → languages → git → ssh</span>
<span class="c"># → vscode → iterm2 → wm → macos → touchid → doctor</span>
<span class="ok">✓</span> every step is idempotent, and re-runnable alone with --only &lt;step&gt;</pre>
          </div>

          <p>A <strong>Brewfile</strong> lists every tool, app and font. <strong>GNU Stow</strong> symlinks the dotfiles, so editing a config means editing the repo. <code>doctor</code> checks the result, and <code>dump.sh</code> syncs back anything that isn't a symlink. It also handles the dull but important bits: SSH-signed commits, Touch ID for <code>sudo</code> and sensible macOS defaults.</p>

          <h3>The desktop</h3>
          <p><strong>AeroSpace</strong> tiles every window, so nothing overlaps and I move around with <span class="combo"><kbd>⌥</kbd> <kbd>h</kbd><kbd>j</kbd><kbd>k</kbd><kbd>l</kbd></span>. Some workspaces have a fixed job (chat, calendar, email), apps are routed to them automatically, and switching to one that's empty opens its apps. <strong>JankyBorders</strong> outlines the focused window, since tiled windows have no title bar to tell you where your keystrokes will go.</p>
          <p>The menu bar is <strong>SketchyBar</strong>, extended with a small native Swift helper:</p>
          <ul>
            <li><strong>Machine load at a glance.</strong> CPU and memory cards stay out of the way until the machine is busy, then turn amber, then red. With agents building in the background, this is my early warning.</li>
            <li><strong>Meetings that don't sneak up.</strong> A pill appears half an hour before a call, joins it in one click and can start a Granola recording. It reads my calendars and never writes to them.</li>
            <li><strong>Screen-share safe.</strong> Presentation mode hides app, media and meeting titles. A mic/camera card lights up only while either is live.</li>
            <li><strong>Display-aware.</strong> It works around the laptop's notch and switches layout profiles when I dock or undock.</li>
          </ul>

          <h3>The terminal</h3>
          <p><strong>iTerm2</strong>, <strong>zsh</strong> with no framework, <strong>Starship</strong> for the prompt, <strong>mise</strong> and <strong>uv</strong> for runtimes, and modern replacements for the classics: <code>eza</code>, <code>bat</code>, <code>ripgrep</code>, <code>fd</code>, <code>zoxide</code>, <code>fzf</code>. My pull-request queue lives in <strong>gh-dash</strong>, grouped by what needs doing (CI failing, changes requested, ready to merge, needs my review), with a few keybindings I added:</p>

          <table class="keys">
            <tr><td><kbd>i</kbd></td><td><strong>Review in lazygit</strong>The PR is checked out into a throwaway worktree and shown as one staged diff. I leave line comments in a pending GitHub review without touching my own checkout.</td></tr>
            <tr><td><kbd>I</kbd></td><td><strong>Claude reviews first</strong>A fresh-context agent reviews the PR and puts its findings into <em>my</em> pending review. Nobody sees anything until I've edited it and submitted.</td></tr>
            <tr><td><kbd>T</kbd></td><td><strong>Jump to the task</strong>Focus the agent session that wrote this PR.</td></tr>
          </table>

          <h3>The agents</h3>
          <p><code>~/Projects</code> is a <strong>cockpit</strong> for NOVA's codebases: the app, its design system and the services around it. A Claude Code session started there doesn't write feature code: it starts tasks, tracks them and gets them merged. I describe the work, and the cockpit builds an isolated worktree from the latest base branch, registers the task and opens a new tab with a pre-briefed session. Every task starts <strong>read-only, in plan mode</strong>, and works through the same pipeline:</p>

          <ol class="pipeline" aria-label="Task lifecycle">
            <li><span class="step">Ticket</span></li>
            <li><span class="step">Definition of Done</span></li>
            <li><span class="step wait">Clarify</span></li>
            <li><span class="step wait">Plan</span></li>
            <li><span class="step">Implement + tests</span></li>
            <li><span class="step">Gate</span></li>
            <li><span class="step">Self-review</span></li>
            <li><span class="step">Verify in the app</span></li>
            <li><span class="step">Draft PR</span></li>
            <li><span class="step tier">Second review</span></li>
            <li><span class="step wait">Ping me</span></li>
          </ol>
          <div class="legend-row"><span class="wait">waits for me</span><span class="tier">complex &amp; critical work only</span></div>

          <p>How deep a task goes depends on a <strong>complexity tier</strong>. A typo fix gets a plan of a few lines that I can approve at a glance. Standard work gets a fuller plan and a full self-review. Complex work also gets an <strong>independent second review</strong>: a fresh agent sees only the PR and the checklist, never the author's reasoning, and reviews it cold.</p>
          <p>Once a PR is open, a loop in the cockpit watches CI and review comments and sends anything actionable back to the session that wrote the code. When a human review comment applies more widely, it's added to that repo's <strong>review checklist</strong>, which every future task reads while coding and again at self-review. For example:</p>
          <p class="pull">“Fail-open guards: never discard the error from a query used as a guard. A read error must fail the action, not let the write proceed.”</p>
          <p>A lesson written there is a review comment no future PR needs. Around all of this, <strong>clorch</strong> shows which sessions are working, idle or waiting on me. A shared slot pool makes expensive builds and test runs <strong>queue instead of competing</strong>, so the laptop stays responsive while many tasks are in flight.</p>
        </section>

        <!-- ─────────────── 06 ─────────────── -->
        <section class="layer" id="guardrails">
          <h2><span class="n">07 — Professional defaults</span>Guardrails I <em class="s">don't bend</em></h2>
          <p>Speed only counts if you can stand behind what ships. These rules hold for every task, whatever its size:</p>
          <ol class="benefits">
            <li><div><strong>Plans before code.</strong>Every task starts read-only, in plan mode, and nothing changes until I've approved its plan. For a small fix that plan is a few lines, so approving it takes a glance.</div></li>
            <li><div><strong>Isolation by default.</strong>Each task has its own branch, worktree and environment. Nothing touches my main checkout, and nothing touches another task.</div></li>
            <li><div><strong>Evidence, not claims.</strong>A PR carries its Definition of Done, test output, a self-review summary and proof from the running app. "Done" is something you can check.</div></li>
            <li><div><strong>A human reviews every diff.</strong>Agents open drafts. I review, my teammates review, and people decide what merges.</div></li>
            <li><div><strong>Lessons get written down.</strong>Review feedback becomes a checklist, so the same comment never has to be made twice.</div></li>
            <li><div><strong>The machine is a shared resource.</strong>Expensive work queues for a slot rather than grinding the laptop for everyone, me included.</div></li>
          </ol>
        </section>

        <!-- ─────────────── 07 ─────────────── -->
        <section class="layer" id="who">
          <h2><span class="n">08 — People &amp; fit</span>Who it's <em class="s">for</em></h2>

          <p>First, who does what. The setup only works because the split is explicit:</p>
          <div class="cards">
            <div><h3 class="card-title"><em>Me</em></h3>Decide what's worth building, shape the ticket, approve plans, review every diff, merge.</div>
            <div><h3 class="card-title">The <em>agents</em></h3>Turn tickets into tested, verified draft PRs, fix red CI and respond to review comments.</div>
            <div><h3 class="card-title">My <em>team</em></h3>The Experience Team and the wider engineering team review and approve, like always. Agents open drafts; people decide what ships.</div>
          </div>

          <p>And who should copy it:</p>
          <div class="fit">
            <div class="yes"><div class="kicker">A good fit</div><ul>
              <li>Product engineers with a steady stream of well-scoped features and bugs.</li>
              <li>People who already review more code than they write, and are fine with that.</li>
              <li>Anyone who likes owning their tools. It's all plain scripts, JSON and markdown you can read and change.</li>
            </ul></div>
            <div class="no"><div class="kicker">Not a fit</div><ul>
              <li>Exploratory work where you don't know what you want yet. Agents amplify clarity; they don't create it.</li>
              <li>Anyone who won't read the diffs. The whole system assumes a human reviews everything.</li>
              <li>Teams without CI or tests. The gates are what make handing work off safe.</li>
            </ul></div>
          </div>
        </section>

        <!-- ─────────────── 08 ─────────────── -->
        <section class="layer" id="benefits">
          <h2><span class="n">09 — Benefits &amp; trade-offs</span>What it <em class="s">buys</em> me</h2>
          <ul>
            <li><strong>More in flight, less in my head.</strong> Several tasks move in parallel, and each one keeps its own context, so picking one up again is <kbd>T</kbd>, not ten minutes of rereading.</li>
            <li><strong>Time back for the product work.</strong> More of my day goes to deciding, shaping tickets, reviewing and talking to people: the parts of the job I think matter most.</li>
            <li><strong>Quality that compounds.</strong> Checklists and cold reviews mean the same mistake isn't made twice, by me or by an agent.</li>
            <li><strong>Reproducible by default.</strong> One command rebuilds the Mac, and every config change lands in a repo as I make it.</li>
            <li><strong>Focus.</strong> The desktop gets out of the way. Load, meetings and the PR queue are visible without going looking for them.</li>
          </ul>

          <h3>The honest trade-offs</h3>
          <ul>
            <li><strong>The bottleneck moves to review.</strong> Reviewing well is most of the job now, which is why review has its own tooling.</li>
            <li><strong>Agents are only as good as the ticket.</strong> Vague in, vague out. That's why every task starts by writing a Definition of Done.</li>
            <li><strong>It's my own glue.</strong> Shell scripts, skills and config. That's powerful, but when something breaks, I'm the maintainer.</li>
            <li><strong>Compute and tokens aren't free.</strong> That's why depth scales with complexity, and heavy multi-agent work is kept for the few tasks that need it.</li>
          </ul>
        </section>

        <!-- ─────────────── 09 ─────────────── -->
        <section class="layer" id="iterating">
          <h2><span class="n">10 — Always iterating</span>Never <em class="s">finished</em></h2>
          <p>This setup is never done, and I've stopped wanting it to be. It changes on two rhythms:</p>
          <div class="recap">
            <div><div class="kicker">Every week</div>I fix whatever got in my way: a keybinding, a script, a smarter default. Review feedback flows into the checklists on its own, so the system also gets better while I'm not looking.</div>
            <div><div class="kicker">Every month</div>I step back, look at where the time actually went, and change the structure: a new stage, a new guardrail, a new repo onboarded, or something removed because it stopped earning its place.</div>
          </div>
          <p>What makes that sustainable is the same discipline I'd expect from any production system. Everything lives in versioned files. Decisions sit next to the config they explain. And a change doesn't count until it's written down somewhere the next task, or the next version of me, will read it.</p>
          <p>Next month, parts of this post will be out of date. That's the point.</p>

          <h3>Standing on shoulders</h3>
          <ul class="credits">
            <li><a href="https://github.com/nikitabobko/AeroSpace">AeroSpace</a><span>tiling window manager</span></li>
            <li><a href="https://github.com/FelixKratz/SketchyBar">SketchyBar</a><span>the menu bar</span></li>
            <li><a href="https://github.com/FelixKratz/JankyBorders">JankyBorders</a><span>focus outline</span></li>
            <li><a href="https://github.com/dlvhdr/gh-dash">gh-dash</a><span>PR dashboard</span></li>
            <li><a href="https://github.com/jesseduffield/lazygit">lazygit</a><span>git &amp; review UI</span></li>
            <li><a href="https://github.com/androsovm/clorch">clorch</a><span>agent session monitor</span></li>
            <li><a href="https://claude.com/claude-code">Claude Code</a><span>the agents</span></li>
            <li><b class="pair"><a href="https://starship.rs">Starship</a> &amp; <a href="https://mise.jdx.dev">mise</a></b><span>prompt &amp; runtimes</span></li>
          </ul>
          <p>If you want to take something, start with <a href="https://github.com/birrejan/mac-setup">mac-setup</a>. It's public and it's meant to be forked. If you're building your own agent workflow, I'd love to compare notes.</p>
        </section>]]></content:encoded>
      <category>Setup</category><category>Career</category><category>AI agents</category>
    </item>
  </channel>
</rss>
