Running agents is an operations job
One morning in September I sat down to twelve agent sessions, and every one of them had the same name.
Each session starts in the same home repository and fans out from there - one into a leagues club’s point-of-sale mapping, one into research for a civil engineering firm’s drawing reviews, one fixing a deploy script, one drafting a proposal. The workspace manager labelled them all after the folder they started in. Twelve tabs, one word, repeated. I genuinely could not tell which session was doing what, or where any of them was up to.
It was a small thing. It also sums up most of what the last three months have taught me.
On a normal day now I have somewhere between eleven and twenty agent sessions open at once, each in its own workspace, and I click between them rather than watching them all. Overnight a separate build factory works through issues and turns them into pull requests. Nothing merges and nothing deploys without me.
I expected the hard part to be the prompting. Getting the brief right, getting the model to understand the work. That turned out to be the easy part, and it keeps getting easier.
The hard part is operations. Visibility, data handling, isolation, capacity, restarts, change management. The same discipline I spent twenty years applying to servers and networks, now applied to a fleet of things that write code. My background is systems, not software, and I’ve never been more glad of it.
From one terminal to a fleet
For the first part of this year, working with AI meant one Claude Code session in one terminal on my laptop. One conversation, one task, and when the lid closed, the work stopped.
In March I moved the sessions onto a small cloud server and ran them inside tmux, the terminal multiplexer sysadmins have leaned on for decades. Sessions survived the laptop closing, and I could check on one from my phone. That worked for three or four sessions. Past that, tmux has no idea what’s running inside each window. It’s a grid of terminals, and every one looks the same until you go in and read it.
In late July I moved to herdr, a terminal workspace manager built specifically for AI coding agents. Each agent gets its own workspace, and the sidebar shows every agent’s state - working, idle, or waiting on me. I attach to it from my Mac and the sessions keep running when I disconnect. Since the 0.9 release it records which conversation lives in which pane, so after a restart the sessions resume themselves.
The part that changed how I work is that herdr is built for agents to drive as much as for me to watch. Everything I can do in the interface, an agent can do from the command line: create a workspace, start a session in it, send it a prompt, wait for a particular line of output, read what’s on another pane’s screen.
So one session can start others. I had a Claude Code session audit the versions running across every client server. It wrote up the findings and a shared brief - which session owns which issues, work in your own git worktree, test box first, merges and deploys are mine - and then spun up four more sessions, one per client and one for the shared LLM gateway, each told to read the brief and start on its own list. Roughly:
herdr workspace create --cwd ~/git/tessellant --label "client-a:fleet-wave" --no-focus
herdr pane run <pane-id> "claude 'Read the fleet-waves brief, then work the client-a issues'"
Three clients and a shared service, upgraded side by side, each in a labelled workspace I can click into, none of them treading on the others. In tmux that would have been me opening windows and pasting briefs by hand. Here it’s one instruction to one agent.
The agents multiplied alongside it. Claude Code is the daily driver. I use pi sometimes, and Claude Code on the web or phone for when I’m away from a keyboard. The bigger change was going from using an agent to running several, and running several is a different job.
A box of its own
Also in late July, everything moved onto a dedicated machine: an ex-government Dell small-form-factor desktop off eBay, about $500 all up, sitting in the garage rack. It runs my home automation as well, so it earns its keep twice.
Moving off the laptop and off rented servers changed more than I expected.
It runs any time, from anywhere. The box is always on, so nothing depends on my laptop being open. Scheduled work runs as ordinary system timers - bank feeds pull overnight, idle headless browsers get reaped every few minutes, the build factory works its queue while I sleep. I reach all of it from the Mac, the iPad or my phone, and a push notification only arrives when something genuinely needs me.
It’s contained. The disk is encrypted and unlocks from the machine’s own security chip at boot. The box isn’t on the public internet at all; it’s only reachable over a private Tailscale network. The build factory runs as its own unprivileged user with its own session server, so its runs can’t touch my sessions and mine can’t touch its. Containers run rootless under Podman.
Over the top of every agent sits a permission layer. Merging to main, deploying, writing to a client’s systems, editing its own settings - all blocked. When one of those is needed, the agent does the preparation and hands me a single short command to run myself. I drew that line properly after an agent booked a lunch on the wrong day, twice, and I learnt to insist on short after a long one-liner wrapped on paste and ran as several separate commands.
It’s a test bench. An agent can stand up a complete copy of a client’s platform in Podman - database, services, the lot - test against it, and tear it down. A disposable copy of the real thing is the single most useful thing you can give an agent. It also leaves a mess. In early September the disk hit 96% and the factory, correctly, refused to start new runs. The culprit was 481 leftover container volumes, ten of them in use. One prune freed 24 GB. Nothing cleans those up unless you tell it to.
Capacity is the constraint, so plan for the restart
The box started with a single 16 GB stick of RAM. Plenty for a home server, I thought.
By September it was carrying twenty agent sessions, a handful of test stacks and the home automation VM. On 8 September it ran out of memory and hard reset. The biggest single culprit was a browser-automation tool I’d configured globally. It started its own helper stack of around 270 MB for every session, whether or not that session ever opened a browser. Seventeen sessions, about 4.5 GB, before anything had happened. Nothing on the box capped agent memory.
The fixes went in the same day: swap from 4 GB to 16 GB, a memory policy that kills the greedy process instead of the whole server, and the browser tool scoped to the one project that needs it. Three days later a second stick went in. 32 GB, dual channel, and with eleven agents running there’s now 22 GB free where there used to be seven.
A footnote for anyone buying RAM for an office desktop: the new stick was an enthusiast kit that only reaches its rated speed through an XMP profile, and the Dell BIOS doesn’t do XMP. Both sticks dropped to the lowest common speed. Still a net win, but buy plain JEDEC-rated memory for business-class machines.
The crash cost more than the reboot. When the box came back, the workspace manager restored the layout as bare shells, and closing them overwrote the saved layout with a single empty workspace. Twice I resumed around eleven sessions by hand. The obvious command - “continue the last session in this folder” - is ambiguous when five sessions share one folder. Resuming by exact session id is what works. Before the RAM upgrade I captured the map of which pane held which session, and all eleven came back exactly where they were.
Then the root fix, found embarrassingly late: the workspace manager had native session restore all along. The per-agent integrations had never been installed. Installed now, and sessions come back on their own.
The principle is state over memory. An agent session is a running service with state. If that state isn’t written down somewhere durable, the next reboot decides what you keep.
If you can’t see it, you aren’t supervising it
Back to the twelve identical names. Supervision only works if you can tell, at a glance, which session needs you and which one is fine to leave. Without that, you’re clicking through tabs reading scrollback, and the session that actually needed a decision waits twenty minutes while you check on one that didn’t.
The tool had no way to label a workspace from what the agent was doing, so I built one. A small skill reads each session’s transcript and renames its workspace to the client and the task - client work gets a client prefix and a short description, internal work gets a t: prefix, personal work gets none. Sessions waiting on input turn bold red in the sidebar. Clients get a colour tint.
Even the titles agents give themselves drift. One session was still titled after the server error it started on, hours after it had moved on to mapping a client’s sales data. Anything that labels a session once and never again is going to lie to you by lunchtime.
The factory had the same problem. A build run once sat in what looked like a blank pane for twenty minutes, because each stage only printed its output when it finished. I assumed it had hung. It hadn’t. The fix was a live journal viewer next to every run, so there’s always a visible heartbeat.
You can’t operate what you can’t observe. Make sure each agent tells you what it’s doing now, and that the one waiting on you is the loudest thing on the screen.
Transcripts are a data store, so treat them like one
Every agent session writes a transcript to disk - plain text, every message, every tool output. Those transcripts turn out to be one of the most-read files on the box, because agents mine them constantly. “Check the previous sessions for how we set this up” is one of the most useful instructions I give.
In September I had an agent sweep about a hundred recent transcripts. It found API keys and a VPN private key that I had pasted into chats on two separate days. They were sitting in plain text, readable by any later agent searching old sessions for context. The same day, an agent reading another agent’s live screen to work out what it was doing pulled a service password into its own context.
The model did nothing wrong in either case. What failed was data classification. I had been treating transcripts as conversation history. They’re a data store holding whatever passed through the session, with no classification, no retention policy and a very curious set of readers.
What changed: anything that reads another session now takes only the message text - no tool output - and redacts anything shaped like a secret before it goes anywhere. The exposed keys were rotated. And the habit changed: secrets go in the secret store and the agent gets pointed at it, not handed the value in chat.
If you run agents, go and look at where your transcripts live, and search them for anything that looks like a key. You will probably find something.
Blast radius is set by the environment, not the instructions
In mid-September, agents working on a clean-up task ran a client project’s test suite on the shared box. Three times. What nobody noticed was that a local preview database happened to be running, and a dozen test modules quietly fell back to it when no test database was configured. The tests overwrote and deleted rows.
The local copy only - production was never touched. But it looked like two flaky tests rather than data loss, which is the worst way for a problem to present.
The agents’ instructions were fine. Nobody told them to touch that database. Nothing stopped them either. The suite connected to whatever was up, and something was up.
So the fix went into the environment: the test setup now refuses that database outright, a guard test fails if it ever tries, and every agent gets told explicitly until that change is everywhere. Instructions describe what you intend. The environment decides what’s possible.
The same goes for what “success” means. More than once a deploy has reported a clean rollback while leaving new code running. Green is a status, not a fact. Verify the thing, not the report about the thing.
Playbooks rot like code
My agents run on playbooks - skills and instruction files that tell them how to dispatch a build, check a box, write in my voice. I treated these as documentation. They behave like code, and they break like code.
Upgrading the workspace manager from 0.7 to 0.9 silently changed a command two factory skills depended on. Nothing errored. It would have failed the first time a build was dispatched, and I only caught it by reading the changelog against every place the skills used that command. For weeks, a content skill pointed at a voice file that didn’t exist, so the voice layer silently failed to load on every run. The output was generic, and I put that down to the model.
The bigger surprise came from measuring. In late September I had agents read seven weeks of my own transcripts - 540 sessions, just under 3,000 things I’d typed, nearly 12,000 shell commands run on my behalf. Thirty-five skills had fired 126 times in all. Eight had never fired once. I never call a skill by name. I just say what I want, and the skill’s description either matches how I actually talk or it doesn’t.
So the library got rebuilt from the evidence: seven new skills for the work that actually fills my days, eleven rewritten so their descriptions match my real phrasing, four retired. Then each new one was checked back against the transcripts, which caught three more misses in the new work itself.
A playbook is only as good as its last real run. Put lessons in the file the agent reads, not a note beside it. Check dependency changelogs against usage. Test skills by running them, and measure whether they fire at all.
Better output, not just more of it
Speed stopped being the problem early. For most of these three months the question has been how to make what comes out better.
The factory had a second review stage from the start, and for a while I trusted it. Then I noticed every “second opinion” was the same model in a fresh session, reading the same brief and reaching the same conclusions. So I gave the reviewer a different job description: start clean, read the history of the change, and try to prove it’s broken - build a throwaway copy and reproduce the failure, or withdraw the finding. Across one batch it overturned five of the seven changes the pipeline had already passed. One would have let the builder quietly rewrite the tests it was about to be graded against. The model hadn’t changed; only the brief had. The panel has since moved to a different vendor too, and on its first real run it vetoed a change the original reviewer had passed - for less than a cent.
The other lever was routing. Early on every build ran the full test-first process, however small. A change that’s one database query and a rendered page doesn’t need that; a change to authentication does. The factory now sends work down a light lane or a full lane depending on its shape, which made the small things faster and left the care for the changes that deserve it.
And the whole chain got stricter as it went. Every client repository has tests that run on every pull request, and a deploy re-runs them on the box and rolls itself back if they fail. That’s standard practice anywhere software is taken seriously. The difference is that the thing opening the pull requests is now usually an agent.
If you’re running three
You don’t need twenty sessions for any of this to apply. Three is plenty.
- Once more than one agent is running, give them a home that isn’t the laptop you close at night, and a tool that knows they’re agents.
- Name every session by what it’s doing now, and make the one waiting on you the loudest thing on screen.
- Find your transcript files. Search them for secrets, rotate what you find, and stop pasting keys into chats.
- Isolate anything destructive at the environment level - separate users, separate test databases, guard tests, permissions - rather than trusting instructions.
- Know what each session costs in memory, and have something that kills the greedy process instead of the box.
- Write down which session is which before any restart, and resume by id.
- Put hard-won lessons in the file the agent reads, and test playbooks by running them.
- Give the reviewer a different brief from the builder.
If someone else runs them for you
Most people reading this won’t run agents themselves. They’ll pay someone who does, or approve a team that wants to. You don’t need to know what a worktree is to ask the questions that matter:
- Where do the agents actually run, and what happens to the work when that machine restarts?
- Where are the conversation logs kept, who can read them, and have they been checked for passwords?
- What are the agents not allowed to do on their own? If the answer is “nothing”, you have your answer.
- When an agent says a change worked, what checks it independently?
Anyone running agents seriously will have a boring, specific answer to each of those. Vague answers are the tell.
None of this is new. It’s the ninety per cent that was always there in running systems, arriving on your own desk. The other half of the three months - what got built, and the pile of what didn’t - is in The backlog.
The agents do the work. Running them is the job.