Ask an agent to do the same thing twice and it may find three different ways to do it. Sometimes that's useful. When the task is running the Rails tests, I'd rather it use the procedure we already know works.

That's what I've been trying to wrangle: consistent execution without expecting a model to become predictable. Not removing agents from the work. Removing the parts they shouldn't have to figure out again.

A well-equipped kitchen isn't breakfast

I like to think of this as a kitchen. A prompt is the order ticket. AGENTS.md or CLAUDE.md is the cookbook on the counter—you don't want it buried in sticky notes. Skills are recipe cards. Slash commands are the procedures taped inside a cabinet door, the things you always reach for and never quite remember. Subagents are sous chefs with their own counter space.

These are useful ways to organize work. Give the agent a narrower task, a better recipe, and only the context it needs. Don't show it the whole kitchen when it only needs the frying pan.

But a well-equipped kitchen doesn't mean the chef actually follows the recipe. Picture a robot with ridiculous long arms reaching through a window to get to the stove, breaking the glass and knocking over the milk. It technically reached the frying pan.

A long-armed robot outside a kitchen reaches through broken windows to fry an egg while spilling milk onto the counter.

We can add another instruction: work from inside the kitchen. Don't break windows. Then another. Eventually we're documenting all the ways the agent has surprised us instead of giving it a reliable way to fry an egg.

Skills and commands help an agent choose what to do. They aren't the operation itself.

We already know how to do this

We remove repetition from code by extracting functions, methods, and shared components. We do something similar with prompts when we turn repeated instructions into skills. The next step is to extract the procedure those instructions describe.

Instead of asking an agent to find the right test command, run it, interpret what happened, and decide whether to continue, code can run the declared command and record its result. The agent can still diagnose a failure. It doesn't need to rediscover how this repository runs tests.

Give the robot a machine that cooks eggs and another that pours milk. It still chooses what to ask for, but it doesn't need to invent the mechanism each time.

A robot presses a button on an orange machine whose mechanical arm cooks an egg, beside a separate button-operated machine pouring milk into a glass.

A failed command is different from a passing command. No tests running is different from tests passing. Those distinctions shouldn't depend on how confidently an agent describes its work.

This is what my first ADW implementation did: Ruby scripts wrapped CLI calls to Claude Code, Codex, and Amp. Agents wrote and reviewed code. The runner controlled the phases and checks. Hooks can enforce specific boundaries too. TDD and Ralph-style build-and-check loops are other ways to discipline the work; none requires building an entire factory first.

The scripts helped. Then I had little programs scattered across repositories, each with its own inputs and outputs. I was still telling agents, "Go look at how we solved this over there and do that again." I'd stopped repeating some instructions without really making the pieces easy to reuse.

Swamp gave the pieces a shared shape

Swamp gives me typed models with methods, workflows that connect those methods, versioned data about what ran, and vault references for credentials. Community extensions let me reuse someone else's capability before building my own.

A method owns an operation. A workflow defines how operations fit together. Their results become data that the next step can consume, or that I can inspect later. This was the missing part of the scripts: a common way to compose the work and see what happened.

The agent still needs to discover the right tool. Repository guidance and skills can tell it to search installed types and community extensions first. That instruction isn't a guarantee either. But once the agent invokes the method, the method executes its defined procedure instead of asking the agent to improvise it.

Deterministic here describes the encoded procedure, not a promise of identical results. APIs fail. Repository state changes. An embedded agent call still uses judgment and tokens. The useful boundary is knowing which parts are code and which parts are model decisions.

PR triage started as another repeated request

I kept asking agents to help manage a stack of PRs: find failing CI, read review comments, check the code, suggest a fix, draft a response. That's a lot of coordination to repeat in a conversation.

With granular API tools, the agent has to coordinate those calls and carry the intermediate context. An MCP connection gives it tools, but doesn't necessarily give it this whole process. I wanted to define the plumbing once and spend model turns where judgment actually helps.

The review-comment flow now reads the stack's recorded thread data, selects open feedback, skips fresh drafts, and invokes scoped read-only agents to assess comments against the code. It records the verdicts, reply drafts, and suggested fixes. Posting a reply or applying a change is a separate action.

Recorded PR feedback
  → select comments that need attention
  → scoped agent assessment + reply draft
  → record suggestions
  → my decision

From the outer agent's perspective, that can be one invocation followed by a response. It isn't one model call in total: the review agents inside still consume tokens. The savings I'm aiming for are the repeated discovery and coordination turns, not free reasoning.

One review comment caught a real mismatch: a project's image count included collaborator photos, while the eligible-photo selection only allowed the company's own photos. The suggested change checked the actual eligible count before choosing a project. That's where I want the model involved—understanding the comment and its consequences—not figuring out how to fetch every thread again.

These pieces grew into a dashboard. It joins recorded PR, CI, and ticket state, shows what needs me, and surfaces reply drafts and held fixes. Some narrowly eligible mechanical CI fixes have their own automated path; changes that need my decision stay held. The dashboard isn't the invention. It's a useful surface over the recorded work.

The background work doesn't need a chief-of-staff agent constantly watching my PRs. Scheduled methods refresh the recorded state, and downstream steps reuse it. Agents enter when there's a comment to assess or a fix to prepare. I don't need a model to keep reasoning about whether it's time to poll GitHub.

The factory is the same idea, across more steps

A software factory is our normal development process encoded into executable steps: plan, build, test, review, deliver. Agents generate code and help review it. Code runs the checks and retains the handoffs. People set direction and decide what to approve.

You can coordinate that process with saved prompts, an orchestrator, an implementer, reviews, and human checkpoints. My iteration has been on how much of the repeated coordination to put into code. The process doesn't need to be invented again just because we're using Swamp to express parts of it.

My current factory works toward an integrated draft that I can try and refine. It uses scoped planner, checker, implementer, and reviewer calls. Runtime methods execute declared checks and record whether the evidence passed, failed, or is missing. A visible draft isn't a release certificate. When the direction settles, I can keep one PR or package the work into a useful stack.

The reuse is what makes this worth more to me than another isolated script. My cli-agent extension runs bounded agent jobs in the factory, PR-comment triage, and CI repair flows. Review-panel and review-rubric extensions support other review pipelines. Worktree lifecycle operations used by the factory also support the dashboard's fix path.

Improve one building block and several projects can benefit. Yes, I could build all of this without Swamp. I had already built a lot of it without Swamp. I like having pieces I can compose across repositories instead of another button attached to another one-off machine.

A reasonable question: if we already have a command that sets up a development environment, what are we actually gaining by putting Swamp around it?

Not much if we're just giving the command another name. The useful abstraction is the process around it: establish the environment, check that it's healthy, run the operation, and retain what happened. Existing commands can still do the underlying work.

One example is a QA flow that looks for an available PR environment and invokes a browser agent for a smoke test. Environment selection and setup belong in the repeatable procedure; exploring the app is the bounded agent job. A team can share those capabilities without everyone adopting my entire factory.

Repeatable can still be wrong

Agents compound our tendencies. That includes my tendency to get a few things working, get excited, and crank the machinery up to eleven.

I track my merged PRs over time alongside automated review grades. Throughput increased as I moved from manual work and Copilot to Claude Code and heavier agentic workflows. Then I slowed down a little and got more intentional.

When the quality appeared to dip, my first question was whether the agents were producing worse code. Looking back across earlier work, I recognized a pattern I'd had before using them: go fast, don't come back to clean up, and ship work I wouldn't be happy with on a slower day. More output made that tendency easier to see.

The grades use OOP and extensibility criteria informed by Sandi Metz, my preferences, and feedback I've received at work. They aren't an objective measure of product quality, and the rubric changed over time. But the recorded reviews gave me something concrete to revisit rather than just blaming the tools or trusting my impression of how things were going.

I wrote about scaling the factory back because I was spending too much time getting work through its process. A roughly 240-line PR took about two days. A 17-PR stack of about 2,000 lines took a week. Line count doesn't establish how difficult the work was, but the friction made me question what I'd built.

I'd overindexed on receipts, contracts, and gates before useful feedback. Checking receipts was becoming the work.

There was also a check that rejected a planner's trailing punctuation when it was supposed to verify intent. It did the wrong thing very consistently. Code can check required fields. Matching strings isn't the same as understanding whether a plan satisfies the request.

So the goal isn't to make everything deterministic. It's to put mechanics in code, keep judgment where it belongs, and get useful work in front of me sooner. Recorded results make mistakes easier to investigate; they don't make the process correct by definition.

The next time I explain the same procedure to an agent for the third time, that's a candidate for extraction. Start with a skill or command if that's enough. If the repeated steps need stronger execution guarantees, make a script, a Swamp method, or a workflow. Search for existing extensions before building another one.

One useful operation. Test it against the original task. Make it discoverable for the next run. Let the next actual problem tell you what to add.