Development WorkflowBranch B
Every Dead Rule Has a Cause of Death. Living Ones Only Prove They Haven't Died Yet.
Four ways a rule file dies, the retro that decides, and the two ways the retro fooled itself.
The ledger
If you develop with an AI you end up with rule files. CLAUDE.md, agent definitions, whatever Markdown sits under .claude/. I’ll call all of it rule files. Skills stay skills.
Over a few months, 41 skills existed in our monorepo. 26 were alive at the same time at the peak. Today there are 4. Counting rule files and skills together, at least 47 were deleted, and every one of them was wired into a live run when it died: a 400-line skill for teaching the tester its job, a rules file that was meant to apply itself by path and was never read once, one development flow per stack that later folded into a single /develop.
Nearly every deletion carries its reason in the commit message. The few that don’t, I remember. So this is archaeology, mostly.
Why bother. Two reasons, both about cost. A contradictory pair of rules does not throw an error. A human reading “you may read this” next to “never read this” would stop and ask. The model reconciles the two silently, picks one, finishes the task and says nothing. We had exactly that pair, and I’ll get to it. The second reason is duller: every line an agent loads is tokens it spends reasoning about that line, in every dispatch, for as long as the line exists.
The trunk post of this chapter sorted dead rules into three rough piles. The commit log shows four causes of death. Here they are, so you can skip writing them.
Death 1: knowledge in the wrong layer
When the project started, I wanted dispatched agents to stop exploring from zero, so the root CLAUDE.md described the current state of things. One line:
default to hardcoded English, use translation keys only if component already has i18n setup.
Alongside it: the whole directory tree, and a row of version numbers, React 19, TypeScript 5.x, MUI v7. All of it described what the frontend looked like that week.
A few months later the frontend got i18n. The line became false and did not disappear. Every agent that read it did as told. Same for the tree and the versions: a rule describing the present goes stale the moment the code moves, and the person moving the code has no reason to think about a file three directories up. Some didn’t know it existed.
And because this is a monorepo, the backend read it too. The frontend was reading stale facts. The backend was reading facts that were never its business. Two symptoms, one cause: the knowledge lived too far from the code that invalidates it. The people who would notice it go stale couldn’t see it. The stacks that could see it had no use for it.
The fix was a move, not a rewrite. Directory layout, test commands, coding style went into each project’s own CLAUDE.md, next to the code that changes them. The root keeps what applies to everyone.
Today every CLAUDE.md in the repo adds up to 768 lines. Under the old layout an agent read all of it. Now a frontend run reads the root, 33 lines, plus the frontend’s own, 31 lines. Everything else loads when a task needs it.
Death 2: hardcoding what the outside world decides
The first-generation skills opened with model: claude-opus-4-6. The reasoning was simple. I want the strongest model, I don’t trust the default, pin it.
Then the model upgraded and it became claude-opus-4-7[1m]. Then I got tired of editing version numbers and made it model: opus. Then a new model family appeared and every skill needed the same edit again. We pulled all of it and root got a ban:
Flow files never name a specific model: normal dispatches inherit the session’s model.
Model names, versions, paths. Everything the outside world decides belongs to the same class: it changes faster than you edit files. Hardcode one and you have taken on a debt that comes due at the next release. Since the ban, model upgrades have cost us zero edits.
Death 3: asking the agent to check itself
Every time a run went wrong we added a sentence. Afraid of a skipped step, we added a checklist. Afraid of an agent overstepping, we told it to ask itself whether it had authorization. The tester skill got to 400 lines partly this way. My favourite:
Do not tell the user “done” until every row checks. Rows 1, 4, 5 are the most common failure modes.
It did not work. The agent itself admitted it: the checklist was right there and rows 1, 4 and 5 got skipped anyway. So we appended “don’t skip.” The second generation’s frontend flow grew from 130 lines to 571 this way.
A rule asking the model to reflect on itself works about as well as a sign saying “no running.” Both rely on the runner judging the runner. Nothing outside the agent can verify whether the reflection happened. Whether it actually asked itself the question, only it knows.
So we split every such rule in two. Which part is fact, which part is judgment.
Facts went into JSON. The verify command, the test scope, what to clean up, which roles loop. Fields. The agent copies them and runs them. There is no room for “I think this one doesn’t need it.”
Judgment stayed in the rule file, but not as “remember to ask yourself.” It became a condition you can look up in a file:
planned→go requires a non-empty
go:line; developing→dev-done requires every work item done-with-evidence.
Before acting, look at the file. No value, no action. If the agent acts anyway, the plan file is the evidence and the retro catches it. Both halves are the same move: swap silence for something observable.
I only understood the general form of this later. Every gate in this flow is enforced by the agent the gate constrains. The coder checks whether the coder may proceed. That is why appending “don’t skip” never worked and never will. A sentence in a rule file is a request. A tool permission is a wall.
The one rule that has never needed a “don’t skip” is the reviewer’s tool list. It has no Edit and no Write. It cannot quietly fix what it flags. Not because it is disciplined, because it has no hands.
Bash can write too, so that wall had a gap. Hooks closed it. They came back declared in each role file’s frontmatter, so the boundary rides with the role. The reviewer’s Bash is now an allowlist, and anything outside it, sed -i included, exits with an error. Four of our seven role files carry a hook. The charter that governs our flow text says it plainly: a path- or command-shaped boundary on one subagent needs no failure first. Its hook replaces the prose.
Two things about walls I didn’t expect.
They can filter content, not just remove capability. We had a rule about what a code comment may contain. Half of it is mechanical: no ticket ids, no pointers to docs, English only. All three are a regex. That half became a hook and its prose was deleted. The other half, what the code and the specs do not show, requires judgment and stayed as one sentence in root. One rule, cut down the middle. Half mechanism, half request.
And walls have doors. The comment hook fails open if its script errors, and it exempts Markdown and the docs directories. I’m writing that down rather than pretending the wall is solid.
Where it stands: four roles fenced, three still on prose. The team lead’s own rule, never read source files yourself, is prose, and retros have caught it reading anyway.
Death 4: a patch that outlived its premise
A subagent once cited code that did not exist and the team lead made a decision on it. We patched:
If there is no Read tool call covering the file you are about to cite, STOP. Read first.
Correct at the time. Weeks later the architecture changed. The team lead’s context became the most expensive resource in a run, and a lead that reads source files is a lead filling its own head with noise. The same rule flipped completely:
never Read source files or full diffs yourself
Verifiable facts now come from a research dispatch that brings back the citation. Same problem, two opposite fixes, both right, different premises. This is the pair from the opening. Both were true, just not at the same time. Had the first not been deleted cleanly, they would have been live together, and the model would have reconciled them without a word.
Who decides: the retro
The four deaths above were diagnosed by hand, at first. Then we built the thing that does it for us.
At the end of every /develop run the agent writes a retro. Three things go in: what the user corrected, which gate failed and retried, which rule it read and then did something else. One report, into git. The report changes nothing by itself. It accumulates, and every so often a human reads the pile at once, decides which lessons become rules and which rules die, makes the edits, and deletes the reports. We call that a harvest. The human makes the final call.
What the retro bought us, and what it didn’t.
Deleting got cheap. I wanted to delete rules and worried about deleting the wrong one. “It costs almost nothing to keep” is the default every engineer grew up with, and for code it is often true. For text a language model reads on every dispatch it is not. So I wrote the trade into the charter. Delete a rule I shouldn’t have, and the next run trips on it, the user corrects it, the retro logs the correction, and I put it back. One recorded friction. Keep a rule I shouldn’t have, and nothing trips. It dilutes everything next to it, forever, and nobody files a report. Delete first, patch back if something breaks. Flow files are down more than forty percent from their peak.
Friction became mechanism, not more words. The reflex after a bad run is to add a sentence: last time this went wrong, so don’t. One friction showed up three times. Git on Windows converts line endings on checkout, and our workflow scripts fail to start with Windows endings. Three runs stuck, three retros, one of which dutifully proposed a rule telling the model to convert the file before running it. Only at harvest, with the reports side by side, was it visibly the same bug three times. The fix was one line in .gitattributes forcing that directory to Unix endings. The rule files gained zero words. The retro’s job was not the proposal. Its job was writing it down three times, so a human could see the repetition and pick a better fix than the one proposed.
The numbers lied, politely. I’m an engineer, so I wanted KPIs: how many redo rounds, how many agents, how many design changes. Every report came back looking fine. Wrong dispatches: zero, now and then one. Work items with evidence: one hundred percent, in nearly every report. Same report, two paragraphs down, the user was correcting the agent’s understanding and asking for the design to be redone. I had measured whether some steps happened, not whether the flow was any good. Five revisions of the definitions later I deleted them.
The second attempt was worse. I added a table asking the agent to mark each rule it read as load-bearing or no difference, so quiet but useful rules would have evidence at the next cleanup. A month in, twenty-odd reports had marked nearly two hundred rows load-bearing and a dozen as no difference. At the next harvest, load-bearing rules got deleted anyway and not one of the dozen was deleted because of the table. Ask a model whether a rule helped it and it will say yes. It has no way to tell, and neither do you from its answer. Thirty days, then gone.
Cost is the one thing worth recording. Now a script runs at the end of each development and copies out the token usage of every agent in the run, input, output, cache reads, wall time, into a table. No judgment involved, just numbers off the session log. It can’t say whether the flow is better. No two tasks are the same size. It can say where the tokens went. Across seven runs, the model’s output was under one percent of them. Over ninety percent were cache reads: agents loading the same rule files and source code into context, again and again. The main context’s peak varies threefold between kinds of task. By token count, the model barely writes. It reads. And every sentence in a rule file is read once per dispatch.
What is allowed to stay
Faster, more accurate, cheaper. That is a wish, not a criterion. Add a reminder and call it accuracy. Drop a document and call it speed. Every edit has a justification, so arguments go nowhere.
The charter now applies a text test to every sentence in a flow file. To stay, it must be one of three things.
A Fact the model cannot derive at the moment it reads. Not the directory tree or the tech stack, those are in the repo. Something like “this folder is a reference copy of the old service, ignore it for new work”, which is written nowhere else.
A Decision between reasonable options, recorded so it isn’t re-decided every run. Independent work items run in parallel, for instance. As models get better at picking, this class dies first.
A Correction where the model’s default differs from what we want, backed by at least one run where it actually did the thing. Not “I bet it might do X, so forbid X.” A recorded correction.
Everything else is derivable and gets deleted on sight. The burden of proof is on retention.
Alive only proves not dead yet
Every dead rule above has a cause of death on file. A living rule can only show it hasn’t died.
Cheap rules drift toward immortality. One line, never in the way, nobody ever has to decide to delete it. The tenth line of our first skill was a single word: ultrathink. Months of retros and not one evaluated whether it did anything. It sat there, maybe useful, maybe not. I deleted it from every file recently. Nothing happened.
The default used to be: prove it’s useless and I’ll delete it. Now it’s: prove it’s used, or it goes.
Is this step still necessary?
We had an agent whose only job was writing development notes. After every run it recorded why the code was written the way it was, so the next agent wouldn’t have to guess. Seventy-odd notes over several months. Sixty-odd retros mentioned it. Not one asked whether it should exist. Neither had I.
At a harvest I held it against the three classes. The facts it recorded could be read off the code. Every code change meant a note to update, one more layer to maintain. Every run waited for it to finish writing. Zero for three. The agent was removed. The decisions worth keeping moved into one or two lines of comment next to the code they explain, and the behaviours QA can observe moved into the acceptance criteria. The step is gone and nothing has gone wrong.