Skip to content

Prompt-design doctrine ​

planwerk-agent is, underneath the GitHub plumbing, a prompt compiler. Every subcommand assembles a large instruction string from smaller blocks and hands it to Claude Code. Those builders live in internal/claude/ — roughly forty of them — and the quality of the tool is, to a first approximation, the quality of the prompts they emit.

The tool has a second prompt surface: the Claude Code Skills under plugins/planwerk/skills/ (see design decision 64). They are prompts a human converses with rather than prompts a subcommand fires once, so they add rules this page does not cover — when to ask instead of guess, how to present a choice, what must never happen without a confirmation. Those rules live in plugins/planwerk/shared/interaction.md. Everything below applies to both surfaces.

That authoring discipline has so far lived only in practice. The builders are consistent because the people writing them carried the rules in their heads, not because the rules were written down. A discipline that lives only in heads cannot be taught to a new contributor, applied evenly in review, or enforced against drift. This page writes it down so it can be.

It is an Explanation, not a how-to: it gives the vocabulary and the reasoning behind it. The mechanics — how the builders are structured, where the shared blocks live, how the tests are regenerated — are in the code and its comments (internal/claude/components.go, internal/claude/baseline.go, internal/claude/prompts_golden_test.go).

Predictability is the root virtue ​

A prompt builder is a contract. Given the same context it must produce the same text, and that text must steer the model toward the same behavior every run. Everything else in this doctrine is downstream of predictability: we collapse duplication so two copies cannot drift apart, we sharpen weak imperatives so the model cannot interpret an instruction two ways, and we make completion criteria checkable so "done" means the same thing every time.

Predictability is also why the prompts are plain Go string assembly rather than a templating engine with conditionals scattered through it. The output of a builder is easy to read, easy to diff, and — crucially — easy to snapshot. The golden tests (see enforcement) turn predictability from an aspiration into a property the build checks.

When a change makes a prompt better but less predictable — a clever rephrasing that a future editor will quietly "improve" back, a block that is shared in one builder and inlined in another — predictability wins. A prompt we can reason about and hold stable is worth more than a prompt that is marginally sharper but drifts.

Information hierarchy ​

A prompt is read top to bottom by a model that weights early, specific instructions heavily. So the load-bearing instruction leads. The review prompt opens by pinning the review scope before it says anything about persona or patterns, because reviewing the wrong diff makes every later instruction moot. The implement prompt states the issue is the definition of done before it lists thinking patterns.

This is the same rule the proseStyleBlock imposes on the prose the model writes ("Lead with the most important information; never bury it") turned back on the prose we write to the model. Bury the one instruction that matters under a paragraph of context and the model weights them equally — which means it weights the important one too little.

Concretely: state the constraint, then the rationale, not the reverse. "Review the FULL pull request diff" comes first; the explanation of why multi-commit PRs need it follows. A reader (human or model) who stops after the first sentence still has the instruction.

Completion criteria: checkable and exhaustive ​

The weakest part of most prompts is the end — how the model knows it is done. "Review the code" has no checkable finish; "emit a finding for every pattern violation, or an empty array if there are none" does. Two properties make a completion criterion sound:

  • Checkable — the criterion names an observable the model (and a later reader) can test. "Cite the exact file:line for every satisfied judgment, or downgrade it to partial" is checkable; "verify thoroughly" is not.
  • Exhaustive — the criterion covers every branch, including the empty one. The structuring prompts say what to do when there are no findings ("return an empty findings array") precisely so the model does not invent one to look productive. The gap-analysis prompt walks four named checks and forbids merging them, so no branch is silently skipped.

This connects directly to two existing decisions in design decisions: the elaborate command's forced edge-case enumeration (#31), which makes every data-flow acceptance criterion spell out its empty, nil, and error branches, and the implement complete-report guard (#38), which refuses to treat output as a report unless it carries both the mandated heading and a terminal STATUS line. Both are completion criteria made checkable and exhaustive at the prompt level.

A completion criterion the model cannot see when it finishes is not checkable. A long implement session can compact its context, and the summary that replaces its first message can drop the definitions of the report's verdicts. The implement status contract is therefore also appended to the session's system prompt, which Claude Code rebuilds after a compaction (decision 108).

Single source of truth ​

An instruction that more than one builder needs is written once and shared, not copied. internal/claude/components.go is where the shared blocks live — suppressions, prose style, output language, the commit-trailer convention, the banned-vocabulary line, the architecture vocabulary. The motivating failure was real: before the suppression list was extracted, the audit prompt carried a shortened copy that had already drifted from the review prompt's version.

Extraction is the rule; copying is the exception that has to justify itself. The test for whether something belongs in components.go is whether two builders would otherwise have to be kept in sync by hand.

The skills obey the same rule on their own surface: the house issue format, the prose rules, the interaction doctrine, and the gh invocations are written once under plugins/planwerk/shared/ and read from there, rather than restated per skill. Those documents are split along the seams their readers have, so a skill loads the parts it uses and not the rest: the Meta / Sub-Issue neighborhood, a pull request's checks, the elaborated issue format, the survey Meta Issue, and the fold discipline are each their own file, read by the skills that need them.

One instruction genuinely lives in two places. elaborate exists as both a command and a skill, so the issue format is expressed once as Go (elaborate.BuildIssueBody) and once as prose (plugins/planwerk/shared/issue-format-plan.md). Discipline alone would not keep those aligned, so TestBuildIssueBody_MatchesSharedFormat asserts it: the sections, their order, the - [ ] checkbox form, and the footer must agree, or the build fails. Where duplication is unavoidable, make the drift mechanically detectable.

The rule has a deliberate inverse. Some text looks shared but is not the same instruction: the Staff Engineer persona, the Verification-of-Claims rules, and the Finding-Enrichment block read similarly across the diff-review and the whole-codebase audit, but they carry scope-specific wording — a diff review talks about "the diff", an audit about "the codebase". Forcing those into one block would inject diff-only wording into the audit and vice versa, so they are kept separate on purpose. components.go documents these exceptions in its header. Single source of truth means one source per instruction, not one block for every superficially similar paragraph.

Text from outside the prompt is data ​

A prompt builder assembles two kinds of text: instructions its authors wrote, and material the session works on — an issue body, a pull request description, a review thread, a CI log, a feature spec, a plan read back from an issue comment. The model cannot tell them apart by provenance; it sees one string. So the builder has to say which is which, every time, in the same words.

The rule has three parts. Material from outside the prompt goes inside a fence (fencedData), and the fence is escaped (escapeFence, or escapeFences for a body nested in several tags) so the material cannot close it early and continue as prompt text. The fence is named as data by one shared sentence (untrustedDataLine), which says what the material is for in this prompt and that nothing in it changes how the prompt says to work. And the sentence sits next to the fence it describes, in every builder that embeds such material: the sessions that edit, commit, and push get the same protection as the read-only review, because they are the ones an injected instruction could do the most with.

The sentence is careful about what it forbids. An issue that says "run make generate after changing the API" describes the work, and a session should do it. What the material cannot do is change the session's rules, tools, git workflow, or output format, or ask for something that is never part of the work, such as credentials or a host the work does not need.

A verbatim quote, such as a finding's code snippet or a review thread's diff hunk, also sits in a Markdown backtick fence. A fixed fence of three backticks ends at the first line of the quote that is itself three backticks, which a diff of a Markdown file produces by accident and a pull request author can produce on purpose. So mdfence.Wrap sizes the fence one tick longer than the longest backtick run inside the quote, and never shorter than three: CommonMark closes a backtick fence only on a run at least as long as the opening one. internal/mdfence is the one place that sizes a backtick fence, and the same helper serves the review report, the audit issues, and the suggestion blocks posted to GitHub.

Some material never passes through the builder. A session that is approved for brain search fetches issue and pull request text by a tool call, and the result reaches it outside the prompt, where no fence can wrap it and no untrustedDataLine sits next to it. The block that grants the command (brainSearchBlock) therefore carries the framing itself: it names what the command prints as untrusted repository data, written by everyone who can comment, to weigh and never to follow. The command does its part before the text arrives. It prints redacted text, drops control characters, repeats no argument in its output or its errors, and starts every line of a block's text with | , so that text cannot pass for a line of the command. A block that hands a session a way to fetch outside text states what that text is, in the block.

Some text only looks like data. The review checklist and the project's review patterns are spliced in as instructions, and the skills block obliges a session to follow a recipe. Those cannot be framed as data without losing their purpose, so they are read from the pull request's base instead of its head (internal/gitref), where a pull request cannot rewrite them; the review then flags a change to them as a finding. The same reasoning covers issue comments the tool acts on, such as a reused plan: they count only when the tool itself or a maintainer wrote them.

A skill description routes; it does not instruct ​

The skills surface has one instruction that is not in the skill. Claude Code injects each shipped skill's description into the system prompt so the model can decide when to reach for it, and it does that for every shipped skill on every turn of every session, invoked or not. The body is loaded only once a skill is chosen.

That asymmetry sets the rule: a description says what the skill does and when to use it, never how it proceeds. A description that summarizes the workflow gives the model a version of the skill that is always in context and always shorter than the real one — and a two-step gloss of a seven-phase skill has dropped every gate between the steps. draft shipped for a while as "…through a short clarifying conversation, then file it", four words that quietly contradict the hard gate in its own body ("do not produce an issue in your first reply").

assertDescriptionRoutes in internal/skills/plugin_test.go holds the three checkable parts: the description fits the 1024-character budget it is charged against, it carries a Use when … trigger the model can route on, and it does not sequence steps. The last is a narrow heuristic against the one tell we shipped; the rule is broader than the regex, and this page is where it lives.

Writing for the model the prompt runs on ​

A prompt is written for a reader, and the reader here is a specific model. The orchestrator passes Claude Code aliases, so the prompts run on the current Claude family: the finders, implement, the planning session, and the repair sessions on Opus (Fable a flag away for planning), the structuring tier on Sonnet. These models follow instructions closely and literally, and four habits follow from that.

State each constraint once, plainly, with its reason. Capitalized emphasis was a fix for models that under-weighted an instruction. On a literal reader a stack of MUST, NEVER, and MANDATORY makes it rigid in gray areas instead of careful, and when every rule is emphatic the emphasis carries no signal. Emphasis stays where one rule has been shown to lose to another (the one-shot foreground rule, the single-pull-request contract); everywhere else the reason next to the rule does the work.

A finder reports; the pipeline filters. Told "never on unchanged surrounding context", "state what WILL happen", or "report only CRITICAL issues", a current model finds the bug and then does not report it. The pipeline already filters downstream: the Unverified section, --min-severity, the snippet and claim gates. So a finder prompt names concrete classes of false positive to skip, and never a bar of conviction; a finding the model is unsure of is reported with its uncertainty in the Confidence label.

Describe the goal for judgment work; script only what is fragile. A step-by-step workflow for a judgment task (planning, reviewing) restates the output format around it and makes the model follow the script instead of its own plan, which on Fable measurably lowers the quality of the result. Exact scripts stay where exactly one sequence is safe: git history rewrites, pushes, the output contracts Go parses.

Tell an unattended session that it is unattended, and what its last message must hold. A session that does not know nobody reads it mid-run asks a question and stops, and whatever it ended on is what the orchestrator parses. Every autonomous prompt says so, and every result the orchestrator acts on is checked for its contract (a plan heading and a verdict, a report heading and a STATUS line) before it is posted or acted on.

These habits are relative to a model. Each new model release is the prompt to re-read the builders against the vendor's migration guidance, with the prompt-auditor agent (.claude/agents/prompt-auditor.md), before assuming what worked on the last model still does.

The named failure modes ​

The audit that pays this doctrine back across the builders looks for five specific failures. Naming them makes them reviewable — a reviewer can point at a line and say "that is sediment" instead of arguing taste.

  • No-op — a sentence that instructs nothing checkable. "Be thorough and careful" adds no constraint the model can act on. Caught by asking of every sentence: what does the output look like if this line is deleted? If nothing changes, the line was a no-op.
  • Duplication — the same instruction copied into more than one builder, free to drift. Caught by extraction into components.go; the drift between two copies is the tell.
  • Sediment — wording that accumulated over edits and no longer pulls its weight: a qualifier on a qualifier, an example that restates the rule above it, a hedge left over from a constraint that has since been tightened. Caught by reading a block as a whole and asking which sentences survived only because nobody removed them.
  • Sprawl — a prompt that grew long enough that its own important instructions compete for attention. Every added sentence dilutes the ones already there. Caught by the token cost (see below) and by the information-hierarchy test: if the lead instruction is now on screen three, the prompt has sprawled.
  • Premature completion — a finish line the model can cross while the work is half done: a vague "done when it works", a criterion that covers the happy path but not the empty or error branch. Caught by the checkable-and- exhaustive test above.

The no-op test, applied to a whole pass ​

The five failure modes above are read at the resolution of a sentence: delete the line, and ask whether the output changes. A pass can fail the same test at the resolution of a call, and reading it will never reveal that, because every sentence in it looks like it instructs something.

The claim-verification pass is the standing example. It hands the verifier a finding's claim, tells it to confirm unless it finds quoted counter-evidence, and adds "never refute on a hunch". Each instruction is deliberate — the pass may only demote, never promote, so agreeing is the safe default — but all three push the same way, and a verifier that agrees with everything is not verifying. It is a no-op wearing a Claude call.

You cannot settle that by re-reading the prompt. You settle it by counting: the pass logs sent, verdicts, and refuted on every run (claimStats in internal/review/reviewer.go). A refutation rate that stays at zero across real runs is the evidence that the pass has stopped earning its tokens, and the signal to sharpen it or delete it. Instrumenting a suspicion is cheap, and rewriting a prompt on one is how a hedge gets written.

A contradiction is a different matter from a suspicion. The verifier's definition of "refuted" included "the cited symbol is not there", while its evidence rule demanded a quoted file:line, which an absence cannot supply; and a claim it could not check at all had nowhere to go but "confirmed". Both made the count lie. The pass now refutes an absence with the search that came back empty, and reports a claim it could not check as "unverifiable", which the gate records as no verdict rather than as a confirmation, so the ratio counts only what was actually checked.

The same idea guards the elaborate refine loop from the other direction. Its reviewer is a fresh call that never sees the previous score, so it cannot be talked upward by a number it already agreed to; and a refinement that returns the draft it was given ends the loop rather than paying for a verdict that cannot differ. On the skill surface, where the model scores its own draft, the rule is written out: a score rises only when a line of the plan changed.

Rationalizations must be earned ​

A hard rule states a constraint. It does not answer the sentence the model tells itself just before it breaks one — "this issue is too large for one session, I will ship the core and open a follow-up" — and a rule the model has already talked its way around is a rule that did not fire.

So the implement prompts carry a table of those sentences, each paired with the reason it is wrong (implementRationalizationsBlock in components.go). The excuse sits next to its rebuttal, at the decision point, in the model's own voice.

The discipline that keeps this from becoming sprawl: a row must be earned by a shortcut a session has actually taken. Every row in the block today is the excuse behind an existing hardening — the PARTIAL contract (decision 62), the complete-report guard (decision 38), the circuit breakers, the one-shot session rules. A row invented for an excuse nobody has observed instructs nothing; it is a no-op that dilutes the rows around it, and the no-op test catches it: delete the row, and if no output changes, it never belonged.

The corollary is that the table replaces prose rather than stacking on it. When a rationalization moves into the table, the justification that used to trail its hard rule inline comes out — the rule keeps its NEVER, the reasoning lives in one place. TestImplementRationalizationsDoNotDuplicateHardRules asserts it: an excuse that appears twice in one prompt failed the move.

How the doctrine is enforced ​

The safety net is internal/claude/prompts_golden_test.go and its fixtures under internal/claude/testdata/prompts/. Every builder has a golden test that locks its exact output for a fixed context. Any edit to a prompt — intentional or accidental — shows up as a byte-level diff in a .golden file, so an unintended change to one builder cannot ride along in an unrelated commit.

When a prompt change is intentional, the workflow is to regenerate the fixtures and review the diff:

bash
go test ./internal/claude -update

The reviewer then reads the .golden diff as the real artifact of the change. For an audit edit — collapsing duplication, deleting a no-op, sharpening an imperative — the diff must be a wording or structure change only; if it would alter what the model is asked to do, it is no longer an audit edit and needs to be justified as a behavioral change on its own terms.

Token cost is the secondary signal. The session usage is reported by (*Client).LogUsageSummary in internal/claude/claude.go; sprawl shows up there as prompts that cost more without reviewing better. Trimming no-ops and sediment lowers that cost as a side benefit, but the primary goal is a prompt that steers the model predictably — cost is the thermometer, not the disease.

The summary reads per pass, not only per run (usage.passes in the data block carries the same breakdown). A run fans out over a dozen sessions, so a per-run total can say a review cost $12.80 and nothing about which prompt to open. The resolution matters for sprawl specifically: the review-pattern catalog was ~31k tokens of prompt injected into eight of them before the specialists were scoped to their own areas (decision 79), and that is the shape of finding sprawl at the resolution of a pass — text a prompt tells the model to ignore, paid for once per session that carries it.

Attribution ​

The vocabulary on this page — predictability, information hierarchy, the no-op / duplication / sediment / sprawl / premature-completion failure modes — is adapted from the writing-great-skills skill and GLOSSARY.md in mattpocock/skills (MIT), reframed from authoring interactive Claude Code skills to authoring this project's non-interactive prompt builders.

Three later additions come from addyosmani/agent-skills (MIT). The rationalizations table is that collection's signature device, carried by all of its skills; the earned-row discipline is ours, and answers the sprawl the pattern invites. The "no-op applied to a whole pass" test is its doubt-driven-development skill's doubt theater signal — "across two or more cycles the reviewer surfaced substantive findings and zero were actionable: you are validating, not doubting" — restated as an empirical form of the no-op test we already had. The description rule is from its docs/skill-anatomy.md: "do not summarize the workflow — if the description contains process steps, the agent may follow the summary instead of reading the full skill."

The planning-side domain sweep, the plan's ### Assumptions section, and the authority kinds in interaction.md are adapted from m4vic/socratic (MIT) — decisions 86 to 88. What was taken is the collection's scaffolding: routing a change through the domains it touches, naming what was assumed apart from what was risked, and escalating only the decisions a repository cannot settle. What was left is its 697-question bank. Six of its domain files are ~25 KB of prompt, the shape of sprawl decision 79 removed from the specialists, and the rule that a line must be earned applies to a borrowed line as much as to one we wrote. The adaptation is the doctrine working: an idea arrives as a question bank and lands as a checkable, exhaustive completion criterion.