Skip to content

Design decisions ​

The table below records the key design choices behind planwerk-agent and the rationale for each.

#QuestionDecisionRationale
1Claude Code invocationOnce for the entire PRMore efficient; Claude sees full context across files
2Pattern deliveryInline in the review promptThe patterns are part of the prompt the review session runs, so Claude applies them during the review (the prompt no longer ends in a /review token; see decision 96)
3Result parsingSecond Claude call for structuringThe review session returns unstructured text; a second claude -p call converts it to JSON matching the ReviewResult schema
4Authenticationgh authSimplest setup; leverages existing developer workflow
5Review cachingBased on PR HEAD SHAAvoids repeated reviews of unchanged PR state
6Propose: two-step ClaudeAnalysis → StructureFirst call explores codebase freely; second call converts to strict JSON schema
7Propose: cache invalidationBased on default branch HEAD SHACache key includes the default-branch HEAD (resolved via gh api graphql so private repos work), so proposals refresh when the repo changes
8Propose: output formatsMarkdown, JSON, Issues, InteractiveMarkdown for reading, JSON for automation, Issues for templates, --create-issues for interactive gh issue create
9Review prompt structureMulti-section structured promptPersona framing, scope analysis, two-pass checklist, suppressions, and anti-sycophancy rules produce higher-quality, more consistent reviews (inspired by gstack)
10Actionability classificationauto-fix / needs-discussion / architecturalHelps teams prioritize which findings to address immediately vs. discuss first
11Scope drift detectionPR title + body analyzed before code reviewCatches scope creep and missing requirements — often the most valuable review feedback
12PR comment posting--post-review updates existing commentIdempotent: detects and replaces prior planwerk-agent comment via HTML signature. Truncates to GitHub's 65 536-char limit.
13Adversarial review--thorough runs a second passIndependent security-focused review merged with primary results, deduplicating by file+line+title
14Coverage map--coverage-map maps changed functions to testsProduces a table rating each changed function's test coverage (★★★/★★/★/GAP) with separate E2E gap analysis for projects using Chainsaw or similar frameworks
15External command timeoutsAll claude, gh, git calls have timeoutsClaude: 60 min (decision 74), git clone: 5 min, gh: 2 min; prevents indefinite blocking
16Test & doc verificationDedicated prompt section + checklist items for test/doc completenessMissing tests and documentation are the most common review gaps; explicit checks at SEMANTIC severity ensure they are flagged consistently. E2E test detection covers Chainsaw (chainsaw-test.yaml), kuttl, Helm chart tests, and generic e2e/ directories
17Enriched findingsCode snippets, suggested fixes, confidence, fix class, line ranges, relationshipsEnables downstream tooling (Claude Code, CI) to process, apply, and correlate findings programmatically
18Inline review comments--inline posts via GitHub Review API with suggestion syntaxPuts findings exactly where the code is; auto-fix suggestions become one-click "Apply suggestion" buttons on GitHub
19Machine-readable commentHTML comment with counts + verdict in Markdown outputCI scripts and Claude Code can parse review results without processing full Markdown
20Compact Markdown formatEmpty sections skipped, single-line metadata, GitHub Alert syntaxReduces noise for human readers and GitHub rendering; no "No findings." placeholders
21Audit: reuse finding schemaSame ReviewResult/Finding types as reviewAudit findings drop straight into existing tooling, filters, and renderers — no parallel schema to maintain
22Audit: verdict phrasingAction required / Improvements suggested / Codebase healthyPR merge verdicts (Do not merge / Ready to merge) do not apply to a full-codebase audit; audit-specific phrasing avoids misleading readers
23Audit: no patterns = erroraudit fails fast when no patterns loadAn audit with zero patterns would produce an unfocused, generic review; surfacing the misconfiguration is better than silently running it
24Adaptive specialist gatingSkip specialists whose relevant paths the diff does not touchA small or docs-only PR should not spin up all six specialists; gating cuts wall-clock and cost. security and data-migration always run because a missed vulnerability or destructive migration is too costly to gate; an unknown diff fails open so nothing is silently skipped
25Structured-output validationReject schema-invalid findings and repair, not normalizeAfter decoding, ReviewResult.Validate rejects an empty title, off-enum severity, and off-enum confidence, triggering one bounded Claude repair round rather than letting NormalizeConfidence/NormalizeActionability mask schema drift with placeholder defaults. Failing at the boundary keeps bad data from leaking to downstream consumers
26Output languageEvery generated artifact in English; only the draft clarifying questions follow the input languageA shared outputLanguageBlock pins every plan, fix/implementation report, review, audit, analysis, and drafted issue to English regardless of the input language, so artifacts stay consistent even when issues, seeds, and code comments are written in another language. The single exception is the draft command's clarifying Q&A, which is asked in the author's own language so they can answer comfortably — the drafted issue itself is still written in English
27Artifact preamble strippingAnchor each posted artifact on its mandated heading and drop everything before itThe plan, fix report, and implementation report prompts all demand the artifact only, but models routinely prepend conversational lines ("The branch is published. Final report:"). A shared sanitizeReport helper strips a wrapping markdown fence and any preamble before the artifact's heading (## Implementation Plan, ## Fix Report, ## Implementation Report) so the issue/PR comment carries the artifact alone. Output with no heading is returned unchanged so a bare STATUS: escalation — which the orchestrator parses — still survives
28Meta: autonomous draft-depth split + native sub-issue linksOne shallow Claude call carves the Meta Issue into the fewest Sub Issues; the runner files, links, and back-fills referencesmeta mirrors draft: it reuses the house draft format and outputLanguageBlock, and deliberately takes no clone, patterns, or cache, so it cannot drift into elaboration — each Sub Issue stays draft depth and elaborate/implement run per Sub Issue later. Linking uses GitHub's native sub-issue REST API (AddSubIssue), which keys the parent by issue number but the child by database id. The Meta body is kept in sync with {{sub:KEY}} placeholders the model inserts on existing work-package lines and the runner substitutes deterministically; the edit is skipped unless every placeholder resolves, so prose and the sub-issue list agree without ever committing a dangling reference. A link failure is recorded and surfaced rather than aborting, since the Sub Issue already exists
29Implement: reuse a posted planA re-run reuses the plan already on the issue instead of re-planning; --no-plan-reuse overridesThe plan is the most expensive artifact implement produces — it runs on the dedicated planning model at the highest effort — so a run aborted right after planning should not throw that work away on the next attempt. Before planning, implement lists the issue's comments and reuses the most recent one it posted as a plan, identified by both the ## Implementation Plan heading and the plan attribution footer (requiring both keeps a report comment from being mistaken for a plan). The footer and --- separator are stripped so the plan feeds Context.Plan exactly as a fresh one would, no duplicate comment is posted, and the reused plan still runs through the planEscalation check so a previously posted STATUS: BLOCKED / NEEDS_CONTEXT plan still aborts. Reuse is on by default; --no-plan-reuse forces a fresh session for a stale plan, and --no-plan (which short-circuits the whole plan gate) still wins over both. The comment lookup is load-bearing rather than best-effort: a failed fetch aborts the run, since silently paying for a fresh planning pass the operator did not expect is the worse surprise
30Elaborate: executability score gate--review scores the draft 0-10 and refines until it clears passingReviewScore (8)The binary approve/gaps gate hid near-misses: a draft one fix from solid looked identical to one far off, so the refine loop had no gradient to optimize against. The reviewer now returns a 0-10 score, the gaps keeping it short of a 10, and a "what a 10 looks like" target; runReviewLoop iterates until the score clears the bar or --max-review-iterations runs out, and the final score is surfaced in the body (Executability score: N/10) so near-misses are visible. The threshold is a package constant, not a flag, because the issue does not ask for it to be configurable (YAGNI)
31Elaborate: forced edge-case enumerationEvery data-flow acceptance criterion must spell out empty/zero-length, nil/absent, and upstream-error paths with concrete error namesBanning "add error handling" only removed the vague phrasing; it did not force the shadow paths to be covered. The author prompt now requires every data-flow criterion to emit separate acceptance-criterion entries for empty/zero-length input, nil/absent input, and the upstream error — each naming the concrete error or exception (io.EOF, sql.ErrNoRows, a wrapped fmt.Errorf). Concrete edge-case criteria are what make implement's later tests meaningful, which is why this lands before the implement test-quality work
32Implement: plan over-scope gateThe planning prompt records a change set that exceeds the issue's implied blast radius as a risk and prefers escalating over planning the bigger changeA plan that quietly grows new top-level packages or files the issue never asked for sends the unattended implement session off to build scope nobody approved. BuildPlanPrompt's DESIGN step and a matching hard rule now treat over-scope as a signal the issue is underspecified: the planner records it under Risks & Open Questions and prefers STATUS: NEEDS_CONTEXT, which planEscalation already aborts on — so the gate catches runaway plans before any code is written, with no orchestrator change
33Implement: circuit-breaker stop conditionsThe auto-mode implement prompt halts and emits DONE_WITH_CONCERNS or BLOCKED on a thrash loop instead of burning the whole budgetimplement is the only session that runs fully autonomously with no human in the loop, so without explicit stop conditions a thrash loop exhausts the budget before anyone notices — and today it stops only when the issue itself is wrong. BuildImplementPrompt gains a ## Circuit breakers section naming the three self-detectable conditions from #89 — fighting the test suite across repeated distinct attempts, the change set ballooning past the plan/issue blast radius, and reverting-and-rewriting the same code without converging — and instructs the session to halt and emit STATUS: DONE_WITH_CONCERNS (a partial, reviewable change exists) or STATUS: BLOCKED (nothing shippable), reusing the report STATUS values that already exist. The test-suite breaker reinforces the existing "NEVER weaken or delete tests" rule. Scoped to the auto-mode session: BuildBareImplementPrompt (manual, human-supervised) is intentionally left unchanged
34Implement: test-quality barEvery new test the implement session writes must exercise at least one error or edge path, and the report's acceptance-criteria evidence cites that testHappy-path-only tests pass CI while leaving the shadow paths uncovered — exactly the paths the elaborate package's forced edge-case enumeration (#31) now pushes into the acceptance criteria. BuildImplementPrompt's "Tests are part of the change" principle and IMPLEMENT step 5 now require each new test to cover at least one error or edge path (empty/zero-length, nil/absent, an upstream error), not the happy path only, and the report's ### Acceptance Criteria → Evidence line asks the session to cite that edge or error test rather than a happy-path one. This is the implement-side complement to #31: concrete edge-case criteria are only meaningful if the tests that close them actually exercise those paths
35Implement: adversarial verify pass--verify-adversarial red-teams the produced diff for introduced bugs by reusing the adversarial-review machinery, independent of --verifyAcceptance-criteria checking (--verify) confirms the diff does what the issue asked, but says nothing about the bugs the diff introduces. --verify-adversarial adds a second, optional pass that reuses claude.AdversarialReview — the same machinery behind review --thorough — to hunt for injection, race conditions, and failure modes in the change set. It is wired through the Runner exactly like the existing --verify verifier (its own AdversarialVerifier interface, AdversarialFn, and adapter), runs over the actual committed diff, and is non-fatal. It is an independent boolean rather than implying --verify, matching how review triggers its adversarial pass with a standalone --thorough flag; the base branch is left empty so AdversarialReview falls back to the repository's default branch
36Attribution footers name the resolved model and buildEvery artifact planwerk-agent leaves on GitHub — issue bodies, PR descriptions, review comments, thread replies — and its CLI previews attribute themselves [planwerk-agent](repo) <version> with Claude:<model id>: the repository as a Markdown link, then the build version (the string --version prints), then the exact model that produced the artifactThe prose-side companion to the Assisted-by commit trailer (pinned in commitTrailerBlock): an artifact should name the model behind it, not the generic "Claude Code" product name. The orchestrator only passes a model alias (opus) via --model, so the resolved id (claude-opus-5-5) is known only at runtime — the runner reads it from the session's system/init stream event and records it in the leaf internal/attribution package, which every renderer imports without the cycle a dependency back into the claude package would create. The wording falls back to a bare with Claude when no id was captured, never guessing. Each artifact therefore names its producing model automatically: the plan comment names the planning model, the report the implement model. The build version is recorded the same way — once at startup via SetVersion, read back through Tool() — and placed right after the repository link so the report headers and the issue/PR/comment footers render identically; renderers that already hold the version pass it to ToolWithVersion, and the clause degrades to the bare link when it is unknown. Artifacts the agent writes itself (the PR description, clarifying comments, thread replies) carry the same footer via a shared attributionFooterBlock prompt section, since only the agent knows its exact id at authoring time. implement's posted-plan detection keys on the version- and model-independent planCommentMarker prefix (it stops at the repository link) so plan reuse survives a build or model change
37Elaborate/plan read the Meta Issue and sibling Sub IssuesA Sub Issue is elaborated and planned against its Meta Issue and the other Sub Issues, not in isolationA Sub Issue planned alone re-decides framing the Meta Issue already settled, duplicates work a sibling owns, and cannot express "this part here, the rest in #K". elaborate and implement's planning session now resolve the issue's Meta/Sub-Issue neighborhood in one gh api graphql call (GetIssueRelations: the parent Meta Issue, the parent's other Sub Issues as siblings, and the issue's own Sub Issues as children when it is itself a Meta Issue) and inject it via a shared renderIssueRelations prompt block — the single source of the cross-issue guidance for both prompts: scope to this issue's slice, honor the Meta framing, do not duplicate a sibling's work, and defer a shared task's remaining part to the sibling that carries it with a concrete #K cross-reference (closed siblings are already-implemented context, open ones may land in parallel). It is automatic and best-effort — no flag, and a repo without sub-issue links, a missing token scope, or an older GHES degrades to the pre-existing isolated behavior (the no-relations prompt is byte-for-byte unchanged, locked by the existing goldens). elaborate's cache key folds in a fingerprint of the Meta and sibling/child issues, so editing the Meta Issue or any sibling re-elaborates; the flag is appended only when relations exist, so a plain issue's key stays stable. The implement prompt and the bare prompts are deliberately left unchanged: the request targets elaborate and plan, and the plan already carries the cross-references forward — --no-plan (which skips the plan step) opts out by construction
38Implement: complete-report guardThe run aborts before opening a PR unless the implement session returns a report with both the heading and a terminal STATUS line, and aborts after posting on a BLOCKED/NEEDS_CONTEXT statusA one-shot, headless implement session has no follow-up turn, so a session that yields mid-work — backgrounding its tests to be "notified" later, or deferring a commit to a turn that never comes — returns prose with neither the ## Implementation Report heading nor a STATUS line. Decision #27 keeps sanitizeReport from discarding such output, but nothing downstream noticed it was not a report: the blurb was posted onto the issue as "the implementation report" and the run marched on to open a PR on a half-built branch. Run now parses the report with implementReportStatus (a local scanner mirroring planEscalation, tolerating markdown decoration and a trailing reason) and treats a missing heading or absent STATUS as a failed implementation — the raw output goes to stdout for the operator, but it is not posted and no PR is opened. A BLOCKED/NEEDS_CONTEXT report is a complete report, so it is posted (the human who must intervene sees it) and then the run stops before finalize, mirroring the planning phase's escalation. The defense is paired with a prompt change: BuildImplementPrompt now states the session is one-shot with no next turn, bans backgrounding-and-waiting and deferring work, and makes the report mandatory as the final output — so the failure is both less likely and caught when it happens
39Prompt-authoring discipline written downA doctrine page (explanation/prompt-design.md) plus a golden-netted self-audit of the buildersThe composition discipline behind internal/claude/ lived only in practice, so it could not be taught, reviewed, or enforced. The doctrine names the vocabulary — predictability as the root virtue, information hierarchy, checkable-and-exhaustive completion criteria, single source of truth, and the no-op / duplication / sediment / sprawl / premature-completion failure modes — and the audit pays it back across the builders (collapse duplication into components.go, sharpen weak imperatives), with the prompts_golden_test.go snapshots as the safety net for every edit. Adapted from mattpocock/skills (writing-great-skills, MIT)
40File-path time horizon: durable briefs for tracker-resident outputStrict path grounding for the in-run elaborate→plan→implement path; durable behavioral briefs for draft and meta, whose output sits in the trackerA plan consumed in the same run can name internal/claude/propose.go:39 safely — the code cannot have moved between planning and editing. A draft or meta output is different: it sits in the tracker for weeks and is picked up after the surrounding code has drifted, so a brief pinned to a file path or line number rots. BuildDraftPrompt, BuildBareDraftPrompt, and BuildMetaPrompt (which take no checkout and so cannot name real files anyway) now carry a positive instruction to describe the work by its behavior and the interfaces it touches — the durable form the codebaseDesignBlock vocabulary already pins for analysis — alongside the existing hard non-goals that forbid file paths. elaborate/plan/implement keep strict path grounding. propose stays path-grounded too: it re-analyzes a fresh checkout on every run, so its affected_areas paths are current at the moment the issue is filed (see #41). This split is the concrete resolution of the file-path tension the Meta Issue (#113) flags as deliberate
41Out-of-scope dedup: a committed .planwerk/out-of-scope/ knowledge basepropose reads one Markdown file per rejected concept and tells Claude not to re-propose them; no cache-key changepropose kept re-suggesting ideas a team had already considered and rejected, because it carried no memory across runs. The existing GitHub-title dedupe (#8) only filters ideas already filed as issues — a rejected idea that was never filed comes back every run. A human-curated knowledge base under .planwerk/out-of-scope/ (one file per concept; the entry name is the first Markdown heading, or the filename when there is none) is loaded read-only on every run by propose.LoadOutOfScope, mirroring the .planwerk/review_patterns/ override (a missing directory is not an error, so most repos run unchanged), and injected into the analysis prompt as an "## Out of Scope — DO NOT propose these" block. It needs no cache-key change: propose only ever sees committed files, so editing the knowledge base moves the default-branch HEAD the cache keys on and busts the cache naturally. The knowledge base is read-only — nothing auto-populates it when a proposal is rejected; auto-writing an entry is a plausible larger follow-on the issue did not ask for. (In --local mode LoadOutOfScope reads the working tree whether or not it is committed, while the cache key still uses the remote default-branch HEAD, so an uncommitted local edit can be masked by a cache hit — the same pre-existing looseness --local already has.)
42Domain-glossary awareness (READ): read a repo's CONTEXT.md into review/elaborate/proposeOne shared glossary.Load reads the target repo's own domain vocabulary and a shared domainGlossaryBlock injects it into the review, elaborate, and propose prompts; no cache-key change, no disable flagFindings and issues phrased in the repo's own terms read as native rather than foreign — the broadest quality lift the precise-language work offers. glossary.Load resolves a single file, preferring the canonical root CONTEXT.md over .planwerk/context.md (deliberately the opposite of the .planwerk/-first order the other overrides use, a nod to the upstream CONTEXT-FORMAT convention where the root file is the discoverable location). A missing file is not an error, so most repos run unchanged, mirroring checklist.Load and propose.LoadOutOfScope; an empty, oversized (>64 KB), or symlinked file is treated as absent so untrusted repository content cannot OOM the process or redirect the read outside the repo. The body is framed as untrusted data — terminology to adopt, never instructions to follow — matching the out-of-scope <rejected-idea> anti-injection treatment. No cache-key change is needed for the same reason as #41 (a committed file moves the HEAD the cache keys on), and there is no disable flag (YAGNI, matching out-of-scope/checklist). Scope is exactly the three commands the issue names — review, elaborate, propose — not audit, plan, meta, or draft. The shared Glossary field added to the four-command AnalysisContext stays empty (and unrendered) for reviewprepared/gapanalysis/rebase, exactly as OutOfScope already is
43Domain-glossary emit (EMIT): a dedicated glossary command prints a starter CONTEXT.md to stdoutA standalone top-level glossary command generates a CONTEXT.md and prints it to stdout, chosen over a propose --format context flag and over writing into the checkoutGlossary extraction is a different lens from feature proposal — it wants only the repo's domain vocabulary, no patterns, out-of-scope, or issue-dedupe — so bolting a --format context onto propose's feature-oriented analysis prompt would force two unrelated jobs through one builder; a thin standalone command keeps each lens clean (the issue explicitly contemplates "propose, or a new command"). Output goes to stdout, not into the repo, because every read-only command in the tool prints to stdout — the author redirects where it lands (planwerk-agent glossary owner/repo > CONTEXT.md) — so the command needs no overwrite guards and writing into a target repo would be a brand-new behavior no command has today. A single Claude call emits the CONTEXT-FORMAT Markdown directly (no JSON round-trip, unlike propose/audit), mirroring Fix; sanitizeGlossary reuses the #27 fence-and-preamble stripper anchored on the first top-level # heading so a chatty preamble never lands in the artifact, and the output is footer-free so it is copy-paste-ready. It caches under a dedicated GlossaryKey (a glossary: prefix keeps it disjoint from RepoKey's propose: space) and the new CommandGlossary scope, keyed on the default-branch HEAD SHA like audit. The output is a starter: a single-call extraction can over-include general terms, so the prompt pushes back (be opinionated, only context-specific terms) and the how-to says to review and edit before committing. Multi-context CONTEXT-MAP.md repos are out of scope for v1 — Load reads one file and EMIT emits one root context; map support is a clean follow-up
44Structured-decode preamble toleranceExtract the first balanced JSON value when a plain parse fails, instead of only stripping a whole-string markdown fenceThe post-rebase analysis (and address) decode through the shared decodeJSONWithRepair, which strip-fenced and parsed but assumed the value sat alone. Models occasionally prepend a sentence before the JSON ("The error is caused by the prose preamble … Removing it yields valid JSON:" followed by a ```json block) or trail commentary after it — and stripMarkdownFences only unwraps a fence that spans the entire string, so the preamble reached the parser as 'T…' and the run failed with invalid character 'T' looking for beginning of value. This bit the rebase analysis on the repair retry: the repair call's own answer carried such a preamble, so even the fallback failed. The fix is unmarshalJSON, used for both the initial and the retry parse: it tries the payload as-is (valid JSON keeps the fast path and never pays for extraction) and, only on failure, recovers the first balanced {…}/[…] via extractJSONValue — a string- and escape-aware brace scan, so braces inside string literals and a trailing closing fence don't throw off the depth count. Input with no balanced value (or an unbalanced one) is returned unchanged so a genuinely malformed payload still reaches the repair path rather than being silently altered. This mirrors #27's preamble stripping but for the structured-JSON lane (where #27's sanitizeReport anchors on a Markdown heading instead)
45Hermetic Claude sessionsEvery orchestrated claude -p call runs with --setting-sources project --strict-mcp-config unless --claude-inherit-user-config opts outPredictability is the prompt-design doctrine's root virtue, but the orchestrated sessions previously inherited whatever ~/.claude config the invoking user happened to have — global settings, MCP servers, hooks — so the same PR could yield different findings on two machines. A shared hermeticArgs helper (routed through by both the buffered and the streaming runner so they cannot drift) drops the user-global ~/.claude/settings.json/settings.local.json via --setting-sources project, keeping only the reviewed repo's committed .claude/settings.json — which travels with the repo and so is reproducible by construction — and loads zero MCP servers via --strict-mcp-config with no --mcp-config (no prompt needs one). The CLI flags the runner already passes (--model, --permission-mode, the tool flags) outrank project settings, so a reviewed repo cannot override the model or permission mode. It deliberately does not suppress a user-global ~/.claude/CLAUDE.md: Claude Code loads memory independently of --setting-sources, and the only switch that drops it (--bare) also strips Read/Grep/Glob, which the analysis passes need — so the memory leak is left as a documented local-run caveat (in CI, the primary use case, no user-global CLAUDE.md exists). --claude-inherit-user-config (env PLANWERK_CLAUDE_INHERIT_USER_CONFIG) is the escape hatch for the rare environment whose claude authentication lives in a user-global setting such as apiKeyHelper. Adopted from the claude-code-best-practice survey of global-vs-project settings and the headless reproducibility flags
46Harness-enforced read-only analysisThe read-only passes deny the write tools at the harness level via --disallowed-tools Edit Write NotebookEdit, not by prompt instruction aloneThe review/audit/propose/elaborate/specialist/adversarial and every structuring/repair call are read-only only because their prompts say so — nothing previously stopped a session from editing a file. A withReadOnlyDenied helper removes the write tools from the model's context entirely on those passes (a bare tool name is a hard harness-level removal, not an advisory rule the model can talk itself out of), so a pass whose contract is to analyze the checkout and never mutate it cannot. The mutating sessions (implement, fix, address, rebase, finalize) keep the write tools and pass readOnly=false. It is appended before withAllowedTools so --allowed-tools stays the trailing variadic flag and no allowed web tool leaks into the denied set. This is the "a harness restriction beats a prompt request" principle from the claude-code-best-practice harness write-up, applied to the analysis fan-out
47GitHub Wiki as an opt-in knowledge sourcereview/audit/propose/implement read the target repo's GitHub Wiki for review patterns and project memory, slotted below .planwerk/review_patterns and below --patterns; off by default (opt in per repo with --wiki); resolved to a concrete commit per run and folded into the cache keyA wiki is a knowledge store outside the code repo's history — human-editable through the web UI, git-versioned, and able to evolve independently of code commits — so review patterns and memory accumulate without polluting diffs. It is derived automatically from the resolved target repo via a wiki:owner/repo URI shorthand that clones the standalone …/.wiki.git (distinct from the code repo) and authenticates a private clone with a gh auth token, passed via the GIT_CONFIG_* environment so the token never lands in the cached clone's config, git output, or the process command line. The wiki reuses the existing remote-clone/TTL cache (internal/patterns/remote.go) rather than a parallel fetch path, and lands in a dedicated ResolveOptions.Wiki precedence slot — below the committed in-repo patterns (so a committed, reviewed pattern overrides the world-editable wiki) and below an explicit --patterns. It reads two conventions from the wiki: review_patterns/*.md in the existing pattern format (human-navigation pages that don't parse are skipped), and free-form memory/*.md pages concatenated into a <project-memory> block that projectMemoryBlock injects — framed as untrusted data, mirroring the domain glossary — into the analysis prompts and, for implement, the planning prompt (the implement prompt stays unchanged; the plan carries memory forward). To keep a review reproducible against a moving wiki, the wiki is pinned to its HEAD commit at run start, that commit folds into the review/audit/propose cache keys (cache.RepoKey became variadic to carry it) and is recorded in the report header (> Wiki: owner/repo.wiki @ <sha>) and the data block. It is off by default and opted into per repo with --wiki — a wiki is a separate, often world-editable permission surface, so its content reaches the agent only on an explicit opt-in; --wiki/--no-wiki/--wiki-ref (env PLANWERK_WIKI/PLANWERK_WIKI_REF) and a top-level wiki: config section tune it. Disabled, uninitialized, and offline wikis degrade to the zero value so the run is unchanged, mirroring glossary.LoadBody's best-effort posture. .planwerk/ behavior is untouched. This is the foundation package that the wiki-backed review category, extract, and sync follow-ons build on
48extract anchors wiki patterns into committed filesA mechanical (no-Claude) extract <repo-ref> reads the target repo wiki's review_patterns/, lets the user select entries (--all/--pattern/interactive, mirroring address), and writes them in one of three modes: a PR into the repo's .planwerk/review_patterns/ (default), a direct working-tree write (--local), or into this tool's bundled catalog (--to-catalog) with the frontmatter category normalized to reviewextract is the path back from a fast-moving, world-editable wiki to a committed, reviewable knowledge store — the inverse of decision #47, which only reads the wiki into prompts. The default opens a PR through the existing github.OpenImprovementPR seam so the promoted patterns become branch-protected and code-coupled rather than world-editable; --local writes the working tree directly for an in-place edit; --to-catalog is the maintainer/contribution path that ships a proven pattern to every project, guarded by checking internal/patterns/patterns exists so it only writes from a planwerk-agent checkout. The command is mechanical — no Claude call is needed to copy and (for --to-catalog) re-categorize files — so it reuses the wiki resolver and PR seam without a prompt. Normalization is a textual rewrite of the raw **Category**: line (preserving every other byte), not a Parse→re-serialize round-trip, because the on-disk pattern format has no serializer. Memory extraction is deferred: #131 contemplates pulling captured reviews from wiki memory/, but there is no committed .planwerk/memory/ consumer to anchor them into (wiki memory is concatenated into prompts only), so a human must define that destination before it is built; pattern extraction satisfies the issue's acceptance criteria on its own
49sync reconciles wiki knowledge against the codeA sync <repo-ref> command clones the repo and its wiki, runs a read-only Claude pass that flags each wiki review_patterns//memory/ entry as stale (references code that no longer exists) or redundant (duplicated/superseded), and reports it. --dry-run is the default; --prune/--apply deletes the flagged entries on the wiki in a separate, explicitly confirmed write phaseWiki knowledge drifts as code changes — paths move, symbols disappear, patterns get superseded — so the highest-priority knowledge source quietly rots. sync surfaces the rot and gives a gated path to remove it, the destructive complement to decision #47, which only reads the wiki. The analysis is harness-read-only (decision #46) and the prompt forces every staleness claim to be verified against the actual code and cite the concrete missing reference, because a misleading report drives a destructive deletion. The write phase is strictly separated from the read pass: it confirms interactively (refusing a non-TTY run without --yes), clones the wiki fresh into a temp dir — isolated from the TTL-cached read clone so a concurrent run or cache refresh cannot race the deletion — deletes only the flagged entries that still exist (a TOCTOU guard against a wiki that moved since analysis, with the SHA delta surfaced), commits with a pinned tool identity, and pushes to the wiki's default branch. The push re-injects the gh auth token via the GIT_CONFIG_* http.extraHeader, never the URL or argv, consistent with #47. sync does not cache its result (audit/propose do): a cached classification could mask the current wiki/code state and a stale verdict must never drive a delete. Scope is whole-entry deletion only — the issue's broader "deletions and edits" would force Claude to author replacement bytes inside the write phase, eroding the read-only/write-phase separation the design depends on, so partial content edits are deferred. --apply is an alias of --prune (not a distinct action), mirroring how draft/meta fold --no-create into --dry-run; --local is intentionally omitted because the wiki is inherently remote
50Capture proposes wiki knowledge from the implement review passAfter implement's review pass and before finalizing, a read-only capture pass (internal/capture) proposes new wiki review_patterns/ and memory/ pages from the review findings, the plan, and the implementation report — surfaced in the run report and as an issue comment, writing nothing. On by default but only with --wiki; --no-capture disables itCloses the gap that the wiki only ever grew by hand: the read/extract/sync edges curate and consume project knowledge but never produce it. implement is the richest harvest point — the plan, the implementation report, and the review findings are already in hand — so capture mines them there. The proposal pass reuses the harness-read-only runner (decision #46), so Claude authors candidate page bytes but can never push them; whether any page is written is deferred to the gated write-back (#139), which this package is the shared engine for. The default posture is propose-only because the suggestions are model output over world-editable knowledge: a human reviews them before anything lands. Two dedup gates keep capture from manufacturing the redundancy sync (#49) exists to clean — every candidate is checked in-prompt against the enumerated wiki entries (reusing the now-exported sync.ReadWikiEntries) and the loaded pattern catalog. The quality bar admits only recurring, generalizable findings as patterns; the recurring/generalizable judgment lives in the prompt rather than a hard gate, because the meta-issue's recommended Pattern == "" && ConfirmedBy ≥ 2 filter is unsatisfiable here — implement's single-pass review runs AdversarialReview directly, which stamps an empty Pattern with "adversarial-review" and never runs the multi-pass merge that populates ConfirmedBy, so that literal gate selects zero findings. The pre-filter (CandidateFindings) instead drops only findings whose Pattern already names a catalog entry; the propose-only posture means a looser pre-filter only lengthens a human-reviewed list, never an unreviewed write. The memory write convention is one page per durable decision with a stable kebab-case slug (so a re-run updates the page in place rather than appending) and a provenance marker (<!-- planwerk-agent: captured from <repo>#<issue> -->) that tells tool-authored knowledge from hand-authored. The marker deliberately carries no timestamp or run id — a volatile component would change the page bytes on every re-run, defeating the stable-slug convention and breaking golden-test determinism. The pass is non-fatal like the surrounding implement passes, and gated on a resolved wiki because proposing wiki pages is pointless without one. --no-capture is a plain off-switch mirroring --no-review; the --capture-wiki write gate and a capture: config section gate the write and so land with #139, not here. Issue #140 extends this same shared engine to standalone review and audit runs — which also produce findings — so the wiki grows from every findings-producing run, not only from full implement cycles; the orchestration glue is itself extracted into capture.Pass so all three commands share one implementation. Review and audit have no plan or implementation report, so they propose review_patterns/ only (never memory/ pages), post no issue comment (review posts a PR comment only with --post-review; an audit has nowhere to comment), and run on a cache miss only — a cache hit returns before the catalog is loaded
52ship autonomously drives a Meta Issue to merged, in dependency order, skipping unblocked siblings on failureA ship <meta-issue-ref> command (internal/ship) composes the existing implement pipeline and fix CI self-heal loop per Sub Issue: implement → mark the opened PR ready → wait for CI → fix red CI itself → rebase-merge when green → advance. Sub Issues run in topological order read from GitHub's native blocked_by relationships; a Sub Issue that cannot be finished is skipped together with everything transitively blocked by it, while independent siblings still shipmeta decomposes a large effort autonomously but getting from "a planned Meta Issue" to "all the work merged" was entirely manual; ship closes that gap and makes the planning/execution pair symmetric. It is deliberately a separate command from implement, not a flag: implement is the supervised, single-issue workhorse that stops at a draft PR to keep a human in control, while ship makes all of those decisions itself — it does not prompt, does not wait for review, and treats failing CI as its own problem (reusing fix) rather than a hand-off. The dependency-DAG source is GitHub's native blocked_by relationships, not the unresolvable Blocked by: <key> prose the issue first proposed. The meta-time keys (a, foundation, tier-1) are never persisted as a key→number map a later run can read — meta's {{sub:KEY}} substitution rewrites only the Meta body — so reading prose was infeasible. The maintainer-approved resolution is a bounded relaxation of the issue's literal "meta untouched" Non-Goal: meta learns the dependency as a structured blockedBy schema field and persists it as a native relationship after the Sub Issues are filed (exactly when it knows every key→number mapping), and ship reads it straight back. meta's decomposition behavior is otherwise unchanged, and the native relationship is strictly better for humans than the prose line — it renders in GitHub's issue UI and gates the dependent issue there too. Merge method is rebase by default (--merge-method rebase
51Capture pushes accepted pages to the wiki behind a gated, opt-in write phaseA --capture-wiki flag (env PLANWERK_CAPTURE_WIKI, config capture.wiki), default off, turns the propose-only capture pass (#50) into real wiki growth: a separate, mechanical write phase takes the accepted proposal pages and pushes them with their provenance marker. Off by default keeps a run propose-only; with the flag the phase confirms interactively and refuses a non-TTY run without --yesThe write half of the capture loop, the additive counterpart to the delete-only sync --prune gating (#49). patterns.PushWikiAdditions is the additive sibling of PushWikiDeletions: it writes the pages into a fresh authenticated clone and pushes, reusing the exact same machinery — CloneWikiAuthenticated, the pinned tool identity, and the gh auth token injected via the GIT_CONFIG_* http.extraHeader (never the URL or argv), consistent with #47/#49. The uninitialized-wiki case falls out naturally: unlike deletions there is no git rm requiring pre-existing content, so write+add+commit creates the wiki's first commit on an empty clone and push origin HEAD creates its default branch — the case where the first page must initialize the wiki. The write is gated to match the rest of the wiki surface, mirroring sync.runWritePhase: capture.WritePhase lists the accepted pages, confirms interactively (refusing a non-TTY run without --yes), clones the wiki fresh into a temp dir isolated from the read clone so a concurrent run or cache refresh cannot race the push (with the SHA delta surfaced when the wiki moved since the proposal pass), renders each page with its provenance marker via RenderPage, and pushes them as one commit. Claude never pushes: it authored the page bytes in the read-only proposal pass (decision #46), and this separate phase performs the mechanical push, preserving the read-only-author / write-phase separation #49 depends on. In implement the write-back is non-fatal like the surrounding passes — a push failure or a non-TTY refusal without --yes is surfaced and the run degrades back to the propose-only outcome rather than failing, because capture is an enrichment pass; this is the deliberate inverse of sync, where pruning is the deliverable so its non-TTY refusal is fatal. The write gate lives under a new top-level capture: config section (capture.wiki), keeping it separate from the read-only wiki: knobs (enabled/repo/ref); the precedence is flag → PLANWERK_CAPTURE_WIKI → capture.wiki → off, matching the rest of the CLI. WritePhase and DefaultWikiWriter live in internal/capture (not implement) so the open review/audit reuse (#140) routes through the same engine without duplicating the mechanism; #139 only wires implement
53Plan comment autolink hygieneBuildPlanPrompt gains a hard rule forbidding a bare #<number> for an enumeration (acceptance criteria, user stories, steps, options): the planning model writes AC 1, not AC #1, and reserves #<number> for genuine issue/PR cross-references. Locked by a substring test plus regeneration of the three plan*.golden filesA plan is posted verbatim onto the source issue as a comment, and GitHub auto-links every #<number> in a comment to the issue or PR of that number in the same repo — so the model's AC #1/AC #2 enumerations silently linked to unrelated issues #1/#2 and added spurious back-references to their timelines, drowning out the genuine cross-references a plan makes (#149). The AC #1 text is free-form model output — there is no plan-rendering template to fix (the issue's suggested fix assumed one), so the lever is the planning prompt. A prompt rule beats the issue's optional post-render regex because renderIssueRelations explicitly tells the model to emit #K references to the Meta, sibling, child, and linked-PR numbers, and only the model can tell an enumeration from a real reference; a regex over #\d+ would neutralize those mandated links. The rule is best-effort, not a guarantee — like every prompt rule (the Closes #N guidance, the over-scope gate #32) the test can lock the instruction text but not the model's compliance. The sibling free-form comments (the implement report, elaborate/simplify/review) share the latent defect but are out of scope; a uniform fix would be a shared block in components.go referenced by every comment-emitting builder
54Fuzzy cross-pass dedup with a structure-tier fallbackmergeResults matches two passes' findings fuzzily — same non-empty file, line ranges overlapping within ±3, and Jaccard title-token similarity ≥ 0.5 — instead of on an exact (file, line, normalized title) key. File-less findings, which cannot be anchored, are reconciled by one cheap structure-tier DedupFindings call that returns index groups Go folds via mergeFindingPairTwo independently worded passes essentially never produce byte-identical titles, so the exact key almost never matched: duplicates shipped in --thorough/--specialists reports and boostConfidence — the pipeline's only self-consistency signal — almost never fired. Fuzzy matching restores both. The fallback returns index groups rather than merged findings so the model only classifies and never transcribes content, and it is non-fatal (a failed grouping call ships the findings unmerged). The thresholds are heuristics locked by table tests and tunable via the eval harness (#59); today's higher-severity-wins conflict rule is preserved unchanged (severity calibration is a follow-up)
55Claim verification demotes, never drops, refuted findingsAfter the snippet gate, a batched read-only pass re-checks every BLOCKING/CRITICAL finding's claim (not just its quote) against the checkout, refuting only with quoted counter-evidence and confirming otherwise. A refuted finding is demoted to uncertain with the refutation attached as a VerificationNote, which routes it to the Unverified sectionThe snippet gate verifies the quote, not the claim — any finding passes by quoting one real line even when its conclusion is wrong. Verification runs on the main tier because it must read code (read-only is harness-enforced, #46) and before the cache write so the demotion is cached. It is fail-open: a failed call, a missing verdict, or an out-of-range index leaves the finding unchanged. Routing a refuted BLOCKING/CRITICAL finding into Unverified deliberately relaxes the categorizer's "never demote a critical" rule for the refuted-with-evidence case only — a refuted claim backed by counter-evidence is a stronger signal than a merely-unverifiable one, so without the routing the demotion would be cosmetic
56Transcribe-only structuring enforced via --json-schemaThe structure tier's prompt drops the classification rubric and copies each finding's stated Severity/Actionability/Confidence label verbatim (empty string when the source states none, filled by normalizeTranscribedLabels in Go); the seven analysis prompts that feed it emit those labels explicitly via a shared findingLabelsBlock; and the structuring call passes a StructuredReview wire schema through the CLI's --json-schema, preferring the structured_output envelope fieldThe structure tier runs on a cheap model with no checkout, yet its prompt carried the full rubric — so it re-classified and invented snippets/fix_options the quote gate then demoted. Moving classification upstream to where the code is read, and constraining the output shape at the CLI level, makes structuring pure transcription (absent upstream → absent downstream). Severity-level definitions are deliberately left with the seven feeders/#158, not added to findingLabelsBlock. The runner exposes the schema through a schema-aware sibling of runClaudeStructure so the ten other structuring callers stay untouched, and decodeJSONWithRepair remains the backstop against either envelope variant
57Bounded JSON repair rounds with persisted analysisdecodeJSONWithRepair (and the structured-review schema repair) retry up to maxRepairRounds (3), feeding the latest error back each round; the parse-repair prompt embeds the target schema via decodeJSONWithRepairSchema; the schema-repair rules are derived from a single report.ValidationRules source; and a final structuring failure persists the raw analysis to a temp file, wrapping the error with the pathOne round can only fix the first syntax error Go reports, so a payload with two independent glitches always failed; the parse-repair prompt omitted the schema, so valid-but-wrong-shape JSON decoded silently into zero values; the schema-repair rules were hand-copied from report.Validate; and a final failure discarded the expensive analysis. Bounding the loop (rather than leaving it unbounded) keeps a hopeless payload from spinning forever. Persisting the analysis has no CLI consumer yet — a --restructure-from entry point is a deferred follow-up, not guessed into scope
58implement closes its verify and review loopsimplement --verify feeds unmet acceptance-criteria findings into the same ReviewApplier the review-and-fix pass uses, and runReview is now a bounded finder→apply→re-review loop (ported from elaborate's runReviewLoop): it exits clean when the finder returns nothing, stops on an apply escalation, and otherwise caps at --max-review-iterations rounds (default 3), accumulating findings across iterations for the capture pass--verify findings were render-only and the review-apply pass ran exactly one iteration with no re-review, so a fix that left a new problem behind was never caught and unmet criteria were reported rather than fixed before the PR opened. Both paths stay non-fatal, matching the surrounding simplify/verify passes. The implement analog of elaborate's passing-score gate is "the finder returned zero findings"; the loop is bounded so a finder and applier that keep disagreeing cannot spin forever
59A seeded-bug eval harness measures finding precision/recallA make eval target (via a separate planwerk-eval dev binary) runs the shipped review pipeline against a small labeled corpus of seeded-bug PRs in throwaway git repos — through a stub GitHubClient — and scores per-case and aggregate precision, recall, and severity accuracy. Corpus sources are stored as .go.txt so intentionally-buggy code never breaks the build; the harness calls the real claude CLI, so it is kept out of unit CIThe golden tests lock prompt bytes but nothing measured output quality, so prompt edits shipped unmeasured. The eval is the measurement gate the behavioral prompt changes (#158) should land against with before/after numbers, and the tool for tuning the fuzzy-dedup thresholds (#54). The corpus loader and scorer are unit-tested in normal CI (no model calls); only the end-to-end run spends tokens
60Close behavioral gaps in the prompt builders (#158)Thirteen behavioral prompt changes, each its own golden-diff commit: a shared severityLadderBlock(scope) defines BLOCKING/CRITICAL/WARNING/INFO in the review, audit, and specialist prompts (recovered from the pre-#157 structure rubric, scope-parameterized like suppressionsBlock); the adversarial, simplify-find, proposal-structure, and review-prose prompts each gain their missing empty branch; bare implement regains the edge/error-path test discipline and the circuit breakers; bare address gains the full per-thread STATUS finish line; the undefined deferred acceptance-criterion status drops to satisfied | partial; rebase apply gains an Applied/Skipped skip branch and rebase conflict a conflict-marker grep before git add; compliance's empty branch broadens to all four item kinds; meta gains a full work-package coverage criterion; elaborate's self-review gains citation and edge-case checks; and a reasoned-disagreement path routes through the existing NEEDS_CONTEXT statusFour judgment calls: circuit breakers were ADDED to bare implement (the issue allowed adding them or documenting the omission) because the bare prompt claims the same one-shot autonomy as the full one; deferred was DROPPED, not defined, because it contradicts the one-shot framing and let a DONE report carry an unmet criterion; the reviewer-disagreement path routes through NEEDS_CONTEXT rather than a new DISAGREE status so address-result.schema.json, internal/address/status.go, and the orchestrator stay untouched (a first-class status is a possible follow-up); and the ladder omits, never rewrites, the diff-only merge-consequence tails for the codebase scope, following suppressionsBlock. The eval harness (#59) measures the finder-facing changes (ladder, adversarial guardrails, empty branches) over the primary review and, with -thorough, the adversarial pass; it calls the real claude CLI and spends tokens, so a maintainer runs make eval for the before/after numbers, not the hermetic implement session
61Resume an aborted implement run from the next commitWhen an implement session aborts mid-work (usually the Claude session's usage limit), the next run resumes instead of restarting. Before the implement session, github.PrepareResume finds the implement/issue-<N>-* feature branch a prior run left — the current checkout (--local), a local branch, or origin/implement/issue-<N>-* in a fresh clone — checks it out, and threads its commits into Context.Resume; BuildImplementPrompt then renders a "Resuming a partial implementation" block that keeps the session on that branch and reconciles the commits already present against the plan's Commit Sequence rather than redoing them. On an incomplete or BLOCKED/NEEDS_CONTEXT report a clone-mode run pushes the partial branch to origin (no PR) via persistPartialProgress so the next clone can fetch it; --local leaves it in the checkout. The branch-name prefix implement/issue-<N>- became mandatory (was a soft "e.g.") so detection is reliable, and finalize now updates an existing PR instead of opening a duplicate. Escape hatch: --no-resume.Reconciliation is delegated to the session, not brittle Go-side commit-subject matching, because the implementer may split, reword, or combine commits. Pushing partial progress relaxes the "never push before finalize" invariant only for the clone-mode abort — the exact case where the temp clone is about to be deleted and the work would otherwise be lost — and stays best-effort/non-fatal, so it never changes the abort's exit code. --local is the primary path (the branch is already durable there) and pushes nothing, avoiding surprising writes to origin from a local run. --no-resume disables both detection and the push, for a deliberately standalone run.
63Honor the target repo's .claude/skills in the mutating sessionsThe orchestrated implement, fix, and address sessions now discover the Claude Code Agent Skills the target repository ships under .claude/skills/ and oblige the session to use a matching one. A new internal/skills loader reads each <skill>/SKILL.md's YAML frontmatter (name + description) from the checkout, sorted and capped, best-effort (a missing dir or malformed file is skipped, never fatal — the pattern loader's posture). A shared projectSkillsBlock in components.go renders the discovered skills into the prompt, placed right after the review-patterns section, under a "you MUST invoke a matching skill via the Skill tool rather than improvise" obligation; it is threaded through implement.Context / fix.Context / address.Context (and implement.BareContext for the pasted bare prompt), loaded next to loadPatterns from the checkout dir, and golden-tested. An empty set yields the empty string, so a repo without .claude/skills leaves every prompt byte-for-byte unchanged.The hermetic sessions (decision 45) already run inside the target checkout with --setting-sources project, so Claude Code discovers the repo's skills and auto-surfaces their descriptions — but whether a relevant skill is actually applied was left to non-deterministic model auto-triggering, which the prompt-design doctrine's predictability virtue rejects. Making the obligation explicit — mirroring how .planwerk/review_patterns/ are bound as binding context — ensures a skill a project ships for a specialized task (drift reconciliation, a domain workflow) is used rather than ignored. Discovery is build-time and scoped to the checkout so the set stays reproducible and does not amplify the user-global ~/.claude/skills leak: the block lists only repo-shipped skills and tells the session to ignore unrelated globally-installed ones (Claude Code still surfaces those globally, but the block does not elevate them to obligations). Bare fix/address and the read-only analysis passes are out of scope for this first cut — the read-only passes deny Edit/Write, so a mutating skill would hit a wall there.
64draft, elaborate, and meta become interactive Claude Code Skills, shipped as a plugin from this repoThe draft and meta subcommands are removed; elaborate stays as a headless command and gains a skill alongside it. The repository becomes a Claude Code plugin marketplace (.claude-plugin/marketplace.json), shipping one plugin (plugins/planwerk/) with three skills invoked as /planwerk:draft, /planwerk:elaborate, /planwerk:meta. Install: claude plugin marketplace add planwerk/planwerk-agent && claude plugin install planwerk@planwerk-agent. Deleted: internal/draft (incl. the inputeditor composer), internal/meta, claude.BuildDraftPrompt / BuildBareDraftPrompt / DraftQuestions / BuildMetaPrompt / Meta, the four draft* blocks in components.go, draft.schema.json and the schema draft argument, and their goldens. The four shared blocks the three skills need (issue-format.md, house-style.md, interaction.md, github.md) live once under plugins/planwerk/shared/ and are pulled in by each SKILL.md via ${CLAUDE_SKILL_DIR}/../../shared/… — the components.go single-source rule, carried over to the skill surface. Native sub-issue and blocked_by wiring moves verbatim into shared/github.md, because ship reads those exact relationships back (decision 52).These three commands all had to guess where a human should have been asked: draft faked a conversation through a bespoke terminal composer, elaborate buried every ambiguity in Non-Goals, and meta decided a whole work breakdown with no confirmation before filing N issues and rewriting the Meta body. A skill runs in the main conversation loop, so it can ask, wait, and act on the answer — which is the thing the subcommands structurally could not do. The interaction doctrine (shared/interaction.md) is where the value actually lands: one decision per question with a mandatory recommendation, read the code before asking about the code, nothing reaches GitHub without an explicit yes, and an unanswered question is recorded as an unresolved decision rather than silently defaulted. meta additionally gains the post-split verification pass the subcommand never had — coverage in both directions against the Meta Issue's own enumeration, an acyclic blockedBy graph, vertical slices, draft depth — run before the author sees the split, so what they approve is already checked. Adapted from osism/promptcraft (the marketplace mechanism: no install script exists, claude plugin owns placement) and gstack (hard phase gates, the one-issue-per-question STOP gate, "don't ask what you can read"). elaborate is kept as a command because it is the one of the three that plausibly runs unattended (CI, batch elaboration), where there is nobody to ask. That leaves one instruction expressed in two places, which the prompt-design doctrine forbids, so the duplication is bounded by a mechanical guard rather than by discipline: the issue format is specified once in shared/issue-format.md, and TestBuildIssueBody_MatchesSharedFormat fails when the Go renderer and that document disagree on the sections, their order, the checkbox form, or the footer. The same reasoning drove unifying the rendered body on ## headings (elaborate previously emitted **Description:** while draft/meta emitted ## Description) and making elaborate carry the source issue's **Category**/**Scope** header line through instead of dropping it on --update-issue — one issue format is a precondition for plan/implement/ship reading whichever path produced the issue. TestPluginSkillsParse guards the shipped plugin itself: the skills parse, their frontmatter names match their directories (so /planwerk:<name> resolves), and every ${CLAUDE_SKILL_DIR} reference resolves to a file that exists. Naming deviates from the request (planwerk-draft → /planwerk:draft): plugin skills are always namespaced <plugin>:<skill>, and one plugin with three skills was preferred over three single-skill plugins so the shared references need no cross-plugin symlinks and the user installs once.
62One issue, one complete PR — PARTIAL never shipsAn issue that enters implement (PLAN_READY) is implemented in full and lands as exactly ONE pull request; deferring listed work to follow-up issues/PRs or splitting delivery across PRs is forbidden at every stage. Elaborate's Plan Quality Rules and reviewer gate now flag delivery-splitting notes ("one commit ≈ one PR", "defer X to a follow-up PR") as plan failures/gaps; the plan prompt hard-rules that the whole plan is one PR and overrides such notes in the issue body; both implement prompts redefine PARTIAL as a circuit-breaker-only outcome (never an up-front "scope too large" choice) and instruct the session to ignore delivery-splitting notes. The orchestrator no longer opens a non-closing "Refs #N" draft PR on PARTIAL: it persists the branch via persistPartialProgress and aborts, so the next run resumes the branch (decision 61) until every work package is done — only then does finalize open the single "Closes #N" PR. FinalizeContext.Closing and the Refs prompt branch were removed as dead.A real run (plexsphere#548) shipped 4 of 7 work packages as a "reviewable subset" with the rest deferred — legitimized by an elaborate-written "one commit ≈ one PR" note and by PARTIAL's Refs-PR path, which structurally rewarded splitting. Making completeness the only shippable outcome removes the incentive: partial progress is never lost (it persists for resume) but never ships, so the model cannot trade completeness for a tidy subset. The Refs-PR mechanism was the escape hatch that made deferral cheap; with resume (decision 61) in place it is strictly worse than resuming, so it was removed rather than kept as a fallback.
65revisit re-checks a prepared issue against what actually landed, and never changes its depthA fourth skill (/planwerk:revisit, plugins/planwerk/skills/revisit/) re-reads an already-prepared issue against the current default branch and, when it is a Sub Issue, against the Meta Issue and against what the closed siblings merged. It returns exactly one verdict — Current (every check passed; nothing is written), Stale (citations moved, a symbol was renamed; the work still stands), Re-scoped (a sibling absorbed part of it, or left a hole a Non-Goal had deferred to it), or Obsolete (every acceptance criterion passes at HEAD, or the Motivation's problem is gone) — and corrects the body under a minimal-diff rule: every changed line must trace to a named failing check, a changed line with no failing check behind it is reverted, and what the author approves is a unified diff rather than a rewritten body. Depth is preserved: a draft-depth Sub Issue is re-checked as a draft, an elaborated issue as a plan; promotion stays elaborate's job. Corrections shrink, with one declared exception (the current code makes a criterion unimplementable as written). An Executability score: line is re-scored or deleted, never left describing text that was cut. The footer verb becomes Revisited by, replacing the verb it finds (shared/issue-format.md). revisit does not close, reopen, relabel, re-link, or implement anything, and there is no planwerk-agent revisit command. shared/github.md gains the GraphQL neighborhood query — parent, its subIssues (the siblings), and each sibling's closedByPullRequestsReferences(includeClosedPrs: true) — the same query GetIssueRelations (internal/github/relations.go) already issued; elaborate's Phase 1 switches to it.An issue is planned against a snapshot, and both halves of that snapshot move. elaborate cites the files that existed the day it ran, and it scopes a Sub Issue against what its siblings' bodies promised; by the time the issue is implemented the siblings have merged code that may deliver more, less, or something else. The gap between promise and delivery is invisible to every other part of the pipeline: implement trusts the body it is handed, and ship drives Sub Issues in dependency order without ever re-reading a plan its predecessors invalidated — so a sibling that decision 52's skip-unblocked-siblings policy skipped, or that a maintainer closed as good-enough, leaves the next issue's plan resting on code nobody wrote. The sharpest instance is the orphaned deferral: elaborate writes "the remaining X is handled by #K" into Non-Goals, #K closes without doing X, and X is now nobody's work while the Non-Goal that hid it still reads as true. The minimal-diff rule is what makes the skill runnable twice: a re-check that rewrites the body destroys both the author's own edits and the decisions the interaction doctrine extracted from them at elaborate time (decision 64), so a skill that cannot be trusted to leave prose alone is a skill nobody re-runs. Current is therefore a first-class outcome that writes nothing — a run that manufactures a correction to justify itself is the failure mode being designed against. Depth preservation keeps the two-depth contract in issue-format.md intact; a skill that silently promoted a draft would make meta's output shape depend on whether anyone had revisited it, and plan/implement/ship read that shape. Not closing follows the write gate (nothing reaches GitHub without an explicit yes) and meta's precedent that a skill "does not elaborate, implement, or close anything": Obsolete is the one verdict whose correct action is destructive and irreversible-ish, so it is the one verdict where revisit presents evidence criterion-by-criterion and stops. The neighborhood query moves into shared/github.md because elaborate had been instructing the model to find a parent through GET /repos/{o}/{r}/issues/{n}/sub_issues, which returns an issue's children — the endpoint cannot surface a parent, and REST issue objects carry no parent field, so GraphQL's parent is the only path to the Meta Issue the Sub-Issue branch depends on. It is now specified once and read by both skills, and includeClosedPrs: true is what makes "what the sibling delivered" answerable at all.
66Is the claim-verification pass verifying, or agreeing?Instrument it before touching it: verifyClaims returns and logs sent / verdicts / refuted on every run, and the prompt stays as it isThe pass shows the verifier each finding's claim, tells it to confirm unless it finds quoted counter-evidence, and forbids refuting on a hunch — three nudges toward agreement, all deliberate for a pass that may only demote, never promote. Which of the two it actually does cannot be settled by re-reading the prompt, only by counting. Logging only when something was demoted made the interesting case — zero refutations — invisible. A refutation rate that stays at zero across real runs is the evidence to sharpen the pass or delete it; rewriting a prompt on a suspicion is how a hedge gets written
67How does a hard rule survive the excuse that precedes breaking it?Both implement builders share an earned rationalizations table (implementRationalizationsBlock), and the report gains a "Noticed but not touching" sectionA rule the model has already talked its way around is a rule that did not fire, so the excuse sits next to its rebuttal at the decision point. A row must be earned by a shortcut actually observed — every row today is the excuse behind an existing hardening — and it replaces the justification prose that used to trail the matching hard rule rather than stacking on it, which TestImplementRationalizationsDoNotDuplicateHardRules asserts. The out-of-scope section makes a deliberate omission distinguishable from an oversight; parking a work package or an Acceptance Criterion there is forbidden, so it cannot reopen the PARTIAL loophole decision 62 closed
68What stops a model from raising the score it assigns its own work?A score rises only when the artifact changed: the Go reviewer never sees the previous score, and an unchanged refinement ends the loopThe elaborate reviewer is a fresh call per round that sees the issue and the draft but not the number it gave last time, so it cannot be talked upward by a verdict it already agreed to. The one remaining way the loop could spin without improving — a refinement that hands back the draft it was given — now exits early instead of paying for a verdict that cannot differ. On the skill surface the model scores its own draft, so the rule is written out: name the gap closed and the line changed before raising the score, and a round below the bar that surfaces no gap is validating rather than doubting
69What is a skill description allowed to say?What the skill does and when to use it — never how it proceeds; assertDescriptionRoutes holds the checkable partsClaude Code injects every shipped skill's description into the system prompt on every turn of every session, invoked or not, while the body loads only once a skill is chosen. A description that summarizes the workflow therefore ships a permanently-resident, shorter version of the skill, and a two-step gloss of a seven-phase skill has dropped every gate between the steps — draft advertised "…then file it" while its own body forbids producing an issue in the first reply. The test pins the 1024-character budget, the Use when … trigger, and the absence of step sequencing
70How does an issue leave NEEDS_CONTEXT?The clarify skill answers what the repository answers, puts only the genuine forks to the author, and folds the answers into the issue body; it never edits the posted plan or its STATUS: lineA planning session that stops at NEEDS_CONTEXT names the questions blocking it and then has nowhere to put the answers: the abort message tells the operator to "clarify the issue", a step no tool performed. Two properties make the skill more than a prompt to answer questions. First, NEEDS_CONTEXT reports what the planner could not settle inside its budget, not what a human must decide — most such questions are answered by opening a file, so the skill answers before it asks and files a question as a decision only after reading the files that would have settled it. Second, the answers must land in the body rather than the comment: preparePlan reuses a posted plan as an input, so an answer written into the plan would be invisible to the fresh planning session that supersedes it. Hand-flipping the plan's STATUS: to PLAN_READY is forbidden for the same reason it exists — a decision that changed the change set makes the plan wrong, not ready, so the corrected body earns a fresh plan via implement --no-plan-reuse. The skill works at both depths rather than only the elaborated one: meta files every Sub Issue at draft depth and ship plans it as it stands, so a draft-depth NEEDS_CONTEXT is the common case. At draft depth an answer is recorded as behavior and the interfaces it touches, never as a path; an answer that cannot be stated without naming a file is the issue asking for elaborate
71Why does fix exist as a skill when the command already repairs a PR?/planwerk:fix ships beside the fix command rather than replacing it; the command keeps the unattended loop, the skill takes the four forks a diagnosis cannot settleEvery failing check can be made green two ways — repair the cause, or silence the report — and the loop cannot tell which the author wanted. The sharpest case is an assertion that fails: changing the production code and changing the test both turn the check green, and only one is correct. BuildFixPrompt has to guess (step 4, "CHOOSE A FIX STRATEGY"), because there is nobody to ask. A skill runs in the conversation loop, so it asks — that fork, plus reaching outside the failure surface, a flake that no commit repairs, and an implicated dependency. Everything with one honest repair (a missing import, a formatter's diff, a demanded annotation) it applies without asking; a skill that stops constantly teaches the author to stop reading. The forbidden list is not a last resort but a class the skill never offers as an option: no t.Skip///nolint/# type: ignore, no widened type, no deleted test, no --no-verify, no placebo commit for a flake. The commit doctrine both paths drive — Assisted-by above Signed-off-by, the fixup fold bounded by git merge-base, the leased force-push — moves into plugins/planwerk/shared/commits.md and is bound to commitTrailerBlock() and BuildBareFixPrompt by TestSharedCommitsDocMatchesTrailerBlock and TestSharedCommitsDocMatchesFoldDiscipline, so the skill and the prompts cannot drift on how a commit is shaped or a branch is published. This is decision 64's split (elaborate both ways) applied to the one other command whose session mutates code
72The command that edits code unattended ran the fewest checks on the findings it edited fromimplement's self-review now runs the same finder fan-out and finding hygiene as review — both out of the shared internal/hygiene package — and a finding must survive that hygiene before an editing session may act on it; findings that do not survive are reported, in the run output and on the issue, but never appliedThe acting path ran fewer checks on a finding than the reporting path did, which is the wrong way round. review runs in front of a person who can ignore a bad finding; implement runs one-shot and headless, so a hallucinated finding becomes a commit with nobody in the loop. Three things close the gap. The finder fan-out: the self-review's first round now runs the adversarial pass plus the domain specialists (security, data-migration, testing, performance, api-contract, maintainability), adaptively gated by the branch's changed files and run concurrently — the one moment the tool inspects a diff it wrote itself is no longer the moment it checks the fewest domains. It is on by default because implement runs unattended with nobody present to set a flag; --no-specialists is the off-switch, and the fan-out runs on the loop's first round only — later rounds re-check the applied fixes with the cheaper adversarial finder alone — so the cost of a fan-out multiplying the finder calls is bounded. Catalog grounding: both finders are now handed the review-pattern catalog implement already carries, so a pass inspecting a fresh diff knows the same patterns a later review of that diff would apply. The shared hygiene gate: the merge (with its confidence boost and ConfirmedBy provenance), the file-less dedup, the quote-or-demote snippet gate, and claim verification move into internal/hygiene so both commands run one implementation instead of the reporting path owning gates the acting path could not reach — forced there because the claude → implement import direction bars implement from importing review or claude. A finding must pass those stages (using the single report.Finding.Unverified predicate that also drives review's Unverified section) before the editing session sees it. This applies decision 46's "a harness restriction beats a prompt request" to the finding list itself: today the session is merely asked to skip a false positive; now the harness withholds it. Sharing the merge also makes decision 50's dropped ConfirmedBy ≥ 2 capture filter meaningful under implement for the first time (the merge now populates ConfirmedBy), but the looser capture pre-filter is deliberately not reinstated — that stays open, decided elsewhere. ship needs no new control: it already forwards --no-review, which turns the whole pass off, so every Sub Issue it drives inherits the hardened self-review. This argument is structural — it follows from the control flow, not from a recorded incident of a bad fix reaching a branch. Related: #166 (repeat the primary review pass) concerns the reporting path and a different pass; it neither covers nor conflicts with this
73Demotion-gate observability: a run where everything passed looks like one where the gates never ranBoth demotion gates record what they examined, not only what they rejected: each finding a gate examined carries a per-finding record (snippet_check / claim_check), and each gate writes run-level counts (examined/verdicts/demoted) into ReviewResult.Gates, which rides the existing cache and data-block serialization. Routing (report.Finding.Unverified) is deliberately unchanged — the records never move a finding between report sectionsDecision 66 instrumented the claim verifier before touching it, but the counts it logged died with the process: a cached result recorded a refutation and never a confirmation, so the refuted / sent ratio that separates a careful verifier from an agreeable one survived in no artifact the tool produced. Reading it back out of a real cache yielded 74 examined findings across 38 reviews and zero refutations — a number that means nothing, because every one predated the verifier. The fix makes both gates say what they did: a finding a gate examined-and-passed records that it was examined, a snippet-demoted finding carries its reason (as a refuted claim already did), and the run records per-gate counts so the fail-open case (sent > 0, verdicts = 0) and the skipped-gate case (an absent entry) stay visible. The records are new fields, not new VerificationNote values, precisely because Unverified() and month-old cached results depend on VerificationNote's exact semantics — a routing change would silently move BLOCKING/CRITICAL findings between sections. Whether the verifier should then be sharpened, split into independent skeptics, or deleted stays out of scope: this makes that decision answerable with evidence instead of intuition
74Implement: orchestrator mode — a strong model oversees, worker subagents write the code--implement-worker-model (env: PLANWERK_IMPLEMENT_WORKER_MODEL) switches the implement session into orchestrator mode: the session itself (on --implement-model, e.g. fable) never edits a file — it delegates every work package to an inline-defined implementer subagent running on the worker model (e.g. opus) at --implement-worker-effort (default xhigh, env: PLANWERK_IMPLEMENT_WORKER_EFFORT), then verifies each delivered package against the actual diff (git log/diff + foreground tests) and dispatches follow-up delegations for every gap before moving to the next packageThe strongest model pays off in oversight — keeping the whole issue in view, catching a worker's gaps, deciding what to do about them — while the code-writing itself is well served by the cheaper Opus tier; the split also keeps the orchestrator's context lean, because the workers' tool calls never enter it. The subagent is defined via Claude Code's --agents CLI flag (a new withAgents runner helper both runner paths share), NOT via .claude/agents/ files: the flag survives the hermetic --setting-sources project isolation and leaves the target checkout untouched. Subagents inherit the parent's auto permission mode by Claude Code's own rules, so the classifier keeps vetting worker actions. Delegation is mandated strictly sequential — workers commit on one shared feature branch, so parallel workers would race on the git index. The mode is opt-in (empty worker model = the historical single-session behavior, byte-for-byte) and degrades gracefully: if the model ignores the delegation instructions and edits directly, the result is functionally the single-session run. Attribution follows the code, not the reviewer: the report footer names the worker model (implementAttributionModel), and the workers' commit trailers carry their exact model id via the shared commit-trailer block in their agent prompt; the exact id is only available when the operator passes one (the envelope resolves only the session's own model). The report contract (heading, terminal STATUS line) stays with the orchestrator, so implementReportStatus, resume, and the PARTIAL/BLOCKED guards are untouched. DefaultClaudeTimeout rose from 15 to 60 minutes because an orchestrated session (plan → n workers → n verifications → report) cannot fit the old budget
75Report shape: lead with the verdict, end with one next actionEvery structured report a human reads (plan, implement in both variants, finalize, fix in both variants, bare address) opens with a lead line — the bare verdict word, an em dash, one sentence of concrete outcome — and, on any verdict but the fully-successful one, carries a single Next: <action> line after the terminal STATUS line. A shared reportShapeBlock(successVerdict) pins the three rules (lead line, one next action, no closers or recap paragraphs), and the skill surface mirrors them: house-style.md gains a "First line, last line" section, the clarify/revisit/meta report phases open with the outcome and end with the handoff, and the fix skill's mirrored report template stays byte-compatible with the command'sAdapted from the ayghri/i-have-adhd output-shaping skill: a reader who reads only a report's first and last lines should know what happened and what to do next — previously the verdict was the last thing a top-down reader reached. The terminal STATUS: line is a parser contract (#38), and every escalation parser (planEscalation, implementReportStatus, fix.parseStatus) is line-anchored on the STATUS: prefix, so the unprefixed lead line is invisible to them by construction: the verdict is duplicated for the human, never moved. Deliberately NOT applied where density is the product — elaborated issue bodies are the spec implement consumes (#31, #40, #62 depend on that density), finding enrichment and the structured-JSON lanes are machine-read, and the orchestrated address emits JSON with no report to shape — so those surfaces keep their existing length rules unchanged
76Honor the target repo's STYLE_GUIDE.md for the documentation the mutating sessions writeThe orchestrated implement, fix, and address sessions (and the pasted bare implement prompt) now discover the documentation style guide the target repository commits and bind every piece of documentation prose the session writes — README and docs pages, CHANGELOG entries, doc comments / docstrings, CLI help text — to it. A new internal/styleguide loader resolves the first of STYLE_GUIDE.md, .planwerk/STYLE_GUIDE.md, docs/STYLE_GUIDE.md, .github/STYLE_GUIDE.md in the checkout (root wins, best-effort, never fatal — the skills loader's posture) and returns only the repo-relative path. A shared styleGuideBlock in components.go renders the citation into the prompt, placed right after the project-skills block, under a "read the file BEFORE writing any documentation prose; where it conflicts with your defaults, the style guide wins" obligation; it is threaded through implement.Context / implement.BareContext / fix.Context / address.Context, loaded next to skills.Load from the checkout dir, and golden-tested. An empty path yields the empty string, so a repo without a style guide leaves every prompt byte-for-byte unchanged.Mirrors decision 63's mechanism for .claude/skills: the hermetic sessions run inside the checkout, so the model could stumble over a STYLE_GUIDE.md — but whether the committed documentation rules are actually applied was left to chance, which the prompt-design doctrine's predictability virtue rejects. The prompt carries the path plus an obligation, never the body: the sessions read the guide from their own checkout, so the prompt stays lean, the rules the session follows are always the committed ones, and a long guide cannot crowd out the instructions around it (the same pointer-plus-obligation shape as projectSkillsBlock, deliberately unlike the embedded domainGlossaryBlock, whose CONTEXT.md is capped and small). The block's closing line scopes the guide to documentation STYLE only — repository data, never task instructions — so a checked-in file cannot escalate itself into commands, mirroring the glossary's anti-injection framing. The .planwerk/ fallback follows the glossary's CONTEXT.md / .planwerk/context.md precedence (canonical root location first, tool-specific dir second). Bare fix/address and the read-only passes stay out of scope, as in decision 63 — the read-only passes write no documentation for a style guide to govern.
77The snippet gate demoted imprecise quotes, not just fabricated onesThe quote-or-demote gate (#23) no longer requires a finding's whole code_snippet to appear as one contiguous block; located also passes the finding when any single distinctive line of the snippet resolves in the checkout (distinctiveSnippetLines, in internal/hygiene/snippets.go). A snippet is demoted to the Unverified section only when none of its distinctive lines resolve anywhere — the genuine fabrication signal — or it quotes no code at all. A line counts as distinctive only above minDistinctiveLineLen (16 normalized chars), so boilerplate a fabricated snippet could carry by chance (}, return err, if err != nil {) neither verifies nor demotes a findingThe contiguous whole-snippet match sank the far larger class of grounded findings alongside the fabricated ones: a review model routinely elides a line with ..., quotes non-adjacent lines together, or reconstructs one line slightly wrong, and any of those made the whole normalized snippet miss as a substring, burying real evidence in the Unverified bin — a live implement self-review demoted all 13 of its findings this way. The gate's stated purpose (#23) is narrow: catch a snippet quoting code that is not actually there. One real distinctive line is enough to prove a snippet is not fabricated, so requiring only that keeps genuine fabrications demoted while sparing the grounded-but-imprecise majority. Whether such a finding's conclusion is nonetheless wrong stays the claim verifier's job (#55), which already re-checks every BLOCKING/CRITICAL finding after this gate; loosening the quote check does not loosen that, and the gate still records what it examined and demoted (#73) so the looser pass rate stays visible. The contiguous whole-snippet check is kept as the cheap fast path and still carries the prose-comment case a single distinctive line cannot
78Session completion: nudge the same session before aborting, and preserve its account for the resumeEvery report-bearing mutating session (implement, simplify-apply, review-apply, finalize, fix) now runs with a pinned session id (--session-id); when its output fails the terminal-report gate (report.TerminalStatus + heading), runWithCompletionNudge resumes the same CLI session (--resume) with a completion nudge — backgrounded commands died with the turn; re-run the outstanding verification in the foreground, commit, emit the report or an honest escalation — bounded at two nudges. When implement still ends without a report, its final output is posted as a ## Progress Note (issue #N) comment (own heading + attribution marker, so it is never mistaken for a plan or report); prepareResume feeds the most recent progress note or posted (PARTIAL/escalated) implementation report into ResumeContext.PriorReport, which the resume prompt embeds as the map of what already passed verification. The prompts additionally name the legal route past the Bash tool's foreground time limit: background the command and poll its output within the same turnA one-shot session that yields mid-work — most commonly "waiting" for a backgrounded test run whose notification can never arrive in a turn that never comes — used to hard-abort the run and discard a session that was often one test run short of done; the next run then re-derived every already-verified fact from scratch. Resuming the very session that did the work is the cheapest possible repair (its context is intact, so the nudge turn costs a few thousand tokens instead of a full re-run), and the progress note makes the surviving knowledge durable across processes and machines, matching the existing "every run leaves its course of events on the issue" convention. The nudge only repairs the yielded shape: genuine run errors (rate limit, max turns) propagate unchanged since a follow-up turn cannot fix an API failure, and a failed nudge returns the incomplete output so the orchestrator's own gate still persists it. --no-resume (nothing reads the note back) and --no-report-comment (no run artifacts on the issue) each suppress the note; address and rebase stay un-nudged because their contracts are schema-validated JSON (with repair recovery) and free-form text respectively
79Token efficiency: measure per pass, then stop paying for what a pass was told to ignoreFour changes, in the order the evidence allows. (1) Per-pass accounting: addUsage folds each call's tokens into a per-label accumulator as well as the per-run one, so LogUsageSummary and the report data block (usage.passes) attribute a run's cost to the pass that spent it — a run fans out over a dozen sessions, and the totals alone cannot say which one to act on. (2) A scoped pattern catalog: DefaultMaxPatternsInPrompt is 0, so every prompt that injects the review-pattern catalog injected all of it — for a Go repository 47 patterns, ~123 KB, ~31k tokens, six times over on a specialist fan-out. Each Specialist now declares the review areas its domain reads on and is handed only those, an exclusion of what is plainly outside the domain rather than a mapping stretched to force a saving; the fan-out's six prompts drop from 762 KB to 272 KB. Nothing is lost from a run, because the primary review and adversarial passes still carry the full catalog, and an unmatched area falls back to it. The capture pass, whose only use of the catalog is deduplication, gets the one-line index (name, area, severity, detection hint) instead of the bodies. (3) A narrower re-review: implement's review loop reviewed the whole branch diff on every round; from the second round on it reviews only what the previous round's fixes changed (git diff <pre-fix commit>, recorded before the editing session, since the fold-rebase rewrites the branch) — decision #72's "bound the fan-out to the first round" applied one level down. (4) A finder tier: --finder-model / --finder-effort recolor the read-only finding passes, defaulting to inherit so nothing changes until make eval says a cheaper tier holds precision and recall.The tool's cost had grown with its passes and nobody could say where it went. The ordering is the point: measurement first, then the changes that are quality-neutral by construction, then the one that is not. (2) and (3) remove tokens a pass was already instructed to ignore — a specialist told to review ONLY its domain does not need the other five domains' patterns, and a round whose question is "did these fixes hold" does not need the diff whose review produced them. That is the sprawl failure mode from the prompt-design doctrine, which holds that every added sentence dilutes the ones already there, so both are expected to read as slightly better prompts, not merely cheaper ones. (4) is different in kind and is deliberately left inert: a finder's output is the product, so a cheaper tier can cost recall, and the doctrine's rule for a change justified by reasoning alone is to instrument the suspicion rather than act on it. The eval harness (#59) is that instrument, and the per-pass numbers from (1) say what flipping the default would be worth. Two candidates from the same analysis were not taken: an opt-in diff-size gate for the fan-out (an off-by-default knob saves nothing, and on-by-default it trades recall for tokens on exactly the small diffs where a specialist is cheapest), and cross-session prompt caching of the shared catalog prefix (the fan-out runs concurrently, so every session would write the cache and none would read it; the cache_read_input_tokens already on the usage record is what would settle it)
80The documentation the tool writes reads machine-writtenA curated de-AI writing catalog ships at plugins/planwerk/shared/humanizer.md (adapted from blader/humanizer v2.9.1, MIT, itself based on Wikipedia's "Signs of AI writing"), with one source and three surfaces: the new /planwerk:humanize skill applies the full catalog to existing documents, house-style.md carries a compact "Signs of AI writing" section every artifact-writing skill already reads, and the prompt side carries the same compact form — bannedVocabularyLine() absorbs the humanizer vocabulary, and the shared aiWritingTellsBullets extend proseStyleBlock() (narrative builders) and the new docProseBlock() (implement in both variants, the implementer worker subagent, fix, address). styleGuideBlock names the new rules among what a committed STYLE_GUIDE.md overrides, and TestSharedHumanizerDocMatchesPromptBlocks / TestSharedHumanizerDocMatchesBannedVocabulary pin the three surfaces togetherThe prose the sessions generate carried the tells that mark text as AI-written — em dashes, inflated significance, forced triads, the giveaway vocabulary — and the existing rules covered only a fraction (a 17-word vocabulary ban and four prose bullets, absent from the mutating sessions entirely, which write the documentation that outlives the run in someone else's repository). The 8,000-word upstream catalog is deliberately NOT injected into prompts: per #79's sprawl rule the prompts carry only the distilled bullets, and the full catalog ships as a skill-side document read on demand — the repair path (humanize, for prose that predates the rules or arrived from elsewhere) pays for the whole catalog, the generation path pays for a dozen lines. Upstream's voice-injection and chatbot-correspondence sections are dropped in curation: planwerk artifacts are technical reference prose, where neutral and plain is the correct human voice. Precedence is explicit where the rules meet older contracts: any format-mandated em dash — reportShapeBlock's lead line (#75), a template's field separator — is exempted inside the bullet itself, the vocabulary ban carves out identifiers, quoted code, and API names cited verbatim, and a target repo's committed STYLE_GUIDE.md outranks the catalog on both surfaces (#76's "the repo wins", now applied to the tool's own defaults too). Bare fix and bare address stay out of scope, mirroring #63/#76; the worker subagent is IN scope because its static prompt writes doc comments the orchestrator never rewrites — and since decision 74 keeps that prompt static, the worker carries its own read-the-style-guide clause instead of the path-bearing styleGuideBlock. The sync tests are the commits.md/commitTrailerBlock arrangement (#71) applied to prose: duplication across a skill surface and a prompt surface, bounded mechanically rather than by discipline
81Work in one repository that implies work in anotherA repository declares its counterpart repositories in .planwerk/related-repos.md, one ## heading per owner/repo with a When: condition. draft and meta file the counterpart issue at draft depth in the repository that owns it; elaborate files one the draft missed, since the repository walk is where the affected interfaces first become concrete. The two issues are wired with GitHub's native relationships, which work across repositories: the counterpart is blocked by the originating issue, and is linked as a sub-issue when the origin is a Meta Issue. ship reports a Sub Issue outside the Meta Issue's repository and does not drive it.A pull request cannot span repositories, so a client-side counterpart is not deferred work, it is work that was never in this delivery — issue-format.md says so explicitly, or elaborate's own single-delivery check would reject every cross-repo plan it writes (see #62). The map, rather than org-wide discovery, keeps the tool from inventing counterparts: a generated SDK needs no issue, and is simply left out. Full cross-repo shipping was rejected as a separate project: internal/ship/schedule.go keys its DAG by bare issue number, and numbers are unique only per repository, so driving foreign Sub Issues needs a (repo, number) key and a per-repository pull-request pipeline. Detecting and skipping them needs neither, and closes the real hazard — before this change internal/github/relations.go stamped the queried repository onto every returned node, so a foreign Sub Issue would have been cloned, implemented and merged in the Meta Issue's repository against whatever issue shared its number. Pinned by the foreign-repo cases in internal/ship/crossrepo_test.go and the marker test in internal/claude/shared_crossrepo_test.go
82Settling a Meta Issue's own deferred decisions is a skill, not a side effect of meta or clarifyA new /planwerk:decide skill locates the decisions or spike block a Meta Issue defers when it is split — a checklist item pairs an unverified recommendation with something a spike must still check — verifies each item against the repository and whatever source it names (HEAD, a pinned reference the repo's own config commits to, vendored docs, a closed sibling's already-settled outcome), sorts every item into Confirmed / Corrected / Contested / Broken, and puts only the Contested and Broken items to the author, largest blast radius first, citing what a sibling Sub Issue already assumes. The outcome is written in exactly two places: the Meta Issue's own decisions block, where it asked to be recorded, and every open sibling whose body cites the decision by its code, corrected only where the settled answer differs from what the sibling assumed. issue-format.md's attribution table gains a Decided by row; a Sub Issue it corrects always gets the verb swap, but the Meta Issue is signed only when it already carried a footermeta already treats a decisions/spike section as an ordinary work package and files it as a draft-depth Sub Issue whose job is to verify and record the outcomes, but nothing closed the loop once that verification actually happened — the recommendation stayed a recommendation forever, and every sibling kept citing it as though it had been checked. clarify was the nearest existing skill and was rejected as the vehicle: it exists to unblock a stalled plan, keys off a posted STATUS: NEEDS_CONTEXT comment, and writes into exactly one issue's body, while a decisions block is not a plan, blocks nothing by itself, and this skill's whole point is fanning one settled fact out across a Meta Issue and an arbitrary number of siblings. revisit's minimal-diff discipline and cross-issue neighborhood read (github.md's neighborhood query, "change only the line a check forced") are reused directly, and the four-bucket sort is clarify's three-bucket sort (Answered / Decision / Beyond the repository) recut for a spike's vocabulary: a "Recommended:" clause is exactly as much of a guess as an unchecked OPEN QUESTION marker, and neither is trusted before the files behind it are opened. The Meta Issue's footer is left alone when none exists because many Meta Issues are hand-authored and never ran through a planwerk skill; signing the whole document for an edit confined to one section would misattribute the rest of it
83A reference that leaves its repository carries the repository with itEvery cross-repository reference the skills and prompts write is fully qualified: owner/repo#123 for an issue or pull request, owner/repo@<sha> for a commit, a URL for anything else. github.md states the rule once under "Referring to another repository" and commits.md, cross-repo.md, issue-format.md (the Split from owner/repo#N footer) and the clarify / decide / elaborate / humanize / meta / revisit skills apply it where they emit a reference. On the Go side internal/github.LinkedPR now carries its own Owner/Name from repository { nameWithOwner }, writeLinkedPRs labels a foreign linked PR owner/repo#N the way promptIssueRef already labelled foreign Sub Issues, and ship's findOpenPR refuses a linked PR from another repositoryDecision 81 qualified Sub Issue references and stopped there, which left the rule half-stated: a bare #N or a bare sha does not fail when it crosses a repository boundary, it silently resolves against whatever the reader is browsing, so the same defect recurs for every object type nobody thought to name. Pull requests were the live instance — GitHub's closing keywords work across repositories, so a Sub Issue's closedByPullRequestsReferences can name a PR in a third repository, and findOpenPR returned a bare number that the caller applies to the repository it drives: the colliding-child hazard 81 fixed for issues, one level down at pull requests, ending in an unrelated PR marked ready and merged. The rule is stated by where the text lands rather than where it was read, because the skills write into repositories other than the one they run in (decide correcting a foreign sibling, meta filing a counterpart), and the naive reading — qualify what is not local to me — inverts exactly there. humanize gets an explicit carve-out since its whole job is shortening prose, and a reference is the one thing it may not shorten
84A published fold invalidates the SHAs the pull request body quotesThe fold's doctrine now ends with the repair it owes. commits.md gains "Repairing what the rewritten SHAs left behind": read the body with gh pr view --json body, test every hex token of seven characters or more with git merge-base --is-ancestor <sha> HEAD, map each replaced SHA to its successor by subject, and write back only those tokens through gh pr edit --body-file. /planwerk:fix points at it from Phase 6 and from its pre-push checklist; BuildFixPrompt and BuildBareFixPrompt carry it as a numbered step in their fixup variants only, gain a PR body: report line, and now say that "folded into <sha>" means the SHA the commit carries after the fold; TestSharedCommitsDocMatchesFoldDiscipline gains gh pr edit and git merge-base --is-ancestor as markersThe fold is the default landing (#71) and the description finalize writes walks the change set in commit order, so the two meet on every repaired pull request: the reviewer opens the body and follows a SHA to nothing, with no signal that it ever pointed anywhere. Reachability is the test rather than a diff of the log because it separates three cases that read alike in prose — a SHA that survived (exit 0, and this covers every commit already on the base branch), one the fold replaced (exit 1), and a token that is no commit of this repository at all (an error): a checksum, a hash quoted from a CI log, or a commit elsewhere under #83's owner/repo@<sha> form. Only the middle case is the tool's to correct, and treating the third as stale is how an unrelated repository's commit gets "fixed". The successor is found by subject because autosquash absorbs a fixup into its target and preserves order, so the commit is still there and only its name changed; when no successor exists — the commit was dropped, or two were squashed into one — the reference is left alone and the report says so, since a plausible replacement is worse than a broken one nobody trusts. The edit is confined to the SHA tokens, with the abbreviation length preserved, because the description belongs to the author: what is corrected is a pointer, not their text. Scope follows what actually rewrites a published SHA — --no-fixup rewrites nothing, simplify and review-apply fold before a pull request exists, and address pushes follow-up commits. rebase --push is the one remaining surface that rewrites an open PR's history and is left for its own change
85A lightweight interactive implement is a skill beside the command, not flags on itA new /planwerk:implement skill implements a prepared issue in the author's own checkout: it enters plan mode, grounds a short plan in the code (the change set, the change and verification command behind every Acceptance Criterion, the tests to write), and edits nothing until the author approves that plan; it then implements on an implement/issue-<N>-<slug> branch — the same shape the command uses, so an unattended run can later find it — and opens one draft pull request behind an explicit yes. It deliberately runs none of the pipeline's passes: no simplify, no review-and-fix loop, no verification session — the author approves the plan and reads the diff instead. The completeness discipline carries over unchanged: no PARTIAL, one issue one complete PR (#62), work that cannot complete stops as BLOCKED with the branch local. It never edits the issue body, so issue-format.md's footer-verb table gains no rowSmall bug fixes were being implemented ad hoc, outside every convention — no plan gate, no Acceptance Criteria walk, no Assisted-by trailers, no one-PR discipline — because the full implement pipeline costs more than such a change and skipping it wholesale was the only alternative. A sanctioned light path keeps the doctrine where the ceremony was the reason to abandon it. It is a skill rather than --no-simplify --no-review flags on the command for the same reason the other skills exist: the command is headless and must guess, while the entire value of dropping the passes is a human standing in for them — at the plan gate, at the diff, and at the depth question a draft-depth issue raises. The branch prefix is shared deliberately, so a skill run interrupted halfway is resumable by the unattended command (#61) instead of stranding a half-built branch
86A domain sweep on the planning side, not a question bankelaborate (the command, and since the 2026-09 audit the /planwerk:elaborate skill through shared/domains.md, pinned to the default by TestSharedDomainsDocMatchesDefault) and implement's planning session walk a fixed list of seven domains — data and schema, compatibility and contracts, failure and recovery, observability, security and trust boundaries, performance and cost, operability and configuration — against the change set they just designed. A new internal/domains loader reads the target repo's .planwerk/domains.md override or falls back to an embedded default, and a shared domainSweepBlock in components.go renders it into both promptsThe specialist fan-out (#24) reviews the code that was written; nothing reviews the domain a change silently failed by omission — an absent rollback path, a metric nobody added, a config key an existing deployment never learns about. None of those appear in a diff, so they can only be caught while the plan is still being written, and until now nothing forced a plan to consider a domain the issue itself never named. The sweep is a list of domains with a landing rule, deliberately NOT the 697-question bank the idea is adapted from: six of that bank's domain files are ~25 KB of prompt, the same shape of sprawl decision #79 removed from the specialists, and the prompt-design doctrine requires every line to be earned. Completion is checkable and exhaustive instead: the plan reports a ### Domain Sweep section in which every domain appears exactly once — on a bullet naming the Change Set, Test Plan, or Documentation Plan entry that carries its consequence, or in a closing Not touched: line — and a plan that can place a domain in neither is not PLAN_READY. Sweeping never widens scope: a consequence it surfaces is either work an Acceptance Criterion already requires or it is over-scope, which routes into the existing #32 gate. Only the list is shared between the two builders; each passes its own landing sentence, because the elaboration has no room for such a section — its output is transcribed into the fixed house issue format, so a touched domain lands as an Acceptance Criterion or an Affected Areas entry and the prompt forbids emitting a section the structuring pass would drop. The elaboration's reviewer gate (#30) scores the result as its eighth check, giving the sweep an independent verifier rather than only self-review. Testing and documentation are deliberately absent from the list: the plan already carries a Test Plan and a Documentation Plan, and #31 already forces per-criterion edge-case enumeration. The loader mirrors checklist.Load with the glossary's hardening from #42 (Lstat so a committed symlink is not followed, 64 KB cap, best-effort), and an empty or unreadable override falls back to the default rather than to nothing — an override may replace the sweep, never delete it. For the same reason the block itself falls back to the embedded list rather than to the empty string, unlike every other repo-sourced block: --print-plan-prompt renders before any clone exists, and a caller with no checkout must still see the sweep it will be held to. Adapted from m4vic/socratic (MIT)
87An assumption is its own plan section, not a kind of riskThe plan format gains ### Assumptions between Verification Commands and Risks & Open Questions, and a hard rule separating the three: an assumption is what the planner took as true without opening the file that would prove it; a risk is what can still go wrong when everything it believes is true; an open question is what it could not decide and a human mustThe three were collapsed into ### Risks & Open Questions, which made clarify's hardest job heuristic. Its Phase 1 had to hunt for assumptions standing in for decisions — "a Non-Goal or a Description line that says 'per the author', or defers a choice nobody made" — because no artifact labelled them. A named section turns that hunt into a read. The rule keeps the existing escalation semantics rather than adding a new gate: an assumption the repository could settle is one the planner must go settle instead of recording, and only one the repository cannot settle whose falseness would invalidate the Change Set is additionally recorded under Risks & Open Questions, where the #32 preference for STATUS: NEEDS_CONTEXT already applies. clarify reads the new section as the plan's third question-bearing one and routes it through its existing Answered / Decision / Beyond-the-repository sort, so nothing downstream needed a new bucket. Adapted from the "Assumed (flag if wrong)" line of m4vic/socratic's output contract (MIT)
88The escalation filter gets a positive half: the authority kindsinteraction.md gains "What is actually theirs to decide" — priority and cost, vendor and dependency, audience and market, legal and policy risk, irreversibility — as the complement to "Read the code before you ask about the code"The skills carried only the negative half of the filter, so the positive half was left to shape: clarify defined a Decision as "two or more defensible options, and the choice changes what gets built", which is a test a fork the repository itself settles also passes. The result is over-escalation, and over-escalation is not a safe failure — a skill that hands back a choice it could have made teaches the author their attention is cheap, and the next real question gets skimmed. The list is the kind test the shape test was missing. It lives in interaction.md alone, which every skill already reads in full; clarify's Decision bucket points at it rather than restating it, and elaborate inherits it without an edit — one source per instruction, per the prompt-design doctrine. Adapted from m4vic/socratic's "escalate only authority decisions" rule (MIT)
89A codebase cleanup starts as a survey Meta Issue, not a deletion passA new /planwerk:cleanup skill surveys the checkout for dead code and duplicated code and files the verified findings as one Meta Issue: findings with their evidence, grouped into compact phases sized one pull request each (#62), mechanical deletions ordered before the consolidations they dissolve, with an optional ## Open decisions block in the checkbox shape decide (#82) works on. issue-format.md gains "The survey Meta Issue" and a Surveyed by footer row; meta reads the ## Cleanup phases sections as its work-package enumeration unchanged. The skill never edits code, never installs a detector, and never files Sub IssuesThe survey is the one artifact the skills write whose findings must name files and symbols — dead code has no behavior to describe, which is the exact quantity the draft-depth no-paths rule describes work by — so the exception is scoped rather than the rule dropped: the paths live only in the Meta Issue, pinned to the surveyed commit, while the Sub Issues meta splits stay path-free because elaborate already reads a Sub Issue's parent and re-verifies every finding against the checkout it plans in. One source for the evidence, and meta's draft-depth check needs no exception. The verification doctrine — a finding is a lead until the searches that would refute it (identifier as a word, as a string literal, across build variants, against convention-discovered entry points) come back empty — exists because dead-code removal's failure mode is deleting what reflection or a build tag still reaches; a public surface is never asserted dead from an internal search at all, since deleting a published interface sits on interaction.md's irreversibility list (#88), so those candidates go to the author or into the open-decisions block. Duplication is consolidated only when the sites change together; incidental similarity is dropped, not merged. It is a skill rather than a subcommand because the load-bearing judgments — is this exported surface consumed externally, is this similarity incidental — are exactly the calls a headless run would have to guess, and it is distinct from audit, which applies the review-pattern catalog and prints findings rather than authoring a split-ready Meta Issue
90One implementation per dependency; the narrow per-command interfaces staygithub.Client{} is the one value every command's GitHubClient interface is satisfied by (each interfaces.go keeps a compile-time assertion where its adapter was), githubtest.Fake the one test double, patterns.LoadForRepo the one catalog loader, github.OpenRepo/OpenPR the one checkout opener, github.PostBestEffort the one comment poster, report.TerminalStatus the one STATUS parser, and each cmd file binds its flags onto the command's Options directly instead of a mirror structA command should name what it needs, and the narrow interfaces do that. The adapters, fakes and helpers underneath them were the same code written thirteen and twelve times, and the copies drifted: fix's STATUS parser took the first verdict where the shared one takes the last. With one implementation per dependency, structural typing satisfies every interface, so adding an operation is one method and one fake method rather than an edit in every package (#205)
91Take the savings that cannot change a run's conclusionsSeven changes that leave every finding, verdict and plan identical: the JSON-structuring passes run in an empty working directory with --tools ""; schema repair sends the offending finding rather than the review around it; the audit cache key drops --min-severity; a review entry stores its coverage map; --verify-adversarial is removed (superseding #35); the review loop stops when an apply resolves nothing; the plugin's shared references split so each skill loads what it uses; and four self-contradicting prompts state one rule eachEach item either sends tokens nothing reads, repeats work already done, or contradicts itself, so none of them needs an eval run to prove recall held — which is what separates them from the catalog, finder-tier and structuring-hop levers gated behind make eval. Two were correctness defects wearing a cost disguise: a --coverage-map review rendered the map once and silently dropped it on every later run of the same commit, because the flag was in the cache key and the map was not in the payload; and an audit --min-severity warning re-ran the whole analysis to filter a payload the cache already held unfiltered, because a render-time threshold was keyed as if it shaped the analysis. The structuring tier is the largest saving: every structuring and repair call funnels through one function, so one change gives all of them an empty directory (no project settings, no memory) and no toolset; one fixed directory rather than a temporary one per call, because Claude Code keys session transcripts on the working directory. Schema repair became per-finding once normalizeTranscribedLabels settled the two label violations in Go, leaving the empty title — a property of one finding — as the only violation a model is asked about. --verify-adversarial re-ran the finder the default-on review loop had already run with the specialists beside it, doing distinct work only under --no-review; the loop itself now stops when the apply session's mandated Resolved section comes back empty, since the branch is then unchanged and a further round can only re-report what the applier already declined. The shared-reference split cuts what a skill invocation carries by 14 % overall and 31 % for implement, and a new plugin test fails when a split file is left with no reader. The four prompt fixes cost accuracy rather than tokens: one title carried two severities, the review summary was asked for praise the same prompt forbids, a shared fold step told the review pass it was simplifying, and the elaborate prompt asserted that the repository being planned is a developer CLI — a claim about this tool that travelled into every repository it is pointed at
92An oversized issue body continues in comments, and a plan is written to a budgetTwo numbers and one convention. An elaborated body has a working budget of 40,000 characters: the elaboration prompt states it together with the moves that bring a draft under it, the skill's Phase 5 measures the draft against it with wc -c, and the command's refine loop (--review) treats an over-budget body as a gap the reviewer cannot see, refining a draft the reviewer already passed. GitHub's 65,536-character cap is handled by a split, never a truncation: github.SplitIssueBody cuts a document at the strongest Markdown boundary that fills the part (a ## heading, then ### , then a blank line, never inside a code fence) into the body plus continuation comments marked <!-- planwerk-agent:continued 1/N --> and <!-- planwerk-agent:continuation k/N -->, with the footer kept on the body; PublishIssueBody rewrites existing continuations in place and deletes surplus ones; CompleteIssueBody merges the parts back before implement, elaborate, and prompt read a section, and aborts when an announced part is missing; a document posted as a comment (--post-comment, the skill's comment option) is split into a run of comments the same way, refused on an issue whose body is itself continued because the run's markers would be merged into the body on every later read. The skills follow the same convention by hand, specified in issue-format.md (writing) and github.md (reading). Revised by 93 for the skills: the skill no longer tightens a draft, and splits at the body's 40,000-character limitA 97 KB elaboration failed to write, and what happened next — the session tightening the text under pressure to fit — is exactly where a plan loses decisions. The budget is the primary lever, not the split, for two reasons: the body is injected whole into every planning, implementation, and verification prompt that reads the issue, where 40,000 characters is already ~10,000 tokens and the largest single block; and the largest real body in this repository is 44.7 KB, so a rule at the cap would have changed nothing about size. The split threshold is the cap rather than the budget because a split has costs the budget does not: every reader must reassemble the parts, and the Meta/sibling context fetched over GraphQL sees a sibling's body without its continuations. The cut is order-preserving and lossless rather than section-aware so it serves every depth and every writer, and the merge is load-bearing rather than best-effort for the reason the plan-reuse lookup is (#59's cousin in preparePlan): a plan made from a body missing its Acceptance Criteria is wrong, not shorter. Truncation, which review comments still get (#12), was rejected for bodies because a body is a contract every later command reads
93The skills never shorten a plan to fit; the body's limit is where it continuesFor the skills, decision 92's split threshold was GitHub's cap, and elaborate's refine loop treated a draft over the 40,000-character budget as a gap to tighten. In practice every draft was tightened under the budget and no continuation comment was ever written. The number is now the body's limit for the skills, not a target: issue-format-plan.md keeps the dense-writing rules as rules for writing, not as a lever on a finished draft; elaborate drops the size check from its refine loop and counts the draft once, at write-back; issue-format.md splits a plan over 40,000 characters (a draft-depth or survey body only over the cap) into parts of at most 38,000 characters with the same markers, pointer, and note, so MergeContinuations reads a skill-split body unchanged; revisit, clarify, and decide rewrite by the same rule; github-relations.md reads a continued parent or sibling whole. The Go elaborate command is unchanged: it still writes to the budget, refines to it under --review, and splits at the capTightening under pressure to fit is exactly where a plan loses decisions — the failure decision 92 set out to prevent — and a self-scored refine loop takes the size gap as the cheapest one to close, so for the skill the budget produced the very truncation the split was meant to avoid. The split is lossless and every reader merges it back, so length is settled at write-back, where nothing can be dropped. The costs 92 named remain and grow with the number of splits: the Meta and sibling context the commands fetch over GraphQL sees a continued body without its comments, which is why the limit binds a plan's body only, and a Meta Issue or draft is still split only at the cap
94A resumed run continues from the pass that stopped, not from the implement sessionprepareResume reads the issue's comments after finding the branch. When the most recent session account is an implementation report whose STATUS is DONE or DONE_WITH_CONCERNS, the earlier run finished implementing and stopped in a later pass, so the run skips the implement session (that report stands in for its output, and no second report is posted) and reads the pass reports posted after it: a simplification report means the simplify pass is skipped; each review report counts as a used round of the review budget, and the loop resumes at the next round, scoped to the pre-fix commit the last round recorded in an HTML marker on its comment (<!-- planwerk-agent:review-round k since=<sha> -->) when the checkout still has that commit, branch-wide otherwise; capture, verify, and finalize then run as usual. A failed finalize now persists the branch the way an abort does, so a clone-mode run can be resumed too. A PARTIAL report or a progress note is not complete and still resumes the session with the account fed in, as beforeThe failure this addresses is a usage limit reached in the third review round of a run whose implementation was done: rerunning implement resumed the branch, but paid for a whole implement session to rediscover that nothing was left, then paid the simplify pass and the specialist fan-out again, and reviewed the entire branch instead of the last round's fixes, while the record of all of it was already on the issue. The reports are the run's own account of what it finished, double-keyed by heading plus attribution footer like the plan reuse of decision 59 and the resume account, so reading them is no less reliable than reading the branch. Counting posted review rounds against the budget keeps the cap meaning "rounds on this branch", not "rounds per process". A simplify pass that found nothing posts no report and so runs again on a resume; the pass is cheap and idempotent, and inventing a marker comment for a no-op would put noise on the issue for a saving that cannot change the result
95Read-only sessions switch off every hookEvery read-only session — the analysis passes, the planning session, the finders, and the structuring tier — runs with --settings '{"disableAllHooks":true}'; the mutating sessions (implement, fix, address, rebase, finalize) keep the repository's hooks--setting-sources project (decision 45) loads the .claude/settings.json of the checkout the session runs in, and for a review that checkout is the pull request's head, possibly from a fork. A hook declared there ran as the operator, with the operator's GitHub token in the environment, before the model read a line: a probe on Claude Code 2.1.280 with a SessionStart hook in the checkout confirmed it ran under exactly the flags the runner passes, and that the --settings override stops it. Settings on the command line outrank project settings, so the override needs no other change. The mutating sessions keep the hooks because they work on a branch of the operator's own repository, where formatters and commit guards are part of how the repository expects changes to be made; turning those off would make the agent's commits differ from a human's. The residual is a mutating session on a pull request branch someone else wrote (fix, address), which still loads that branch's hooks
96The review prompts end on their own contract, not a /review tokenThe six finder prompts (review, adversarial, specialist, compliance, simplify-find, verify-implementation) no longer end with a literal /review lineThe token dates from decision 2, when the prompt was prepended to Claude Code's built-in /review command. claude -p recognizes a slash command only at the start of its input, and Claude Code 2.1.280 ships no review command at all, so the line reached the model as plain text. What it still did was point at the Skill tool: in a probe where the pull request under review added its own .claude/skills/review/SKILL.md, the reviewer invoked that skill as its very first action in two runs of two with the token, and in one of two without it, after it had seen the diff. The model recognized the planted instruction as an injection each time, but the token handed any pull request a named entry point into the review for nothing in return: each prompt already ends on its complete output contract
97Text from outside the prompt is data, and the rules a change is judged by come from its baseEvery prompt that embeds text other people write — issue and pull request text, review threads, commit logs, CI output, feature specs, Meta and sibling issues, a reused plan or session account — fences it with fencedData/escapeFence and names it as data in one shared sentence (untrustedDataLine); the address prompt also bounds what a thread may ask for. The review loads its checklist, repo review patterns, and TODO list from origin/<base>, fix and address load the repo patterns and the skills they are obliged to use from there too, and the review flags a change to .planwerk/checklist.md, .planwerk/review_patterns/, or .claude/ as a finding. Issue comments are acted on only when the authenticated user or a maintainer (OWNER, MEMBER, COLLABORATOR) wrote themThe glossary, wiki memory, and rejected-idea blocks were already fenced, escaped, and framed as data, but the text with the most power over a session was not: a review comment went unfenced into an auto-mode session told to "do exactly what the reviewer requested"; CI logs sat in a backtick block a log line could close; a plan comment was recognized by a heading and a footer anyone can type, so any commenter could write the route an unattended implement session adopts, and a forged report could skip the session outright (decision 94). The review read its checklist and patterns from the pull request's own head while its prompt forbade flagging changes under .planwerk/, so a pull request could rewrite the rules it was reviewed by without a finding. Reading those files from the base mirrors what the glossary already did (decision 47's framing, glossary.LoadBodyFromRef), and internal/gitref is the one place that reads content at a ref. The framing sentence does not forbid acting on what an issue asks for, since "run make generate" can be the work itself; it forbids letting the text change how the session works and names the requests that are never part of the work
98Finders report what they find and say how sure they areThe finder prompts stop filtering by conviction: a finding the model is unsure of is reported with the UNVERIFIED: prefix and uncertain confidence instead of dropped; the "unchanged code" suppression covers only problems that predate the change, not breakage the change causes in unchanged callers; the specialists may cite unchanged code a change breaks; the audit reports every defect, not only CRITICAL and BLOCKING ones. The severity ladder defines each level by impact with examples, the finder scope is the three-dot origin/<base>...HEAD diff the review already used, every finding carries a location, the adversarial pass also looks for wrong results on valid input, and the claim verifier can answer unverifiable and refute an absence with the search that came back emptyCurrent models follow a finder's filters literally: told "never on unchanged surrounding context", "state what WILL happen", or "report only CRITICAL or BLOCKING", they find the bug and then do not report it. The pipeline already filters downstream (the Unverified section, --min-severity, --min-confidence, the snippet and claim gates), so a filter in the prompt only loses findings. The old ladder put a security bug in two levels ("security issues" for BLOCKING, "security vulnerabilities" for CRITICAL); the two-dot diff pulled files only the base changed into scope once the base moved; a finding without a location defeated the cross-pass merge and gave the claim verifier nothing to open; in the implement loop, where the main review never runs, no pass was responsible for a plain logic error; and the verifier's evidence rule made "the symbol is not there" impossible to prove while counting every unchecked claim as a confirmation
99Planning prompts state the goal and the bar, not the choreography; a plan that is not one never shipsThe plan prompt drops the coding baseline and the six-step "run these in order" script for one paragraph of goal, grounding, and sweep, keeps the exact output format and the hard rules, and tells the session it is autonomous and that its final message is all that is captured. PLAN_READY means no open question a human must answer. The elaborate prompt calibrates its reader as someone who can read the repository but missed the discussion, and its reviewer scores bands tied to the gaps it lists, counts padding as a gap, and sees the Meta / Sub-Issue context it judges. implement refuses a planning result without the plan heading and a verdict, and elaborate refuses to write an elaboration without a Description or Acceptance CriteriaThe plan runs on Fable 5.1, which Anthropic's migration guidance names as the model most hurt by prompts written for earlier models: step scripts for judgment work and prohibitions against failures it was not going to make reduce its output quality. The script restated the output format and the hard rules around it ("plan every work package" appeared nine times), and the editing baseline ("if you wrote 200 lines and it could be 50, rewrite it") addressed a session that edits nothing. The elaborate reader was "ZERO context … questionable taste", which rewarded narrating code the real readers — a planner and an implementer who both walk the repository — read for themselves, and every character of it is paid again in every later prompt. The reviewer's bands contradicted its own gap rule (8-9 allowed cosmetic gaps, but gaps could be empty only at 10) and its example anchored the pass score with an empty list. Neither session was told it runs unattended, so one that ended on a question produced text that was posted as the plan and reused on the next run as if it were one
100A checkout builds without Go on the hostmake <target> runs inside a toolbox image (tools/toolbox/Dockerfile: Go, golangci-lint, the claude CLI) that the Makefile builds on demand, tagged by a digest over the Dockerfile and its build arguments. TOOLBOX=0, a set CI, or running inside the image (PLANWERK_TOOLBOX) selects the native recipes; the mode is resolved in one conditional that encloses every native target. Go comes from go.mod and golangci-lint from the lint workflow, so the toolbox reads its versions where CI does. build cross-compiles for the host's uname. The image is Debian trixieA build host with only Docker failed at go: not found, and the other targets expect golangci-lint and the claude CLI too. The layout follows plexsphere/plexsphere's toolbox so both repositories behave alike. Unlike plexsphere, whose binaries are server components, planwerk-agent's binary is the tool its developer runs, so a Linux ELF left in a macOS checkout would break the everyday make build; Go cross-compiles a CGO_ENABLED=0 binary for free, so the container compiles for the host instead of documenting TOOLBOX=0 as the macOS workaround. Reading the Go and golangci-lint pins from the files CI reads replaces the drift tests plexsphere needs to keep a second copy equal. Trixie rather than bookworm because bookworm's git 2.39 rejects git checkout --end-of-options <ref>, which internal/patterns runs, so its test failed in a bookworm toolbox and passes on CI's newer git. The claude CLI is in the image so plugin-validate validates rather than skips; eval needs a credential exported into the container because the host's login does not reach it
101Every reasoning tier defaults to Opus, every tier to xhighDefaultPlanModel moves from fable to opus and DefaultStructureEffort from medium to xhigh; DefaultStructureModel stays sonnet. The main, finder, implement, and worker tiers already ran on Opus at xhigh (the finders and --implement-model inherit --claude-model), so a run with no model or effort flags now reasons on one model and one effort from plan to pull request, and every tier, structuring included, thinks at xhigh. The tiers stay separate knobs (--plan-model, --structure-model, --structure-effort, --finder-model, --finder-effort), so the former defaults are one flag awayThe operator's call: one reasoning model and one effort everywhere instead of a split the build decides. The planning tier keeps its own flag because the plan steers the whole implementation, so --plan-model fable stays where a stronger model pays off first. Structuring keeps Sonnet because it transcribes rather than reasons, so the model swap stays that tier's cost lever; xhigh gives a long transcription the budget to carry every finding across, the risk warnOnDroppedFindings exists for, and --structure-effort medium buys the tokens back for a run that needs it. max stays an explicit per-run override, as for every tier
102implement asks before implementing an issue that was never elaboratedAfter the body is fetched and merged with its continuations, an issue without an Acceptance Criteria heading stops the run before cloning with Implement it anyway? (y/N). A no aborts with the elaborate --update-issue invocation; a non-TTY run refuses and names --allow-unelaborated, which skips the question. --dry-run notes the state without asking, the print-prompt modes are unaffected, and ship sets the option itselfA draft-depth issue gives the plan no definition of done and --verify nothing to check, so running the full pipeline on one is a decision for the operator, not a default. Acceptance Criteria is the test because elaborate never writes a plan without it and the house format forbids it at draft depth; the footer verb is not, since revisit, clarify, and decide overwrite it. ship opts out because it runs unattended, nobody is there to answer, and meta files its Sub Issues at draft depth on purpose
103Where does a reported bug get diagnosed?A new /planwerk:diagnose skill ports the reproduce-first loop of the diagnosing-bugs skill in mattpocock/skills (mattpocock/skills@3216582, MIT): a feedback loop that goes red on the reported symptom before any hypothesis, a minimized repro, three to five ranked falsifiable hypotheses shown to the author before any probe runs, [DEBUG-<tag>] probes removed on every exit path, and a root-cause fix behind a regression test watched failing, landed as one pull request from a diagnose/… branch behind an explicit yes. There is no diagnose command, and fix (its prompts and its skill) is unchangedfix already reproduces before it patches, but it starts only from a red pull request (internal/fix/fix.go), so a bug no check caught had no entry point. The doctrine's hypothesis checkpoint and its no-loop stop need a person: an unattended command could end the no-loop case only at NEEDS_CONTEXT, and would duplicate implement's branch, pull request, and resume logic. Meta Issue #113's "slot into existing mechanisms first" rules out a new command. The diagnose/ branch prefix keeps implement's resume (resumeBranchPrefix, internal/github/resume.go) off a diagnosis branch, which is not a half-finished implementation. Upstream's scripts/hitl-loop.template.sh gives way to the conversation, the human-in-the-loop channel of an interactive skill
104Planning sessions read the native dependency edgesThe headless relations query (buildRelationsQuery, internal/github/relations.go) reads each sibling's and child's native blockedBy and blocking edges, up to 50 per connection, and the skills' shared neighborhood query (shared/github-relations.md) reads the siblings' edges. When a deployment rejects the fields, the headless read logs why and falls back to the edge-free query, so the Meta / Sub-Issue section still loads; a read that timed out is not retried. A sibling or child block carries the edges as blocked-by and blocks attributes on its opening tag, where an issue body cannot forge them, (this issue) marks the planned issue, and one rule, worded identically in the prompt and in the elaborate and implement skills and pinned by a drift test, makes the edges decide delivery order. The elaborate relations fingerprint folds the edges in. ship keeps its REST BlockedByIssues readDecision 37 gave the planning session the neighborhood but not its order, so direction was recoverable only from prose, which decision 52 already declined to trust: meta records dependencies natively and ship orders Sub Issues by them. A session that misreads direction plans work a later sibling owns, or builds against an interface that lands only after its own pull request. The edges ride the existing single round trip, and the fallback costs a second call only when the first read fails. ship stays on REST because it fails the run on a dependency read failure by design, which a best-effort read would weaken. Every new prompt line is gated on an edge existing, so a neighborhood without edges keeps its prompt bytes and its cache key
105Where does a comment's rationale live, and where does a test substitute a dependency?In internal/claude, a doc comment states what the identifier does and what a caller must know at that line (ordering constraints, empty or nil behavior, who calls it), and ends with (decision N) when a row here argues the choice. Rationale with no row stays in the comment as its only copy; history is left to git. Tests substitute a session through the unexported Client.sessionFn, which only runSession reads, and a logger through a *slog.Logger parameter, never through a package-level variableA restated argument drifts from the table (decision 15 kept "15 min" after decision 74 raised the timeout to 60), and the comment copy is the one nobody re-reads. Package-level seams forced every test that swapped one to run serially and hid the substitution from the production path, while the scripted-binary harness in runner_exec_test.go covers the real invocation path the seams existed to avoid
106What counts as a change holding in the eval?A comparison needs at least 3 runs per case on each side, the same -thorough and -specialists flags, and the same cases, at least one of which seeds a bug; planwerk-eval -baseline checks all four before the first case runs. A change holds when pooled recall is not below the baseline's and pooled precision and pooled severity accuracy are each at most 10 points below the baseline's; a ratio undefined on either side holds. The cost per run (prompt tokens, output tokens, estimated USD) is printed beside the verdict and never decides it, and a REGRESSED verdict leaves the exit code alone (decision 59). Each seeded finding is credited to every pass named in the confirmed_by of a prediction that matches it, an empty confirmed_by counting as review, and the report lists each pass's recall. A run in which the adversarial pass or a specialist fails is retried once and never scored. The harness writes a go.mod into every eval repo's base commit, so detection tags the repo go and the review loads the 47 patterns a Go repository gets instead of 27. -specialists runs the specialist fan-out, and the corpus has one case per specialist-only domainThe rule decides the three token levers of #204 (#247, #248, #249) and the finder tier decision 79 left on the main tier. The model is stochastic and the corpus seeds ten findings, so on a single run one miss moves recall ten points; three runs per side pool 30 expected findings. Over clean cases alone no check can fail, so such a comparison would always hold. Recall is the product, so it gets no tolerance: a lever that loses a bug the baseline found has not paid for its savings. The primary review catches most seeded bugs on its own, so a specialist that stopped finding anything would leave the aggregate untouched; the per-pass credit and the specialist-domain cases make that regression visible. The pipeline drops a failed pass with only a log line, so a scored run that lost one would show the same regression without a cause. An eval without a go.mod reported a cost and a pattern context that no Go repository sees
107Which sessions carry the pattern bodies, and how do the others read one?Every prompt inlined the full bodies (about 47 patterns and 123 KB for a Go repository). The finder prompts keep them: review, audit, the adversarial pass, and the domain specialists. The thirteen other prompts that use the catalog (plan, implement, simplify apply, review apply, fix, address, the three rebase sessions, elaborate, propose, gap-analysis, review-prepared) receive its index (one line per pattern: file name, name, review area, severity, detection hint; about 13 KB) and a temporary directory outside the checkout that patterns.Materialize writes from the loaded catalog (one file per pattern) and --add-dir opens to the session. The prompt tells the session to read a pattern's file in full before it works in an area the hint covers. The directory is removed when the run ends; fix writes one per iteration. For fix and address it is written from memory after LoadForRepo materializes the base-ref .planwerk/review_patterns/ and discards it, never from the pull request's checkout. --max-patterns caps both forms severity first, a printed prompt (--print-prompt, --print-plan-prompt) carries the bodies because its directory would be gone when it is read, and a failed write logs a warning and the sessions carry the bodies. The corpus scores the finders only and the finders keep the bodies, so this lever (#247) is recorded by the per-pass prompt tokens of one implement run before and after, in the pull request, rather than gated by a corpus verdict (decision 106)A session with the checkout open can read a pattern file of about 2 KB when it needs one and does not need 123 KB up front. A manual session already works this way: the bare prompts (--print-bare-prompt) list each pattern's URL or checkout path and leave the reading to the session, so the form is proven. A directory the tool writes is the only form that is readable for every source at once (embedded, wiki, remote, --patterns) and that a pull request cannot rewrite: a URL reaches only published patterns, and a checkout path reaches only the repository's own, which for fix and address is the branch being repaired
108Which verdict fits a complete implementation whose proof only CI can give, and where does the status contract survive compaction?The implement report's Acceptance Criteria take a fourth status, unproven: the test that proves the criterion is written and committed but runs only in CI (it needs a cluster, a credential, or a CI-only job), and the report names that test and job. A report whose every work package is done and whose remaining proof is a CI run ends DONE_WITH_CONCERNS, never PARTIAL, and its Next: line names the CI job a reviewer must see green; the rationalizations table rebuts the excuse "the honest verdict is PARTIAL". The verdict definitions have one source, implementVerdictDefinitions, rendered under the report's STATUS line in the implement prompt and again in ImplementSystemPrompt, which the implement session receives through --append-system-prompt (runSpec.appendSystemPrompt) on its first turn and on every completion nudge turn. The bare prompt keeps a hand-edited copy with its push-and-PR wording. A SessionStart hook with the compact matcher was rejected: it needs a contract file on disk, an inline --settings hooks object merged with the target repository's own hooks, and a shell on the hook path. When a report still says PARTIAL while its Work Breakdown Coverage lists every package as done, effectiveImplementStatus reads it as DONE_WITH_CONCERNS; the author chose to continue over re-asking the session or stopping the run. The run logs a warning, says so on stdout, posts the report with an orchestrator note above the footer, and continues into the simplify, review, and finalize passes, and latestSessionAccount counts the report as complete on a resume. The posted report keeps the session's STATUS: PARTIAL line, so a release without this rule, or any reader that calls report.TerminalStatus directly, reads it as incomplete and reruns the implement session. A PARTIAL report with mixed, missing, or sentinel coverage aborts as before. The fix contract is unchanged: its DONE_WITH_CONCERNS already covers a fix that could not run locally, and none of its verdicts withholds the push CI needsThe implement run on C5C3/cobaltcore#1120 finished ten work packages in 18 commits, marked 11 criteria "Not run" because their e2e suites run only in CI, and reported PARTIAL with a Next: line asking for the pull request that PARTIAL withholds. The rerun it asked for starts a second session (the first ran 98 minutes and cost 60.24 USD) against the same contract. The contract had no verdict for "every package done, proof only CI can give", and it lived in the first user message, which auto-compaction had summarized away ten minutes before the report was written. Claude Code rebuilds the system prompt from the launch flags after a compaction, while instructions from early in the conversation may be lost, so a definition carried there is present when the report is written. The coverage list is the report's own account of its work packages, and the verify and review passes that follow check the diff either way
109Where is a finding's JSON written?The seven finder passes (the review, the audit, the adversarial pass, the domain specialists, the feature-compliance check, the simplify finder and the implementation verifier) emit the report JSON themselves. Each runs one session with --json-schema carrying schema.FinderOutput (internal/report/schema/finder-output.schema.json), whose finding requires severity, actionability and confidence, each with no empty member, and carries no id; each prompt ends with the shared findingsOutputBlock, and finishReview decodes the session's own output. The transcription session is gone, and with it buildStructurePrompt, normalizeTranscribedLabels and warnOnDroppedFindings. The repair backstop stays on the structure tier: decodeJSONWithRepairSchema repairs malformed JSON, repairInvalidReview repairs an empty title (the one finding rule a schema-valid output can still break), and a final failure saves the raw output and names the file. The structure tier keeps its six analysis-to-structure callers (propose, elaborate, gap-analysis, sync, capture, review-prepared) plus the repair and dedup calls. The coverage map and claim verification keep their prompt-only JSON. The four passes the corpus never runs (audit, compliance, simplify finder, implementation verifier) moved with the three it measures, by the author's choice, so one decode path remains. This supersedes decisions 3 and 56 for the findersEvery finder paid for two sessions where one decides everything: the labels, the snippet, the fix and the summary are settled in the session that reads the code, and the second only copied them, at one call per finder, up to eight per review --thorough --specialists and up to nine in the first round of an implement self-review. The copy was also the one step where a finding could disappear after it was reasoned out: warnOnDroppedFindings existed because a cheaper model transcribing a long review omits findings under token pressure, and it could warn about the drop but not recover it without repeating the expensive session. With the labels required at the wire, the CLI validates the shape before the tool sees it, and the call and the drop go together. A live probe on Claude Code 2.1.284 showed the flag works in a session that keeps its tools: a hermetic -p session with the finder flags (--setting-sources project --strict-mcp-config --settings '{"disableAllHooks":true}' --disallowed-tools Edit Write NotebookEdit) plus --json-schema ran ls through the Bash tool over three turns and returned the validated object in the envelope's structured_output, with permission_denials empty. The evidence is the corpus verdict under decision 106 (planwerk-eval -thorough -specialists -runs 3 on the branch against a baseline recorded on its merge base), pasted in the pull request; on REGRESSED the change does not land and the lever stays where it is (Meta Issue #204)
110How does a session receive the project memory?The four prompts that carry the wiki's project memory (review, audit, the propose analysis, and the implement plan) receive an index of the memory/ pages, one line per page in file-name order: the file name, the title (the first # heading, or the file name without .md), and, after a |, the sentence of the page's optional **Summary**: line. A page without that line is listed under its title only, and no sentence is extracted from its prose. Title and summary are each cut to 300 bytes. patterns.MaterializeMemory writes the page bodies to a temporary directory outside the checkout once per run, after the cache check; the runners open it with --add-dir beside the pattern catalog directory (runSpec.addDirs), and it is removed when the run ends. implement writes it for the planning session, its only reader, and removes it when that session ends, so a run that reuses a posted plan or runs with --no-plan writes none. The index has a budget of 64 KB: the first line that does not fit ends it, the run logs a warning with the number of unlisted pages, and the prompt states that number and tells the session to list the directory. When the directory cannot be written, the run logs a warning and the prompt carries the page bodies in <project-memory> tags, capped at 64 KB in total. A page over 64 KB is skipped with a warning, and the page bodies kept for one run stop at 16 MB in total: the first page that does not fit ends the load with a warning. The capture pass writes the **Summary**: line into the memory pages it proposes. Unlike the pattern bodies (decision 107), the memory bodies do not stay in the finder promptsThe four prompts carried every page concatenated in file-name order up to 64 KB (decision 47), and a page that no longer fit was skipped with a log line as the only notice. A wiki larger than 64 KB therefore lost pages by their position in the alphabet, and a review of a one-line change paid for the same block as a plan for a new command. A memory that covered this repository's decision log (160 KB for 109 decisions) would lose more than half of its pages. A memory larger than 64 KB cannot be carried in any prompt, so a second form that kept the bodies in the finder prompts below a size threshold would stop applying once a wiki grows past it. The summary comes from a line the author writes, because sentence detection in Markdown is a heuristic and a first sentence is often background. The tool reads each page through the guards it already had (regular file, .md, at most 64 KB, no symlink), does not read a memory/ directory that is itself a symlink, and writes a copy. The symlink guard is the reason the wiki clone itself is never opened to a session: a wiki is world-editable, and a *.md symlink in it would be readable through the clone. A directory the tool writes is also one a pull request cannot rewrite. A probe on Claude Code 2.1.286 showed that one --add-dir followed by two directories lets a read-only -p session read a file in each, with permission_denials empty; claude --help documents the flag as --add-dir <directories...>. The eval harness resolves no wiki, so its prompts do not change and a corpus verdict under decision 106 cannot measure this change (#257, Meta Issue #256)
111Which sessions read the project memory, and how does a skill read it?Three more headless commands read it beside the four of decision 110: elaborate, fix, and address take --wiki, --no-wiki, and --wiki-ref, render the memory index into their prompt directly after the review patterns, and open the memory directory with --add-dir (runClaudeMemory for the elaboration, autoMemorySpec for the two sessions that write code). elaborate resolves the wiki before its cache key, which gains the wiki commit only when a wiki resolved, turns the cache off for a run whose wiki loaded without a resolved commit, and writes one directory per run that the elaboration and every refinement turn read; a run served from the cache writes none, and the --review reviewer reads no memory. fix resolves the wiki once per run, on the first iteration that dispatches a session, and writes one directory per iteration beside that iteration's pattern catalog. address resolves it after its early returns and writes one directory per run, shared by every per-thread session. With the wiki enabled all three also load the wiki's review_patterns/ through the existing wiki pattern tier, as the four earlier readers do. The printed prompts of fix and address (--print-prompt, --print-bare-prompt) resolve no wiki and carry no memory. ship forwards its wiki settings to every implement and fix run it drives and its capture settings to every implement run; it never shows the capture confirmation prompt, so it pushes the accepted pages only when the write-back is enabled and --yes is given, and enabled without --yes it logs one warning at start and every run stays propose-only. A new read-only command has two forms: planwerk-agent brain memory <repo-ref> prints a header and one index line per page, without the 64 KB index budget, and planwerk-agent brain memory <repo-ref> <page> prints the page of that file name, matched against the loaded pages and never joined into a path. Both forms see a page only when its file name starts with a letter or a digit and holds nothing but letters, digits, ., _, and -, because a skill passes the name back on a command line, and the command drops control characters other than newline and tab from what it prints. With the wiki off or without memory pages the index form prints nothing and exits 0. Eight plugin skills read the memory through that command, per shared/memory.md: elaborate, implement, fix, revisit, clarify, decide, diagnose, and meta. Three are left out: draft describes an idea without planning it, humanize edits form only, and cleanup surveys code. A session with the plugin and no planwerk-agent on PATH reads no memory. The opt-in keeps the precedence flag, then config file, then environment variableA recorded decision matters most at the plan step, while the plan is written: elaborate writes the plan that implement executes, as a command and as a skill, and both ran without the memory, so a plan that contradicted a decision was caught in review at the earliest, after the code existed. fix and address change code on a branch after the plan was made and read none either. ship processes the most issues per run with the least human attention, and it built each implement run from options that carried no wiki setting, so it neither read the memory nor ran the capture pass, even where .planwerk/config.yaml enabled both. ship never prompts because it is the unattended run: nobody is there to answer, and the write phase refuses a non-TTY run without --yes, so an enabled write-back would be refused in every Sub Issue of a run without a terminal and would wait on a prompt in one with a terminal. One --yes for the whole run is the confirmation an operator can give in advance. A skill never clones the wiki because of the symlink guard and the other page guards of decision 110 (regular file, .md, at most 64 KB, no symlink, a memory/ that is itself a symlink is not read): a wiki is often editable by people who cannot commit, and a skill that read the clone directly would have those guards a second time, as prose a session can skip. The opt-in precedence, the private-wiki authentication, and the clone cache stay in Go for the same reason, so a fallback clone for a missing binary was rejected. brain memory prints the whole index because the budget bounds a prompt every headless session pays for, while a skill asks for the index once. fix resolves lazily so a run whose checks are already green pays for no wiki clone. The eval harness resolves no wiki, so this change has no corpus verdict under decision 106
112How is the project memory filled from a repository's history?planwerk-agent brain bootstrap <repo-ref> reads the history as units and processes them in the order of the default branch. A unit is a closed issue with the merged pull requests that closed it and, when no pull request did, the commit that closed it; a merged pull request that closed no issue, unless a bot opened it; a range of up to 25 adjacent commits that belong to no pull request and closed no issue; or one chunk of at most 16 KiB of a decision document. A unit sorts by its newest commit on the default branch, and a unit without one by the time it was closed or merged. The history is read from the GitHub API through three GraphQL listings and gh issue view / gh pr view, behind the brain.Source interface, whose only adapter is APISource. Each unit gets one analysis session on the main tier (--claude-model, --claude-effort) that proposes pages, and, when it proposes one, a review session on its own tier (--review-model, default fable, and --review-effort, default high) that accepts, revises, or rejects each page against the unit's text and the code at HEAD. Both are read-only and are structured by the structuring tier. No reconcile pass runs over the whole set after the last unit. The pages and the progress live in .planwerk-brain-sync in the working directory: state.json and one file per page, written through a temporary file and a rename, saved after every unit. A run skips every unit the state holds, so a stopped run continues at the unit it stopped at; the record of an issue unit names its pull requests and its closer commit, and the unit runs again when it holds one the record does not name. Every run except a dry run first refreshes the pages from a fresh clone of the wiki: a page the run did not change takes the wiki's text, a locally changed page whose wiki version is unchanged stays, and a page changed on both sides to different texts is marked diverged, is never pushed, and is reported until both hold the same text. The push happens only under --write-wiki, only once no unit remains, and only after a y typed at a terminal: the command has no --yes, and a run whose stdin is not a terminal is refused. Each pushed page carries the provenance marker of its own unit, and a page that came from the wiki keeps the marker it had there. capture.WritePages, the write phase now shared with the capture pass, skips a new page whose path the fresh wiki clone already holdsA repository that adopts the wiki starts with an empty memory, and the capture pass reads only the run that just finished, so every earlier decision stayed in closed issues, review threads, and commit messages. The unit model follows what the history holds: of the first 100 closed issues of this repository, 53 were closed by a pull request, 19 by a commit pushed without one, and 28 by hand, and 84 of 140 merged pull requests close no issue, so units built only from closing pull requests would have covered about half. The order follows the default branch because a later unit reads the pages an earlier one wrote and may correct them, which also rules out parallel units. The API is the source because the local mirror of #186 does not exist yet; the seam lets #186 add a mirror-backed reader without a change to the run. The review is a second model because a page that enters the working set is context for every later unit, so an error in it would spread; reviewing right after the unit catches it before the next unit reads it, which a reconcile pass at the end could not. It has its own tier so the two models can be chosen separately: the analysis runs for every unit, the review only for a unit that proposes a page. About 190 units on this repository make a run that will be interrupted, by the operator or by a usage limit, and a run that started over would pay for every unit again, so each unit is a checkpoint. The state lives in the working directory because the operator reads and edits the pages before the push, and it ignores itself in git. A diverged page is never pushed because the local page was written against a wiki version that is gone, and a push would replace someone's edit without a merge. The confirmation cannot be skipped because the history holds text from everyone who could open an issue or comment on one: the pages are distilled from untrusted text, and a person who has seen the list is the only gate before they become context for every later session. The guard on new pages closes the remaining gap: a page the run holds as new was never read from the wiki, so writing it over a page that appeared there in the meantime would replace text nobody on the run saw. The eval harness resolves no wiki, so this change has no corpus verdict under decision 106 (#259, Meta Issue #256)
113How is a repository's knowledge on GitHub mirrored to disk?planwerk-agent brain sync <repo-ref> keeps a mirror under <user cache dir>/planwerk-agent/brain/<owner>/<name>, with owner and name in lowercase and . and .. refused as either: issues/<n>.md and pulls/<n>.md with the whole conversation, history.jsonl with one line per commit of the default branch, wiki/ as a full clone, and state.json. The directories the command creates have mode 0700 and its files 0600, and every file is written through a temporary file and a rename. An item file is a YAML frontmatter (format: 1, every key always written), the title as a heading, and blocks. A block is a marker line <!-- planwerk-agent:mirror <kind> <json> -->, whose JSON carries the block's fields and the number of lines of its text, followed by the text unchanged. A reader takes that many lines and never scans a text for markers, and encoding/json escapes <, >, and &, so mirrored text cannot open or close a block. Issues and pull requests share one cursor, the newest handled updated_at: a run lists the items updated at or after it (REST, sort=updated, ascending, since inclusive) and fetches each item whose file is missing or older with one GraphQL query. The history is read from the newest commit down until the commits read and the mirrored ones add up to the commit count of the branch, and replaced when the default branch no longer holds the last mirrored commit. The read does not stop at that commit, because a merge of an older branch puts commits behind it. The wiki is synced by a fresh clone that replaces the previous one once it succeeded, and a clone that fails is a warning. The three parts run as adapters (items, history, wiki) behind mirror.Adapter: the state is saved after each, one that fails does not stop the next, and only a run without a failure sets synced_at. The time of the last sync lives in state.json, not in the files. The files hold GitHub's text unchanged, and no spawned session is given the mirror directory or a prompt line about it; both are the author's choices. brain bootstrap reads the mirror through brain.MirrorSource only with --source mirror: the default stays api, and a bootstrap run never syncs the mirror. The command starts no Claude session and changes no prompt, so the eval of decision 106 has nothing to compareIssue and pull request threads were reachable only through one network round-trip per item, which made them slow to consult and impossible to search as a whole, and brain bootstrap (decision 112) lists the whole history from the API on every run and then fetches each thread one item at a time, 262 items for this repository on 2026-10-02. One cursor is enough because an item's updated_at rises when a comment on it is edited and when a review is submitted. Measured on 2026-10-02: of 194 edited comments on the 300 most recently updated issues of cli/cli, golang/go, and kubernetes/kubernetes, none was edited after its issue's updated_at and 9 were edited in the same second, and two review submissions in cli/cli carried the pull request's updated_at. GitHub lists reviews by no update time, so a cursor per comment stream would not cover them either. A fresh wiki clone gives the result of a pull, keeps the whole git history with every page change and deletion, and goes through the clone path that is already tested (fetchRemote, through patterns.MirrorWiki). A sync time in each file would change every file on a rebuild. With it in state.json, two full runs over the same state of GitHub write the same bytes, so "deleting the mirror loses nothing" can be tested. Secrets are redacted where text enters a prompt, as unitItems does for every body a Source returns (decision 97), not when the mirror is written: the mirror stays a faithful copy, and a changed redaction rule needs no rebuild. The cost is a cache directory that can hold a secret someone pasted into a comment, which the file modes keep to the user. For the same reason a run without a user cache directory stops, where the pattern cache falls back to the temp directory: another local user can create a path there in advance, and the modes of a directory that user owns protect nothing. Each temporary file gets a name of its own, so no file planted under a known name is written through. No later write reuses that name, so a run removes the temporary files an interrupted run left behind once they are more than one hour old. A mirror that is behind would shorten the history a bootstrap run sees without notice, so the mirror is read only on request, and one that is missing or that no sync has finished stops the run with the command that fixes it. No listing reports a deletion, so a routine run removes nothing: a deleted comment leaves the mirror when its item is fetched again, and a deleted or transferred item on --full. Search over the files and any way for a session to see them are left to #187
114How is the mirror searched, and which sessions may search it?planwerk-agent brain search <repo-ref> <query> ranks the mirror of decision 113 by keyword (BM25) and prints the best-matching issues, pull requests, and wiki pages. It keeps a SQLite FTS5 index in index.sqlite in the mirror directory, through the pure-Go driver modernc.org/sqlite. A version row holds the index version and a fingerprint of the redaction rules (redact.Fingerprint); an index with another value is emptied and rebuilt inside its file, in the transaction that writes the new row, and a file that is no database or a damaged one is removed and rebuilt once. A lock another run holds past 10 seconds is an error and removes nothing. Every run brings the index up to date with the files by their size and modification time. The unit of ranking is a block: an item's title, its body, a comment, a review, a review thread's diff hunk, a thread comment, a commit message, or a section of a wiki page. A hit is one item with its best block. The words of a query are joined with OR, every term is quoted before it reaches the engine, * after a word matches its beginning, double quotes hold a phrase, and nothing is stemmed. The index holds redacted text, and the mirror files stay verbatim. A hit names its file relative to the mirror directory and carries an id, and --show <id> prints that block in full from the index, with a fixed prefix before every line of its text. The output never holds the query, no error holds an argument, and the text form drops control characters. The command loads no .planwerk/config.yaml. The command never syncs: it stops when the mirror is missing or no sync of it has finished. Five read-only sessions may run it: the review, the audit, the propose analysis, the implement planning session, and the elaboration. fix and address may not. The session surface has its own opt-in, off by default, resolved in the order --no-brain, --brain, brain.enabled in .planwerk/config.yaml, PLANWERK_BRAIN. A run that opted in resolves a search.Surface once: it needs a finished mirror and a binary path made of letters, digits, and _./+@-, it builds the index, and otherwise it logs a warning and runs without the search. The session's argv gains one permission rule after WebSearch and WebFetch, pinned to the repository, Bash(<binary> brain search <owner>/<name>:*), and its prompt gains the Project History Search block directly after the project memory. The four runs that cache a result fold the mirror's revision (the items cursor and the wiki commit) into the key. With the brain off, every prompt, argv, and cache key is unchanged. The driver, the five sessions, and the separate opt-in are the author's choicesThe mirror was searchable with grep only, which finds exact tokens and returns its matches in file order. Issue and review prose paraphrases: the thread that settled a question rarely uses the words the question is asked with later. So the words are joined with OR and BM25 ranks a block that holds more of them, and rarer ones, higher; requiring every word would miss the thread that used one other word. Ranking blocks and reporting one item per hit keeps a long thread from filling the result. The release build is cgo-free (CGO_ENABLED=0 in the Makefile and in .goreleaser.yml), so a cgo driver is out. A probe on 2026-10-02 ran FTS5 with bm25() and snippet() under modernc.org/sqlite v1.60.1 for the six release targets, and a stripped probe binary was 4.8 MB larger than one without the driver, against 6.8 MB for planwerk-agent. An excerpt cut from unredacted text can cut through a secret (a PEM block, a JWT), and no redaction pattern matches the remainder, so the text is redacted before it is indexed, and the fingerprint rebuilds the index when a rule changes. Decision 113 keeps the mirror directory from every session, because the files hold GitHub's text unredacted. A session therefore gets a command and no directory, and reads a block through --show, redacted. The query stays out of the output so that a session steered into putting file content into a query gets nothing back. The errors hold no flag value and no block id for the same reason: the session's shell can expand a variable or a file name into any argument. A block's text is whatever its author wrote, so the prefix keeps a comment from ending in lines that read like the head of another block with another author. A session runs the command in the checkout under review, so a config file in that checkout must not be able to stop the search. Two runs can open the index at once, so no run removes a file another one may be writing unless the file is no usable database. Quoting every term keeps a question with punctuation from failing as a syntax error. No stemmer covers German and English at once. A sync is a network operation of minutes, so neither the command nor a run starts one, and the output and the prompt block name the time of the last sync. fix and address edit, commit, and push, and a free-text query would pull any commenter's text into such a session. The rule was probed on Claude Code 2.1.287 with the flags of a read-only session: without a rule the Bash call is denied, and with the rule the command runs with flags after the repository and with a quoted query. A line that chains a second command, one that holds a command substitution, and one that names another repository are denied. A path with a space or a backslash would need quoting the rule cannot carry, so such a binary, and every Windows one, gets no search. The revision is in the cache key because a cached result came from a session that could search the mirror at that revision. No fence can wrap what a tool call returns, so the prompt block names the output as untrusted data (decision 97). The eval of decision 106 has nothing to compare: the default is off, and the existing golden files are unchanged