Commit Graph

18 Commits

Author SHA1 Message Date
Tom Boucher
aa7697fe97 fix(#2997): include phase_id_convention in resolved config (#3098)
* fix(#2997): include phase_id_convention in resolved config

_baseConfig in config-loader.cts is an explicit allowlist of keys copied
from parsed into the resolved config. phase_id_convention was in
VALID_CONFIG_KEYS (manifest line 83) and survived the unknown-key filter,
but was never copied into _baseConfig — silently dropped on a clean read.
The milestone-prefix validation check could only be activated via the
ROADMAP frontmatter fallback, not the documented project-config surface.

Added phase_id_convention: get('phase_id_convention') ?? null to _baseConfig.
3 tests: survives resolution, null round-trips, absent resolves to null.

* chore(#2997): backfill changeset PR number 3098

---------

Co-authored-by: sim <sim@local>
2026-08-05 21:26:38 -04:00
Tom Boucher
9a76ca6783 fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter

extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.

Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.

The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.

Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (3eb1cede2), making it
the only non-text file under src. file(1) reported it as data and text tools
silently skipped it, defeating the audit rule that says to search the authored
source; tsc passed because the bytes sat inside a comment, so no gate caught it.
It is live on next.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): pin unterminated-frontmatter detection and its negative space

Covers the discriminator on both sides. The positive rows are the issue's own
repro (LF and CRLF) plus the key-count boundary 0/1/2 around the ">= 1 parsed
key" threshold. The negative rows are the documents that reach the same branch
and must stay silent -- above all a Markdown thematic break at byte 0, which is
how this class of check has previously shipped a false positive on valid
Markdown.

Deduplication is tested on both halves of the composite key: a repeat of the
same (path, cause) is suppressed, a genuine second failure in a different file
is not, and a Windows and POSIX spelling of one path resolve to a single key.
The reset seam is asserted to actually clear -- #2674 is the precedent where a
reset that silently failed to clear made every later dedup assertion a vacuous
pass, and the cases only passed because each happened to pick an unused key, so
every case here uses a path unique to itself.

Assertions are on typed surfaces throughout -- the frozen reason enum and the
dedup-set size -- never on diagnostic prose. The one CLI-level case asserts a
differential between two runs (whether stderr is empty) rather than matching a
message, and is the wired user-reachable surface for this fix. Stream failure is
injected by overriding process.stderr.write and restoring it, never chmod 0o000,
which root bypasses.

Two properties guard the ~50 call sites of the changed function: the new
optional path argument is inert with respect to the parsed value, and LF/CRLF
spellings of a document still parse identically.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): raise the truncation threshold and repair the dedup key

Isolated adversarial review found the one-key discriminator false-positives on
ordinary Markdown: a thematic break above a single labelled line -- `Note:`,
`Author:`, `TODO:`, `See:` -- parses as exactly one key and was reported as
corruption, which is the precise failure the design claimed to prevent and the
changeset promised was fixed. The threshold is now two keys. A file truncated
after exactly one key becomes a false negative; that is the same
precision-over-recall direction already taken at zero keys, and every GSD
artefact this guards carries two or more frontmatter keys.

Three dedup-key defects, each of which could silently swallow a real diagnostic:

- Backslash normalization is removed. A backslash is a legal filename character
  on Linux and macOS, so folding it to a forward slash made two genuinely
  different files share one key. Two spellings of one Windows path may now
  report twice; two distinct files can never silence each other. Lost signal is
  the worse failure.
- The key namespaces are tagged so a file literally named like the unnamed
  digest fallback can no longer collide with a path-less caller whose content
  hashes to that digest -- computable for any predictable content, no brute
  force needed.
- The source identity is computed once rather than hashed twice per emission.

Corrects the previous commit. The two NUL bytes in src/config-loader.cts were
NOT in a JSDoc comment as that message claimed; they were deliberate separators
in the live dedup key, and stripping them degraded it to bare concatenation.
They are restored as escape sequences -- byte-identical runtime string, and the
file is text again so grep can see it. The diagnostic script that misled me
indexed a character-offset string with a byte offset.

Also threads sourcePath through the STATE.md and PLAN.md readers so the two
artefacts epic #1879 is actually about name their file rather than reporting
under a content digest.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): correct fixtures and assertions left behind by the review fixes

The previous commit changed two behaviours deliberately and the suite still
encoded the old ones, so gsd-test came back red with six failures across both
lanes -- all of them mine.

Fixtures carrying a single frontmatter key no longer clear the two-key
truncation threshold, so the CLI differential and the two path-less dedup cases
were asserting a diagnostic that is now correctly withheld. They now carry two
keys, which is what a real interrupted write of a GSD artefact looks like.

The Windows/POSIX case asserted that two spellings of one path collapse to a
single key -- the exact folding that was removed because it also collapsed
genuinely distinct POSIX files whose names contain a backslash. Inverted to
assert they now report separately, with the reasoning recorded inline so the
trade is not silently reversed later: mild duplicate noise on one Windows path
is acceptable, a swallowed diagnostic is not.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): name the file at every read site, and report each file once

The diagnostic reached only the four frontmatter CLI verbs, so ~47 of 53 call
sites reported a truncated file under an anonymous content digest instead of
naming it. Since naming the file is the whole point -- it is what an operator
can act on -- that was a gap in the deliverable, not a scoping choice. 43 of 53
sites now pass the resolved path.

Closing it surfaced a defect the original design missed. A single truncated
STATE.md is parsed twice in a normal run: once by the read wrapper, which holds
the path, and again by a pure core downstream, which is handed only the string
and cannot know it. Those two parses keyed separately, so one file produced two
diagnostics -- and wiring more sites made the collision more likely, not less.
Every emission now registers both identities the input could be known by and
checks both before writing, so whichever caller arrives first speaks and the
other is suppressed. Distinct files with distinct content still report
separately, which is the property ADR-1411 actually requires; two files whose
truncated content is byte-identical collapse to one report, which stays the
documented limit.

Ten call sites deliberately keep no path. Two are frontmatter's own round-trip
checks during set and merge, where passing a path would report on every write.
The other eight are the state-transition pure cores, which ADR-1769 defines as
(content, intent, deps) -> newContent with injected I/O; threading a path
through them would contradict that recorded decision, so it is surfaced rather
than taken unilaterally. With the widened key they no longer double-report, and
in the normal flow the named parse runs first, so the file is still named.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): inject the STATE.md path into the transition cores

The six state-transition cores parsed STATE.md frontmatter without knowing
which file it came from, so a truncated STATE.md reached the operator as an
anonymous content digest on exactly the artefact epic #1879 is named for.

ADR-1769 section 3 shapes these as (content, intent, deps) -> newContent with
injected deps, and deps is the seam for precisely this: something the core
cannot derive without doing I/O. It already carries roadmapProvider and a
phase-inventory provider on that basis, each documented as injected rather than
imported so the core stays pure and testable without disk access. A resolved
path is data, not I/O, so an optional sourcePath member extends the established
pattern rather than contradicting it, and every existing stub keeps compiling
because the member is optional.

updateCore and reconcileCurrentPosition take no deps and are left alone. With
the widened dedup key they cannot double-report, and in the normal flow the read
wrapper has already named the file by the time they run.

Also regenerates gsd-core/bin/lib/state-transition.cjs. That artifact is tracked
rather than gitignored, unlike most of its siblings, so leaving it stale would
have shipped a runtime without this change to anyone reading the repo without
building. tsc had skipped the re-emit because its incremental build info still
recorded an emit that had since been reverted, so the stale output survived a
clean build; clearing tsconfig.build.tsbuildinfo forced it. The
compiled-artifact-sync gate is what surfaced the drift and now reports all nine
tracked artifacts matching their source.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): stop the widened dedup key from hiding a second file

The previous commit widened the dedup guard so one file parsed twice -- once by
a read wrapper holding the path, once by a pure core holding only the string --
reported once instead of twice. It did that by checking BOTH keys before
emitting, which silently traded one defect for a worse one: two DIFFERENT files
whose truncated content happened to be byte-identical now collided on the shared
content digest, and the second file's diagnostic was swallowed. That is the
over-coarse keying ADR-1411 explicitly forbids, reintroduced while fixing
something else.

The guard now checks only the key matching what the caller actually knows -- a
named read checks its path key, a path-less read checks its digest key -- while
still recording every key the input could later be identified by. The redundant
path-less re-parse of an already-named file stays silent, and two distinct files
always both report.

Verified across all six orderings: same file named-then-anonymous reports once;
two different files with identical content report twice; two different files
with different content report twice; the same path twice reports once; two
path-less parses of identical content report once; two path-less parses of
different content report twice.

The suite caught this -- twenty failures, all in the unusable-input tests that
reuse one truncated fixture across different paths. The local probe written
alongside the broken change did not, because it compared two files with
different content and could therefore only confirm the expected behaviour.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): count diagnostics emitted, not identities interned

The suite measured the size of the dedup set as a stand-in for "how many
diagnostics were emitted". That held only while one emission recorded exactly
one key. Once an emission began recording every identity the input could later
be matched by -- a path key and a content key for the same file -- the set grew
by two per write and twenty assertions read 2 where they expected 1.

The production behaviour was correct throughout; the proxy was not. Set size
counts identities, which is an implementation detail of the guard. The
behavioural claim these tests exist to make is how many diagnostics an operator
actually saw, so the module now exposes that directly as an emission counter and
the suite asserts on it. The set-size accessor stays for assertions genuinely
about key shape.

The local probe written alongside the change did not catch this because it
counted process.stderr.write calls -- the right thing -- while the suite counted
set growth. Verification now asserts both and requires them to agree, so a
future divergence between the counter and real writes fails immediately rather
than being discovered a bench run later.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): retire two assertions that outlived the behaviour they described

Both tests encoded assumptions the dedup fix invalidated, and both were caught
by the suite rather than by the probe written alongside the change.

The forged-path case asserted that a file named like the anonymous digest
fallback must not suppress a later path-less report. That premise is gone: an
emission now records every identity its input could be matched by, so ANY named
report of some content silences the anonymous re-parse of that same content --
which is the same-file guard working as intended, and has nothing to do with the
crafted name. The property still worth defending is that a crafted filename can
never silence a real file reported under its own path, so that is what the test
now asserts, with the deliberate suppression documented beside it.

The reset-seam case ended by reading the size of the dedup set and expecting 1.
Set size counts interned identities, not diagnostics written, and one emission
now interns two. It asserts the emission counter for the event and keeps a
weaker set-size check for the interning.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): close the review findings on the discriminator, dry-run and counter

Three orthogonal review passes ran against the final diff. Their findings:

A labelled preamble under a leading rule was still misreported. Raising the key
threshold to two only moved the boundary, because two colon-labelled lines are
as common in ordinary prose as one -- a document opening with a rule over an
Author and a Reviewed-by line, then prose, was called corrupt. Key count alone
cannot separate the two. What does is what follows: a write interrupted part way
through a frontmatter block ends mid-block, so every line of the region is still
frontmatter-shaped, whereas a document merely opening with a rule goes on to
prose. Both conditions are now required, and each closes a false-positive class
the other leaves open. Nested list values and indented continuations stay
frontmatter-shaped, so legitimate truncations are unaffected.

`state rebuild --dry-run` reported a truncated STATE.md anonymously. The write
path is named only because readModifyWriteStateMd parses with the path first;
the dry-run branch reads the file directly and never did. Dry-run is the
read-only mode an operator reaches for first when they suspect corruption, so it
is the one that most needed to name the file. reconcileCurrentPosition takes the
path as an optional argument now and rebuildCore passes it down. That function
was previously left alone on the grounds that a read wrapper always names the
file first -- this is the flow that disproves it.

The emission counter counted write attempts rather than writes, so on a broken
stderr it claimed a diagnostic had reached the operator when nothing had. It is
incremented only after a write that completed, and the broken-stderr test now
asserts the count as well as the return value.

Two documentation defects. The module described a guarantee it does not keep:
one file yields one diagnostic only when the named read comes first. The reverse
ordering emits twice, and that is deliberate -- a path-less caller cannot
identify its file, so suppressing the later named report would also suppress a
genuine second failure in a different file whenever two files share identical
truncated bytes, which ADR-1411 ranks the worse failure. The comment now states
the asymmetric guarantee and a test pins it. Separately, the CONTEXT.md glossary
entry still described backslash normalization that a later commit removed, and
asserted the opposite of what the tests pin; no lint checks prose against code,
so nothing caught it.

Also converts three body-level try/finally blocks to t.after(), per
CONTRIBUTING.md's rule that try/finally belongs only in helpers with no test
context -- the file's own emissionsDuring helper already did this correctly.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#1882): tell the operator what the truncated-frontmatter warning means

A user who has just seen the new warning is acting, not studying, so this lands
in the How-To quadrant beside the other "if you see X" branches in
debug-a-failed-execution, not in reference or explanation. It gives them what
the warning means for this run, three steps to restore the file, and the fact
that the warning changes no return value or exit code.

It also states the case that matters more than the warning itself: silence does
not prove the file is intact. GSD says nothing when the partial block carries
fewer than two fields or reads as prose, because a Markdown document opening
with a horizontal rule is indistinguishable from one of those. A reader chasing
missing metadata needs to know not to treat quiet as clean. Why that threshold
exists is explanation and deliberately stays out of a how-to.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1882): backfill changeset pr number to 2712

* test(#1882): constrain each branch of the frontmatter-shape check

CI's mutation gate came in at 61.56 against a threshold of 62, and the surviving
mutants were concentrated in isFrontmatterShaped -- the function added last, in
response to review, and the only one never given tests of its own. It was
exercised solely through extractFrontmatter, which covers the composite decision
but leaves each branch of the predicate unconstrained: drop the blank-line
filter, or any one of the three shape alternatives, and every existing assertion
still passed.

Four cases now pin the halves independently. A blank line inside an interrupted
block must not disqualify it, which constrains the filter and its comparison. An
unindented list item and an indented folded-scalar continuation each exercise one
shape alternative that no other case reaches on its own -- the folded line is
neither a key nor a list item, so it is the only input that distinguishes the
indented branch. And two keys followed by prose must stay silent, which is the
negative half: it fails if the predicate is ever mutated to accept everything,
and it is the case that proves key count alone was never sufficient.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): register the unusable-input suite with the frontmatter mutation shard

The mutation gate reported an identical 61.56 across two runs whose only
difference was four added tests. That is the tell: the tests were never
executed. The frontmatter shard runs a fixed file list in stryker.config.mjs and
scripts/mutation-matrix.cjs, and tests/unusable-input.test.cjs was in neither, so
the entire suite covering the new unterminated-fence branch was invisible to the
gate while passing perfectly well in the normal run.

So the score was not measuring weak tests, it was measuring absent ones: #1882
added mutants to frontmatter.cjs and no test in the shard covered them. Both
lists gain the file; the config already notes they must stay in sync.

This is a registration ripple a new test file carries when it covers a
mutation-tracked module, alongside the .gitignore, eslint, inventory, glossary
and size-baseline ripples a new module carries. Nothing warned about it, which
is why two runs were spent before the identical score gave it away.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 16:50:12 -04:00
Tom Boucher
3eb1cede26 fix(#1880): distinguish a corrupt config from an absent one (epic #1879 Phase 1) (#2688)
* test(#1880): prove corrupt config is indistinguishable from absent

Failing-first. Encodes the issue's runtime repro: a trailing comma in
.planning/config.json currently yields source:builtin-defaults with
degraded:false - byte-identical to the file not existing - and the user's
entire configuration is silently discarded.

Asserts on the typed surface (CONFIG_REASON, _warnedUnusableConfig) rather
than diagnostic prose, per the ADR-1411 amendment's test-methodology clause
and CONTRIBUTING.md's raw-text-matching rule. IO failure is injected by
monkeypatching fs.readFileSync and restoring in t.after(), never chmod 0o000
(root bypasses mode bits).

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1880): distinguish a corrupt config from an absent one

loadConfigResolved wrapped the read, the JSON.parse and the entire config
build in one try with one catch, so ENOENT, EACCES and SyntaxError all fell
through to the same defaults and the branches returned degraded:false -
actively asserting health over discarded configuration. A single trailing
comma in .planning/config.json silently replaced the user's whole config,
reporting source:builtin-defaults degraded:false, byte-identical to having
no config file at all.

ConfigResolution now carries a machine-readable reason. Genuine absence keeps
degraded:false / not_configured; a file that exists but cannot be used sets
degraded:true with config_unparseable or config_unreadable. The same split
applies to the root config and to ~/.gsd/defaults.json.

Control flow is deliberately unchanged. preflight_check reports cyclomatic
141 / cognitive 196 and 93 dependents on this function, with the guidance
that small edits beat one big one, so faults are CAPTURED at the existing
read sites and stamped onto the returns rather than the try/catch being
restructured.

Also carries the ADR-1411 amendment's wiring clause: loadConfig returns
.config alone to ~51 call sites and would never see the new field, so an
unusable file emits a deduplicated stderr diagnostic keyed on resolved path
plus errno. Without it the reason would be an unreachable field and the user
whose config was discarded would still get no signal - the actual defect.

Registers the config-loader seam in lint-resolution-provenance, which until
now guarded only agent-skills.

Caller audit: ConfigResolution.degraded has exactly one consumer outside this
module, cmdAgentSkills (src/init.cts:2259), which destructures
{config, source, degraded} - adding a field does not break it. Its --json IR
now reports degraded:true for a corrupt config, which is the intended fix and
the one observable behavior change.

Closes #1880

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1880): degrade when any config on the path is unusable, not just the last

Two defects found by isolated adversarial review of the first cut.

BLOCKER: the success-path return did not consult configFault. A corrupt ROOT
config whose workstream override happened to parse returned degraded:false /
reason:resolved - the root's settings silently dropped, which is the exact
failure this issue closes, reappearing for any project using workstreams. The
stderr diagnostic fired, so the out-of-band half worked while the in-band half
reported a clean resolve; a --json consumer saw health.

MAJOR: reason was derived from Object.keys(parsed) - the root+workstream MERGE
- so an empty workstream file inheriting a non-empty root reported resolved
despite carrying no settings. Emptiness is now judged on the file actually
read, snapshotted before normalizeLegacyKeys mutates it.

Also: corrects the ConfigResolution JSDoc, which still described the pre-#1880
degraded contract; adds a fast-check property asserting a PRESENT file is
never reported not_configured whatever its bytes (CONTRIBUTING.md parser
rule); and asserts the literal enum values so the provenance lint's
configured_empty/not_configured markers check real assertions rather than
incidental prose in test titles.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1880): reject valid JSON that is not a config object at the read seam

The fast-check property added in the previous commit failed on both node
lanes: a config.json containing 0, "str", [], null or true is valid JSON, so
it parsed "ok", then threw downstream in normalizeLegacyKeys, and the outer
catch reported not_configured - a PRESENT file reported as absent, which is
precisely the collapse this issue exists to close. The property asserts a
present file is never not_configured, and it caught it.

_readConfigFile now validates shape, not just parseability (ADR-227: check
the semantic shape at a trust boundary, not merely the type). A non-object
JSON document is an unusable config, reported config_unparseable.

Adds named regression cases for each non-object form alongside the property,
so the class is documented and not only randomly sampled.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1880): backfill changeset pr number (pr:0 -> 2688)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2452): record a fetch-time shallow failure instead of crashing

This guard failed CI on ubuntu-24 while passing on ubuntu-22 and
windows-24 for the same commit, and passed on other PRs. Not a flake and not
caused by the change under test - a real fragility in the test.

runnerDiff ran the base fetch OUTSIDE its try and only guarded the diff, so it
assumed the failure mode is always 'fetch succeeds, diff reports no merge
base'. At a shallow boundary that lands short of the merge base, git can
instead fail during the FETCH ('unable to parse commit' - the boundary
commit's parent is not available). Which stage git fails at is version and
transport dependent, so on some runners the error escaped runnerDiff and
crashed the test rather than being recorded as the ok:false the assertions
expect. Both stages mean the same thing for what this guard protects: a
shallow base ref cannot resolve the three-dot diff.

Also drops two assert.match calls against git's stderr prose. 'no merge base'
and 'unable to parse commit' are the same condition reported at different
stages, and CONTRIBUTING prohibits raw text matching on subprocess output.
The typed outcome (ok === false) is the contract; the tests now assert that
plus the presence of a cause.

Found while investigating the red lane on #2688; fixed here per the no-defer
rule rather than filed.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 00:09:13 -04:00
Tom Boucher
c3958018dd docs(#2674): amend ADR-1411 — corrupt is not absent (epic #1879 Phase 0) (#2678)
* docs(#2674): amend adr-1411 with the corrupt-is-not-absent house pattern

ADR-1411 reasons only about a resolution miss. It is silent on input that
is present but not usable, which is how five engine read paths (#1879) could
fold an unusable input into the value meaning 'genuinely absent' without
contradicting an Accepted ADR.

Read together, ADR-1411 and ADR-227 converge and do not license throwing as
the cluster's answer: ADR-227 requires malformed input to be coerced rather
than propagated and carves out only genuinely-fatal fields, while ADR-1411
already permits a fallback provided it is 'a visible value, not a silent
substitution'. The defect in these five sites is therefore not that they fall
back but that they fall back invisibly.

Records the pattern that follows: every current return value is preserved, and
the cause is made visible in-band where the result already carries a
provenance envelope, or out-of-band via a deduplicated stderr diagnostic where
it returns a bare value it cannot extend. Throwing stays confined to ADR-227's
genuinely-fatal carve-out, decided per call. Also names the per-applier caller
audit and the lint-resolution-provenance registry gap.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2674): prove the warning-state reset misses the unknown-key dedup set

The two existing cases in this suite only pass because each picks a key
name no other case reuses, so neither can observe whether the reset the
beforeEach calls actually runs.

Failing-first: asserts the exported _warnedUnknownConfigKeys is empty after
_resetRuntimeWarningCacheForTests(). It is not - the helper clears only
_warnedConfigKeys despite documenting itself as resetting per-process
warning state.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2674): reset the unknown-key dedup set with the runtime warning cache

_resetRuntimeWarningCacheForTests documents itself as resetting per-process
warning state but cleared only _warnedConfigKeys, leaving
_warnedUnknownConfigKeys populated across cases. The suite that exists to
test that set - 'loadConfig - unknown-key warning dedup' - calls the helper
in beforeEach expecting exactly this, so the reset was a silent no-op for
it; both cases passed only because each picked a key name the other never
reused. Any later case reusing a key would have had its warning suppressed
by leaked state.

Found while amending ADR-1411, which names this dedup guard as the pattern
five downstream PRs (#1880-#1884) will adopt - shipping the ADR without the
fix would have propagated the footgun to each of them. Folded in here per
CLAUDE.md's no-defer rule rather than filed.

RED verified on 3c4895841 (test only, no fix): linux-node22 reported
'FAIL tests/config-loader.test.cjs - the documented per-process
warning-state reset must clear the unknown-key dedup set too'.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2674): document src/ in the changeset-lint trigger list

CONTRIBUTING.md presented the Changeset Required trigger list as bin/,
gsd-core/, agents/, commands/, hooks/, sdk/src/ - omitting src/, which
scripts/changeset/lint.cjs has in USER_FACING_PREFIXES. src/ is the
TypeScript source of truth compiled into gsd-core/bin/lib/*.cjs, so it is
the most-edited user-facing path in the repo and the omission sends any
contributor who touches it into a CI failure the doc says cannot happen.

Also documents that the lint reads GITHUB_BASE_REF, which only CI sets, so
running it bare locally reports success without evaluating the branch. This
PR hit exactly that: a local run said ok_fragment_present and CI failed
fail_missing_fragment on the same diff.

Found while opening this PR; folded in per the no-defer rule.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2674): add Fixed changeset for the src/ trigger-list and reset fixes

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#2674): restore the round-2 review corrections to the amendment

These edits were made in response to the second isolated review pass but
never staged: later commits used targeted `git add <file>` for the test and
the source fix, so the two markdown files stayed dirty and shipped nothing.
The branch carried the round-1 text, including the ADR-227 misquote the
reviewer raised as a blocker.

Restores: the unconditional-diagnostic clause (ADR-227's GSD_DEBUG opt-in
was never implemented, so citing it as the precedent was wrong), the dedup
key, #1882 folded into the out-of-band mechanism instead of a fourth
mechanism-less category, the narrowed caller-audit rationale, and the
test-methodology clause.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 20:25:39 -04:00
Tom Boucher
46ba02acde feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs

* fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions

* fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests

* chore(#2630): backfill changeset pr to 2661
2026-07-26 01:42:47 -04:00
Tom Boucher
0bbbca2a46 fix(#2069): forward model_policy, model_profile_overrides, runtime from global defaults (#2442)
* test(#2069): add fail-first regression for global-defaults dropped keys

Adds four failing-first regression cases to tests/defaults-json-fallback.test.cjs:

- model_policy forwarded from ~/.gsd/defaults.json
- model_profile_overrides forwarded from ~/.gsd/defaults.json
- runtime forwarded from ~/.gsd/defaults.json
- parity: model_policy survives identically whether it lives in the global
  defaults or in a project's .planning/config.json

All four fail on unfixed code (Branch D of loadConfigResolved builds
_globalBaseCfg from a whitelist that omits these three keys). The
project-config path at config-loader.cts:602-604 already forwards them,
so the global path should too.

* fix(#2069): forward model_policy, model_profile_overrides, runtime from global defaults

The _globalBaseCfg whitelist in Branch D of loadConfigResolved previously
omitted three keys that the project-config path forwards parsed['…']:

  - runtime
  - model_profile_overrides
  - model_policy

so ~/.gsd/defaults.json silently dropped them. A machine-wide model policy
(or runtime / profile overrides) was honored inside a project (where
.planning/config.json carries it) but ignored for out-of-project runs —
resolve-model fell back to the profile default with no warning.

Adds the three entries to _globalBaseCfg in the same (globalDefaults['…']) || null
shape as the sibling keys and the project-config path, so global defaults
honor them identically.

Regression tests in the prior commit (#2069 fail-first) demonstrate the
fix on the same suite that previously failed.

* test(#2069): extend parity test to all three previously-dropped keys

Code review (subagent) flagged that the parity test only asserted
model_policy shape-parity between global-defaults and project-config paths.
A future regression breaking just runtime or just model_profile_overrides
shape (e.g. someone changing parsed['runtime'] to ?? null in the project
path) would slip a single-key test.

Extends the parity test to assert deepStrictEqual / strictEqual across
all three keys: model_policy, model_profile_overrides, runtime. Same
two-dir setup, three cheap assertions.

* chore(changeset): backfill pr:2442 in .changeset/sturdy-seals-fly.md
2026-07-19 22:44:08 -04:00
Tom Boucher
b2961c3f69 fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022) (#2336)
* test(#2070): fail-first tests for adaptive model_profile and models tier validation

Encodes the three acceptance criteria from #2070 plus the boundary cases the
resolver silently ignores today (non-string values, empty string, mistyped
phase-type key), and pins VALID_TIERS to a catalog-derived set.

Red phase: these fail against current src/ by design.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022)

W004 sourced its profile list from a hand-maintained literal that predated the
adaptive profile, so `"model_profile": "adaptive"` was false-flagged. It now
reads VALID_PROFILES, which model-catalog.cts derives from model-catalog.json.

models.<phase_type> was validated nowhere: the resolver's tier gate silently
drops unknown values, so a typo like `"planning": "opuss"` was an undiagnosable
no-op. A new W022 flags unknown phase-type keys and invalid tier values
(including non-string values, which the same gate also drops).

VALID_TIERS moves from a function-local literal in model-resolver.cts to a
catalog-derived export, so health and the resolver cannot disagree by
construction rather than by parity test. Object.values(adaptiveTierMap) is
['opus','sonnet','haiku'] plus 'inherit' — identical to the previous literal,
so resolution behavior is unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2070): changeset for validate health adaptive profile + W022

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* fix(#2070): close review findings — malformed models, tier-list duplication, changeset gate

Review of the initial fix surfaced three real defects, folded in per the
no-defer rule:

1. verify.cts: the W022 guard skipped a top-level `models` that is present but
   not a plain object (`[]`, `"opus"`, `5`, `true`). The resolver ignores those
   identically, so they were the same undiagnosable no-op #2070 targets — just
   one level up. They now warn; absent/null/{} stay silent.

2. config-loader.cts: RUNTIME_OVERRIDE_TIERS was a second hardcoded copy of the
   tier vocabulary this change had just de-hardcoded elsewhere. It now derives
   from the catalog via ADAPTIVE_TIER_VALUES (no 'inherit' — runtime overrides
   resolve to a concrete tier). Byte-equivalent to the old literal.

3. scripts/changeset/lint.cjs: USER_FACING_PREFIXES omitted `src/`. Post-ADR-457
   the product source is src/*.cts compiled to a gitignored gsd-core/bin/lib,
   so the `gsd-core/` prefix is dead coverage for library code and a src/-only
   PR could merge with no release note — including this one. Adding `src/`
   closes the gate; tests/ stays non-user-facing.

Also corrects a false docstring in the VALID_TIERS test: value-equality cannot
detect a re-hardcoded literal, so the test no longer claims it does.

Two review findings were rejected with evidence rather than actioned:
- W021 double-allocation is governed by ADR-612 ("W021 renumber -> void ...
  kept, message-disambiguated"), not a defect.
- Global-defaults validation would be a false-positive generator: config-loader
  reads ~/.gsd/defaults.json only on the "no .planning/" branch, and health
  early-returns E001 without .planning/, so those values provably never affect
  resolution in any context health can run.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* test(#2070): regenerate install goldens for the changeset-lint change

scripts/ ships as an installed artifact, so scripts/changeset/lint.cjs's content
hash is pinned in all 18 runtime golden fixtures. Adding 'src/' to
USER_FACING_PREFIXES changed that hash and tripped every golden parity check.
Regenerated via `npm run gen:golden`; the only delta is the lint.cjs hash.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

* docs(#2070): backfill PR number 2336 into changeset

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 06:48:04 -04:00
Tom Boucher
880fbd963a fix(#2206): strip trailing slashes in isGitIgnored to avoid CRLF check-ignore quirk (#2235)
* fix(#2206): strip trailing slashes in isGitIgnored to avoid CRLF check-ignore quirk

isGitIgnored was called with a trailing slash (`.planning/`) in config-loader.
git check-ignore has a longstanding quirk: a CRLF .gitignore with blank lines
falsely reports any path WITH a trailing slash as ignored. This silently set
commit_docs=false on Windows repos (where CRLF .gitignore is the norm), skipping
all planning-doc commits. Normalize trailing slashes inside isGitIgnored so every
call site is protected.

Closes #2206

* docs(#2206): add changeset fragment

* docs(#2206): backfill PR number
2026-07-13 13:04:16 -04:00
Tom Boucher
6addeccd19 feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates
an external API/SDK/service can no longer seal without a decided coverage matrix.

- src/api-coverage.cts: deterministic detector (compound verb+noun signal +
  <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix
  parse/validate/render with field-length caps.
- check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as
  a token under .planning/phases/ only (traversal-neutralized); validates
  COVERAGE.md or blocks iff a strong integration signal is detected and no
  matrix exists; fail-closed when phases tree exists but phase unresolvable.
- capabilities/ai-integration: workflow.api_coverage_gate config key (default
  true), plan:pre contribution, blocking verify:pre gate. Data-driven.
- gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch.
- Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e.
  Code+security review findings fixed (stopword FP, scope containment, pipe/cap
  rejection, prompt-injection message hygiene).
- Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs.

Closes #1562
2026-07-07 15:11:12 -04:00
Rezolv
75552f7ea0 fix(#1533): prototype-pollution guard in _deepMergeConfig (#1534)
* fix: prototype-pollution guard in _deepMergeConfig (audit M4)

The root↔workstream config merge iterated Object.keys(overlay) with no
__proto__/constructor/prototype guard, while four sibling paths in the same
file (lines ~315/319/331/341/549) guard them. A workstream/root config.json
with {"__proto__": {...}} could pollute the merged object's prototype chain
and spoof unset config flags (per-object, not global Object.prototype).

Adds the same three-key continue guard at the top of the overlay loop plus a
regression test for __proto__/constructor/prototype overlay keys.

Closes a gap missed by the closed config proto-pollution hardening
(#751/#1406/#663).

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

* chore(changeset): Fixed fragment for #1534 (config proto-pollution guard)

Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-21 22:46:34 -04:00
Tom Boucher
e7855bc217 fix(#1459): user-owned consent store gates third-party capability activation; env/cwd in disclosure; loader validator parity (#1473) 2026-06-20 01:59:03 -04:00
Tom Boucher
353f63d170 feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2)

Promote the registry from a frozen data file to loadRegistry({includeInstalled}),
composing the first-party registry with a validated installed overlay (ADR-1244 D2):

- Extract the conformance validator to a shared runtime-callable module
  (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it
  verbatim, guarded by a generative-parity test (no build-time/runtime drift).
- capability-loader.cts: loadRegistry({includeInstalled}) composes first-party
  ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and
  <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party
  always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic-
  prefixes); full merged-set cross-capability validation; engines.gsd load-time
  re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path
  escapes rejected.
- semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed.
- Wire surface/state + loop to the overlay; loop injects a blocking gate for each
  skipped gate-kind overlay (fail-closed).
- cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd)
  + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/
  config-set call (never eager at module load, never wrong-cwd); first-party path
  unchanged with no cwd.
- run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test
  hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact);
  capability-validator.cjs stays linted (#551 migration coverage).

Closes #1431

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1431): add changeset for runtime capability registry overlay

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52)

The cwd-aware overlay config-key federation added to config-schema.cts
(_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd
threading) introduced mutable surface uncovered by config-schema's mutation
test set, dropping its score to 39.58% (below the 52 break threshold). Add a
real-overlay-fixture describe block exercising every branch (cwd guard, overlay
loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker
score 39.58% -> 77.08%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 14:41:20 -04:00
Tom Boucher
484a5b7b86 fix(#1415): loadConfig provenance + agent-skills diagnostic (Resolution Provenance P2) (#1424)
Implements ADR-1411 P2 / #1415, closing #1366.

Part 1 — config-loader.cts:
- Adds `loadConfigResolved(cwd, options) → ConfigResolution { config, source, degraded }`
  with six tagged branch return paths:
  A1 ws+wsconfig → source:'workstream', degraded:false
  A2 no-ws+config → source:'root', degraded:false
  B  ws requested, wsconfig absent → source:'root', degraded:true (intercepts recursive call)
  C  .planning/ exists, no config → source:'builtin-defaults', degraded:false
  D  no .planning/, global defaults readable → source:'global-defaults', degraded:false
  E  no .planning/, no global → source:'builtin-defaults', degraded:false
- `loadConfigResolved` calls `findProjectRoot` at entry to anchor resolution to the
  nearest .planning/ ancestor (cwd-drift fix, heuristic 4 from P1)
- `loadConfig` becomes a one-line delegation: `return loadConfigResolved(cwd, options).config`
- Exports `loadConfigResolved` in the `export =` block

Part 2 — init.cts cmdAgentSkills:
- Imports `findProjectRoot` from `./project-root.cjs`
- Anchors to project root before loading config (fixes #1366 cwd-drift)
- Uses `loadConfigResolved` for provenance; passes projectRoot to buildAgentSkillsBlock
- Computes `configured` + `reason` (AgentSkillsReason enum): 'resolved' |
  'not_configured' | 'configured_empty' | 'configured_unresolved'
- configured_empty and configured_unresolved emit stderr WARNING; not_configured is silent
- --json IR gains: configured, reason, source, degraded (in addition to existing
  agent_type, block, skills_count, warnings)

Tests (TDD):
- tests/config-loader.test.cjs: 8 new provenance tests (RED before impl, GREEN after)
- tests/agent-skills.test.cjs: 7 new diagnostic tests (RED before impl, GREEN after)
- All 306 tests across 4 suites pass (config-loader:33, agent-skills:72, init:105, workstream:96)

Docs:
- CONTEXT.md: Config Loader Module entry updated with loadConfigResolved interface;
  Resolution Provenance entry notes P2 is now implemented
- docs/CLI-TOOLS.md: --json field reference table added for agent-skills

Closes #1366
Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:05:15 -04:00
Tom Boucher
8c3d934a90 refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete) (#1295)
* refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete)

After T0–T6 nothing imports core, so retire the spine and its scaffolding:
- delete src/core.cts (and the gitignored gsd-core/bin/lib/core.cjs artifact;
  remove its .gitignore + eslint-ignore entries)
- delete scripts/lint-core-spine-imports.cjs + its allowlist; drop it from the
  package.json lint:ci chain
- regenerate docs/INVENTORY-MANIFEST.json (drops the core.cjs surface)
- sweep stale references: CONTEXT.md glossary back-compat clauses (spine retired,
  callers import the leaf directly), planning-config.md CONFIG_DEFAULTS owner,
  and false present-tense core.cjs claims in leaf-module docstrings

The ADR-857 decomposition is complete: the former Core god-module is fully
dissolved into its leaf modules; no re-export spine remains. No behaviour change.

Closes #1294

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1294): migrate the computed-path core.cjs importers the literal grep missed

bin/install.js used require(path.join(_gsdLibDir, 'core.cjs')) (a computed
path, and bin/install.js was never in the convergence lint's scan roots), and
~8 test files referenced core.cjs via path.join/readFileSync/existsSync/FILE_ARG
forms the literal-string migration grep missed. Route install.js's symbols to
their leaves (RUNTIME_PROFILE_MAP->model-catalog, resolveTierEntry/EFFORT_SET->
model-resolver) and repoint/adjust the test references to the leaves. Recovers
the 161 'Cannot find module core.cjs' failures from the spine deletion.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 18:50:46 -04:00
Tom Boucher
b10e56818b feat(#1169): complete ADR-857 phase 6 — migrate features to Capabilities, revive dead gates, harden conformance gate (#1183)
* test(#1168): make phase-6 gate un-gameable — reject empty stubs + require loop shrink

The migration assertion previously checked only role==feature, so a registration-only stub (empty hooks, logic left inline) would turn the gate green while phase 6 stayed incomplete — the exact false-completion pattern this gate exists to prevent. Strengthen it: each ADR-named feature must OWN its behavior (>=1 hook, or a command family); and plan-phase.md/execute-phase.md must shrink strictly below their frozen pre-phase-6 sizes (94519/93166 LF bytes), which also defeats double-run gaming (declare a hook but keep the inline block -> file does not shrink -> red).

Gate now 5 pass / 4 fail (orphaned execute:wave:post, empty/unregistered features, config-key leaks, no shrink). Green is now reachable only by REAL migration. Refs #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate gap-analysis to a Capability (plan:post gate)

First real ADR-857 phase-6 migration (pattern-defining tracer). gap-analysis moves from an inline post_planning_gaps branch in plan-phase.md to a real plan:post gate Capability:

- capabilities/gap-analysis/capability.json: role:feature, plan:post gate (when=workflow.post_planning_gaps, blocking:false advisory), OWNS workflow.post_planning_gaps (federated out of central schema). - plan-phase.md: inline config-get + gsd_run gap-analysis block replaced with a plan:post render-hooks call site dispatching the gate; file shrinks 94519->93279. - src/check-command-router.cts: cmdGapAnalysisPlanPost runs the real gap analysis via gap-checker. - post_planning_gaps removed from central manifest; resolves via federated config (default true preserved). - tests/post-planning-gaps-2493: re-pointed to assert capability ownership.

Verified: gate 5 pass / 4 fail (gap-analysis cleared from migration, plan:post-orphan, config-leak, and plan-phase shrink checks); loadConfig still returns post_planning_gaps=true; check command runs real analysis; 392/392 in the config/registry/federation/router net. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate profile-pipeline to a command-family Capability

ADR-857 Decision 7: profile-pipeline becomes a command-family Capability (like audit/intel/graphify). capabilities/profile-pipeline/capability.json declares an 8-command family (scan-sessions, extract-messages, profile-sample, write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md) backed by a new gsd-core/bin/lib/profile-pipeline-command-router.cjs; the inline case arms are removed from gsd-tools.cjs. Owns profile-pipeline.enabled (federated).

Verified: registry shows role:feature with commands.length=8; scan-sessions/profile-sample run live via the family; gate cleared profile-pipeline from the empty-stub failure (only tdd/schema-gate/drift remain); 296/296 registry+inventory+gsd-tools tests; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1167): wire execute:wave:post + implement ui.safety-gate check

Revives the second dead gate from #1167: ui.gates@execute:wave:post was declared but never dispatched AND its check.query (ui.safety-gate) was unimplemented. Adds the per-wave execute:wave:post render-hooks call site in execute-phase.md (fires after each wave's merge/cleanup, before the next forks) and implements cmdUiSafetyGate (frontend + UI-SPEC aware, mirrors cmdUiPlanGate) in check-command-router. +17 regression tests.

Verified: phase-6 orphaned-points conformance test now PASSES (gate 6 pass / 3 fail); ui-safety-gate routable in dot+hyphen forms; check-ui-safety-gate 17/17, check-ui-plan-gate 18/18; lint 0 errors. Refs #1167, #1168.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate drift (schema + codebase) to execute:wave:post gates

Removes the inline schema_drift_gate + codebase_drift_gate steps (77 lines) from execute-phase.md; drift becomes a Capability with two execute:wave:post gates (verify.schema-drift blocking, verify.codebase-drift advisory) dispatched via the per-wave render-hooks call site. check-command-router routes verify.schema-drift / verify.codebase-drift to the real detectors. Federates workflow.drift_threshold / drift_action / schema_drift_gate out of central.

Also fixes the execute:wave:post dispatch prose to run NON-blocking (advisory) gates too — the prior version only ran blocking gates, which would have silently dropped the codebase-drift advisory after its inline step was removed. Behavior preserved.

Verified: gate 7 pass / 2 fail (drift cleared from stub + config-leak; execute-phase.md 92297 < 93166 frozen -> shrink passes); both drift checks run real detection; loadConfig defaults preserved (threshold=3, action=warn, gate=true); drift-detection 56/56 + schema-drift 34/34; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate tdd to a Capability (plan:pre contribution + execute:post gate)

tdd becomes a real Capability: a plan:pre contribution injects the <tdd_mode_active> planner guidance (rendered from PLAN_PRE_HOOKS_JSON like security's contribution), and an execute:post gate (tdd.review-checkpoint, advisory) runs the real end-of-phase RED/GREEN review via a new check-command handler. Inline tdd_mode reads + the inline planner block + the tdd_review_checkpoint step are removed; workflow.tdd_mode is federated out of central. The MVP+TDD per-task RED-commit gate is preserved — TDD_MODE is now derived from the execute:post hooks (capId==tdd active), not an inline config-get.

BEHAVIOR CHANGE (documented, not silent): the --tdd CLI flag now persists workflow.tdd_mode=true via config-set instead of being per-invocation. Rationale: tdd is now a config-toggled Capability, and env vars do not persist across the workflow's separate bash blocks (config does), so an ephemeral override isn't cleanly achievable; --tdd therefore enables the tdd capability, consistent with how all capabilities are toggled.

Verified: gate 7 pass / 2 fail (tdd cleared from stub + config-leak; plan-phase + execute-phase both < frozen sizes); contribution injection + execute:post gate dispatch wired; MVP+TDD gate preserved; tdd.review-checkpoint runs real review; full unit suite 556/0; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): migrate schema-gate to a plan:pre contribution Capability

The plan-time schema-push detection (former plan-phase.md §5.7) becomes a schema-gate Capability: a plan:pre contribution (into:planner, when:workflow.schema_push_detection) whose fragment carries the full ORM-detection + [BLOCKING] schema-push-task injection logic, rendered into the planner via the existing plan:pre render-hooks dispatch. The inline §5.7 block is removed (plan-phase.md 94519->90445). workflow.schema_push_detection is a new capability-owned (federated) key, default true. (The execute-side schema-drift gate was migrated separately into the drift capability.)

Verified: registry inlines the fragment (len 2704) so it is actually delivered at plan:pre; gate 8 pass / 1 fail — all 5 ADR-named features now real Capabilities, only the config-leak test remains (intel/security, next unit). Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(#1169): close the 3 capability config-key leaks — phase-6 gate now GREEN

Removes the last inline config-get reads of capability-owned keys from plan-phase.md. security_asvs_level/security_block_on now flow through the security plan:pre contribution via a new loop-resolver configValues mechanism (resolves declared config keys with the same 4-level precedence as activation and attaches them to the rendered hook); the §5.55 banner reads them from PLAN_PRE_HOOKS_JSON. intel.enabled becomes a real intel plan:pre step (ref.command: intel api-surface) dispatched via render-hooks; the inline intel branch is gone. gen-capability-registry now validates ref.command as a third dispatch shape.

Verified: phase-6 capstone conformance gate is FULLY GREEN (9/0); 3 leaks gone (grep=0); security configValues resolve to {2,medium}/default {1,high}; intel step present only when enabled; loop-render-hooks 62/0, capability-registry 287/0, capability-state/federated-config 113/0; lint 0 errors. Closes the migration half of #1169. Refs #1139, #1167, #1168.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): address adversarial review — restore schema-drift block, generic planner injection, uniform gate contract

Adversarial review caught 2 real regressions the green gate missed: (1) schema-drift no longer blocked — the execute:wave:post dispatch read GATE_RESULT.block but verify.schema-drift emitted drift_detected/blocking, and onError:skip wrongly bypassed positive blocks; (2) only tdd's plan:pre contribution was injected into the planner, dropping schema-gate's schema-push detection and security's threat-model guidance.

Fixes: (A) every gate check returns a uniform boolean 'block' under --raw (the dispatch form), with advisory gates (tdd/gap) carrying their report in 'message'; (B) gate-dispatch contract corrected at all sites — onError governs command errors only, a blocking gate's positive block always halts; (C) generic planner injection of all plan:pre contributions where into=='planner' (tdd + schema-gate + security incl configValues); (D) two new conformance assertions: planner contributions injected generically + every gate check.query returns boolean block under --raw.

Verified: gate 11/11; all 6 gate checks return boolean block under --raw; full suite 595/0; lint 0 errors. Refs #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): restore MVP+TDD end-of-phase blocking escalation (2nd adversarial pass)

The migrated tdd execute:post gate is statically blocking:false, but the contract (references/execute-mvp-tdd.md + CONTEXT.md) requires the end-of-phase TDD review to ESCALATE from advisory to blocking when MVP_MODE && TDD_MODE && a TDD plan misses a RED/GREEN commit. The migration prose had downgraded this to a 'strong advisory recommendation' — silent loss of the blocking escalation. Restore it: the tdd-gate dispatch now refuses to mark the phase complete (Phase blocked message) under MVP+TDD when GATE_RESULT.block is true; advisory otherwise.

Also strengthen tests/execute-mvp-tdd-gate.test.cjs: hasBlockingEscalation previously matched any line with 'blocking'+'mvp+tdd' (so 'advisory (blocking: false) ... under MVP+TDD' was a false green); now it requires the real refusal semantics ('refuse to mark the phase complete' / 'phase blocked'). Caught by 2nd adversarial review pass.

Verified: execute-phase.md 92702 < 93166 frozen; mvp-tdd-gate + phase-6 gate 19/0; full suite green; lint 0 errors. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): restore MVP+TDD proceed-block, codebase auto-remap, schema skip-flag (3rd adversarial pass)

3rd adversarial pass found 4 more silent regressions: (1) the tdd MVP+TDD 'refuse to mark complete' was nullified by a downstream 'ALWAYS proceed regardless of gate results' line — proceed is now conditional (stops on an active MVP+TDD block); (2) the test now asserts the proceed is NOT an unconditional override; (3) codebase-drift auto-remap (spawn gsd-codebase-mapper when drift_action=auto-remap) was dropped — the execute:wave:post advisory dispatch now consumes spawn_mapper/directive; (4) GSD_SKIP_SCHEMA_CHECK bypass was lost from the gate path — cmdVerifySchemaDrift now honors the env var (block:false when set).

Verified: no unconditional proceed; GSD_SKIP_SCHEMA_CHECK=true -> block:false; gate 11/11 + mvp-tdd 9/9; full suite 569/0; lint 0; execute-phase.md 93109 < 93166. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): init.cts reads federated config keys from nested path (4th adversarial pass)

Config federation moved tdd_mode/research/nyquist_validation from flat config.<key> to nested config.workflow.<key>, but src/init.cts still read them flat — so init.plan-phase/init.execute-phase emitted tdd_mode:false / research_enabled:undefined / nyquist:undefined regardless of config (a public command-contract regression; the migrated loops use render-hooks so enforcement was unaffected). Read via config.workflow (type-safe Record cast). Now init reflects the same resolved values + federated defaults (research/nyquist default true) as the render-hooks path.

Verified: build clean; init.plan-phase emits tdd_mode:true/research:false/nyquist:false for set config, defaults true for empty; full suite 591/0; lint 0. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1169): add changeset for ADR-857 phase-6 completion (PR #1183)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): complete phase-6 migration fallout — restore TEXT_MODE, fix registry .claude leak, re-point stale workflow-contract tests

The capability migration left real regressions and stale consumer tests that
the per-module unit suite missed but the full cross-platform suite caught (27
failing tests):

Real source regressions (fixed):
- execute-phase.md lost its AskUserQuestion TEXT_MODE plain-text fallback when
  the inline schema_drift_gate step was removed — non-Claude runtimes would
  stall. Restored, and the execute:post gate-dispatch prose de-duplicated to
  cite the execute:wave:post contract (loop body shrinks below the frozen
  pre-phase-6 ceiling while keeping every onError/blocking nuance).
- capabilities/tdd inline fragment hardcoded `@~/.claude/gsd-core/references/tdd.md`,
  baked verbatim into the committed capability-registry.cjs and leaked the
  install path on 11 non-Claude runtimes (registry .cjs is copied, not
  path-converted). Made the fragment path-free; regenerated the registry. The
  phase-6 conformance gate now guards this (no ~/.claude install path in any
  capability source or the generated registry).
- plan-phase.md: removed a §5.7 stub re-added in error and routed Branch 2 to
  step 6 (schema-gate is a plan:pre capability, §5.7 is gone).

Stale workflow-contract tests re-pointed to the capability dispatch they now
must assert (behavior verified preserved in source first, assertions kept
equal-or-stronger): bug-621 + bug-2851 (gap-analysis via gsd_run render-hooks
plan:post + registry binding), feat-2527 (tdd_mode federated out of central),
phase6-planning + plan-phase-ui-redirect (§5.6 bounded by ## 6.),
plan-phase-drift-guard (intel when:intel.enabled skip branch).

profile-pipeline-command-router.cjs un-ignored from eslint (hand-written, no
TS source) + stale disable comments removed. Size baseline regenerated.

Verified: full suite 15140 tests / 0 fail; lint 0 errors; conformance gate green
legitimately. Refs #1139, #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1169): add ADR-857 E2E content-test coverage for the 12 loop points + capability deliverables

Grounds the capability engine in behavioral E2E tests (drive the real
render-hooks/check CLI + the real registry, assert typed result content — no
source-grep), structured around what ADR-857 says to deliver. 207 tests; each
genuineness-checked (flip the expectation, confirm it fails).

Per-loop-point dispatch (7 files): empty-point negative-space across the 6
no-hook points; verify:post 3-step resolution+ordering+onError; plan:pre
contribution/configValues + ui.plan-gate + intel; plan:post gap-analysis;
execute:wave:post drift+ui gates via the check route (schema-drift block/skip,
codebase-drift threshold BVA, auto-remap); execute:post tdd.review-checkpoint
RED/GREEN; ship:pre security gate resolution + frontmatter-get predicate pieces.

ADR-deliverable coverage (4 files): predicate boundary held (edge/prohibition
probes stay core, not off-by-default Feature Capabilities — phase-6 exception);
core loop runs with zero capabilities (all 12 points empty, init bundles
resolve); contribution merge (multiple ordered <contribution from=> blocks);
federated-config key removal on uninstall.

federated-config allowlisted for its 3-file split (unit + integration +
lifecycle). Refs #1139, #1167, #1168, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): remove dead drifted converter dups + address adversarial review

Lint cleanup (root-caused, not waved off): src/runtime-artifact-conversion.cts
carried 11 agent-converter functions (+5 orphaned consts/helpers) that were
never exported, never called, and had silently DRIFTED from the live
hand-authored copies in bin/install.js (one even referenced an undefined
`claudeToCopilotTools`). Deleted the dead duplicates; install.js's live copies
are untouched (it never imported these). Lint now 0 errors / 0 warnings.

Adversarial-review (Codex) findings fixed:
- HIGH: execute-phase.md TDD_MODE used `jq ... || echo false`, silently
  disabling the MVP+TDD blocking gate on jq-less runtimes. Reverted to the
  `node -e` form (node is guaranteed; matches the file's other node-e usages) so
  a missing optional tool can no longer fail-open a blocking safety path.
- MEDIUM: federated-config-key-removal orphan-key test was vacuous (it skipped
  the orphan assertion). Now asserts the removed capability's key is genuinely
  not surfaced/validated after uninstall.
- LOW: phase-6 conformance leak regex broadened to catch absolute-home and
  Windows-backslash `.claude/(gsd-core|commands|agents|hooks)` paths, not only
  `~`/`$HOME` forward-slash forms.
- LOW: bug-2851 plan:post dispatch assertion now requires `--raw` (matched its
  stated contract).
- nit: plan-pre intel-step test duplicate assertion replaced with a distinct
  structured-output check.

Size baseline regenerated (execute-phase.md 93089 < 93166 frozen). Refs #1167, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1169): make runtime-homes-descriptor-drive titles environment-independent

The descriptor-equivalence test embedded the absolute golden config path
(`os.homedir()`-derived) directly in each `test(...)` title, so titles differed
between macOS (`/Users/x/.claude`) and Docker (`/home/gsdtest/.claude`). Every
test PASSES on both platforms (15885/0 leaf tests each), but gsd-test-summary
compares results by title and reported 29+29 false "only in Mac / only in
Docker" discrepancies for tests that actually pass everywhere.

Move the golden path out of the title and into the assertion message (still
shown on failure); titles are now byte-identical across platforms so the
cross-platform comparator matches them. No assertion logic or golden values
changed. Refs #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1169): derive TDD_MODE via gsd_run --active-cap, not node -e (fix prompt-injection CI gate)

The prior fix reverted execute-phase.md:181 from jq to `node -e` to close a
Codex HIGH (jq||echo-false silently disabling the MVP+TDD blocking gate on
jq-less runtimes) — but the CI prompt-injection scanner BLOCKS new `node -e` in
workflow markdown (inline code-exec = injection vector), turning the security
gate red. Both forms were wrong: node -e fails the scanner; jq fail-opens a
blocking safety gate; `config-get workflow.tdd_mode` is forbidden by the
conformance leak gate (tdd_mode is capability-owned).

Correct fix (what Codex recommended): a gsd_run-native boolean. Add an
`--active-cap <capId>` flag to `loop render-hooks <point>` that resolves hooks
the normal way and prints exactly `true`/`false` for whether a capId is active
— scanner-safe (canonical launcher, no inline code), node-reliable (no optional
jq to fail-open), and leak-free (render-hooks resolution, not config-get).
execute-phase.md:181 now `TDD_MODE=$(gsd_run loop render-hooks execute:post
--active-cap tdd)`. +5 behavioral tests for the flag.

Verified: prompt-injection-scan --diff origin/next → 0 findings; conformance
gate 13/13 (execute-phase.md 92934 < 93166); execute-mvp-tdd + tdd-mode +
loop-render-hooks 87/0; lint 0/0. Refs #1167, #1169.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 21:07:55 -04:00
Tom Boucher
607813f5d0 feat(#1136): consume resolved capability state (#1153)
* feat(#1136): consume resolved capability state

* chore(#1136): add capability state changeset
2026-06-12 23:43:07 -04:00
Tom Boucher
6dbd895028 feat(#910): federated config merge in config-loader (ADR-857 phase 3b) (#914)
Build the federated config merge: each capability owns its config-key slice
(ADR-857 decision 3 / ADR-894), and loadConfig merges them defensively. The
registry now emits a full configSchema index ({key:{owner,type,default,
description}}, generator-validated); a new src/federated-config.cts resolves
federated keys defensively (skip central keys -> pending-migration warning,
skip malformed slices -> warning never throw, else type-checked user override
?? default, with nested dotted-path lookup and enum validation); and loadConfig
applies the overlay on every return path.

Wired as a provably-empty no-op channel: every UI-pilot key is still central,
so validKeys is empty and loadConfig returns byte-identical output on all paths
(identity return when the overlay is empty; shared CONFIG_DEFAULTS never
mutated). Registry-only; no key is cut over; nothing in the live loop changes.

Closes #910

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 23:37:41 -04:00
Tom Boucher
a74e71b049 refactor(#885): extract loadConfig cluster into config-loader.cts (#886)
ADR-857 rollout phase 2e — the largest core.cts extraction. Move the
configuration-loading subsystem (loadConfig + _getConfigDefault/
_getNestedConfigDefault/CONFIG_DEFAULTS/_deepMergeConfig, isGitIgnored +
_gitIgnoredCache, _warnUnknownProfileOverrides + RUNTIME_OVERRIDE_TIERS + the
dedup Sets, _resetRuntimeWarningCacheForTests) out of core.cts into a new leaf
module src/config-loader.cts. core.cts re-exports the public surface
(loadConfig, isGitIgnored, CONFIG_DEFAULTS, RUNTIME_OVERRIDE_TIERS,
_resetRuntimeWarningCacheForTests); 12+ callers unchanged.

Cycle-free: config-loader imports only leaves (configuration, config-schema,
planning-workspace, shell-command-projection, core-utils, model-catalog). All
core-internal helpers loadConfig touches moved with it to avoid a cycle. core
keeps CANONICAL_CONFIG_DEFAULTS for its model-resolver functions, which now
resolve loadConfig via the binding — this unblocks the final model-resolver
extraction (2f). core.cts: 1275 -> 792 lines.

Repointed tests/config-field-docs.test.cjs (a docs-parity source check) to read
the CONFIG_DEFAULTS literal from its new home (config-loader.cjs). New-CLI-module
checklist done (.gitignore, eslint, INVENTORY 95->96 + row, manifest,
ARCHITECTURE, CONTEXT.md "Config Loader Module"). Adds tests/config-loader.test.cjs
(27 tests: behavioral + shim-identity + adversarial config fixtures).

Gates: lint, code-review, security-review (prototype-pollution guard confirmed
intact), codex adversarial-review (0 findings; byte-identical move). Mac 4115
pass; clean-build docker: full-suite hit the local mirror's known incremental-tsc
non-determinism on an unrelated re-exported symbol (findPhaseInternal, from
already-merged 2d), but a clean targeted rebuild of the affected file passed
163/0 — CI's clean full matrix is the authoritative gate.

Closes #885

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 15:36:38 -04:00