Commit Graph

3785 Commits

Author SHA1 Message Date
Tom Boucher
1abb0d3427 feat(#1432): capability source resolver + ledger (ADR-1244 Phase 3) (#1443)
* feat(#1432): capability source resolver + ledger (ADR-1244 Phase 3)

ADR-1244 D3 + D4 — additive, testable-in-isolation modules (the /gsd:capability
install command + consent gate are Phase 4). NEW modules only; not yet wired into
bin/install.js.

- src/capability-source.cts → resolveCapabilitySource(spec, opts): one seam, one
  adapter per source kind — local (fs copy), git (execGit clone+checkout, https/ssh/
  git transports only), npm (execNpm pack --ignore-scripts + tar, NEVER npm install),
  tarball (https download + sha512 integrity-before-extraction + tar), registry (stub).
  SECURITY: install never executes capability code (copy/extract only, --ignore-scripts);
  integrity verified before staging; symlink members rejected at interior, source-root,
  AND tar-member (verbose-listing) layers; tar-slip member paths rejected pre-extraction;
  npm specs with shell metacharacters (incl %) rejected (execNpm uses a Windows shell);
  git ext::/file:// transports + leading-dash/metachar refs rejected; atomic staging
  (.staging→rename, restore-on-failure); full Phase 1/2 validator suite + engines.gsd
  pre-check on the fetched manifest. Test seam _setCapabilitySourceHttpGet.
- src/capability-ledger.cts → per-runtime .gsd-capabilities.json install manifest:
  readLedger/writeLedger (atomic via platformWriteSync)/recordInstall (idempotent,
  prototype-pollution-guarded)/removeEntry/reconcile (orphan report, hardened against
  hostile/non-string/'..' files[] members — never throws, never oracles outside runtimeDir).
- 48 new tests (34 source incl. the full security matrix, 14 ledger). Two Codex
  adversarial rounds; all 12 findings fixed + regression-tested.
- .gitignore + eslint.config.mjs: built capability-source/ledger.cjs git+eslint-ignored
  (ADR-457 #551 migration coverage); CONTEXT glossary + INVENTORY rows + manifest regen.

Closes #1432

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1432): add changeset for capability source resolver + ledger

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 16:23:20 -04:00
Tom Boucher
353f63d170 feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2)

Promote the registry from a frozen data file to loadRegistry({includeInstalled}),
composing the first-party registry with a validated installed overlay (ADR-1244 D2):

- Extract the conformance validator to a shared runtime-callable module
  (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it
  verbatim, guarded by a generative-parity test (no build-time/runtime drift).
- capability-loader.cts: loadRegistry({includeInstalled}) composes first-party
  ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and
  <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party
  always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic-
  prefixes); full merged-set cross-capability validation; engines.gsd load-time
  re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path
  escapes rejected.
- semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed.
- Wire surface/state + loop to the overlay; loop injects a blocking gate for each
  skipped gate-kind overlay (fail-closed).
- cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd)
  + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/
  config-set call (never eager at module load, never wrong-cwd); first-party path
  unchanged with no cwd.
- run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test
  hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact);
  capability-validator.cjs stays linted (#551 migration coverage).

Closes #1431

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1431): add changeset for runtime capability registry overlay

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52)

The cwd-aware overlay config-key federation added to config-schema.cts
(_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd
threading) introduced mutable surface uncovered by config-schema's mutation
test set, dropping its score to 39.58% (below the 52 break threshold). Add a
real-overlay-fixture describe block exercising every branch (cwd guard, overlay
loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker
score 39.58% -> 77.08%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 14:41:20 -04:00
Tom Boucher
2421cf1b4a feat(#1430): versioned capability manifest + native stamping (ADR-1244 Phase 1) (#1436)
* feat(#1430): versioned capability manifest + native stamping (ADR-1244 Phase 1)

Make the capability manifest versioned — the data substrate the Capability
Ecosystem (ADR-1244) keys off:

- capability.json gains a REQUIRED semver `version` plus the optional
  ecosystem envelope (`engines.gsd`, `compatVersions`, `integrity`,
  `provenance`); the build-time conformance validator enforces them via a new
  `validateVersionEnvelope()` (exported for the Phase 2 runtime overlay).
- All 32 native capabilities stamped with `version` (= package version,
  lockstep) + `engines.gsd`; `sync-manifest-versions.cjs` gains a glob sweep
  that keeps them in sync, and the issue-844 regression guard is extended.
- Strict SemVer 2.0.0 grammar blocks metacharacter/space/unicode smuggling in
  version strings; range/integrity fields are shape-validated (satisfaction
  and the load-time gate are deferred to Phase 2/4).
- Capability rel-paths emitted forward-slash for cross-platform git correctness.

Closes #1430

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1430): add changeset for versioned capability manifest

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 10:45:53 -04:00
Tom Boucher
8040a6bac0 chore(#1417): add resolution-provenance CI guard (Resolution Provenance P4) (#1428)
Adds scripts/lint-resolution-provenance.cjs — a registry + ratchet CI guard
that locks in the agent-skills configured_empty/not_configured contract tests
so they cannot be silently removed, and establishes a registration point for
future config-interpreting read verbs (ADR-1411 P4).

Design rationale:
- REGISTRY (one entry: agent-skills → src/init.cts → tests/agent-skills.test.cjs)
  is the canonical registration site; new verbs are added here.
- For each registered verb, the guard asserts its test file contains BOTH a
  `configured_empty` assertion AND a `not_configured` assertion — proving the
  configured-empty-vs-not-configured contract is explicitly tested.
- NOT a universal static detector (intractable / false positives) — mirrors the
  no-adhoc-markdown-parsing grandfather pattern.
- Uses scripts/lib/allowlist-ratchet.cjs (assertWithinAllowlist) so stale
  allowlist entries fail (ratchet-down) and novel offenders always fail.
- checkRegistry() is factored as a pure exported function tested in
  tests/lint-resolution-provenance.test.cjs without shelling out.
- Wired into lint:ci (package.json) and lint step name updated in test.yml.
- CONTEXT.md ### Resolution Convention extended with P4 guard sentence.
- Allowlist starts empty ([]) — agent-skills already has its tests.

Closes #1417
Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 08:50:22 -04:00
Tom Boucher
2d406d2f17 fix(#1416): add resolution.cjs to ESLint ignore list (ADR-457 migration-coverage) (#1426)
The new src/resolution.cts (P3) compiles to gsd-core/bin/lib/resolution.cjs, a
tsc-generated artifact that must be eslint-ignored (lint the .cts source, not the
emitted .cjs). The 551-eslint-bin-lib-coverage test enforces this and was missed
in PR #1425 — fixing next.

Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:56:54 -04:00
Tom Boucher
b0c774c2e3 feat(#1416): formalize Resolution convention + agent-skills value envelope (Resolution Provenance P3) (#1425)
Narrows P3 of ADR-1411 (Resolution Provenance, epic #1411) based on an
adversarial fit-analysis that showed a single Resolution<T> envelope adopted
by agent-skills, capability-state, and capability-writer fails the deletion
test: configured/reason are meaningless for capability verbs, and
capability-writer's errors[] (operation-not-applied) cannot fold into
warnings[]. The only genuinely shared seam is warnings: string[].

Changes:

- src/resolution.cts: new pure types+builder leaf — exports Resolution<T>
  {value, configured, reason, warnings}, makeResolution<T>() builder, and
  AgentSkillsValue {block, skills_count}. No other src/ imports.

- src/init.cts: cmdAgentSkills --json IR gains additive value:{block,
  skills_count} field (built via makeResolution). All existing flat fields
  (agent_type, block, skills_count, warnings, configured, reason, source,
  degraded) are retained unchanged for back-compat.

- src/capability-state.cts: doc comment on ResolveCapabilityRuntimeStateResult
  naming it the canonical read-verb envelope. No JSON change.

- src/capability-writer.cts: doc comment on SetCapabilityStateResult naming it
  the canonical mutation-verb result (warnings=advisory, errors=operation-
  not-applied). No JSON change.

- CONTEXT.md: new ### Resolution Convention glossary entry after
  ### Resolution Provenance.

- docs/adr/1411-resolution-provenance.md: P3 narrowing amendment appended.

- tests/resolution.test.cjs: 9 unit tests for makeResolution (new).
- tests/agent-skills.test.cjs: 2 P3 tests for value.block/value.skills_count
  and back-compat of all flat fields.

All 277 tests pass (5 suites). npm run lint clean. All lint checks pass.

Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:44:52 -04:00
Tom Boucher
484a5b7b86 fix(#1415): loadConfig provenance + agent-skills diagnostic (Resolution Provenance P2) (#1424)
Implements ADR-1411 P2 / #1415, closing #1366.

Part 1 — config-loader.cts:
- Adds `loadConfigResolved(cwd, options) → ConfigResolution { config, source, degraded }`
  with six tagged branch return paths:
  A1 ws+wsconfig → source:'workstream', degraded:false
  A2 no-ws+config → source:'root', degraded:false
  B  ws requested, wsconfig absent → source:'root', degraded:true (intercepts recursive call)
  C  .planning/ exists, no config → source:'builtin-defaults', degraded:false
  D  no .planning/, global defaults readable → source:'global-defaults', degraded:false
  E  no .planning/, no global → source:'builtin-defaults', degraded:false
- `loadConfigResolved` calls `findProjectRoot` at entry to anchor resolution to the
  nearest .planning/ ancestor (cwd-drift fix, heuristic 4 from P1)
- `loadConfig` becomes a one-line delegation: `return loadConfigResolved(cwd, options).config`
- Exports `loadConfigResolved` in the `export =` block

Part 2 — init.cts cmdAgentSkills:
- Imports `findProjectRoot` from `./project-root.cjs`
- Anchors to project root before loading config (fixes #1366 cwd-drift)
- Uses `loadConfigResolved` for provenance; passes projectRoot to buildAgentSkillsBlock
- Computes `configured` + `reason` (AgentSkillsReason enum): 'resolved' |
  'not_configured' | 'configured_empty' | 'configured_unresolved'
- configured_empty and configured_unresolved emit stderr WARNING; not_configured is silent
- --json IR gains: configured, reason, source, degraded (in addition to existing
  agent_type, block, skills_count, warnings)

Tests (TDD):
- tests/config-loader.test.cjs: 8 new provenance tests (RED before impl, GREEN after)
- tests/agent-skills.test.cjs: 7 new diagnostic tests (RED before impl, GREEN after)
- All 306 tests across 4 suites pass (config-loader:33, agent-skills:72, init:105, workstream:96)

Docs:
- CONTEXT.md: Config Loader Module entry updated with loadConfigResolved interface;
  Resolution Provenance entry notes P2 is now implemented
- docs/CLI-TOOLS.md: --json field reference table added for agent-skills

Closes #1366
Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:05:15 -04:00
Tom Boucher
d220ba6fd6 fix(#1414): resolve project root from descendant via nearest ancestor .planning/ (Resolution Provenance P1) (#1423)
Add heuristic (4) to findProjectRoot: after heuristics (1)-(3) (sub_repos,
multiRepo, .git+parent-.planning/) are exhausted without a match, perform a
second bounded walk-up within FIND_PROJECT_ROOT_MAX_DEPTH to locate the
nearest ancestor directory containing a .planning/ subdirectory. Returns that
ancestor as the project root, so loadConfig finds the correct config instead
of falling through to defaults when gsd-tools is invoked from a plain
descendant subdirectory of a single-repo project (#1366 cwd-drift gap).

Ordering is load-bearing: the new walk runs AFTER the existing loop so
sub_repos workspaces (where a child sub-repo may have its own .planning/)
still resolve correctly to the parent workspace. The existing own-.planning/
guard (heuristic 0, #1362) and the depth bound (FIND_PROJECT_ROOT_MAX_DEPTH=10)
are both preserved unchanged.

Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 00:50:30 -04:00
Tom Boucher
6f2547610e docs(#1412): ADR-1411 — Resolution must report provenance, not fall open silently (#1413)
Decision record for the Resolution Provenance epic (#1411, P0). Establishes that
context resolution (config loading, project-root anchoring, workstream resolution)
must report its provenance rather than fall open silently to defaults — the
resolution-side analog of ADR-227. Binds the Config Loader Module, Project-Root
Resolution Module, and I/O Module. Adds the ADR, the index/seam-map entry, and the
CONTEXT.md glossary term.

Part of #1411
Closes #1412

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 23:42:17 -04:00
Tom Boucher
dcc7bc0d15 fix(#1406): guard setGsdConfig test helper against prototype pollution (#1407)
CodeQL js/prototype-pollution-utility (alert #40) flagged the setGsdConfig
test helper in tests/git-base-branch.test.cjs: it deep-assigns through a
dot-split key chain with no __proto__/constructor/prototype guard (CWE-915).

Mirror the production guard in src/config.cts (PR #752): inline literal
__proto__/prototype/constructor checks immediately before both write sites,
throwing on a forbidden segment. This is the form CodeQL recognizes; a
Set/pre-loop guard is not. Add a #1406 regression suite asserting the guard
throws on malicious keys and does not pollute Object.prototype, while a normal
nested key still writes. Behavior unchanged for the existing call (16/16 pass).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 21:51:02 -04:00
Tom Boucher
0c9f86d495 fix(#1404): version-only manifest bumps no longer park the back-merge (#1405)
auto-backmerge's needs_review safety net parked the PR whenever a code
file existed only on main. The version manifests (package.json,
package-lock.json, .claude-plugin/plugin.json, gemini-extension.json)
diverge every release by design (next runs a -dev version), so the net
misfired on every release — and that manual-review park is what let the
back-merge sit and go stale (e.g. #1379, which then conflicted with a
later #1777 purity-gate edit to fragments it had deleted).

Exclude the generated package-lock.json outright (it carries a version
per package entry, so a dep bump is indistinguishable from a release
stamp; it only mirrors package.json, still checked). For package.json /
plugin.json / gemini-extension.json, ignore a drop whose main-vs-base
diff touches only the top-level "version" field. A substantive
(non-version) straight-to-main change still parks, preserving the
safety net's real purpose.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 21:02:18 -04:00
Tom Boucher
666d933e16 ci(#1401): add no-adhoc-markdown-parsing rule (fence-strip + section-collect) + grandfather burn-down (#1402)
Tightens the over-broad heading-walk detection: removes heading-walk
entirely and narrows fence-regex to require a multiline body ([\s\S]),
so single-line tests like /^```/ and /^###\s+/ are no longer flagged.
Grandfathers the 10 genuine section-collect sites across state.cts,
milestone.cts, audit.cts, and phase-lifecycle.cts with concrete reasons.
Adds 12 RuleTester tests (3 positive, 9 negative) to eslint-rules.test.cjs.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 20:14:12 -04:00
Tom Boucher
ce8bcb1b95 refactor(#1398): migrate state.cts section-collects onto markdown-sectionizer seam, byte-identical STATE.md (epic #1372 T6) (#1399)
* refactor(#1398): migrate state.cts section-collect regexes onto markdown-sectionizer seam (epic #1372 T6)

Replace ≈18 hand-rolled `/(heading)([\s\S]*?)(?=stop)/` regex splices in
state.cts with direct `tokenizeHeadings` calls that compute the exact
[bodyStart, stopOffset) span, preserving byte-identical STATE.md output.

Non-migratable site left in place: `cmdStateRecordMetric`'s metricsPattern
captures table-header rows in group 1 — not a standard heading+body shape.
Write orchestration (readModifyWriteStateMd / syncStateFrontmatter /
shouldPreserveExistingProgress / #952 no-op guard) is UNTOUCHED.

Verification: t6-headtohead.cjs head-to-head harness runs 25 ops across 6
STATE.md fixture variants (inline, trailing-blanks, CRLF, no-frontmatter,
nested-acc, post-milestone) against origin/next and reports 0 diffs.
No-op guard confirmed: record-session on recorded:false leaves STATE.md
byte-identical. All 150 state tests and 62 milestone/forensics tests pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(#1398): capture T6 state-section-splice characterization; drop throwaway harness

Remove scripts/t6-headtohead.cjs (committed throwaway HEAD-vs-origin/next
byte-compare harness). Capture its coverage as 23 behavioral characterization
tests appended to tests/state.test.cjs, exercising all migrated cmdState*
write-ops across 7 fixture variants (inline, trailing-blanks, CRLF,
no-frontmatter, nested-acc, no-current-pos, post-milestone). Includes the
#952 no-op guard (recorded:false + byte-unchanged assertion) and
CRLF/trailing-blanks edge-case coverage. Also removes the dead
spliceStateSection helper (defined but never called) that was surfacing as
an @typescript-eslint/no-unused-vars warning.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 17:56:17 -04:00
Tom Boucher
308c7be481 refactor(#1396): T5 — migrate uat-predicate + uat onto markdown-sectionizer seam (#1397)
- uat-predicate.cts: replace local _stripFencedBlocks (and its private
  FenceState/StripFencedResult types) with a call to stripFencedCode from
  markdown-sectionizer.cjs (ADR-1372 T5). Both stripFalsePositiveContexts
  step (c) and analyzeMarkdown now route through the seam. The three other
  passes in stripFalsePositiveContexts — frontmatter strip, HTML-comment
  strip, blockquote-line filter — remain caller-side (seam does not do these).
  The unterminatedFence signal consumed by analyzeMarkdown is preserved; it
  is now returned by stripFencedCode (same machine, same contract).

- uat.cts: migrate the ## Current Test, ## Tests, and ## Human Verification
  section-collect patterns onto collectSection/tokenizeHeadings from the seam.
  UAT-specific item parsing (### N. Name blocks, expected/result fields,
  categorization logic) stays caller-side. The HTML-comment strip within the
  Current Test body remains caller-side (UAT document structure, not seam scope).

- tests/markdown-sectionizer.test.cjs: remove the 18-case tautological parity
  guard (DEFECT.GENERATIVE-FIX). Once uat-predicate imports the seam the guard
  compares the seam to itself — removing it is the T5 commitment per ADR-1372.

4-space-indent behavior change (CommonMark correctness improvement): the seam
uses /^( {0,3})/ (CommonMark §4.5 ≤3-space indent); the retired
_stripFencedBlocks used /^(\s*)/ (any indent). A 4-space-indented ``` is no
longer treated as a fence opener (it is an indented code block per CommonMark).
Head-to-head over 9 corpus inputs: 0 diffs on all standard cases; 2 diffs only
on the synthetic 4-space-indent edge cases. No UAT fixture or test in the suite
exercises 4-space-indented fences. The change is a correctness improvement.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 15:44:11 -04:00
Tom Boucher
6e9f8bf50b refactor(#1393): migrate roadmap-parser onto markdown-sectionizer seam (epic #1372 T4) (#1395)
Removes all three inline copies of the fenced-code state machine from
src/roadmap-parser.cts and replaces them with calls to
tokenizeHeadings() from the canonical markdown-sectionizer seam
(ADR-1372 T4).

Changes:
- Drop stripFencedLines() function (copy 1 of 3 — standalone helper)
- Rewrite computeSectionEnd() using tokenizeHeadings() offsets into
  the original content (copy 2 — inline fence loop); the returned
  character offset is preserved exactly for all inputs
- Rewrite getMilestonePhaseFilter() versionOverride fence loop using
  tokenizeHeadings() (copy 3 — inline); sectionEnd offset preserved
- Replace stripFencedLines(roadmap) + unanchored phasePattern.exec()
  with tokenizeHeadings(roadmap) filtered by level and phase-heading
  pattern; headings in inline HTML comments (<!-- ## Phase N: -->)
  are no longer mis-counted (anchored ATX detection is correct)
- Add import { tokenizeHeadings } from ./markdown-sectionizer.cjs

Offset preservation: tokenizeHeadings() records h.offset as the
character index of '#' in the ORIGINAL content; computeSectionEnd()
and the versionOverride path both use h.offset directly as the
section-end character offset — no stripping, no shift.

Corpus head-to-head: 29/30 slots are byte-identical to origin/next.
The 1 diff (getMilestonePhaseFilter phaseCount for the HTML-comment
fixture: 3→2) is an improvement: the old unanchored regex counted
"## Phase 998:" embedded in "<!-- ## Phase 998: ... -->" mid-line;
tokenizeHeadings() correctly requires '#' at line-start (ATX rule).
No existing test asserts on that count; feat-3594 passes unchanged.

All 122 roadmap/milestone/phase tests green.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 15:13:14 -04:00
Tom Boucher
3a1961ebae refactor(#1390): migrate check-command-router + gap-checker onto markdown-sectionizer seam (epic #1372 T3) (#1392)
- check-command-router: stripCommentsAndFences delegates fenced-code
  stripping to seam's stripFencedCode; HTML-comment stripping stays
  caller-side. extractPlanDesignatedSections replaces hand-rolled
  split(/\r?\n/) + /^#{1,6}\s+/ heading walk with collectSections
  driven by DESIGNATED_HEADINGS_RE.

- gap-checker: parseRequirements checkbox-bullet detection migrates
  to iterateBullets (checkbox markers); **ID** extracted caller-side.
  Table-row path and separator-row skip stay caller-side.

- Adds refactor-1390-t3-characterization.test.cjs (43 behavioral
  tests) that were green before and remain green after.

- T1 fail-loud / could-not-parse gate semantics unchanged (verified
  by decisions.test.cjs 59/59 and direct gate invocation).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 14:32:43 -04:00
Tom Boucher
ef2c0a2b7b fix(#1389): exempt auto-backmerge PRs from the Require Issue Link gate (#1391)
The required "Issue link required" check failed on automated back-merge
PRs (chore/backmerge-main-to-next-<sha>), which legitimately map to no
issue — a `Closes #N` would pollute the released CHANGELOG. When such a
PR parks for manual review (needs_review), the maintainer could only
merge via --admin.

Carve them out at the failing step's `if:` (step-level, so the required
check still reports SUCCESS rather than a branch-protection-blocking
"skipped"), keyed on the workflow-authored branch name AND same-repo
identity so a fork PR cannot forge the exemption.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 14:24:57 -04:00
Tom Boucher
75c2e259d0 refactor(#1387): migrate adr-parser onto markdown-sectionizer seam (epic #1372 T2) (#1388)
* refactor(#1387): migrate adr-parser onto markdown-sectionizer seam (T2)

Replace hand-rolled parseSections (split/heading-regex walk/body accumulation)
with collectSections(content, () => true) from the canonical seam. Replace
hand-rolled splitEntries bullet-strip regex with iterateBullets from the seam,
preserving the plain-text-line fallback for byte-identical output. Removes the
last inline heading/bullet scanning from adr-parser.cts; normalizeAdrHeader and
all ADR-specific classification logic are unchanged. 218/218 tests pass before
and after.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#1387): keep splitEntries flat (iterateBullets changed its contract); seam adoption stays in parseSections

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#1387): drop dead preamble reconstruction; add targeted adr-parser mutation tests

parseSections' preamble reconstruction block (heading: null entry) was dead code:
both consumers (parseAdrMarkdown and parseStatusFromSections) skip heading: null
sections immediately on entry. Confirmed via analytical trace and zero-diff corpus
head-to-head across all 44 docs/adr/*.md files.

Adds 11 targeted behavioral tests to kill cheap surviving mutants:
- pushUnique intra-values dedup (kills seen.add removal mutant)
- body split/join round-trip with multi-line prose and entries
- parseStatusFromSections [0] indexing (only first line determines status)
- classifyHeader equality vs prefix-match boundary (exact match, prefix match,
  synonym+letter non-match)
- goal section prose vs entries distinction (bullet markers preserved in context)
- normalizeAdrHeader non-word char removal (parens and slash behavior)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 13:45:05 -04:00
Tom Boucher
917d903330 Merge pull request #1379 from open-gsd/chore/backmerge-main-to-next-350fba48
chore: back-merge main → next (350fba48)
2026-06-17 13:33:48 -04:00
Tom Boucher
cfaa3908cd Merge remote-tracking branch 'origin/next' into chore/backmerge-main-to-next-350fba48
# Conflicts:
#	.changeset/924-claude-flat-skill-layout.md
#	.changeset/happy-finches-travel.md
2026-06-17 13:22:28 -04:00
Tom Boucher
e58b5e1721 fix(#1364): decisions adopt markdown-sectionizer seam + fail-loud coverage gate (epic #1372 T1) (#1386)
* test(#1364,#1365): add decisions regression tests (fail-first proof)

Adds tests/decisions.test.cjs with:
- #1364 recall tests: parseDecisions from markdown-header + em-dash bullets
  (these FAIL on pre-T1 code, proving the bug is present before the fix)
- #1365 fail-loud tests: check.decision-coverage-plan must return passed:false
  for decision-shaped but 0-extracted content (FAIL pre-T1, gate silently passed)
- extractDecisions outcome enum tests (could-not-parse/none-present/parsed)
- Parser QA matrix: CRLF, unicode headings, fenced-code suppression, both bullet forms
- Boundary/threshold tests at limit-1 (0), limit (1)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1364,#1365): adopt markdown-sectionizer seam in decisions.cts; add fail-loud gate

#1364 — Recall: decisions.cts now uses the seam's extractTaggedBlocks and
collectSection for the markdown-header fallback path. Em-dash bullet form
(- **D-NN — title** body) is now recognised alongside the existing colon form.

#1365 — Fail-loud: adds extractDecisions() returning a typed DecisionExtraction
{ decisions, outcome } where outcome is 'parsed' | 'none-present' | 'could-not-parse'.
The blocking gate (cmdDecisionCoveragePlan) now treats could-not-parse as
passed:false with a format-mismatch reason instead of the prior silent passed:true/skip.
gap-checker runGapAnalysis surfaces 'extracted 0 of N — possible format mismatch'
for could-not-parse instead of 'No requirements or decisions to check'.

parseDecisions remains a thin delegate over extractDecisions, so all existing
callers are unaffected.

Seam adoption: stripFencedCode (seam), extractTaggedBlocks(content,'decisions') (seam),
collectSection(content, /decisions?/i, {levelBounded,stripFences}) (seam).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1364,#1365): tighten could-not-parse, parse-miss fail-loud, curly-quote discretion, gap-checker FIX D

FIX A: empty <decisions> scaffolds and all-prose sections no longer return
could-not-parse; outcome is none-present unless the block/section contains
a \bD- token or a parse-miss, preventing false blocks on legitimate phases.

FIX B: parseDecisionLines now tracks parse-misses (D-NN-shaped bullets that
fail both regexes); extractDecisions returns could-not-parse when parseMisses>0
even if some decisions parsed — silent drops no longer mask format errors.

FIX C: curly-quote normalization regex now includes actual U+2018/U+2019
characters so '### Claude's Discretion' (curly apostrophe) correctly yields
trackable:false (regression vs pre-T1 behavior).

FIX D: gap-checker runGapAnalysis surfaces the decision could-not-parse
format-mismatch signal independently of whether requirements items exist —
previously masked inside `if (items.length === 0)`.

Adds 14 behavioral regression tests (fail-first verified manually before fixes).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1365): fail-loud gate on parse-miss regardless of covered decisions

Change the `could-not-parse` guard in `cmdDecisionCoveragePlan` and
`cmdDecisionCoverageVerify` from `decisions.length === 0 && outcome ===
'could-not-parse'` to fire on `outcome === 'could-not-parse'` alone.

Previously a CONTEXT.md with a valid D-01 (covered by the plan) plus a
malformed D-02 (parse-miss) would skip the guard (length === 1), proceed
to coverage, find D-01 covered, and silently return passed:true — hiding
the D-02 parse-miss entirely.

Adds a gate-level fail-first test that places D-01 into a ## Must Haves
section (DESIGNATED_HEADINGS_RE match) so coverage of D-01 would pass on
its own, proving the only path to passed:false is the parse-miss fix.
Also adds the matching verify-side advisory assertion.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1364,#1365): add Fixed changeset (pr:0 placeholder)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1364): backfill changeset PR number (1386)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 12:35:38 -04:00
Tom Boucher
afd95a15a9 fix(#1384): scan live changeset fragments in the #1777 purity gate (#1385)
The product-name-purity gate scanned CHANGELOG.md only, never the
.changeset/*.md fragments that render into it. An impure fragment passed
PR review, sat dormant, and re-introduced a forbidden parenthetical
product description at the next release — even after CHANGELOG.md had been
hand-fixed. This is the recurrence vector behind the 1.5.0 back-merge (#1379)
failure.

- Purify the two live fragments to the already-accepted forms:
  - happy-finches-travel.md: "Claude Code (background dispatch …)"
    -> "Claude Code; background dispatch …"
  - 924-claude-flat-skill-layout.md: "Claude (`~/.claude/…`)"
    -> "Claude at `~/.claude/…`"
- Extend the #1777 gate to also scan live .changeset/*.md fragments,
  reusing one shared detection helper. Archived fragments never re-render
  and are intentionally out of scope.

Test-only + changeset-prose change; no production behavior change.

Closes #1384

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 12:14:15 -04:00
Tom Boucher
94ce20089a chore: purify CHANGELOG parenthetical product descriptions (#1777)
The 1.5.0 release section rendered two product-name parentheticals that
the product-name-purity gate (#1777) forbids:

  Claude Code (background dispatch is kept ...)
  Claude (`~/.claude/skills/gsd-ns-<router>/skills/<stem>/SKILL.md`)

Rewritten to the already-accepted forms (semicolon clause; "Claude at
`path`") so the back-merge into next passes the gate next enforces.
No code or behavior change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 11:38:05 -04:00
Tom Boucher
6210621494 test(#1377): pin real agent-skills IR contracts in two vacuous tests (#1380)
The empty-IR test asserted `parsed === '' || typeof parsed === 'object'`,
which always passes (typeof null === 'object'). Pin the real contract:
no agent type → `output('', raw, '')` → the --json IR is the empty string.

The nonexistent-skill-path test asserted only `block === ''`. Since #1376
added a warnings[] field to the --json IR, also assert warnings[] names the
skipped path so the test guards the silent-drop regression it is named for.

Test-only; no product behavior change.

Closes #1377

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 11:20:21 -04:00
Tom Boucher
7ca8011cd9 refactor(#1373): add canonical markdown-sectionizer seam (epic #1372 T0) (#1381)
* refactor(#1373): add markdown-sectionizer seam (ADR-1372 T0)

Establishes the canonical markdown-structure parsing seam per ADR-1372.
No existing parsers are modified; this is the foundational T0 tier only.

- docs/adr/1372-markdown-sectionizer-seam.md: Accepted ADR defining the
  seam interface, the tiered migration plan (T0-T7), and the prohibition
  enforcement approach (no-adhoc-markdown-parsing ESLint rule in T7).
- src/markdown-sectionizer.cts: Pure module, Node built-ins only.
  Exports: stripFencedCode (CommonMark-correct state machine ported from
  uat-predicate.cts _stripFencedBlocks, CRLF-safe, unterminatedFence
  signal), tokenizeHeadings (ATX headings outside fenced blocks),
  collectSections (line-by-line predicate-driven section collection),
  collectSection (single named section, levelBounded stop, optional
  stripFences), iterateBullets (dash/checkbox/numbered + continuation).
- tests/markdown-sectionizer.test.cjs: 54-test behavioral suite covering
  the parser QA matrix (LF/CRLF, Unicode headings, headings-inside-fences,
  unterminated fences, nested levels, all bullet markers, continuation
  lines, empty/non-string input) plus 4 fast-check property tests
  (idempotence, output shape, never-throws, length monotonicity).
- CONTEXT.md: Markdown Sectionizer glossary entry added (PR review gate).

Tests: 54 pass, 0 fail. Existing adr-parser + uat-passed tests: 22 pass.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(#1373): add extractTaggedBlocks + replaceSection to seam; register inventory

- src/markdown-sectionizer.cts: extend Section type with bodyStart/bodyEnd offsets;
  add extractTaggedBlocks(content, tagName) (inner text of <tag>…</tag> blocks,
  tagName regex-escaped, caller decides fence-stripping) and replaceSection(content,
  section, newBody) (pure character-offset splice for read-modify-write callers);
  update collectSections/collectSection to populate bodyStart/bodyEnd.
- tests/markdown-sectionizer.test.cjs: add 33 new behavioral tests for
  extractTaggedBlocks, replaceSection, and a DEFECT.GENERATIVE-FIX parity guard
  that asserts stripFencedCode and uat-predicate's _stripFencedBlocks agree on a
  shared 9-item corpus; documents the known 4-space-indent divergence.
- docs/adr/1372-markdown-sectionizer-seam.md: list extractTaggedBlocks and
  replaceSection in §"The seam".
- CONTEXT.md: update ### Markdown Sectionizer glossary entry with the two new exports.
- docs/INVENTORY.md: add markdown-sectionizer.cjs row (alphabetically between
  loop-resolver and milestone).
- docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest --write.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1373): clear no-unsafe-assignment + unused-var lint in markdown-sectionizer

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1373): correct section offset/round-trip + CommonMark heading/fence edges; register eslint coverage

FIX 1 (CRITICAL): Enforce content.slice(bodyStart,bodyEnd) === body invariant in both
collectSection and collectSections. bodyEnd is now bodyStart + body.length instead of
the raw stop-line offset, eliminating the trailing-newline overcounting that caused
replaceSection to drop separator newlines (## A\nbody## B gluing bug).

FIX 2 (MED): tokenizeHeadings now accepts ≤3-space indent (CommonMark §4.5) and empty
ATX headings (## / ##   ), text=''. 4-space indent correctly excluded.

FIX 3 (MED): collectSection gains stopAtLevel option — stops at the next heading whose
level ≤ stopAtLevel, independent of the opener's level. Enables state.cts ## sections
that also stop at ### without abusing levelBounded.

FIX 4 (MED): Backtick fence opener info string must not contain a backtick (CommonMark).
Applied in both stripFencedCode and tokenizeHeadings fence state machines. Tilde fences
unaffected.

FIX 5 (LOW): "byte offset" → "character (string-index) offset" in HeadingToken / Section
doc comments.

FIX 6 (LOW): extractTaggedBlocks doc comment documents nested-tag non-support; test locks
the non-greedy close-at-first-</tag> behavior.

FIX 7: Add gsd-core/bin/lib/markdown-sectionizer.cjs to eslint.config.mjs ignores so
tests/551-eslint-bin-lib-coverage.test.cjs passes (3/3).

Tests: 107 pass / 0 fail (was 87; +20 new tests for FIX 1–4, 6).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1373): gitignore tsc-built markdown-sectionizer.cjs (ADR-457 build-at-publish)

The seam's compiled artifact must be a build-at-publish output like every other
src/*.cts->bin/lib/*.cjs module (decisions, core, state, ...), not a committed
file. Add it to the ADR-457 ignore list and untrack it; build:lib/CI regenerate it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-17 11:14:17 -04:00
Tom Boucher
1d7c16a1d0 chore: bump next to 1.5.1-dev.0
Move next onto the -dev prerelease stream after the v1.5.0 release per
ADR-660 (next must not rest at the last-released version).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 10:31:51 -04:00
github-actions[bot]
0dd5194c8f chore: sync next package version to 1.5.0 2026-06-17 14:28:46 +00:00
github-actions[bot]
6a2b6729bc chore: back-merge main into next (350fba48) 2026-06-17 14:28:45 +00:00
Tom Boucher
350fba48b5 Merge pull request #1378 from open-gsd/release/1.5.0
chore: merge release v1.5.0 to main
2026-06-17 10:28:24 -04:00
github-actions[bot]
e67b1d7f91 chore: promote CHANGELOG for v1.5.0 2026-06-17 14:23:45 +00:00
github-actions[bot]
ee3bc368be chore: finalize v1.5.0 2026-06-17 14:17:30 +00:00
Tom Boucher
66085d0080 fix(#1374): surface diagnostic when configured agent skills all fail to resolve (#1376)
* fix(#1374): surface diagnostic when configured agent skills all fail to resolve

buildAgentSkillsBlock returned '' (only ad-hoc per-path stderr warnings) when an agent configured via agent_skills had paths that all failed to resolve — missing SKILL.md, unsafe path, invalid global name, OR a malformed (non-string/non-array) value. query agent-skills --json reported skills_count>0 with an empty block and no machine-readable signal, so a fully-dropped configuration was indistinguishable from a resolved one.

Thread an optional diagnostics collector through buildAgentSkillsBlock: route every skip warning through a warn() helper (stderr + collector), flag truthy-but-malformed config values, emit an aggregate warning when configured paths resolve to zero skills, and surface the collected reasons in a new warnings[] field on the query agent-skills --json IR. Empty arrays and falsy values stay silent (skills_count is honestly 0). skills_count semantics unchanged. Docs updated for the new IR field.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1374): backfill changeset PR number (#1376)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 10:16:51 -04:00
Tom Boucher
120f85164b feat(#1355): detect-and-warn guard for claude-code agent-teams (#1371)
* feat(#1355): detect-and-warn guard for claude-code agent-teams

GSD's multi-agent orchestration can stall under claude-code's experimental
agent-teams (a subagent's completion fails to route to the orchestrator). Per
the maintainer decision, the accepted scope is a read-only detector + one
non-fatal warning — NOT the declined run_in_background/TaskOutput conversion.

- New Teams Status Module (src/teams-status.cts → gsd-core/bin/lib/teams-status.cjs):
  pure resolveTeamsStatus({runtime, env}) + thin CLI cmdTeamsStatus reusing
  resolveRuntime. active = strictly-truthy env flag AND runtime === 'claude'.
- Wire `gsd-tools query teams-status [--active]` (read-only; no capability
  registration needed — conformance gates govern features, not query commands).
- One non-fatal warning in plan-phase.md before the first Agent spawn, gated on
  `query teams-status --active`; zero behavior change on non-claude/teams-off.
- Hermeticity: clear CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS in run-tests.cjs +
  SESSION_ENV_KEYS. Docs reference + CONTEXT.md glossary. Built lib gitignored.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1355): add changeset for teams-detect guard

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1355): bump plan-phase.md workflow size baseline (+407B for teams warning)

The non-fatal agent-teams warning block added to plan-phase.md grew it
92759 → 93166 bytes, past its committed per-file baseline ratchet. The growth
is small, deliberate, and still well under the workflow tier hard cap. Regenerate
the baseline via `npm run size:baseline` (only plan-phase.md changed).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1355): register teams-status.cjs in the inventory manifest

The new teams-status CLI module is a tracked surface; regenerate
docs/INVENTORY-MANIFEST.json (cli_modules family) via
gen-inventory-manifest.cjs --write so the inventory-manifest-sync gate passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 08:52:23 -04:00
Tom Boucher
08e42b0ef1 fix(#1356): rewrite bare ~/.claude paths in the Cursor install branch (#1368)
* fix(#1356): rewrite bare ~/.claude paths in the Cursor install branch

The Cursor branch of _applyRuntimeRewrites only rewrote the trailing-slash
.claude forms, so bare ~/.claude / $HOME/.claude references survived into
installed Cursor artifacts (skills/gsd-surface, skills/gsd-graphify,
workflows/plan-phase, workflows/autonomous), tripping the post-install
"unreplaced .claude path reference(s)" audit. Same regression class as
#983/#2418/#2545 — every other affected branch was patched; cursor was missed.

- Add the three bare-form rewrites (~/.claude, $HOME/.claude, ./.claude) to the
  cursor branch, mirroring cline/trae/augment/codebuddy. They run after the
  slash forms (no double-replace) and use (?![\w-]) so .claude-plugin /
  .claudeignore are not corrupted.
- SKILL.md content also passes through this stage, so no stage-1 converter
  change is needed. Verified: 40 bare refs in the real leaking files → 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1356): add changeset for Cursor bare-path rewrite fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 08:51:48 -04:00
Tom Boucher
c03f3cc6af fix(#1348): canonicalize Codex hooks.json writes to the nested { hooks } shape (#1363)
* fix(#1348): canonicalize Codex hooks.json writes to the nested { hooks } shape

reconcileCodexHooksJsonEvent preserved whatever shape it read, so on an empty,
absent, or legacy top-level hooks.json it wrote top-level event keys
(`{ "SessionStart": [...] }`) that current Codex (deny_unknown_fields) rejects,
instead of the canonical `{ "hooks": { "SessionStart": [...] } }`.

- Lift any top-level event arrays (legacy, empty, or mixed nested+top-level)
  into the nested `hooks` table, merging same-named events so no user/legacy
  entry is dropped and no stray top-level event key survives. Mirrors
  reconcileCursorHooksJson.
- Collapse an empty hook table back to `{}` so removal on an absent file does
  not write a spurious `{ "hooks": {} }`.
- Read path still tolerates both shapes; dedup/removal unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1348): add changeset for Codex hooks.json canonicalization

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 08:51:11 -04:00
Tom Boucher
07adeb50a0 fix(#1342): scope worktree-path-guard to GSD executor runs; fail open for no-repo targets (#1361)
* fix(#1342): scope worktree-path-guard to GSD executor runs; fail open for no-repo targets

The PreToolUse worktree-path-guard fired for any Write/Edit in any linked git
worktree, with no check for active GSD work — so Claude Code plan-mode writing
~/.claude/plans/<slug>.md from a manually-created worktree was hard-blocked.

- Gate enforcement on the GSD isolated-executor branch namespace
  (^worktree-agent-[A-Za-z0-9._/-]+$, per worktree-branch-check.md #2924); the
  guard is a no-op in non-GSD linked worktrees.
- Fail open when a target resolves to no git repository (e.g. ~/.claude/plans/)
  instead of blocking — that is not the #260 main-repo vector. A target inside
  a .git directory still blocks (git rev-parse --is-inside-git-dir).
- The #260 different-git-root hard block (escape to the main repo) is preserved.

Detached-HEAD executors no-op the gate; this is accepted because they are
fail-closed by worktree-branch-check.md (exit 42) before committing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1342): add changeset for worktree-path-guard scoping fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1342): build dot-dot traversal path portably (Windows drive-letter fix)

The traversal test built its file_path by stripping a leading slash from an
absolute externalDir and path.join-ing it after a `..` chain. On Windows the
drive letter (C:\) is not a leading slash, so it survived and path.resolve
produced an invalid doubled-drive path (C:\C:\Users\...), which resolves to no
git repo — the hook failed open (exit 0) and the test expected a block (exit 2).

Use path.relative(worktreeDir, externalTarget) + string concat so the file_path
carries literal `..` segments that resolve to externalTarget on both posix and
win32 (no drive doubling). Verified with path.win32/path.posix.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 08:50:28 -04:00
Tom Boucher
c53fd1f654 fix(#1324): resolve glued phase tokens (#1353) 2026-06-16 21:55:58 -04:00
Tom Boucher
c4161735dc fix(#1325): scope update backup detection (#1354) 2026-06-16 21:55:55 -04:00
Tom Boucher
aab26c7bf4 fix(#1343): parse decision bullets with text before the colon (#1358)
* fix(#1343): parse decision bullets with text before the colon

parseDecisions() silently dropped any `- **D-NN ...:**` decision bullet
whose header had freeform text (a parenthetical, em-dash, or prose) before
the `:**`, so the blocking check.decision-coverage-plan gate computed
coverage over a narrowed set and reported a false pass.

- Broaden bulletRe to tolerate a freeform run before the colon while
  preserving the optional [bracket] tag capture (drives `trackable`).
- Add a parse-miss guard: a line that looks like a D-NN bullet but still
  fails the regex flushes the current decision and warns instead of
  vanishing — the gate-integrity floor.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1343): add changeset for decision-coverage false-pass fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1343): relocate decision-parser regression into owning module test file

CI's lint-regression-test-names bans new bug-NNNN-*.test.cjs files. Move the
9 regression cases from tests/bug-1343-parsedecisions-drop.test.cjs into the
owning parser test file tests/post-planning-gaps-2493.test.cjs (which already
exercises parseDecisions) and delete the banned file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-16 21:55:51 -04:00
Tom Boucher
f3c06f59df fix(#1326): stop emitting Codex agents/openai.yaml sidecars; clean up stale ones (#1360)
* fix(#1326): stop emitting Codex agents/openai.yaml sidecars; clean up stale ones

Codex installs wrote an agents/openai.yaml sidecar under every managed gsd-*
skill dir. Recent Codex builds index both SKILL.md and the sidecar, so each
GSD skill appeared twice in autocomplete (canonical gsd-* name + humanized
display_name).

- Replace writeCodexSkillMetadataFiles / generateCodexSkillMetadataYaml with
  cleanupCodexSkillMetadataSidecars: Codex-only (if isCodex), removes stale
  managed gsd-*/agents/openai.yaml and prunes the now-empty agents/ dir.
- Preserve user-owned dirs (gsd-dev-preferences), non-empty agents/ dirs, and
  non-gsd dirs; lstat-guard against symlinked agents/ so a delete can never
  escape the skills tree; fail-open per directory.
- Codex relies on SKILL.md alone for /skills discovery.
- Update USER-GUIDE/FEATURES docs and rewrite the #774 emission tests into
  cleanup tests.

Scope: the sidecar duplicate only. The separate multi-root (~/.agents/skills
shared-skills) duplicate facet is a distinct concern, not addressed here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1326): add changeset for Codex sidecar cleanup

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-16 21:55:49 -04:00
Tom Boucher
1a678eb0e9 fix(#1359): migrate workflows off deprecated TaskOutput to Read(outputFile) (#1362)
* fix(#1359): migrate workflows off deprecated TaskOutput to Read(outputFile)

The map-codebase and docs-update workflows collected background sub-agent
results with the deprecated Claude Code `TaskOutput` tool using `block: true`,
which has a confirmed main-session hang after the agent completes
(anthropics/claude-code#20236).

Migrate the collection steps to the upstream-recommended pattern: keep
`run_in_background=true` on the Agent spawn, then `Read` each agent's
`outputFile` (from the `async_launched` result) once it reports completion.
Completion-marker contracts and on-disk verification are unchanged, and the
non-Claude runtime fallbacks (sequential_mapping / sequential_generation) are
preserved byte-for-byte. docs-update's timeout note no longer references the
unwired `workflow.subagent_timeout` key (it kept a literal before).

Regression coverage folded into tests/subagent-timeout.test.cjs (the owning
module for background-subagent collection). Workflow size baseline regenerated
for the justified prose growth.

Refs #1355 (same latent hang surface).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1359): add changeset for TaskOutput migration

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-16 21:55:45 -04:00
Rezolv
1698e83336 Merge pull request #1314 from davesienkowski/feat/1279-fail-first-prover 2026-06-16 18:12:13 -04:00
Rezolv
00c05eb717 Merge branch 'next' into feat/1279-fail-first-prover 2026-06-16 17:35:55 -04:00
Tom Boucher
a0dbf8bbdf fix(#1319): use portable Claude skill effort (#1352) 2026-06-16 15:30:17 -04:00
Tom Boucher
c20d741dc9 fix(#1316): preserve prose STATE phase names (#1351) 2026-06-16 15:11:23 -04:00
Tom Boucher
284dc7bc44 fix: resume UAT checkpoint from paused placeholder (#1350) 2026-06-16 14:47:17 -04:00
Tom Boucher
9540fe43b9 fix: have executor self-report worktree metadata (#1349) 2026-06-16 14:19:08 -04:00
Dave
56d4a1bf39 enhance(#1279): project check_violation_fixture scalar — #1278 locate + #1279 proof compose end-to-end (#1346)
Delivers option (a) from the #1314 maintainer review: thread a fourth flat
scalar check_violation_fixture through the projection so a prohibition authored
at spec-phase machine-proves fail-first and greens through the deterministic
path alone — zero hand-authoring at verify time.

- src/probe-core.cts: Prohibition gains check_violation_fixture?; projectProhibitions
  emits it (both kinds) ONLY for a well-formed descriptor and ONLY when non-empty
  (blank/absent -> projects absent -> producer hard-gates, never a partial green).
- src/prohibition-enforcement.cts: descriptorFromProjection reads it back into
  violationFixture via the same numeric-coercion-safe scalar() normalizer.
- Tests (RED-first, proven non-vacuous by reverting both src edits): CHK-02(#1346)
  projection emit, CHK-08(#1346) read-back, CHK-03(D) example round-trip, the
  fast-check round-trip property extended to the 4th scalar (the contract trek-e
  blocked #1301 on), and a real-subprocess COMPOSE capstone greening end-to-end
  through project -> descriptorFromProjection -> default prover+runner.
- Docs flipped from 'hard-gates until #1346' to 'composes end-to-end': verify-phase.md,
  prohibition-probe.md, spec-phase.md authoring, ADR-550 addendum, changeset.
  #1346 now tracks only the node-test causation residual.

190 affected-suite tests green; eslint + tsc clean; size baseline regenerated.
2026-06-16 14:00:49 -04:00
Tom Boucher
6e242bd76a fix: allow quick worktree parent plan base (#1347) 2026-06-16 14:00:06 -04:00
Dave
91bc49c9f1 test(#1279): regenerate workflow size baseline for verify-phase.md (Major 2 prose) 2026-06-16 13:15:31 -04:00