f9d9dfb4bc445fe1ee3670b9b0aa83cae3028591
233 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0c4d570541 |
fix(#1628): type-safe config-set validation — close JSON-coercion enum bypass + enforce capability schema
Three related defects in cmdConfigSet, all 'config-set stores invalid values silently': 1. Missing guards: workflow.security_block_on (enum) and workflow.security_asvs_level (integer 1-3) had no store-time validation. 2. Systemic JSON-coercion bypass: every string-enum guard used VALID_X.includes(String(parsedValue)). Because the value is JSON-parsed before validation, String(["member"]) === "member" let a JSON array slip through and an array was stored in a scalar key. Reproduced on human_verify_mode, statusline.context_position, context_guard_mode, fallow.scope/profile, source_grounding_authority, drift_action, context. 3. Unvalidated capability keys: 32 capability-registry-owned keys (4 enum, 25 boolean, 2 number, 1 string) had no hardcoded guard, so any value — including coerced arrays/objects and out-of-enum strings like code_review_depth=garbage — was stored silently. Fix: a type-safe assertEnumValue() helper (requires typeof === 'string' before membership), routed through all nine central string-enum guards (messages preserved byte-for-byte); plus a generic capability-registry validation block that validates every capability key against its declared type/values (enum via the registry's values — single source of truth — boolean, number, string). Behavioral regression tests cover every central enum key and representative capability keys (array + object coercion rejected, out-of-enum rejected, valid accepted) with boundary coverage for the security keys. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c63fa35b0a |
fix(#1615): allowlist windsurf-conversion.test.cjs in prompt-injection-scan
The commandName validation tests legitimately contain real injection payloads (newline + system-role override phrases, fake [SYSTEM] tags, jailbreak strings) to prove the validator rejects them. The scanner cannot distinguish a test fixture asserting rejection from an actual injection attempt, so CI failed on the test that adds the security control.
Added tests/windsurf-conversion.test.cjs to scripts/prompt-injection-scan.sh ALLOWLIST with a comment citing the defect class.
Also added DEFECT.PROMPT-INJECTION-SCAN-COLLISION-WITH-TESTS to CONTEXT.md so the pattern is documented. Initial draft of that predicate ITSELF triggered the scanner (it quoted the literal injection phrase as an example) — reworded to use descriptive references ('scanner-matching payload', 'instruction-override phrase') since the scanner scans CONTEXT.md too. That meta-collision is now called out in the fix-forward and prevention subkeys.
|
||
|
|
652142521b |
enhance(#1549): validate PR-title issue-ref convention at open time (#1576)
* enhance(#1549): validate PR-title issue-ref convention at open time The release changelog is title-driven: release.yml generates "What's Changed" from PR titles, then format-github-release-notes.cjs buckets each line by its conventional-commit prefix and relies on a `(#<issue>)` in the title to render the issue link. Both rules were enforced only socially, so titles like `fix(core): ...` (no issue link) and `[security] fix(...): ...` (leading tag defeats the `^fix` bucket anchor -> mis-filed under Enhancement) silently broke the changelog, landing on the maintainer as release-time cleanup. Extract the title matcher into one shared module consumed by BOTH the changelog classifier and a new PR-title CI gate, so a title that passes the gate cannot mis-bucket in the changelog (single source of truth). - scripts/lib/conventional-title.cjs (new): classifyBucket + evaluatePrTitle + the anchored regexes. One matcher, two consumers. - scripts/release-notes/format-github-release-notes.cjs: classifyTitle now delegates to classifyBucket (behavior preserved; existing tests green). - .github/workflows/pr-title-validator.yml (new): runs evaluatePrTitle on pull_request opened/edited/reopened/synchronize, for ALL authors (the drift came from member PRs). Trusted base-ref checkout; WARN_ONLY knob for rollout. - tests/conventional-title.test.cjs (new): bucket + gate cases incl. the leading-tag mis-bucket (backfills the untested classifyTitle case) and a cross-check that the classifier delegates to the shared matcher. - CONTRIBUTING.md: document the `type(#<issue>):` rule and no-leading-tag. Claude-Session: https://claude.ai/code/session_01UMV5Qr3H4oFikbuiEauGQk * fix(#1549): check out the PR in pr-title-validator so the new matcher resolves The workflow checked out the base branch (next) as a trusted policy source, but the shared matcher (scripts/lib/conventional-title.cjs) is introduced by this PR and does not exist on next yet — so require() failed and validate-title errored on its own introducing PR. Check out the PR's merge ref instead: the matcher under review is present, the check is self-consistent, and a fork pull_request runs read-only with no secrets, so running the PR's own pure-string regex is safe. * fix(#1549): move conventional-title.cjs out of installed scripts/lib/ bin/install.js bundles every file under scripts/lib/ into the user-installed payload (the changeset CLI's dependencies), and install.test.cjs (#935) asserts that exact set. The new matcher is release/CI tooling that must NOT ship to users, so placing it in scripts/lib/ both broke the install manifest test and would have shipped dead code. Relocate it next to its consumer in scripts/release-notes/ (which the installer does not copy) and update the three require paths (classifier, workflow, test) + the CONTRIBUTING reference. install.test.cjs now 125/125; conventional-title + release-notes suites green; lint:ci clean. * fix(#1549): load title matcher from trusted base ref, not PR code Addresses review (Solvely-Colin + trek-e): the gate checked out the PR merge ref and require()'d evaluatePrTitle from PR-controlled code, so any future PR could edit conventional-title.cjs to return { valid: true } and wave its own malformed title through — a self-bypassable required check. Load the matcher from a base-branch checkout instead (ref: github.event.pull_request.base.ref), the same trusted-policy-source pattern pr-target-validator.yml already uses. The PR can change its title but not the ruler that measures it. An existsSync bootstrap guard skips the check when the matcher isn't on the base branch yet (the introducing PR); every PR after merge is fully gated. This keeps the single shared matcher (#1549's whole point) rather than forking the regex into the workflow. Also per review: - add tests/conventional-title.property.test.cjs (fast-check): any `type(#n): summary` round-trips to valid; evaluatePrTitle/classifyBucket are total functions (never throw). - pin the `fix(#):` zero-digit boundary as missing-issue-ref. Claude-Session: https://claude.ai/code/session_01VqUHNQCh71pEqjo96zkgQL --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
da4a86d8c1 |
feat(#1596): ship GSD skills via .claude-plugin/plugin.json
Phase B-provide of epic #1258. Adds a build-generated skills/ dir + a skills manifest field so plugin-installed GSD exposes gsd-core:<skill> the native Claude Code way. Closes the gap where plugin-only installs lacked the skill surface because bin/install.js never ran. - scripts/gen-plugin-skills.cjs: build step converting commands/gsd/*.md to skills/gsd-<stem>/SKILL.md via convertClaudeCommandToClaudeSkill - .claude-plugin/plugin.json: add "skills": "./skills/" - package.json: add skills to files, gen:plugin-skills to build chain - tests/issue-766-plugin-manifest.test.cjs: Section H conformance (manifest field + dir + frontmatter + count parity) + C2 skills symlink - docs/adr/766-*.md: dated amendment adding skills surface row - .changeset/rapid-bears-hum.md: type Added - skills/: 69 generated gsd-<stem>/SKILL.md files (build-committed) Closes #1596 |
||
|
|
e12a2abfd8 |
feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed), enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way to browse or audit parked seeds on demand. This adds a read-only listing, following the established --list → workflow pattern (per the approved scope on - gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the seeds dir, returns { count, seeds[], summary } JSON with each seed's id, slug, status, scope, trigger_when, planted, title. Optional case-insensitive status filter. User-controlled content is sanitized (sanitizeForDisplay) and every path validated (requireSafePath); read-only. Independent of audit.scanSeeds, which only returns unimplemented seeds for the milestone surface. - /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that renders the seed table. Closes #441 * chore(#441): point changeset fragment at PR #722 * test(#441): allowlist list-seeds test in prompt-injection scan The test asserts that list-seeds neutralizes injection payloads (<system>, [INST]) embedded in seed content, so the fixtures legitimately contain those patterns — same as the sibling security tests already on the allowlist. * fix(#441): use canonical /gsd:capture colon form in list-seeds workflow Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new list-seeds workflow used the hyphen form. * docs(#441): sync help full.md + INVENTORY for --list-seeds Adds the --list-seeds entry to the help reference (help/modes/full.md, per bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json. * docs(#441): add --list-seeds how-to + drop phantom statuses Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers): - USER-GUIDE.md Seeds section (how-to): extend the task to cover auditing parked seeds on demand via --list-seeds, including the status filter — kept task-oriented per Diataxis how-to mode. - CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected from the list-seeds filter vocabulary; the system only produces dormant|active|triggered (src/audit.cts scanSeeds). Reference must be factually accurate and complete. * fix(#441): guard non-scalar status frontmatter in cmdListSeeds A seed with a bare `status:` line (extractFrontmatter yields {}) or a `status: [a, b]` value (yields an array) crashed the whole audit list: `(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string. Coerce every frontmatter read through a `fmStr` helper (mirrors the existing `typeof fm.id === 'string'` guard), so a non-scalar status falls back to dormant and non-scalar scope/trigger_when/title can no longer leak a raw array/object into the JSON contract. Title is now capped symmetrically. Adds regression coverage for empty and array `status:` and non-scalar fields. Refs #441 * docs(#441): align list-seeds workflow status vocabulary The load_seeds step listed `implemented` as an example status filter, but the real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds); `implemented` has no producer. Matches the earlier CLI-TOOLS.md correction. Refs #441 * refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested in-process (review minor #1). No behavior change. Filter comparison now matches the raw lowercased status (both sides already normalized) instead of sanitizeForDisplay(status); sanitization is for output, not matching (review nit #3). * test(#441): add fast-check property coverage and count=1 boundary for list-seeds Adds tests/list-seeds.property.test.cjs with four fast-check properties over deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing (review minor #1). Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2). * chore(#441): sync runtime launcher snippet into list-seeds workflow Propagate the current _runtime-launcher.snippet.sh (with non-Claude runtime home probes) into the new list-seeds.md workflow via scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation. * test(#441): record list-seeds.md in workflow size baseline (#1074) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
fb24d1c7ea |
test(#1496): add behavioral validateCapability check for capability tutorial manifests (#1497)
* test(#1464): add behavioral manifest-validation test for capability tutorial docs Extracts JSON capability manifests from tutorial/reference docs and validates them through the real validateCapability — closing the test gap that let issue #1464's broken tutorial manifest (step missing ref) pass undetected. Adds fix-1464-docs-manifest-validation.test.cjs with: - Suite 1: build/install/reference docs' complete manifests pass validateCapability - Suite 2: adversarial fixtures prove the original #1464 bug shapes are caught (step without ref → "steps[0].ref must be an object…"; id/folderId mismatch) - Suite 3: extractManifests helper unit tests (complete vs. partial block filtering) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#1496): allowlist docs module for 3-file test cluster docs-parity-live-registry, docs-update, and the new fix-1464-docs-manifest-validation sit on the same docs production module; add the docs entry to lint-test-file-count.allowlist.json so the novel-offender CI gate passes. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
33ccf5f89d |
fix(#1367): project-local install uses flat gsd-<cmd>.md layout (fixes /gsd: colon namespace) (#1489)
* fix(#1367): project-local install uses flat gsd-<cmd>.md layout Claude Code project-local installs now write command files as flat gsd-<cmd>.md at .claude/commands/ level instead of commands/gsd/<cmd>.md (subdirectory), so Claude Code registers /gsd-<cmd> (hyphen form) matching hooks, statusline, and all cross-command references. - capabilities/claude/capability.json: local destSubpath commands/gsd → commands - bin/install.js else branch: flat gsd-<stem>.md loop with runtime rewrites - bin/install.js uninstall (1c): remove flat files + legacy subdir cleanup - bin/install.js writeManifest: record flat commands/gsd-<cmd>.md keys - legacy migration: preserves dev-preferences.md across reinstall and uninstall - gsd-core/bin/lib/capability-registry.cjs: regenerated - 6 new regression tests (L0–L5) in bug-1367-*.test.cjs - Updated E suite in bug-3683 + bug-1736, layout + surface + descriptor tests - scripts/lint-regression-test-names.allowlist.json: grandfathered bug-1367 test Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1367): add issue reference to allow-test-rule comment Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
fa1ffb4824 |
fix(#1437): add phase.list-plans to gsd-tools (#1485)
* fix(#1437): add phase.list-plans to gsd-tools Register phase.list-plans in PHASE_COMMAND_ALIASES, implement cmdPhaseListPlans in src/phase.cts (uses findPhaseInternal + scanPhasePlans to return plan_count/has_plans/plans/phase_dir), and wire the handler in phase-command-router. Previously every call produced "Unknown phase subcommand". Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1437): register new test file in lint-test-file-count allowlist Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1437): rename test to fix-NNN convention; update file-count allowlist Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
0d56f544d2 |
feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation (#1458)
* feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation ADR-1244 Phase 6. Adds the capability matrix reference, generated FROM the committed registry so it can never drift from the actual capability set: - scripts/gen-capability-matrix.cjs (--write / --check); --check is a CI drift guard. - tests/capability-matrix-sync.test.cjs (4 tests): drift guard, buildMatrix==committed, every cap present, no placeholders. - docs/reference/capability-matrix.md regenerated from the registry (release-stable: shows engines.gsd, omits the lockstep per-cap version that would churn the file every release). - Consolidated the duplicate trust-model doc: deleted docs/explanation/the-capability-trust-model.md, merged its content into capability-trust-model.md, redirected ~10 references; no stale links remain. - Diataxis verification (now that gsd capability is a real command): corrected the matrix's third-party section — the matrix is the first-party catalogue; the overlay-aware view of installed third-party capabilities is 'gsd capability list', not this generated file. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1435): Added changeset for the capability matrix reference Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1435): address code-review — non-vacuous matrix test + generator polish - capability-matrix-sync.test.cjs: assert the 'security registers a ship:pre gate' precondition unconditionally so the extension-point check can never degrade to a vacuous pass on registry drift. - gen-capability-matrix.cjs: warn (stderr) on an unknown loop point at generation time; rename enginesOf -> fmtEngines for consistency with the other fmt* helpers (output unchanged). - capability-trust-model.md: point the two how-to links at the real files (import-a-capability-from-a-url.md, version-a-capability.md) instead of the bare directory. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1435): backfill changeset PR number → #1458 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
353f63d170 |
feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) Promote the registry from a frozen data file to loadRegistry({includeInstalled}), composing the first-party registry with a validated installed overlay (ADR-1244 D2): - Extract the conformance validator to a shared runtime-callable module (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it verbatim, guarded by a generative-parity test (no build-time/runtime drift). - capability-loader.cts: loadRegistry({includeInstalled}) composes first-party ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic- prefixes); full merged-set cross-capability validation; engines.gsd load-time re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path escapes rejected. - semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed. - Wire surface/state + loop to the overlay; loop injects a blocking gate for each skipped gate-kind overlay (fail-closed). - cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd) + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/ config-set call (never eager at module load, never wrong-cwd); first-party path unchanged with no cwd. - run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact); capability-validator.cjs stays linted (#551 migration coverage). Closes #1431 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1431): add changeset for runtime capability registry overlay Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52) The cwd-aware overlay config-key federation added to config-schema.cts (_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd threading) introduced mutable surface uncovered by config-schema's mutation test set, dropping its score to 39.58% (below the 52 break threshold). Add a real-overlay-fixture describe block exercising every branch (cwd guard, overlay loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker score 39.58% -> 77.08%. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2421cf1b4a |
feat(#1430): versioned capability manifest + native stamping (ADR-1244 Phase 1) (#1436)
* feat(#1430): versioned capability manifest + native stamping (ADR-1244 Phase 1) Make the capability manifest versioned — the data substrate the Capability Ecosystem (ADR-1244) keys off: - capability.json gains a REQUIRED semver `version` plus the optional ecosystem envelope (`engines.gsd`, `compatVersions`, `integrity`, `provenance`); the build-time conformance validator enforces them via a new `validateVersionEnvelope()` (exported for the Phase 2 runtime overlay). - All 32 native capabilities stamped with `version` (= package version, lockstep) + `engines.gsd`; `sync-manifest-versions.cjs` gains a glob sweep that keeps them in sync, and the issue-844 regression guard is extended. - Strict SemVer 2.0.0 grammar blocks metacharacter/space/unicode smuggling in version strings; range/integrity fields are shape-validated (satisfaction and the load-time gate are deferred to Phase 2/4). - Capability rel-paths emitted forward-slash for cross-platform git correctness. Closes #1430 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1430): add changeset for versioned capability manifest Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8040a6bac0 |
chore(#1417): add resolution-provenance CI guard (Resolution Provenance P4) (#1428)
Adds scripts/lint-resolution-provenance.cjs — a registry + ratchet CI guard that locks in the agent-skills configured_empty/not_configured contract tests so they cannot be silently removed, and establishes a registration point for future config-interpreting read verbs (ADR-1411 P4). Design rationale: - REGISTRY (one entry: agent-skills → src/init.cts → tests/agent-skills.test.cjs) is the canonical registration site; new verbs are added here. - For each registered verb, the guard asserts its test file contains BOTH a `configured_empty` assertion AND a `not_configured` assertion — proving the configured-empty-vs-not-configured contract is explicitly tested. - NOT a universal static detector (intractable / false positives) — mirrors the no-adhoc-markdown-parsing grandfather pattern. - Uses scripts/lib/allowlist-ratchet.cjs (assertWithinAllowlist) so stale allowlist entries fail (ratchet-down) and novel offenders always fail. - checkRegistry() is factored as a pure exported function tested in tests/lint-resolution-provenance.test.cjs without shelling out. - Wired into lint:ci (package.json) and lint step name updated in test.yml. - CONTEXT.md ### Resolution Convention extended with P4 guard sentence. - Allowlist starts empty ([]) — agent-skills already has its tests. Closes #1417 Part of #1411 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
120f85164b |
feat(#1355): detect-and-warn guard for claude-code agent-teams (#1371)
* feat(#1355): detect-and-warn guard for claude-code agent-teams GSD's multi-agent orchestration can stall under claude-code's experimental agent-teams (a subagent's completion fails to route to the orchestrator). Per the maintainer decision, the accepted scope is a read-only detector + one non-fatal warning — NOT the declined run_in_background/TaskOutput conversion. - New Teams Status Module (src/teams-status.cts → gsd-core/bin/lib/teams-status.cjs): pure resolveTeamsStatus({runtime, env}) + thin CLI cmdTeamsStatus reusing resolveRuntime. active = strictly-truthy env flag AND runtime === 'claude'. - Wire `gsd-tools query teams-status [--active]` (read-only; no capability registration needed — conformance gates govern features, not query commands). - One non-fatal warning in plan-phase.md before the first Agent spawn, gated on `query teams-status --active`; zero behavior change on non-claude/teams-off. - Hermeticity: clear CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS in run-tests.cjs + SESSION_ENV_KEYS. Docs reference + CONTEXT.md glossary. Built lib gitignored. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1355): add changeset for teams-detect guard Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1355): bump plan-phase.md workflow size baseline (+407B for teams warning) The non-fatal agent-teams warning block added to plan-phase.md grew it 92759 → 93166 bytes, past its committed per-file baseline ratchet. The growth is small, deliberate, and still well under the workflow tier hard cap. Regenerate the baseline via `npm run size:baseline` (only plan-phase.md changed). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1355): register teams-status.cjs in the inventory manifest The new teams-status CLI module is a tracked surface; regenerate docs/INVENTORY-MANIFEST.json (cli_modules family) via gen-inventory-manifest.cjs --write so the inventory-manifest-sync gate passes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bd1f00b1d0 | Merge branch 'next' into chore/1328-chore-remove-orphaned-root-vitest-config | ||
|
|
a13101ee5c |
fix(#1329): existence-filter scoped-CI fallback so a deleted test can't crash the lane
ci-prepare-test-scope.cjs's empty-detection FALLBACK hardcoded tests/core.test.cjs, deleted in #1291. Every scoped lane (scope=targeted| windows) that hit the fallback wrote the stale path into .ci-selected-tests.txt and crashed run-tests with "requested test file(s) not found: core.test.cjs". The full/sharded lanes glob the suite and were immune, so only the scoped lanes went red (e.g. run 27599149212 on #1308). Existence-filter the FALLBACK at write time and fall back to the 'unit' suite sentinel (the #408/#641 path, resolved live by run-tests) when nothing survives, so a stale reference degrades instead of crashing the lane. Detected lists still pass through verbatim (they may carry a suite sentinel and are already filtered by affected-tests-lib). Refactor to an exported, testable resolveSelection(). Add a generative parity guard (DEFECT.GENERATIVE-FIX) asserting every FALLBACK entry resolves on disk or is a known suite sentinel — it fails the instant a refactor deletes a listed file, which #1291 did and CI did not catch — plus resolveSelection unit tests and an end-to-end subprocess test. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c03f97188f |
chore(#1328): remove orphaned root vitest.config.ts left by SDK retirement
vitest.config.ts configured Vitest (not a dependency) to run .ts test files (the repo has none) rooted at ./sdk, a directory deleted when the SDK package seam was retired in #191 (ADR-0174). No npm script, workflow, or dependency references it. Also drop the now-dead sdk/src/*.test.* branch in diff-touches-shipped-paths.cjs isCiGating(), which can never match since the sdk/ tree no longer exists. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1f41a0ce9a |
feat(#1304): add optional activationKey capability manifest field (#1309)
Add an optional activationKey to the feature role of capability.json — the dotted config key that gates the whole capability (e.g. graphify.enabled). gen-capability-registry validates it (non-empty string, reserved-name guard, must be declared in the capability's own config slice, feature-only) and emits it per-capability in the generated registry. Declared on graphify + intel. No runtime consumption yet (resolver wiring lands in #1305). Part of #1302. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8c3d934a90 |
refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete) (#1295)
* refactor(#1294): T-final — delete the core.cjs re-export spine (epic #1267 complete) After T0–T6 nothing imports core, so retire the spine and its scaffolding: - delete src/core.cts (and the gitignored gsd-core/bin/lib/core.cjs artifact; remove its .gitignore + eslint-ignore entries) - delete scripts/lint-core-spine-imports.cjs + its allowlist; drop it from the package.json lint:ci chain - regenerate docs/INVENTORY-MANIFEST.json (drops the core.cjs surface) - sweep stale references: CONTEXT.md glossary back-compat clauses (spine retired, callers import the leaf directly), planning-config.md CONFIG_DEFAULTS owner, and false present-tense core.cjs claims in leaf-module docstrings The ADR-857 decomposition is complete: the former Core god-module is fully dissolved into its leaf modules; no re-export spine remains. No behaviour change. Closes #1294 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1294): migrate the computed-path core.cjs importers the literal grep missed bin/install.js used require(path.join(_gsdLibDir, 'core.cjs')) (a computed path, and bin/install.js was never in the convergence lint's scan roots), and ~8 test files referenced core.cjs via path.join/readFileSync/existsSync/FILE_ARG forms the literal-string migration grep missed. Route install.js's symbols to their leaves (RUNTIME_PROFILE_MAP->model-catalog, resolveTierEntry/EFFORT_SET-> model-resolver) and repoint/adjust the test references to the leaves. Recovers the 161 'Cannot find module core.cjs' failures from the spine deletion. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c76827afbc |
refactor(#1291): T6 — migrate test files off the core spine ahead of deletion (#1293)
The convergence lint only scanned src/ + gsd-core/bin, so ~35 test files still imported core.cjs. Repoint all 33 behaviour importers to the leaf modules directly (same symbol->leaf map as the src migration; leaves are the objects core re-exported by reference), delete the now-meaningless shim-identity describe blocks in the 8 leaf tests, and delete tests/core.test.cjs (forwarded-behaviour coverage now lives at the leaves; resolveWorktreeRoot test relocated to worktree-safety in T0) and tests/lint-core-spine-imports.test.cjs (the lint is removed in T-final). Dropped the stale core.test.cjs entries from the allow-test-rule-refs allowlist; eslint-rules RuleTester fixture path pointed at io.cjs. After T6: ZERO test imports core.cjs. core.cts still builds (now fully unused); T-final deletes it. No behaviour change. Closes #1291 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
76765bc24d |
refactor(#1289): T5 — migrate the final idiom-hard callers off the core spine (#1290)
The last 4 core importers, migrated off non-destructure idioms:
- gsd-tools.cjs: core.{error,ERROR_REASON,setJsonErrorMode,output} -> io.cjs;
core.findProjectRoot -> project-root.cjs (lazy wrapper preserved); inline
resolveWorktreeRoot require -> worktree-safety.cjs
- audit-command-router: DI default `_core ?? core` -> `_core ?? io` (seam preserved)
- intel-command-router: DI default -> `{ output: io.output, timeAgo: coreUtils.timeAgo }` (seam preserved)
- check-command-router: io destructure -> io.cjs; dynamic core['planningDir']
-> planning-workspace, core['findPhaseInternal'] -> phase-locator (typed
imports, dropped the Record-cast bracket hack)
Allowlist is now EMPTY — NO file imports the core spine. core.cts re-exports
are dead weight; T-final deletes core.cts + remaining shim tests + the lint.
No behaviour change.
Closes #1289
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
b108f101b0 |
fix(#1284): grant mcp__perplexity__* to researcher agents + dispatch-table parity guard (#1288)
Adds mcp__perplexity__* to both researcher profiles (generated source-of-truth) and regenerates the agents; adds a generative dispatch-table↔tools parity guard so future provider drift fails CI. Regenerates the agent-size baseline for the +20-byte frontmatter growth. Fixes #1284 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
645601a10d |
refactor(#1286): T4 — migrate 5 large destructure callers off the core spine (batch 3) (#1287)
Migrate the entire core surface of commands (~23 symbols), phase (~17), roadmap, state, template to the leaf modules directly (behaviour-identical — leaves are the objects core re-exports by reference). All 5 now import zero core symbols and are removed from the allowlist (9 -> 4). Dropped a dead `void replaceInCurrentMilestone` from phase.cts; stale core.cjs docstrings fixed. core.cts re-exports untouched (serve the remaining 4 idiom-hard files); teardown is T-final. No behaviour change. Closes #1286 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ec2ecdf28b |
refactor(#1283): T3 — migrate 9 multi-leaf callers off the core spine (batch 2) (#1285)
Migrate 9 files' entire core surface to the leaf modules directly (behaviour-identical — leaves are the objects core re-exports by reference): config, docs, gap-checker, graphify-command-router (namespace core.output -> io.output), init (17 core symbols -> 8 leaves), profile-output, uat, verification, workstream. All 9 now import zero core symbols and are removed from the allowlist (18 -> 9). Stale core.* docstrings corrected. core.cts re-exports untouched (still serve the remaining 9 files); teardown is T-final. No behaviour change. Closes #1283 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a5f213e73e |
refactor(#1281): T2 — migrate 12 single-leaf callers off the core spine (batch 1) (#1282)
Per the T1 design rubber-duck, batch by FILE so each tranche drops convergence-lint allowlist entries. Migrate 12 files' core imports to the leaf modules directly (behaviour-identical — leaves are the objects core re-exports by reference): - io (output/error/ERROR_REASON): agent-command-router, capability-state, capability-writer, frontmatter, gsd2-import, learnings, loop-resolver, task-command-router - roadmap-command-router -> config-loader; workstream-inventory -> core-utils - milestone, verify -> their full leaf sets (both were multi-leaf, not single-leaf as first scoped; migrated completely) All 12 files now import zero core symbols and are removed from the allowlist (30 -> 18). core.cts re-exports untouched (still serve the remaining 18 files); teardown is T-final. Stale core.cjs docstrings in the migrated files corrected to reference io.cjs. No behaviour change. Closes #1281 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
48d9cec6fe |
refactor(#1268): re-home core re-export-spine squatters + migration-convergence lint (#1272)
Re-home the 6 implementation functions squatting in the core.cjs re-export spine (ADR-857) into the modules whose interface they belong to, with core re-exporting them BY REFERENCE so all 32 callers + the shim-identity tests keep resolving unchanged: - worktree-safety: resolveWorktreeRoot, pruneOrphanedWorktrees - git-base-branch (broadened to the Git Query Module): gitWorktreeInfoInternal - agent-install-check (new leaf): getAgentsDir, checkAgentsInstalled - delete the _resetRuntimeWarningCacheForTests wrapper; consumers use a shared resetRuntimeWarningCaches() helper in tests/helpers.cjs Add scripts/lint-core-spine-imports.cjs (migration-convergence lint with a 30-importer allowlist, wired into lint:ci) so the staged spine retirement provably converges: CI fails on any new ./core import. Register the new generated agent-install-check.cjs in eslint-ignore + .gitignore + INVENTORY-MANIFEST.json. No behaviour change. First tranche (T0) of epic #1267. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
cf68841220 |
enh(#1243): consume Claude plugin-provided skills in agent_skills (epic #1258 Phase B) (#1261)
* feat(#1243): consume Claude plugin-provided skills via native Skill-tool directive + grant Skill to agent_skills-consumer agents - Relax global skill name validation to accept namespaced form `^[A-Za-z0-9_-]+(:[A-Za-z0-9_-]+)*$` - Namespaced names (containing colon) on claude runtime emit a Skill-tool load directive instead of a @-include line - Namespaced names on non-claude runtimes are skipped with a warning - Bare unresolved names retain existing warn-and-skip behavior (no promotion to directive) - Grant `Skill` tool to all 22 agent_skills consumer agents; 5 generated agents updated via research-profiles.cjs + regen, 17 hand-authored agents edited directly - Add 16 TDD tests in describe('bug #1243') covering happy/mixed/precedence/negative/cross-runtime/regression/grant cases Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * docs(#1243): document plugin-provided skills in agent_skills Update the Agent Skills Injection reference in CONFIGURATION.md with the three entry forms (project-relative, global:<name>, global:<plugin>:<skill>), the Claude-only runtime behaviour of the namespaced form and the warn-skip on other runtimes, the plugin pre-install prerequisite, and the consumer-agent Skill tool grant. Add docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md with a step-by-step guide for installing the plugin, locating the namespaced skill name, wiring it into agent_skills, and verifying injection. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1243): align agent_skills docs with emitted block format + mixed-block regression test (code-review) - Replace two-section mixed-block example (bogus "Load these plugin-provided skills using the Skill tool:" header) with the actual single-section inline format in CONFIGURATION.md and docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md - Fix quoted warning text in how-to doc to exactly match the emitted string: [agent-skills] WARNING: Plugin-namespaced skill "global:<name>" requires a Skill-tool-capable runtime (claude) — skipping on runtime "<runtime>" - Replace phantom agent slugs (gsd-checker, gsd-researcher, gsd-advisor, gsd-synthesizer) in CONFIGURATION.md Supported Agent Types with real agents/gsd-*.md examples (gsd-plan-checker, gsd-phase-researcher, gsd-code-reviewer, gsd-ui-auditor, gsd-research-synthesizer) - Add byte-identical mixed-block regression test: one path-resolvable global skill + one plugin-namespaced skill on claude runtime → asserts r.ir.block === single-section interleaved block, no secondary header Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(#1243): regenerate agent-size baseline for the Skill-tool grant The 22 agent_skills-consumer agents each grew +7 bytes from adding `Skill` to their tools list; refresh the committed per-agent size baseline (#1074 guard). * chore(#1243): add Added changeset fragment * fix(#1243): traceable allow-test-rule ref + separator-agnostic byte-identical tests (CI) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
b783410815 |
refactor(#1191): inject clock/reset testability seams + handle valid-null settings (#1233)
* refactor(#1191): inject clock/reset testability seams + handle valid-null settings - worktree-safety reapOrphanWorktrees: injectable deps.nowMs clock for deterministic stale-lock boundary tests (mirrors snapshotWorktreeInventory's options.nowMs). - active-workstream-store: _resetControllingTtyCacheForTests() seam clears the memoized controlling-TTY probe cache; test replaces require.cache busting. - gen-capability-registry: export stripGeneratedComment (additive); test imports the real helper + equivalence assertion, keeping the deliberate drift oracle. - install.js readSettings: a successfully-parsed JSON null is treated as empty settings ({}) instead of being mis-reported as malformed; genuine parse failures still warn. readSettings/stripJsonComments exported (GSD_TEST_MODE-guarded require) for real behavioral tests. Closes #1191 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1191): add changeset for valid-null settings fix (#1233) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1191): replace Stryker-incompatible structural reset test with behavioral isTTY-spy The seam-2 reset test read the BUILT active-workstream-store.cjs and grepped for 'didProbeControllingTtyToken = false' — Stryker instruments that file so the literal is absent, failing the mutation DRY RUN. Replaced with a behavioral test that spies on process.stdin.isTTY access count to prove a post-reset probe re-runs (kills the didProbe-reset mutant) without reading source text. Local stryker: dry run passes, score 85.21% >= 80. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1fa7bc594c |
refactor(#1190): extract ADR-230 PR-target branch policy into a tested, fork-safe seam (#1246)
ADR-230's branching-model gate (pr-target-validator.yml) decided allowed/blocked PR targets via inline regex in github-script — untestable. Extracted the decision into committed scripts/pr-target-policy.cjs (classifyPrTarget(base,head)->{decision}), and rewired the workflow to checkout the BASE ref (trusted; fork-tamper-safe) + require the module. Behavior-identical (Codex-verified char-by-char regex equivalence + all side-effects preserved). 70 tests incl. an equivalence oracle battery + hyphen-boundary negatives. Added contents:read for the checkout. Re-attribution: no ADR-230 test references exist (issue's '2 misattributed files' claim not borne out).
Closes #1190
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
00acbc8868 |
fix(#1223): install scripts/fix-slash-commands.cjs so gsd-tools loads (#1240)
* fix(#1223): install scripts/fix-slash-commands.cjs so gsd-tools loads Before this fix, bin/install.js copied scripts/changeset/ and scripts/lib/ into the runtime config dir but omitted scripts/fix-slash-commands.cjs. gsd-core/bin/lib/command-roster.cjs requires this file at module load via require('../../../scripts/fix-slash-commands.cjs'), so every gsd-tools command crashed with MODULE_NOT_FOUND on every installed runtime. Four changes: - bin/install.js copy step: copy fix-slash-commands.cjs into <configDir>/scripts/ with source-missing hard-fail and verifyFileInstalled smoke check - bin/install.js writeManifest: track scripts/fix-slash-commands.cjs (not covered by the changeset/lib subdir loops) - bin/install.js uninstall: best-effort unlinkSync before scripts/ rmdir - scripts/fix-slash-commands.cjs readCmdNames(): wrap readdirSync in try/catch returning [] so skill-based/global installs without a local commands/gsd/ directory do not throw ENOENT Tests added to tests/install.test.cjs (6 new tests): - smoke: install() copies fix-slash-commands.cjs - e2e: spawned gsd-tools.cjs does not crash with MODULE_NOT_FOUND - manifest: writeManifest() tracks the file - uninstall: uninstall() removes the file - readCmdNames unit: export returns an array - readCmdNames spawn: absent COMMANDS_DIR returns exit 0 (no throw) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1223): backfill changeset PR number (#1240) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
cafb874c4a |
fix(#1224): accept --pr 0 placeholder at changeset creation (#1231)
* fix(#1224): accept --pr 0 placeholder at changeset creation The required-field guard `!opts.pr` treated the integer 0 as falsy, rejecting the documented `pr: 0` two-push placeholder with a usage error (exit 2). Non-numeric `--pr abc` (NaN) was also silently accepted before (passes `!NaN === true`... actually `!NaN` is true, so NaN would trigger the guard already). The new explicit checks use `opts.pr === null` for missing flag and `Number.isNaN` for non-numeric input, accepting all finite integer values including 0. The merge-time safety net in parse.cjs (`pr <= 0` → INVALID_PR) is unchanged — a pr:0 fragment is still rejected at lint/render time. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1224): backfill changeset PR number (#1231) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
73b7f45140 |
feat(#1173): wire agent converters into descriptor-driven install path (#1227)
Extends `dispatchKindEntry` in `runtime-artifact-layout.cts` to route agents-kind entries through a converter when the descriptor carries a non-null `converter` field. Adds `stageAgentsForRuntimeWithConverter` to `install-profiles.cts`, expands `VALID_CONVERTER_NAMES` with the 9 agent converter names, and adds a fail-first behavioral test suite (9 tests) proving the new wiring end-to-end. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
22f56f4431 |
ci(#1212): shard windows full-test lane to remove timeout cliff (#1222)
The `full test (windows-latest, *)` lane ran the entire unit suite (~740+ files) in one job whose wall-clock crept against the 20m cap and intermittently CANCELLED (false-negative gate, observed on PR #1207). Prior tactical fixes #869 (15→20m bump) and #1051 (handle-leak) deferred the cliff structurally. Shard the unit suite across 3 parallel runners per OS/node leg so per-job wall-clock is O(total/3) and stays under the cap as the suite grows. - scripts/run-tests.cjs: add `--shard <i>/<n>` — a deterministic, balanced round-robin partition (fileIndex % n === i-1) over the SORTED selected file list. parseShardArg strictly validates i∈1..n, n≥1, integer-only; n=1 is a pure no-op. The 28K Windows argv chunking is preserved within each shard. A legitimately-empty shard (n > file count) exits 0; a selection empty BEFORE sharding still hits the discovery hard error. Composes with --suite and is order-independent (sorted before partition). Exports selectShard/parseShardArg. - .github/workflows/test.yml: test-full becomes the 3 legs × 3 shards = 9-job cross-product (explicit include rows — a base shard dim does not cross-product with include legs, and a nested matrix.leg.os is unresolvable by the H1 shell-policy linter). Unit suite runs sharded; integration/security run once per leg (shard 1). The Required tests fan-in is unchanged: it already needs test-full and checks the matrix-aggregate result, so a failed/cancelled shard fails the gate; the branch-protection check name is preserved. - tests: partition/CLI + pure selectShard contract (completeness, disjointness, balance, determinism, boundaries, fast-check property) + parseShardArg validation, in run-tests-harness.test.cjs; a DEFECT.GENERATIVE-FIX parity guard (per-row shard values 1..N, every leg runs all shards, N == --shard /N denominator) + Required-tests name/needs pin, in ci-test-scope.test.cjs. Closes #1212 Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7edd18fd2b |
feat(#1165): async external_job_waiting half-state + resume/pause contract (#1221)
Core half of #1105: a legal external_job_waiting deferred state so an async-dispatched Execute step (committing a .planning/async-jobs/<job>.json manifest, deferring SUMMARY.md) is not an illegal partial. execute-phase safe-resume, resume-project, and pause-work reconcile against the versioned scheduler-agnostic manifest stability contract without re-dispatching; the producer is the capability half (#1164). Closes #1165. |
||
|
|
1a186013a4 |
fix(#1205): roadmapper applies phase_id_convention to generated phase IDs (#1215)
* fix(#1205): roadmapper applies phase_id_convention to generated phase IDs - Add Phase ID Convention section to <phase_identification> block: documents sequential (default) vs milestone-prefixed forms, and instructs the agent to read phase_id_convention from config.json - Update <output_formats> to show both header and checklist forms for sequential and milestone-prefixed conventions with examples (e.g. ### Phase 1-01: Name, - [ ] **Phase 1-01: Name**) - Add TDD regression test tests/bug-1205-roadmapper-convention.test.cjs (5 assertions, confirmed fail-first then pass after fix) - Update tests/agent-size-baseline.json to reflect legitimate growth - Add .changeset/brave-otters-leap.md (Fixed, pr:0 placeholder) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: backfill changeset pr: 1215 for fix/1205 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1205): move phase_id_convention regression into roadmapper-granularity.test.cjs lint-regression-test-names rejects new standalone bug-NNNN-*.test.cjs files; regression cases must live in the owning module's test file. Move the 5 phase_id_convention assertions (#1205 regression) from the removed tests/bug-1205-roadmapper-convention.test.cjs into tests/roadmapper-granularity.test.cjs as a new describe block, alongside the existing granularity calibration tests. Also update the allow-test-rule comment to cover both #163 and #1205 surface contracts. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1205): fix lint-allow-test-rule-refs for roadmapper-granularity - Add issue ref (see #1205) to allow-test-rule comment in tests/roadmapper-granularity.test.cjs so lint-allow-test-rule-refs passes (new exemptions require #NNN per ADR-456) - Prune stale 'source-text-is-the-product' entry from scripts/lint-allow-test-rule-refs.allowlist.json (ratchet-down; comment now compliant and no longer needs grandfathering) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
9e5d4b266b |
fix(#997): ensure canonical ~/.claude/gsd-core path for plugin installs (#1207)
* fix(#997): ensure canonical ~/.claude/gsd-core path for plugin installs via SessionStart hook Claude Code marketplace plugin installs unpack the package into the version-pinned plugin cache and never run bin/install.js, so ~/.claude/gsd-core/ is never created. Agents, commands, and templates markdown-@-include the canonical ~/.claude/gsd-core/... path (which expands ~ but NOT ${CLAUDE_PLUGIN_ROOT}), so every include resolved to nothing and agents (e.g. the executor) failed. Add a SessionStart hook (hooks/gsd-ensure-canonical-path.js) that, on a plugin install, symlinks the canonical path's immutable subdirs (bin, contexts, references, templates, workflows) to the plugin's bundled gsd-core/ tree. It changes zero @-references, is a no-op in classic installs, preserves user-generated files (USER-PROFILE.md, STATE.md), prunes stale links so it self-heals after `claude plugin update`, uses Windows junctions, and rejects bundled/canonical paths that escape the resolved plugin root (no traversal, no clobber). Registered in HOOKS_TO_COPY (build-hooks), MANAGED_HOOKS, hooks.json SessionStart (runs first, timeout 5), and BUNDLED_GSD_HOOK_FILES. Behavioral regression tests folded into issue-766-plugin-manifest.test.cjs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#997): backfill changeset PR number to #1207 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
98866a0c69 |
feat(#1187): per-module Stryker mutation-score ratchet (ADR-456 80% floor) (#1200)
* feat(#1187): per-module mutation-score ratchet + graduate core-utils ADR-456's 80% mutation floor was unenforceable as a single global break=50: 4 of 6 covered modules sit at 63-79% and forcing them to 80 would require brittle exact-string assertions on equivalent string-literal mutants (a Goodhart's-Law trap). Instead, each covered module declares a minScore floor (locked at its measured score, TARGET 80) enforced per CI shard via stryker --break, ratcheting up over time without brittle tests. - mutation-matrix.cjs: minScore per module + TARGET_MUTATION_SCORE=80, emitted in the matrix; require.main guard + exports for testability. - mutation.yml: per-shard --break <minScore>. - stryker.config.mjs: global break 50->60 as a local backstop (CI uses minScore). - Graduated core-utils (measured 77.5%, floor 75). - context-utilization 79.5->92.3% via behavioral killers (state classification outputs + error-value contract, not exact-string matches) -> minScore 80 (TARGET). - ratchet-integrity guard test (28 cases). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1187): pass mutation break via MUTATION_BREAK env (no stryker --break flag) Adversarial review caught that Stryker 9.x has no --break CLI flag, so the per-shard 'stryker run --break <minScore>' errored out every mutation shard. Read the per-module floor from process.env.MUTATION_BREAK in stryker.config.mjs and set it per shard via env in mutation.yml. Red-green verified: MUTATION_BREAK=99 exits 1, =80 exits 0. Also make the ratchet guard monotonic (RATCHET_BASELINE floors; lowering a floor now fails the guard unless the baseline is edited). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1187): fail closed on bad MUTATION_BREAK + monotonic ratchet baseline Code review: Number(env)||60 failed OPEN — an empty/invalid MUTATION_BREAK (e.g. a future module missing minScore -> matrix expands to '') silently degraded the shard to break 60, letting a high-floor module regress undetected. resolveMutationBreak() now returns 60 only when the env is truly unset (local backstop) and THROWS on present-but-empty/non-numeric/out-of-range (fail closed); stryker.config.mjs imports it via createRequire. Also make RATCHET_BASELINE an equality mirror (=== not >=) so any floor change is explicit in review and no floor can be silently lowered. Tests: 46 (incl resolveMutationBreak cases). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1187): recalibrate config-schema/prompt-budget floors to CI scores First CI mutation run failed two shards: the floors were set from local Stryker runs whose TIMEOUTS were counted as kills (env-variable), inflating scores. CI runs with timeout~0, so the real deterministic scores are lower: - config-schema: local 69.7% -> CI 54.55% (5 local timeouts vanished) -> floor 52 - prompt-budget: local 99.6% -> CI 68.33% (239 local timeouts vanished) -> floor 66 Calibrate floors from CI (the documented source of truth) and record the lesson in the comment so future floors aren't set from timeout-inflated local runs. Baseline updated to match. The other 5 shards passed (deterministic CI scores above their floors). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e9f9ae49c8 |
fix(#1146): single base-branch resolver across forking workflows (#1198)
* fix(#1146): single base-branch resolver across forking workflows Replaces duplicated per-workflow bash detection that silently fell through to :-main on repos where origin/HEAD is unset (git init+remote add+fetch without set-head, most CI checkouts, many worktrees). New CJS module git-base-branch.cjs exposes `gsd_run query git.base-branch` with full precedence ladder: git.base_branch config override → origin/HEAD symref → git remote show origin (authoritative) → local branch presence → "main". All git subprocesses bounded with timeouts; degrades gracefully. Wires execute-phase, quick, ship, complete-milestone, and pr-branch to the single resolver. Removes 14 lines of duplicated detection bash across the five workflows. Includes 7 behavioral tests covering the full precedence ladder including the key regression case (master repo, origin/HEAD unset → must return "master", NOT "main") and an anti-regression guard that fails if any workflow re-introduces the :-main/:-master fallback pattern. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(changeset): backfill PR number #1198 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1146): drop stray PR-body file from branch pr-1146-body.md was committed during changeset backfill but must not be tracked in the repo. Content preserved externally for PR body use. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#1146): add tests for flat base_branch config key and both-branch tie-break Closes two mutation gaps identified in adversarial review: - A2: flat {base_branch: ...} at config root (legacy key form) was covered by code but unguarded against mutation of lines 74-75 in resolver - H: tier-4 tie-break when both main+master exist locally (main wins, per tryLocalBranch JSDoc) was documented but untested 9/9 tests pass. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#1146): allowlist workflow-literal guard as runtime-contract exemption Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#1146): degrade gracefully when gsd_run unavailable in handle_branching bash blocks handle_branching (execute-phase.md) and step 2.5 (quick.md) are extracted and run verbatim by behavioral tests that lack the gsd_run preamble. Adding a || fallback ladder (git symbolic-ref then echo main) keeps the unified resolver as primary in real workflows while letting the test harness succeed without gsd_run defined. Also propagates updated runtime-launcher preamble to pr-branch.md (added in origin/next MemPalace PR) and regenerates workflow-size-baseline.json. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b1e8a74708 |
fix(#1196): wire discuss loop step for capability hooks (#1199)
* fix(#1196): wire discuss loop step for capability hooks discuss was contract-declared (gsd:loop-host marker, in POINT_ORDER and LOOP_HOST_CONTRACT) but structurally unwireable: discuss-phase.md had no `loop render-hooks` dispatch and was absent from the conformance gate's HOST_LOOP_FILES, so capabilities could never wire discuss:pre/discuss:post. - discuss-phase.md: add minimal discuss:pre (before analyze_phase) and discuss:post (after write_context) render-hooks dispatch steps that delegate consumption to a new shared reference (kept under the 32KB #2551 budget; no inline subagent dispatch token). - references/loop-hook-dispatch.md: new canonical, point-agnostic contract for consuming `loop render-hooks --raw` activeHooks (contribution/step/ gate) — single source for hook consumption across host loops. - gen-loop-host-contract.cjs: derive HOST_LOOP_FILES from STEP_WORKFLOWS and export scanWiredPoints()/getWiredLoopPoints() (throws on a missing host file) — one source of truth for the host-loop file + wired-point set. - phase6-capstone-conformance.test.cjs: consume the derived HOST_LOOP_FILES and shared scanWiredPoints (was a hand-maintained duplicate omitting discuss-phase.md + a duplicated regex). - gen-capability-registry.cjs: add validateHooksWired() gen-time guard that rejects a capability hook declared at a valid-but-unwired loop point, with a clear remediation message — failure now surfaces at gen --check/--write time instead of deep in the full conformance suite. - tests (capability-registry.test.cjs): regression + anti-pattern parity guards (every loop-host marker is in STEP_WORKFLOWS/HOST_LOOP_FILES; POINT_ORDER === flattened LOOP_HOST_CONTRACT) so no step can drift into the discuss-class gap again. - docs/INVENTORY*: register the new reference. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1196): backfill changeset PR number (#1199) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5fa4dcd78c |
fix: recover silently-excluded test dirs + test-architecture audit hardening (#1195)
* fix: recurse test discovery so subdir test suites actually run
scripts/run-tests.cjs discovered tests with a flat readdirSync(testDir),
silently excluding tests/observability/ (4 files), tests/dispatch/ (1) and
tests/installer-migrations/ (1) — 94 passing tests — from `npm test` and all
CI lanes. Walk the tree recursively (relative subpaths preserved), classify
suites by basename, and add a fail-on-zero-executed guard for suite/default
runs (escape hatch GSD_ALLOW_EMPTY_SUITE=1) while preserving the empty
--files/--files-from path the CI inert lane relies on.
Unit suite 735 -> 741 files; surfaces ADR-227's observability/dispatch seam.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: retire 5 verified-worthless tests
Adversarial verification confirmed these 5 prove nothing — their coverage is
provided more strictly elsewhere:
- enh-2790 'has a name: field' spot-checks (command-contract enforces /^gsd[:-]/)
- command-routing-hub duplicate construct + duplicate ERROR_KINDS assertions
- no-cjs-sdk-handsync-tooling (guarded files that never existed on main; bug-190
covers the real retired SDK artifacts)
- runtime-artifact-layout cline edge case (subsumed by the explicit-global test
and bug-782-cline-skills-emission)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: add ADR-218 release version-validation coverage
ADR-218 (reject leading-zero versions like 1.01.0; npm duplicate pre-check) had
zero tests — the logic lived only in release.yml bash. Add a test that extracts
the actual rejection regexes from the workflow and exercises them against a
boundary table (leading-zero/malformed rejected, valid accepted) plus structural
wiring assertions. Goes red if the regex is reverted to [0-9]+.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: redesign weak tests into behavioral, deterministic assertions
Per the ADR test audit, rewrite 27 weak test files (test-only, no source
changes) so each can go red for the defect it guards:
- kill pass-always assert.ok(true) placeholders (research-cli, worktree-baseref,
bug-260 security guard, eslint-rules x24, clusters '|| true')
- replace source-text grep with behavioral calls (install Kilo, sh-hook-paths,
plan-review-convergence) and add a repo-layout governance test
- de-flake real-clock/Math.random coupling (phase last_updated, bug-3707 mtime,
context-utilization property, feat-3594)
- fix independence/shared-state violations (bug-492 singleton, issue-844 tmpRoot,
core reapStaleTempFiles, active-workstream TTY, feat-488 GSD_HOME)
- strengthen property/shape-only tests (research-provider/store classification +
collision) and unconditional plugin.json schema validation (issue-766)
Verified: all 28 files run together 1220 pass / 0 fail / 1 skip.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: add no-tautological-assert lint rule, error in test suite
New custom ESLint rule (eslint-rules/no-tautological-assert.cjs) bans asserts
that can never fail: assert(true)/assert.ok(<always-truthy literal>),
'cond || true' inside an assert, and equality asserts comparing two identical
literals. Wired as error on tests/**; full sweep confirmed zero existing
violations so the suite stays green. Prevents the placeholder-assert regressions
the audit redesigns just removed. RuleTester coverage added (6 valid, 8 invalid).
Note: no-only-tests was already enforced via eslint-plugin-no-only-tests, so no
duplicate rule was added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gate new allow-test-rule exemptions to require an issue ref
ADR-456 requires any allow-test-rule exemption added after the ADR to carry a
tracking issue number, but nothing enforced it. New ratchet gate
(scripts/lint-allow-test-rule-refs.cjs, wired into lint:ci) fails when a NEW
allow-test-rule comment lacks a #NNN/URL reference; the 323 existing untracked
exemptions are grandfathered in an allowlist that ratchets down as they gain
refs. Red-green verified (novel untracked offender fails; compliant passes).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: add ADR test-audit evidence report (#1192)
Full risk-first qa-test-architect audit of the ADR portfolio (37 ADRs + 4
platform lenses, adversarial verification of retire verdicts) that drove the
P0 discovery fix, ADR-218 coverage, 5 retires, 27 redesigns, and the two new
lint gates. Filed as point-in-time evidence under docs/issueevidence/, named
for tracking issue #1192.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: replace pre-existing raw NUL byte with escape in feat-3594 fixture
feat-3594's null-byte parser fixture contained a literal NUL byte (pre-existing
on next at
|
||
|
|
ae8bb707bc |
refactor(#1170): remove hand-maintained INVENTORY count scalars (#1179)
* refactor(#1170): remove hand-maintained INVENTORY count scalars The `(N shipped)` heading counts in docs/INVENTORY.md were absolute scalars that collided silently on merge: two branches each bumping the same integer to N+1 produced a clean git merge whose value the merged filesystem (N+2) contradicted, hard-failing inventory-counts.test.cjs on the CI merge commit across all platforms (DEFECT.INVENTORY-MERGE-UNDERCOUNT). - Strip the six `(N shipped)` heading counts + the two prose footnote counts; repoint the intro to INVENTORY-MANIFEST.json as the registry. - Drop the decorative `generated` date from the manifest + its strip-before-compare branch in gen-inventory-manifest.cjs (it conflicted on cross-day merges and is read by nothing). - Delete inventory-counts.test.cjs (scalar-vs-disk gate, the collision source); its drift protection is subsumed by the merge-safe set-membership test inventory-manifest-sync.test.cjs, which stays as the sole gate. - Add inventory-headings-countfree.test.cjs guard (fails if a count is re-added to a heading). - Fix already-broken count-bearing cross-doc anchors to stable count-free slugs in ARCHITECTURE.md + multi-agent-orchestration.md. - Retire the now-impossible DEFECT.INVENTORY-MERGE-UNDERCOUNT + obsolete RULESET.DOC-CONSISTENCY in CONTEXT.md; de-count DEFECT.INVENTORY-DRIFT; correct stale MANIFEST-CANONICAL-KEY (all six families canonical). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1170): backfill changeset PR number (#1179) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
b10e56818b |
feat(#1169): complete ADR-857 phase 6 — migrate features to Capabilities, revive dead gates, harden conformance gate (#1183)
* test(#1168): make phase-6 gate un-gameable — reject empty stubs + require loop shrink The migration assertion previously checked only role==feature, so a registration-only stub (empty hooks, logic left inline) would turn the gate green while phase 6 stayed incomplete — the exact false-completion pattern this gate exists to prevent. Strengthen it: each ADR-named feature must OWN its behavior (>=1 hook, or a command family); and plan-phase.md/execute-phase.md must shrink strictly below their frozen pre-phase-6 sizes (94519/93166 LF bytes), which also defeats double-run gaming (declare a hook but keep the inline block -> file does not shrink -> red). Gate now 5 pass / 4 fail (orphaned execute:wave:post, empty/unregistered features, config-key leaks, no shrink). Green is now reachable only by REAL migration. Refs #1168, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate gap-analysis to a Capability (plan:post gate) First real ADR-857 phase-6 migration (pattern-defining tracer). gap-analysis moves from an inline post_planning_gaps branch in plan-phase.md to a real plan:post gate Capability: - capabilities/gap-analysis/capability.json: role:feature, plan:post gate (when=workflow.post_planning_gaps, blocking:false advisory), OWNS workflow.post_planning_gaps (federated out of central schema). - plan-phase.md: inline config-get + gsd_run gap-analysis block replaced with a plan:post render-hooks call site dispatching the gate; file shrinks 94519->93279. - src/check-command-router.cts: cmdGapAnalysisPlanPost runs the real gap analysis via gap-checker. - post_planning_gaps removed from central manifest; resolves via federated config (default true preserved). - tests/post-planning-gaps-2493: re-pointed to assert capability ownership. Verified: gate 5 pass / 4 fail (gap-analysis cleared from migration, plan:post-orphan, config-leak, and plan-phase shrink checks); loadConfig still returns post_planning_gaps=true; check command runs real analysis; 392/392 in the config/registry/federation/router net. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate profile-pipeline to a command-family Capability ADR-857 Decision 7: profile-pipeline becomes a command-family Capability (like audit/intel/graphify). capabilities/profile-pipeline/capability.json declares an 8-command family (scan-sessions, extract-messages, profile-sample, write-profile, profile-questionnaire, generate-dev-preferences, generate-claude-profile, generate-claude-md) backed by a new gsd-core/bin/lib/profile-pipeline-command-router.cjs; the inline case arms are removed from gsd-tools.cjs. Owns profile-pipeline.enabled (federated). Verified: registry shows role:feature with commands.length=8; scan-sessions/profile-sample run live via the family; gate cleared profile-pipeline from the empty-stub failure (only tdd/schema-gate/drift remain); 296/296 registry+inventory+gsd-tools tests; lint 0 errors. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1167): wire execute:wave:post + implement ui.safety-gate check Revives the second dead gate from #1167: ui.gates@execute:wave:post was declared but never dispatched AND its check.query (ui.safety-gate) was unimplemented. Adds the per-wave execute:wave:post render-hooks call site in execute-phase.md (fires after each wave's merge/cleanup, before the next forks) and implements cmdUiSafetyGate (frontend + UI-SPEC aware, mirrors cmdUiPlanGate) in check-command-router. +17 regression tests. Verified: phase-6 orphaned-points conformance test now PASSES (gate 6 pass / 3 fail); ui-safety-gate routable in dot+hyphen forms; check-ui-safety-gate 17/17, check-ui-plan-gate 18/18; lint 0 errors. Refs #1167, #1168. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate drift (schema + codebase) to execute:wave:post gates Removes the inline schema_drift_gate + codebase_drift_gate steps (77 lines) from execute-phase.md; drift becomes a Capability with two execute:wave:post gates (verify.schema-drift blocking, verify.codebase-drift advisory) dispatched via the per-wave render-hooks call site. check-command-router routes verify.schema-drift / verify.codebase-drift to the real detectors. Federates workflow.drift_threshold / drift_action / schema_drift_gate out of central. Also fixes the execute:wave:post dispatch prose to run NON-blocking (advisory) gates too — the prior version only ran blocking gates, which would have silently dropped the codebase-drift advisory after its inline step was removed. Behavior preserved. Verified: gate 7 pass / 2 fail (drift cleared from stub + config-leak; execute-phase.md 92297 < 93166 frozen -> shrink passes); both drift checks run real detection; loadConfig defaults preserved (threshold=3, action=warn, gate=true); drift-detection 56/56 + schema-drift 34/34; lint 0 errors. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate tdd to a Capability (plan:pre contribution + execute:post gate) tdd becomes a real Capability: a plan:pre contribution injects the <tdd_mode_active> planner guidance (rendered from PLAN_PRE_HOOKS_JSON like security's contribution), and an execute:post gate (tdd.review-checkpoint, advisory) runs the real end-of-phase RED/GREEN review via a new check-command handler. Inline tdd_mode reads + the inline planner block + the tdd_review_checkpoint step are removed; workflow.tdd_mode is federated out of central. The MVP+TDD per-task RED-commit gate is preserved — TDD_MODE is now derived from the execute:post hooks (capId==tdd active), not an inline config-get. BEHAVIOR CHANGE (documented, not silent): the --tdd CLI flag now persists workflow.tdd_mode=true via config-set instead of being per-invocation. Rationale: tdd is now a config-toggled Capability, and env vars do not persist across the workflow's separate bash blocks (config does), so an ephemeral override isn't cleanly achievable; --tdd therefore enables the tdd capability, consistent with how all capabilities are toggled. Verified: gate 7 pass / 2 fail (tdd cleared from stub + config-leak; plan-phase + execute-phase both < frozen sizes); contribution injection + execute:post gate dispatch wired; MVP+TDD gate preserved; tdd.review-checkpoint runs real review; full unit suite 556/0; lint 0 errors. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): migrate schema-gate to a plan:pre contribution Capability The plan-time schema-push detection (former plan-phase.md §5.7) becomes a schema-gate Capability: a plan:pre contribution (into:planner, when:workflow.schema_push_detection) whose fragment carries the full ORM-detection + [BLOCKING] schema-push-task injection logic, rendered into the planner via the existing plan:pre render-hooks dispatch. The inline §5.7 block is removed (plan-phase.md 94519->90445). workflow.schema_push_detection is a new capability-owned (federated) key, default true. (The execute-side schema-drift gate was migrated separately into the drift capability.) Verified: registry inlines the fragment (len 2704) so it is actually delivered at plan:pre; gate 8 pass / 1 fail — all 5 ADR-named features now real Capabilities, only the config-leak test remains (intel/security, next unit). Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1169): close the 3 capability config-key leaks — phase-6 gate now GREEN Removes the last inline config-get reads of capability-owned keys from plan-phase.md. security_asvs_level/security_block_on now flow through the security plan:pre contribution via a new loop-resolver configValues mechanism (resolves declared config keys with the same 4-level precedence as activation and attaches them to the rendered hook); the §5.55 banner reads them from PLAN_PRE_HOOKS_JSON. intel.enabled becomes a real intel plan:pre step (ref.command: intel api-surface) dispatched via render-hooks; the inline intel branch is gone. gen-capability-registry now validates ref.command as a third dispatch shape. Verified: phase-6 capstone conformance gate is FULLY GREEN (9/0); 3 leaks gone (grep=0); security configValues resolve to {2,medium}/default {1,high}; intel step present only when enabled; loop-render-hooks 62/0, capability-registry 287/0, capability-state/federated-config 113/0; lint 0 errors. Closes the migration half of #1169. Refs #1139, #1167, #1168. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): address adversarial review — restore schema-drift block, generic planner injection, uniform gate contract Adversarial review caught 2 real regressions the green gate missed: (1) schema-drift no longer blocked — the execute:wave:post dispatch read GATE_RESULT.block but verify.schema-drift emitted drift_detected/blocking, and onError:skip wrongly bypassed positive blocks; (2) only tdd's plan:pre contribution was injected into the planner, dropping schema-gate's schema-push detection and security's threat-model guidance. Fixes: (A) every gate check returns a uniform boolean 'block' under --raw (the dispatch form), with advisory gates (tdd/gap) carrying their report in 'message'; (B) gate-dispatch contract corrected at all sites — onError governs command errors only, a blocking gate's positive block always halts; (C) generic planner injection of all plan:pre contributions where into=='planner' (tdd + schema-gate + security incl configValues); (D) two new conformance assertions: planner contributions injected generically + every gate check.query returns boolean block under --raw. Verified: gate 11/11; all 6 gate checks return boolean block under --raw; full suite 595/0; lint 0 errors. Refs #1167, #1168, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): restore MVP+TDD end-of-phase blocking escalation (2nd adversarial pass) The migrated tdd execute:post gate is statically blocking:false, but the contract (references/execute-mvp-tdd.md + CONTEXT.md) requires the end-of-phase TDD review to ESCALATE from advisory to blocking when MVP_MODE && TDD_MODE && a TDD plan misses a RED/GREEN commit. The migration prose had downgraded this to a 'strong advisory recommendation' — silent loss of the blocking escalation. Restore it: the tdd-gate dispatch now refuses to mark the phase complete (Phase blocked message) under MVP+TDD when GATE_RESULT.block is true; advisory otherwise. Also strengthen tests/execute-mvp-tdd-gate.test.cjs: hasBlockingEscalation previously matched any line with 'blocking'+'mvp+tdd' (so 'advisory (blocking: false) ... under MVP+TDD' was a false green); now it requires the real refusal semantics ('refuse to mark the phase complete' / 'phase blocked'). Caught by 2nd adversarial review pass. Verified: execute-phase.md 92702 < 93166 frozen; mvp-tdd-gate + phase-6 gate 19/0; full suite green; lint 0 errors. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): restore MVP+TDD proceed-block, codebase auto-remap, schema skip-flag (3rd adversarial pass) 3rd adversarial pass found 4 more silent regressions: (1) the tdd MVP+TDD 'refuse to mark complete' was nullified by a downstream 'ALWAYS proceed regardless of gate results' line — proceed is now conditional (stops on an active MVP+TDD block); (2) the test now asserts the proceed is NOT an unconditional override; (3) codebase-drift auto-remap (spawn gsd-codebase-mapper when drift_action=auto-remap) was dropped — the execute:wave:post advisory dispatch now consumes spawn_mapper/directive; (4) GSD_SKIP_SCHEMA_CHECK bypass was lost from the gate path — cmdVerifySchemaDrift now honors the env var (block:false when set). Verified: no unconditional proceed; GSD_SKIP_SCHEMA_CHECK=true -> block:false; gate 11/11 + mvp-tdd 9/9; full suite 569/0; lint 0; execute-phase.md 93109 < 93166. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): init.cts reads federated config keys from nested path (4th adversarial pass) Config federation moved tdd_mode/research/nyquist_validation from flat config.<key> to nested config.workflow.<key>, but src/init.cts still read them flat — so init.plan-phase/init.execute-phase emitted tdd_mode:false / research_enabled:undefined / nyquist:undefined regardless of config (a public command-contract regression; the migrated loops use render-hooks so enforcement was unaffected). Read via config.workflow (type-safe Record cast). Now init reflects the same resolved values + federated defaults (research/nyquist default true) as the render-hooks path. Verified: build clean; init.plan-phase emits tdd_mode:true/research:false/nyquist:false for set config, defaults true for empty; full suite 591/0; lint 0. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1169): add changeset for ADR-857 phase-6 completion (PR #1183) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): complete phase-6 migration fallout — restore TEXT_MODE, fix registry .claude leak, re-point stale workflow-contract tests The capability migration left real regressions and stale consumer tests that the per-module unit suite missed but the full cross-platform suite caught (27 failing tests): Real source regressions (fixed): - execute-phase.md lost its AskUserQuestion TEXT_MODE plain-text fallback when the inline schema_drift_gate step was removed — non-Claude runtimes would stall. Restored, and the execute:post gate-dispatch prose de-duplicated to cite the execute:wave:post contract (loop body shrinks below the frozen pre-phase-6 ceiling while keeping every onError/blocking nuance). - capabilities/tdd inline fragment hardcoded `@~/.claude/gsd-core/references/tdd.md`, baked verbatim into the committed capability-registry.cjs and leaked the install path on 11 non-Claude runtimes (registry .cjs is copied, not path-converted). Made the fragment path-free; regenerated the registry. The phase-6 conformance gate now guards this (no ~/.claude install path in any capability source or the generated registry). - plan-phase.md: removed a §5.7 stub re-added in error and routed Branch 2 to step 6 (schema-gate is a plan:pre capability, §5.7 is gone). Stale workflow-contract tests re-pointed to the capability dispatch they now must assert (behavior verified preserved in source first, assertions kept equal-or-stronger): bug-621 + bug-2851 (gap-analysis via gsd_run render-hooks plan:post + registry binding), feat-2527 (tdd_mode federated out of central), phase6-planning + plan-phase-ui-redirect (§5.6 bounded by ## 6.), plan-phase-drift-guard (intel when:intel.enabled skip branch). profile-pipeline-command-router.cjs un-ignored from eslint (hand-written, no TS source) + stale disable comments removed. Size baseline regenerated. Verified: full suite 15140 tests / 0 fail; lint 0 errors; conformance gate green legitimately. Refs #1139, #1167, #1168, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1169): add ADR-857 E2E content-test coverage for the 12 loop points + capability deliverables Grounds the capability engine in behavioral E2E tests (drive the real render-hooks/check CLI + the real registry, assert typed result content — no source-grep), structured around what ADR-857 says to deliver. 207 tests; each genuineness-checked (flip the expectation, confirm it fails). Per-loop-point dispatch (7 files): empty-point negative-space across the 6 no-hook points; verify:post 3-step resolution+ordering+onError; plan:pre contribution/configValues + ui.plan-gate + intel; plan:post gap-analysis; execute:wave:post drift+ui gates via the check route (schema-drift block/skip, codebase-drift threshold BVA, auto-remap); execute:post tdd.review-checkpoint RED/GREEN; ship:pre security gate resolution + frontmatter-get predicate pieces. ADR-deliverable coverage (4 files): predicate boundary held (edge/prohibition probes stay core, not off-by-default Feature Capabilities — phase-6 exception); core loop runs with zero capabilities (all 12 points empty, init bundles resolve); contribution merge (multiple ordered <contribution from=> blocks); federated-config key removal on uninstall. federated-config allowlisted for its 3-file split (unit + integration + lifecycle). Refs #1139, #1167, #1168, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): remove dead drifted converter dups + address adversarial review Lint cleanup (root-caused, not waved off): src/runtime-artifact-conversion.cts carried 11 agent-converter functions (+5 orphaned consts/helpers) that were never exported, never called, and had silently DRIFTED from the live hand-authored copies in bin/install.js (one even referenced an undefined `claudeToCopilotTools`). Deleted the dead duplicates; install.js's live copies are untouched (it never imported these). Lint now 0 errors / 0 warnings. Adversarial-review (Codex) findings fixed: - HIGH: execute-phase.md TDD_MODE used `jq ... || echo false`, silently disabling the MVP+TDD blocking gate on jq-less runtimes. Reverted to the `node -e` form (node is guaranteed; matches the file's other node-e usages) so a missing optional tool can no longer fail-open a blocking safety path. - MEDIUM: federated-config-key-removal orphan-key test was vacuous (it skipped the orphan assertion). Now asserts the removed capability's key is genuinely not surfaced/validated after uninstall. - LOW: phase-6 conformance leak regex broadened to catch absolute-home and Windows-backslash `.claude/(gsd-core|commands|agents|hooks)` paths, not only `~`/`$HOME` forward-slash forms. - LOW: bug-2851 plan:post dispatch assertion now requires `--raw` (matched its stated contract). - nit: plan-pre intel-step test duplicate assertion replaced with a distinct structured-output check. Size baseline regenerated (execute-phase.md 93089 < 93166 frozen). Refs #1167, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1169): make runtime-homes-descriptor-drive titles environment-independent The descriptor-equivalence test embedded the absolute golden config path (`os.homedir()`-derived) directly in each `test(...)` title, so titles differed between macOS (`/Users/x/.claude`) and Docker (`/home/gsdtest/.claude`). Every test PASSES on both platforms (15885/0 leaf tests each), but gsd-test-summary compares results by title and reported 29+29 false "only in Mac / only in Docker" discrepancies for tests that actually pass everywhere. Move the golden path out of the title and into the assertion message (still shown on failure); titles are now byte-identical across platforms so the cross-platform comparator matches them. No assertion logic or golden values changed. Refs #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1169): derive TDD_MODE via gsd_run --active-cap, not node -e (fix prompt-injection CI gate) The prior fix reverted execute-phase.md:181 from jq to `node -e` to close a Codex HIGH (jq||echo-false silently disabling the MVP+TDD blocking gate on jq-less runtimes) — but the CI prompt-injection scanner BLOCKS new `node -e` in workflow markdown (inline code-exec = injection vector), turning the security gate red. Both forms were wrong: node -e fails the scanner; jq fail-opens a blocking safety gate; `config-get workflow.tdd_mode` is forbidden by the conformance leak gate (tdd_mode is capability-owned). Correct fix (what Codex recommended): a gsd_run-native boolean. Add an `--active-cap <capId>` flag to `loop render-hooks <point>` that resolves hooks the normal way and prints exactly `true`/`false` for whether a capId is active — scanner-safe (canonical launcher, no inline code), node-reliable (no optional jq to fail-open), and leak-free (render-hooks resolution, not config-get). execute-phase.md:181 now `TDD_MODE=$(gsd_run loop render-hooks execute:post --active-cap tdd)`. +5 behavioral tests for the flag. Verified: prompt-injection-scan --diff origin/next → 0 findings; conformance gate 13/13 (execute-phase.md 92934 < 93166); execute-mvp-tdd + tdd-mode + loop-render-hooks 87/0; lint 0/0. Refs #1167, #1169. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
83e87e3ec0 |
feat(#1180): hard-gate the GSD Version requirement on bug reports (#1181)
Auto-close bug reports opened without a valid GSD Version (Issue Forms enforce required only in the web UI). Bug reports only; version-shaped validation; version-exempt opt-out. Closes #1180 |
||
|
|
aec3374bc2 | feat(#1138): make runtime descriptors authoritative (#1157) | ||
|
|
4ab5c7b3f2 |
feat(#1135): migrate planning hooks to capabilities (#1141)
* feat(#1135): migrate planning hooks to capabilities * chore(#1135): add phase 6 planning capabilities changeset * fix(#1135): satisfy lint for agent hook rendering |
||
|
|
fd01e7a12e |
feat(#1132): complete contribution hook prerequisite
Closes #1132 |
||
|
|
eb051ea696 |
feat(#1123,#1124): enforce duplicate-producer invariant + fail-loud loadCentralConfigKeys in gen-capability-registry (#1131)
Closes #1123 Closes #1124 Refs #857 |
||
|
|
7f1d49935c |
ci(#1104): keep next package.json in sync with the last published release (#1109)
* ci(#1104): sync next package.json version to the last published release next rested on a -dev stream per ADR-660 (1.3.1-dev.0) — a never-published placeholder that leaked to source/dev installs. Make every release type write its exact published version back to next: - finalize/hotfix (push main): auto-backmerge sets next's version to main's released version, folded into the existing back-merge PR (+ pinned setup-node). - rc (no main push): the rc job opens + admin-merges a sync PR after publish. Shared, fail-closed scripts/sync-next-version.cjs stamps package.json + the runtime manifests via the npm version hook and refuses any non-release version. Amends ADR-660 (supersedes the -dev stream decision). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * ci(#1104): harden next-version sync against post-publish failure modes Review hardening (Codex + code-review gates) on the #1104 sync helper and its workflow callers: - release.yml rc Sync step: continue-on-error so a post-publish sync hiccup cannot fail an already-published release (npm immutability would block re-run). - auto-backmerge.yml inline sync: set -euo pipefail + validate VERSION before any shell use (closes a ${VERSION}-in-commit-message injection vector); git add -u instead of -A. - sync-next-version.cjs: reuse an existing open PR instead of failing gh pr create on rc re-runs; regex-parse the PR number and fail loud; discriminate the git diff --cached --quiet exit code (only status 1 == has-diff, else rethrow); git add -u to avoid sweeping runner artifacts into next; tolerate already-merged on admin merge. - tests: +2 (existing-PR reuse, non-diff rethrow); 14/14 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3e836fef0d |
feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/
Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.
Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.
Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
source-checkout-gated build fallback) instead of LLM re-derivation; the engine
capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.
Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.
* test(#550): RED — status×verification re-cut + probe-core engine specs
Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):
status: resolved | dismissed | unresolved (lifecycle, shared)
verification: explicit | backstop | null (only when resolved)
- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
to be extracted — validateResolution(r, validators), validateRequirement,
analyzeCoverage(items, resolutions?, validators), byVerification rollup,
runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
{resolved,backstop}; coverage gains byVerification.{explicit,backstop};
proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
COUNT preserved on every fixture (closed set = resolved+dismissed; doc
line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
rewritten to the two-axis model.
Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).
* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)
Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.
probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
items[] (core never assumes propose is deterministic — edge resolves via
LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
{categories, verification, requiredFieldsByVerification} (ADR-550 #5)
edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.
* chore(#550): register probe-core.cjs artifact in ledgers + inventory
New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:
- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).
probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.
* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]
trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.
Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.
* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard
Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:
- validateResolution now enforces the 'verification is null unless resolved'
invariant for EVERY status (not just resolved): a dismissed/unresolved
resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
(was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
contract; a new test locks that an all-dismissed run is NOT affirmatively covered
(byVerification is the honest gate).
Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.
* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics
Re-review #5 (trek-e) clarity edits:
- Decision 5: annotate that only contract item (a) ships on #584 (the edge
adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
parenthetical — the blessed/implemented semantics are count-preserved = the
CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
carrying the per-tier resolved-status breakdown. The old parenthetical
contradicted the shipped count.
* test(#550): cover runProbeCli structural-guard numeric-count branch
Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.
* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)
The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.
* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)
A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.
* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)
templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.
* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)
The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.
* docs(#550): add how-to for resolving edge-coverage findings (B1)
Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.
* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)
trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.
* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)
trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.
* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)
trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.
* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe
Rebased onto next (
|
||
|
|
e4f0910d62 |
test(#1074): agent-size-budget per-file baseline + line→byte rebase (PR 3/3) (#1097)
Extends the #1074 scheme to tests/agent-size-budget.test.cjs, which used the same assertTightCeiling tier ratchet but was still line-based (never rebased in #717). Completes the migration — the last part of the #1074 epic. - Rebase agent sizing from lines to LF-normalized bytes (#717/#683). - Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests); add a per-agent baseline (tests/agent-size-baseline.json) as the primary anti-creep, and byte hard caps (XL 56 KiB / LARGE 48 KiB / DEFAULT 24 KiB), each above its tier high-water with real headroom. No separate new-file cap: a net-new agent is DEFAULT-tier, already bounded by the DEFAULT cap. - Keep the agent-classification tests verbatim. - scripts/workflow-size.cjs: add generic measureMdFiles(dir, predicate) (workflows + agents share one byte-measurement path); measureWorkflows now delegates to it. - scripts/update-size-baseline.cjs: one 'npm run size:baseline' now regenerates BOTH the workflow and agent baselines (gsd-* filter for agents). Rebased onto next after PR 2/3 (#1096) merged: replicate the scripts/lib/workflow-size.cjs -> scripts/workflow-size.cjs move (PR 1/3) across the generator and the agent test's require; regenerate the agent baseline against current agents (a uniform +170 B preamble drift on all 33 since authoring). Addresses the #1097 review (trek-e): - BLOCKER (acceptance criterion 5): document the agent contract in CONTEXT.md. Adds RULESET.AGENT_SIZE_BUDGET (caps 57344/49152/24576, per-agent baseline, dual size:baseline, shared measureMdFiles seam) and disambiguates it from the separate DEFECT.AGENT-FILE-SIZE-CAP-BREACH 45K-CHAR guard (two units, two purposes). - Docs: now that #1096's docs/TESTING-SUITES.md "Workflow size budget" section is in next, fold in the agent coverage here (renamed to "Workflow & agent size budget"): agent caps + per-agent baseline + the how-to + reference rows, and the disambiguation from the 45K-char guard. - Minor (negative proof): add a boundary-fixture test exercising the hard-cap comparison at cap-1/cap/cap+1 through the real lfByteCount path, so a future threshold/operator edit can't silently neuter a cap. - Nit: align the tier test name wording ("stays within") with the <= operator. Negative proof on a real tracked agent (gsd-planner): baseline catches +10 B; XL hard cap catches 57,516 > 57,344 with the baseline current. Closes #1095 (PR 3/3 child); landing this completes the #1074 epic. |
||
|
|
74d7bc8239 |
test(#1074): add additive per-file workflow size baseline guard (PR 1/3) (#1089)
* test(#1074): add additive per-file workflow size baseline guard (PR 1/3) Introduces a committed per-file size baseline scheme alongside (not replacing) the existing tier anti-creep tests. Green by construction — the baseline records current sizes, so both schemes pass side by side during migration. - scripts/lib/allowlist-ratchet.cjs: add assertFileBaseline (third pure helper, same injected-fail style) — per-file growth/shrink/add/remove diff vs baseline. - scripts/workflow-size.cjs: single source of truth for LF-normalized byte counting (#683) + workflow enumeration, shared by the guard and the generator so they can never measure differently. Lives in scripts/ root (NOT scripts/lib/) because it is dev/CI-only tooling — scripts/lib/ is bundled into the installed runtime, scripts/ root is not, so this keeps it out of the shipped payload. - scripts/update-size-baseline.cjs + npm run size:baseline: regenerate the snapshot (sorted keys, trailing newline, idempotent). - tests/workflow-size-baseline.json: generated snapshot (88 workflows). - tests/workflow-size-budget.test.cjs: import the shared counter (drops the duplicated local byteCount) and add the per-file baseline describe block. - Tests for the helper, the shared module, and the generator (incl. round-trip and fault-injection cases). Refs #1074. Part 1 of 3; PR 2 swaps enforcement, PR 3 covers the agent test. * test(#1074): regenerate workflow baseline after Update-branch merge with next The 'Update branch' merge (652a916b) pulled in next's update.md change (#1090) without regenerating the snapshot, leaving the per-file baseline stale by one file. Re-ran `npm run size:baseline` so the committed baseline matches the merged workflow files. Refs #1074. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |