2085a4ebeba8c720a952ea51597ca508eed53ed8
91 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
195d356d7c | Merge branch 'next' into feat/1346-enhance-verify-phase-project-a-check-vio | ||
|
|
21fe9d3627 |
chore(#1544): test-quality + doc cleanups from the #1507 conversion-module review (#1547)
Follow-up hygiene from the #1507 epic review (release blocker fixed in #1537): - path-replacement.test.cjs: delete the hand-reimplemented computePathPrefix copy (ADR-1508 said to); route all cases (incl. Windows + outside-home) through the real _computePathPrefix so the test can no longer drift from the function. - enh-1511: add an isWindowsHost no-op tripwire characterization test. - enh-1511: add a deterministic (fs-method monkeypatch, root/OS-independent) error-path test for applyRuntimeContentRewritesForCommandsInPlace — asserts the temp dir is rm'd on a read failure with no orphaned gsd-cmd-rewrites-* leak. - runtime-artifact-conversion.cts: @internal note on rewriteStagedCommandBodies (deep-seam companion to rewriteStagedSkillBodies; no production caller today). - ADR-1508 + CONTEXT.md: qualify the "single owner" claim with the deliberate opencode/kilo applyOpencodeFamilyPathPrefix pre-conversion carve-out (#784). No user-facing behavior change (tests + internal docs + one code comment). Refs #1507. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c09c13f295 |
enhance(verify-phase): node-test causation control — prove the RED is content-caused (#1346)
The #1279 node-test machine-proof confirmed a known-bad subject drives the negative test RED, but could not distinguish a genuine content-violation from a deceptive test that reds merely because GSD_PROHIB_SUBJECT is set. Add an optional fifth flat scalar `check_clean_fixture` (-> CheckDescriptor.cleanFixture) threading a KNOWN-CLEAN control subject through projectProhibitions + descriptorFromProjection. When present, the prover also runs the check against the clean subject and requires GREEN, so fail-first is proven only when the check is RED on the violation AND GREEN on the clean subject (content-dependent). Opt-in and additive: absent a clean fixture the prover behaves exactly as post-#1314 (no control, documented residual), preserving the zero-authoring compose path; the lint-rule kind needs no analog (its subject IS the linted file, no env indirection). Coverage: RED-first deceptive case, positive, missing-clean fail-closed, round-trip read-back/emit, fast-check property extended to the 5th scalar, and an end-to-end COMPOSE capstone (honest vs deceptive). Docs: ADR-550 dated addendum, prohibition-probe reference, spec-phase + verify-phase workflows. Closes #1346 Claude-Session: https://claude.ai/code/session_01GsPRb8zvpcT7Eat6vZw8PX |
||
|
|
b65892e6a2 |
docs(#1508): add ADR for Runtime Artifact Conversion Module content-rewrite ownership
Phase 0 of epic #1507. Records the decision to make the Runtime Artifact Conversion Module the single owner of per-runtime content rewriting, flip the dependency direction to installer/layout -> conversion, and close the surface.cts -> bin/install.js getInstallExports relay. Doc-only. Resolves ADR-3660 Initial-Scope deferral; distinct from epic #1258. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD |
||
|
|
e7855bc217 | fix(#1459): user-owned consent store gates third-party capability activation; env/cwd in disclosure; loader validator parity (#1473) | ||
|
|
c866ac1b24 | fix(#1462): fail closed without data loss on a corrupt capability ledger; atomic ledger write (#1469) | ||
|
|
dcceb1a004 |
fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing (#1442)
* fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing resolveConfigHomeFromDescriptor's dot-home-nested probe returned the first bare-existing candidate, so a CLI user (~/.gemini/antigravity-cli) who also had the IDE's ~/.gemini/antigravity dir was silently shadowed to the legacy dir (probed first). Regression from #217 — pre-#217 returned ~/.gemini/antigravity unconditionally. Add an optional probeMarker (gsd-core/VERSION) to the dot-home-nested descriptor: a two-pass probe prefers the candidate GSD installed into, then bare existence, then probe[0]. Behavior is byte-identical when probeMarker is absent (windsurf etc. unaffected). Adds detectAntigravityDirAmbiguity() for installer/operator guidance on already-misinstalled users (auto-relocation ruled out per ADR-0008's single-configDir migration bound). Regression tests fail before / pass after: coexistence + marker-priority cases, end-to-end through the registry descriptor. Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB * chore(#1441): add changeset for antigravity resolver fix Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB |
||
|
|
b0c774c2e3 |
feat(#1416): formalize Resolution convention + agent-skills value envelope (Resolution Provenance P3) (#1425)
Narrows P3 of ADR-1411 (Resolution Provenance, epic #1411) based on an adversarial fit-analysis that showed a single Resolution<T> envelope adopted by agent-skills, capability-state, and capability-writer fails the deletion test: configured/reason are meaningless for capability verbs, and capability-writer's errors[] (operation-not-applied) cannot fold into warnings[]. The only genuinely shared seam is warnings: string[]. Changes: - src/resolution.cts: new pure types+builder leaf — exports Resolution<T> {value, configured, reason, warnings}, makeResolution<T>() builder, and AgentSkillsValue {block, skills_count}. No other src/ imports. - src/init.cts: cmdAgentSkills --json IR gains additive value:{block, skills_count} field (built via makeResolution). All existing flat fields (agent_type, block, skills_count, warnings, configured, reason, source, degraded) are retained unchanged for back-compat. - src/capability-state.cts: doc comment on ResolveCapabilityRuntimeStateResult naming it the canonical read-verb envelope. No JSON change. - src/capability-writer.cts: doc comment on SetCapabilityStateResult naming it the canonical mutation-verb result (warnings=advisory, errors=operation- not-applied). No JSON change. - CONTEXT.md: new ### Resolution Convention glossary entry after ### Resolution Provenance. - docs/adr/1411-resolution-provenance.md: P3 narrowing amendment appended. - tests/resolution.test.cjs: 9 unit tests for makeResolution (new). - tests/agent-skills.test.cjs: 2 P3 tests for value.block/value.skills_count and back-compat of all flat fields. All 277 tests pass (5 suites). npm run lint clean. All lint checks pass. Part of #1411 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6f2547610e |
docs(#1412): ADR-1411 — Resolution must report provenance, not fall open silently (#1413)
Decision record for the Resolution Provenance epic (#1411, P0). Establishes that context resolution (config loading, project-root anchoring, workstream resolution) must report its provenance rather than fall open silently to defaults — the resolution-side analog of ADR-227. Binds the Config Loader Module, Project-Root Resolution Module, and I/O Module. Adds the ADR, the index/seam-map entry, and the CONTEXT.md glossary term. Part of #1411 Closes #1412 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7ca8011cd9 |
refactor(#1373): add canonical markdown-sectionizer seam (epic #1372 T0) (#1381)
* refactor(#1373): add markdown-sectionizer seam (ADR-1372 T0) Establishes the canonical markdown-structure parsing seam per ADR-1372. No existing parsers are modified; this is the foundational T0 tier only. - docs/adr/1372-markdown-sectionizer-seam.md: Accepted ADR defining the seam interface, the tiered migration plan (T0-T7), and the prohibition enforcement approach (no-adhoc-markdown-parsing ESLint rule in T7). - src/markdown-sectionizer.cts: Pure module, Node built-ins only. Exports: stripFencedCode (CommonMark-correct state machine ported from uat-predicate.cts _stripFencedBlocks, CRLF-safe, unterminatedFence signal), tokenizeHeadings (ATX headings outside fenced blocks), collectSections (line-by-line predicate-driven section collection), collectSection (single named section, levelBounded stop, optional stripFences), iterateBullets (dash/checkbox/numbered + continuation). - tests/markdown-sectionizer.test.cjs: 54-test behavioral suite covering the parser QA matrix (LF/CRLF, Unicode headings, headings-inside-fences, unterminated fences, nested levels, all bullet markers, continuation lines, empty/non-string input) plus 4 fast-check property tests (idempotence, output shape, never-throws, length monotonicity). - CONTEXT.md: Markdown Sectionizer glossary entry added (PR review gate). Tests: 54 pass, 0 fail. Existing adr-parser + uat-passed tests: 22 pass. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#1373): add extractTaggedBlocks + replaceSection to seam; register inventory - src/markdown-sectionizer.cts: extend Section type with bodyStart/bodyEnd offsets; add extractTaggedBlocks(content, tagName) (inner text of <tag>…</tag> blocks, tagName regex-escaped, caller decides fence-stripping) and replaceSection(content, section, newBody) (pure character-offset splice for read-modify-write callers); update collectSections/collectSection to populate bodyStart/bodyEnd. - tests/markdown-sectionizer.test.cjs: add 33 new behavioral tests for extractTaggedBlocks, replaceSection, and a DEFECT.GENERATIVE-FIX parity guard that asserts stripFencedCode and uat-predicate's _stripFencedBlocks agree on a shared 9-item corpus; documents the known 4-space-indent divergence. - docs/adr/1372-markdown-sectionizer-seam.md: list extractTaggedBlocks and replaceSection in §"The seam". - CONTEXT.md: update ### Markdown Sectionizer glossary entry with the two new exports. - docs/INVENTORY.md: add markdown-sectionizer.cjs row (alphabetically between loop-resolver and milestone). - docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest --write. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1373): clear no-unsafe-assignment + unused-var lint in markdown-sectionizer Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1373): correct section offset/round-trip + CommonMark heading/fence edges; register eslint coverage FIX 1 (CRITICAL): Enforce content.slice(bodyStart,bodyEnd) === body invariant in both collectSection and collectSections. bodyEnd is now bodyStart + body.length instead of the raw stop-line offset, eliminating the trailing-newline overcounting that caused replaceSection to drop separator newlines (## A\nbody## B gluing bug). FIX 2 (MED): tokenizeHeadings now accepts ≤3-space indent (CommonMark §4.5) and empty ATX headings (## / ## ), text=''. 4-space indent correctly excluded. FIX 3 (MED): collectSection gains stopAtLevel option — stops at the next heading whose level ≤ stopAtLevel, independent of the opener's level. Enables state.cts ## sections that also stop at ### without abusing levelBounded. FIX 4 (MED): Backtick fence opener info string must not contain a backtick (CommonMark). Applied in both stripFencedCode and tokenizeHeadings fence state machines. Tilde fences unaffected. FIX 5 (LOW): "byte offset" → "character (string-index) offset" in HeadingToken / Section doc comments. FIX 6 (LOW): extractTaggedBlocks doc comment documents nested-tag non-support; test locks the non-greedy close-at-first-</tag> behavior. FIX 7: Add gsd-core/bin/lib/markdown-sectionizer.cjs to eslint.config.mjs ignores so tests/551-eslint-bin-lib-coverage.test.cjs passes (3/3). Tests: 107 pass / 0 fail (was 87; +20 new tests for FIX 1–4, 6). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1373): gitignore tsc-built markdown-sectionizer.cjs (ADR-457 build-at-publish) The seam's compiled artifact must be a build-at-publish output like every other src/*.cts->bin/lib/*.cjs module (decisions, core, state, ...), not a committed file. Add it to the ADR-457 ignore list and untrack it; build:lib/CI regenerate it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
56d4a1bf39 |
enhance(#1279): project check_violation_fixture scalar — #1278 locate + #1279 proof compose end-to-end (#1346)
Delivers option (a) from the #1314 maintainer review: thread a fourth flat scalar check_violation_fixture through the projection so a prohibition authored at spec-phase machine-proves fail-first and greens through the deterministic path alone — zero hand-authoring at verify time. - src/probe-core.cts: Prohibition gains check_violation_fixture?; projectProhibitions emits it (both kinds) ONLY for a well-formed descriptor and ONLY when non-empty (blank/absent -> projects absent -> producer hard-gates, never a partial green). - src/prohibition-enforcement.cts: descriptorFromProjection reads it back into violationFixture via the same numeric-coercion-safe scalar() normalizer. - Tests (RED-first, proven non-vacuous by reverting both src edits): CHK-02(#1346) projection emit, CHK-08(#1346) read-back, CHK-03(D) example round-trip, the fast-check round-trip property extended to the 4th scalar (the contract trek-e blocked #1301 on), and a real-subprocess COMPOSE capstone greening end-to-end through project -> descriptorFromProjection -> default prover+runner. - Docs flipped from 'hard-gates until #1346' to 'composes end-to-end': verify-phase.md, prohibition-probe.md, spec-phase.md authoring, ADR-550 addendum, changeset. #1346 now tracks only the node-test causation residual. 190 affected-suite tests green; eslint + tsc clean; size baseline regenerated. |
||
|
|
33b6ee6f0c |
fix(#1279): node-test prover fail-closes on missing violationFixture + honest projection docs
Addresses the #1314 maintainer review (trek-e): - Major 1 (fail-OPEN): defaultProveFailFirst's node-test branch only guarded `if (!fixture)`. A missing/typo'd/stale violationFixture made GSD_PROHIB_SUBJECT point at a missing file; an honest negative test threw ENOENT *inside its callback* (a failing test named distinctly from the file), which isNonVacuousNodeTestRed accepted as proof -> a green forged from a setup crash. Now requires fs.existsSync(path.resolve(cwd, fixture)) before spawning, symmetric with the lint-rule path's file-result guard. Regression test pins it (RED without the guard); a second test pins cwd-relative fixture resolution. - Major 2 (misleading prose): the #1278 projection carries no violationFixture, so the deterministic-locate path always hard-gates (fail-closed) until a check_violation_fixture scalar is threaded through. verify-phase.md and prohibition-probe.md no longer read as if the projected path produces greens; the ADR-550 addendum records both items. Tracked as follow-up #1346. - Documented residual: existence is necessary but not sufficient (a red caused by the env being set vs the subject's content); recorded as a constraint, in #1346. - Nit: stale 'NOT attested fail-first' comment -> 'NOT machine-proven fail-first'. 64 tests pass; eslint + tsc clean; changeset valid. |
||
|
|
49bef1927e |
Merge remote-tracking branch 'origin/next' into feat/1279-fail-first-prover
# Conflicts: # docs/adr/550-spec-phase-probe-contract.md # gsd-core/workflows/verify-phase.md # tests/workflow-size-baseline.json |
||
|
|
7676420d8e | docs(#1279): changeset pr→1314 + node-test non-vacuous-red wording in changeset/ADR | ||
|
|
e50ead7ad2 |
enhance(verify-phase): deterministic auto-locate of the prohibition check descriptor (#1278) (#1301)
* test(1278): RED-first descriptor parity + fail-closed guards + CHK-07 byte-stability (wave 1) - CHK-03 (RED): extend PROB-14 parity in prohibition-probe.schema.test.cjs to carry the flat check_kind/check_target/check_rule scalars through project->write->parseMustHavesBlock; the non-droppable check_kind-presence assertion is the load-bearing RED trigger (fails because projectProhibitions strips check_* on the current build). - CHK-07 (GREEN forward-guard): probe-core.test.cjs pins descriptor-less byte-stability + dispositionForProhibition fail-closed policy, with a t.todo marker forward-locking plan 01-02. - CHK-06 (RED): prohibition-enforcement.test.cjs asserts descriptorFromProjection export + fail-closed on absent/partial/unknown descriptors via the projection adapter (RED until 01-03). - No src/*.cts or .cjs edits; no new test files; lint-test-file-count clean. * feat(1278): add optional flat-scalar check descriptor fields to Prohibition interface (wave 2) - check_kind?/check_target?/check_rule? mirror CheckDescriptor.kind/target/rule (minus caller-attested failFirst, #1279) - optional so existing Prohibition consumers compile unchanged * feat(1278): project check descriptor as flat scalars in projectProhibitions (wave 2) - emit check_kind/check_target (+ check_rule only for lint-rule with a rule) when descriptor well-formed - under-specified/descriptor-less items project byte-identically (CHK-07); flat scalars ride existing parseMustHavesBlock continuation-KV path (no parser rewrite) - add CHK-02 probe-core unit cases pinning the projection - turns CHK-03 parity test GREEN; dispositionForProhibition untouched * feat(1278): descriptorFromProjection read-back adapter feeds fail-closed locate (wave 3) - Add descriptorFromProjection(projected) -> CheckDescriptor | null to src/prohibition-enforcement.cts: renames the projected flat scalars check_kind/check_target/check_rule -> {kind,target,rule?}, or null when the descriptor is absent/non-object (no check_kind key). - failFirst is NEVER sourced from the projection (stays caller-attested; #1279). - rule is set only when check_rule is a non-empty string; the adapter does NOT re-validate kind/target/rule — an under-specified descriptor reconstructs to one the EXISTING runProhibitionEnforcement LOCATE guard rejects (located:false, never green). The merged #1259 guard stays the single source of fail-closed truth. - Turns the RED CHK-06 fail-closed tests (plan 01-01) GREEN end-to-end; CHK-03 / CHK-07 stay green. CheckDescriptor type, locate guard, dispositionForProhibition, and parseMustHavesBlock are unchanged (additive +36/-0). * feat(1278): verify-phase locates prohibition check from projected descriptor (wave 3) - request.check kind/target/rule sourced from projected check_kind/check_target/check_rule via descriptorFromProjection, not verifier invention (CHK-05) - replaces the #1278 author-supplied / tracked-follow-up note with the delivered deterministic-locate behavior - preserves fail-closed routing: absent/partial descriptor -> never green, hard-gate in both modes - failFirst stays a verify-time caller attestation; #1279 bounds the remaining fail-first proof * feat(1278): spec-phase captures wired-check descriptor on test-tier resolution (wave 3) - Step 5.6 'Keep it' / verification: test path captures check_kind/check_target/check_rule, projected onto must_haves.prohibitions for verify-phase deterministic locate (CHK-04) - SOFT capture: a test-tier prohibition without a descriptor is still allowed (no hard authoring block); stays fail-closed/flagged downstream - --auto captures only an unambiguous descriptor, never fabricates a check path - failFirst NOT captured at spec-phase (verify-time attestation; #1279) - PROB-06 soft-gate + text-mode (PROB-09) behavior unchanged * chore(1278): re-baseline workflow size for grown verify-phase + spec-phase prose (wave 3) - spec-phase.md 28438 -> 30343 (+1905), verify-phase.md 35362 -> 36498 (+1136) - regenerated via npm run size:baseline (no hand-picked numbers); growth is the #1278 deterministic-locate + descriptor-capture prose - workflow-size-budget guard green (122/122) * docs(1278): ratify optional check descriptor in dated ADR-550 addendum + type:Changed changeset - Append dated 2026-06-15 ADR-550 addendum ratifying the D3 prohibition-item shape extension (optional flat-scalar check_kind/check_target/check_rule) - Document flat-scalar rationale, deterministic projection/read-back, fail-closed on partial/invalid/absent, #1279/policy out-of-scope - Add .changeset/1278-prohibition-check-descriptor.md (type: Changed) * docs(1278): document optional check descriptor in prohibition-probe reference + FEATURES - Add 'Optional wired-check descriptor (deterministic locate, #1278)' section to the prohibition-probe reference (flat-scalar keys, projection/read-back, fail-closed + backward-compat, failFirst stays attested) - Add deterministic prohibition-check descriptor source entry to FEATURES.md - No CONTEXT.md glossary change: descriptor reuses existing wired-check / verification:test vocabulary, no new glossary term introduced * fix(1278): pass packaging gates — changeset pr field + retired slash-form fix - Add required pr: 1278 to changeset (lint:changeset MISSING_PR hard requirement; plan's 'omit if unknown' was inaccurate — issue number per #1259 convention, updated to real PR number when opened) [Rule 3 - blocking] - Fix retired /gsd-spec-phase -> /gsd:spec-phase at verify-phase.md:83 (wave-3 prose; caught by slash-namespace invariant #3443/bug-2543, blocked CHK-09 full-suite-green) [Rule 1 - bug] - size:baseline + INVENTORY manifest verified in-sync post-build (no diff) * docs(1278): add check descriptor + descriptorFromProjection to CONTEXT.md prohibition glossary * fix(1278): harden descriptorFromProjection round-trip (numeric-coercion + stray-rule) per review - MD-01/LW-01: narrow projected scalars to primitives + String()-coerce, so a numeric-looking check_target (parseMustHavesBlock coerces ^\d+$ to number) reconstructs as a string and locates instead of silently un-locating; no as-string type-lie, satisfies no-base-to-string. - LW-02: attach rule only for the lint-rule kind (drop a stray node-test rule). - LW-03: document the optional check_* keys in the reference Output schema. RED->GREEN tests added in prohibition-enforcement.test.cjs. * chore(1278): set changeset pr to 1301 * test(1278): add fast-check property for the check-descriptor round-trip + fail-closed (trek-e review) RULESET.TESTS.property-based-testing: the projectProhibitions -> render -> parseMustHavesBlock -> descriptorFromProjection chain is a bijective/transformation contract. Adds 2 fc properties to tests/probe-core.property.test.cjs (no new file; ratchet stays at 2 for probe-core): - well-formed descriptors survive the round-trip across the full string domain incl. the numeric-coercion case (target/rule reconstruct as strings); - under-specified/invalid descriptors (absent / target-less / rule-less / unknown-kind) are always fail-closed (never green, flagged, unlocated). Stability is asserted at the descriptorFromProjection layer (the raw parse step is intentionally lossy for numeric scalars; the shared parser is unchanged). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
8cf4716c99 |
docs(#1279): ratify ADR-550 addendum + FEATURES/reference/verify-phase + changeset
- ADR-550 dated 2026-06-15 addendum: machine-proven fail-first REALIZED (D5d CLOSED); ratify violationFixture, GSD_PROHIB_SUBJECT, failFirst demotion - FEATURES REQ-PROHIB-07 reads shipped (machine-proven, not caller-attested) - prohibition-probe reference + verify-phase descriptor shape updated with violationFixture + convention - tag GSD_PROHIB_SUBJECT + violationFixture PROPOSED/renamable (zero live consumers) as a PR-review flag - add .changeset/1279-machine-proven-fail-first.md (type: Changed) |
||
|
|
6a9e6cb0ae |
fix(1259-01): address trek-e review — B1 (fatal/suppressed), B2 (bounded subprocess), M1/M2, minors
Maintainer CHANGES_REQUESTED (reviewed |
||
|
|
31b822b312 |
fix(1259-01): close adversarial review findings — genuine enforcement, honest fail-first scope
Adversarial pre-submission review found the injected-runCheck tests masked a non-functional real runner. Fixes: - BL-01 (false green on vacuous test): the node-test runner now parses the TAP summary and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail) AND a reported test named distinctly from the file — node --test counts an empty file as one passing test, so counts alone could not catch it. - SF-01 (lint anchor never greened): the lint-rule runner now runs the project eslint as --format json and filters by ruleId, so plugin rules (local/*) load via the flat config — bare --rule cannot load a plugin. local/no-source-grep now genuinely greens (covered by a real, non-injected test). - BL-02 (tautological fail-first): the runner no longer echoes the caller's failFirst as if confirmed. failFirst is documented as caller-ATTESTED; the producer requires attestation + a genuine non-vacuous pass. Machine-proven fail-first (needs a violation fixture) is flagged as a tracked follow-up in ADR-550, the changeset, FEATURES, the reference doc, and verify-phase. - SF-02: added real-runner end-to-end tests (no injected runCheck) + pure, exported parse/filter helpers (parseNodeTestSummary, tapTestNames, eslintJsonHasRule, eslintFileResultCount) so the shipping branches are mutation-pinned. - NIT-01/02: LOCATE guard rejects empty-string rule and unknown kinds. - Hardening: spawn checks with NODE_TEST_CONTEXT/NODE_OPTIONS scrubbed so an ambient test-runner context cannot corrupt a verify-time result. - Docs reconciled to the shipped behavior (no 'confirms fail-first' overclaim). |
||
|
|
ce01e1376b |
docs(1259-01): wire verify-phase consumer + ADR-550/FEATURES/reference + changeset
- verify-phase.md: replace test-tier 'fail-closed/deferred' bullet with the check prohibition-enforcement enforcement step (locate -> fail-first -> run -> evidence -> green-or-hard-gate); update determine_status tree - ADR-550 addendum: mark D5d enforcement half LANDED (#1259); cite ADR-857 open-question §147 + D6 (core verify rail) - FEATURES §146: enforcement wording + add REQ-PROHIB-07; keep REQ-PROHIB-06 intact - references/prohibition-probe.md: test-tier enforced + hard-gates via check prohibition-enforcement - changeset (type: Changed) with the D5 '2 no-source-grep invalid cases, not 96' correction |
||
|
|
395fb519e7 |
feat(spec-phase): prohibition probe — surface "must-NOT" constraints (#644) (#1149)
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes #644. |
||
|
|
9a4a448c27 |
docs(#1235): ADR — migrate agent conversion to the descriptor-driven install path (#1253)
Records the design gate for moving agent conversion off the inline bin/install.js loop onto the descriptor path (ADR-3660): the two-(really three-)path problem, the 10 parity behaviors the descriptor agents path must gain (verified against code — the issue's 7 plus Qwen/Hermes branding, the Codex TOML sidecar, and stale-agent cleanup), an AgentConverterContext contract to fix the leaky (content)=>string converter signature, and an incremental per-runtime cutover gated on byte-for-byte golden parity (full + minimal mode). Status: Proposed. Codex-reviewed for technical accuracy. Closes #1235 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f52a7a5f77 |
feat(#1245): add capability ecosystem ADR, PRD, and developer documentation (#1248)
Phase 0 of the Capability Ecosystem epic (#1244): the design record and the third-party-author documentation set, with no runtime or code changes. - docs/adr/1244-capability-ecosystem.md — architecture decision record (amends/extends ADR-857 Decisions 7 & 8) - docs/prd/1244-capability-ecosystem.md — product requirements - Diataxis docs: tutorial, how-to (publish/import/version/remove), reference (manifest schema, /gsd:capability command, capability matrix), explanation (trust model); cross-links added to develop-a-capability.md Refs #1244 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
fbf5d5eab6 |
docs(#1239): ADR — GSD as an embeddable orchestration engine (#1241)
* docs(#1239): ADR for GSD as an embeddable orchestration engine Records the design to invert GSD from a standalone installer that projects onto a host into an embeddable orchestration engine a host loads as a plugin, driven through a negotiated host-integration interface. Unifies ADR-1016 projection (the declarative adapter) with imperative embedding behind one contract; grounded in a 9-host capability survey. Closes #1239 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#1239): de-slash illustrative command placeholders (docs-parity gate) Replace illustrative /gsd:x /gsd-x /gsd.x placeholders with namespace-prefix wording so the docs-parity live-registry check does not parse them as non-live commands. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
bf634b95c3 |
feat(#1213): Capability State Writer — write-side inverse of the resolver (#1225)
* feat(#1213): Capability State Writer — write-side inverse of the resolver Adds src/capability-writer.cts (setCapabilityState + cmdCapabilitySet) and the `gsd-tools capability set` subcommand: the write-side inverse of the capability resolver (ADR-1213). One desired capability state projects onto the substrates — `enabled` drives the runtime surface (canonical on/off), `gates` drive federated config keys (hook granularity), install profile is a read-only floor — then re-resolves and reports divergence (assert-and-report), so "off means off" holds as a write-time invariant. Adds batched setConfigValues; routes gsd:settings capability hook-gates through the writer. Docs: CLI-TOOLS reference, how-to, ADR-1213, CONTEXT.md term. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1213): add changeset for Capability State Writer (#1225) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bd349aa88e |
docs(#1143): add ADR-1143 Claude orchestration capability (Workflow/ultracode) [proposal] (#1155)
Proposes a runtime-gated, default-off BETA capability (post ADR-857) that adopts Claude Code's Workflow tool — the engine behind /effort ultracode — as an optional parallel-execution backend for the GSD loop, and folds the existing ultraplan plan-offload under the same capability. Restores wave parallelism + plan-checker + verifier on Claude Code that #853 currently forces inline. Design-only ADR draft accompanying feature request #1143; implementation is blocked on #857 being released. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0c0a8966ba |
fix(#1151): drive codex sandbox_mode emission from runtime descriptor sandboxTier axis (#1152)
* fix(#1151): drive codex sandbox_mode emission from runtime descriptor sandboxTier axis The sandboxTier runtime-capability axis was cosmetic: declared and validated on all 16 descriptors but read by nothing. The codex per-agent sandbox_mode line was emitted unconditionally from the hardcoded CODEX_AGENT_SANDBOX map, so the descriptor field drove no behaviour (ADR-857 audit finding F10; the "rides along in 5e/5g" promise in ADR-1016 §8 never landed). Make the axis load-bearing: - resolveInstallPlan projects sandboxTier as a 7th InstallPlan axis and fails loud (throws) on a missing/invalid value rather than coercing to 'none'. - installCodexConfig / generateCodexAgentToml gate sandbox_mode emission on sandboxTier !== 'none'. - The per-agent CODEX_AGENT_SANDBOX map is kept: it is GSD agent policy, not a runtime-descriptor property (different layer). Full removal of that map is tracked under #1138 (phase-6 descriptor-residue removal). For codex (sandboxTier === 'codex-agent-sandbox') the emitted TOML is byte-identical to before; for 'none' runtimes sandbox_mode is omitted. Adds leaf, projection, and installCodexConfig threading-seam regression tests; updates the enh-1082 InstallPlan golden master with sandboxTier for all 16 runtimes. Confirmed hypothesis: schema-first vocabulary closure outran consumer wiring, with no conformance gate to catch the orphaned axis. Closes #1151 Refs #857, #1138 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1151): stamp changeset with PR number 1152 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
eb051ea696 |
feat(#1123,#1124): enforce duplicate-producer invariant + fail-loud loadCentralConfigKeys in gen-capability-registry (#1131)
Closes #1123 Closes #1124 Refs #857 |
||
|
|
34bb8ed5f6 |
docs(857): settle phase-6 predicate boundary — verification substrate vs. plug-in tier
Classify edge-probe/prohibition-probe predicate-generation as core verification substrate (not an off-by-default Feature Capability), settled before ADR-857 phase 6 (Migrate) freezes the core/plug-in line. Prompted by @davesienkowski's boundary analysis on #857. - ADR-857: amendment note + top-level decision carve-out + Loop Extension Points exemption + new "Verification substrate vs. plug-in tier" subsection + 2 Alternatives rows + phase-6 rollout exception + Open Questions resolution. - ADR-550: cross-reference pinning the core-substrate classification; ties "exogenous grading" to Decision 4 (judgment-tier) and Decision 5 (test the contract, not the classifier). - CONTEXT.md: glossary — Probe Core / Edge Probe reclassified (probe-core on the contract side, adapters as the generator) + new "Verification substrate (predicate boundary)" term. Decomposition: the verifier<->predicate contract is core/non-toggleable; the generator (probe adapters) is core-default but independently versionable. Docs-only; no user-facing change. Closes #1120 Refs #857 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7f1d49935c |
ci(#1104): keep next package.json in sync with the last published release (#1109)
* ci(#1104): sync next package.json version to the last published release next rested on a -dev stream per ADR-660 (1.3.1-dev.0) — a never-published placeholder that leaked to source/dev installs. Make every release type write its exact published version back to next: - finalize/hotfix (push main): auto-backmerge sets next's version to main's released version, folded into the existing back-merge PR (+ pinned setup-node). - rc (no main push): the rc job opens + admin-merges a sync PR after publish. Shared, fail-closed scripts/sync-next-version.cjs stamps package.json + the runtime manifests via the npm version hook and refuses any non-release version. Amends ADR-660 (supersedes the -dev stream decision). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * ci(#1104): harden next-version sync against post-publish failure modes Review hardening (Codex + code-review gates) on the #1104 sync helper and its workflow callers: - release.yml rc Sync step: continue-on-error so a post-publish sync hiccup cannot fail an already-published release (npm immutability would block re-run). - auto-backmerge.yml inline sync: set -euo pipefail + validate VERSION before any shell use (closes a ${VERSION}-in-commit-message injection vector); git add -u instead of -A. - sync-next-version.cjs: reuse an existing open PR instead of failing gh pr create on rc re-runs; regex-parse the PR number and fail loud; discriminate the git diff --cached --quiet exit code (only status 1 == has-diff, else rethrow); git add -u to avoid sweeping runner artifacts into next; tolerate already-merged on admin merge. - tests: +2 (existing-PR reuse, non-diff rethrow); 14/14 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3e836fef0d |
feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/
Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.
Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.
Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
source-checkout-gated build fallback) instead of LLM re-derivation; the engine
capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.
Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.
* test(#550): RED — status×verification re-cut + probe-core engine specs
Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):
status: resolved | dismissed | unresolved (lifecycle, shared)
verification: explicit | backstop | null (only when resolved)
- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
to be extracted — validateResolution(r, validators), validateRequirement,
analyzeCoverage(items, resolutions?, validators), byVerification rollup,
runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
{resolved,backstop}; coverage gains byVerification.{explicit,backstop};
proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
COUNT preserved on every fixture (closed set = resolved+dismissed; doc
line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
rewritten to the two-axis model.
Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).
* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)
Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.
probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
items[] (core never assumes propose is deterministic — edge resolves via
LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
{categories, verification, requiredFieldsByVerification} (ADR-550 #5)
edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.
* chore(#550): register probe-core.cjs artifact in ledgers + inventory
New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:
- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).
probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.
* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]
trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.
Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.
* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard
Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:
- validateResolution now enforces the 'verification is null unless resolved'
invariant for EVERY status (not just resolved): a dismissed/unresolved
resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
(was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
contract; a new test locks that an all-dismissed run is NOT affirmatively covered
(byVerification is the honest gate).
Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.
* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics
Re-review #5 (trek-e) clarity edits:
- Decision 5: annotate that only contract item (a) ships on #584 (the edge
adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
parenthetical — the blessed/implemented semantics are count-preserved = the
CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
carrying the per-tier resolved-status breakdown. The old parenthetical
contradicted the shipped count.
* test(#550): cover runProbeCli structural-guard numeric-count branch
Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.
* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)
The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.
* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)
A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.
* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)
templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.
* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)
The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.
* docs(#550): add how-to for resolving edge-coverage findings (B1)
Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.
* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)
trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.
* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)
trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.
* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)
trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.
* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe
Rebased onto next (
|
||
|
|
94872662e9 |
feat(#1082): complete phase 5 — descriptor-drive all install surfaces + materialize the InstallPlan — ADR-857/1016/58 (#1080)
* feat(#1077): phase 5f-2 — drive the hookEvents dialect (PostToolUse/AfterTool) from the descriptor postToolEvent (bin/install.js) and preToolEvent (applySettingsJsonHooks in runtime-hooks-surface.cts) now select the event-name dialect from registry.runtimes[id].runtime.hookEvents instead of the hardcoded (runtime === 'gemini' || runtime === 'antigravity') check: hookEvents === 'gemini' → AfterTool/BeforeTool; else → PostToolUse/PreToolUse. hookEvents threaded into the applySettingsJsonHooks opts bag. Equivalence-preserving (Codex-verified): hookEvents 'gemini' is exactly {gemini, antigravity}, 'claude' the rest; undefined → claude dialect (matches the old else). The per-event SET guards (isQwen||claude → SubagentStop/Stop/PreCompact; runtime==='claude' → FileChanged; isGemini → Gemini agent-events) stay HARDCODED — hookEvents (2-value) is too coarse to drive them (the event set differs within hookEvents='claude'); per-event-set drive tracked in #1076. Registry-parity test (enh-1077): asserts BOTH post-tool (AfterTool/PostToolUse) AND pre-tool (BeforeTool/PreToolUse) dialects are a pure function of hookEvents, for gemini/antigravity/claude/augment — non-vacuous (catches a broken hookEvents thread). Closes #1077 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1077): build hooks/dist in before() so dialect-drive test passes in scoped CI hooks/dist is gitignored and absent in scoped/windows CI jobs that do not pre-run build:hooks. Without it, install() finds no hook files and all AfterTool/BeforeTool/PostToolUse/PreToolUse event arrays come back empty, failing every hook-presence assertion. Added an idempotent ensureHooksDist() called in a top-level before() — mirrors the pattern from bug-376-claude-js-hook-gsd-rewriter.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1055): add installSurface/writesSharedSettings/permissionWriter/extendedHookEvents to runtime descriptors Purely additive: four new fields on all 16 runtime capability.json descriptors, validator extended with three new closed-vocab sets, registry regenerated. Test fixtures (VALID_RUNTIME_CAP and makeRuntimeCap) updated to include the new required fields so all 255 capability-registry tests continue to pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1076): drive per-event hook guards from extendedHookEvents descriptor Replace hardcoded runtime-name checks (isQwen||runtime==='claude', runtime==='claude', isGemini) in applySettingsJsonHooks with a single descriptor-driven extendedEvents array derived from the new opts field. Remove isQwen and isGemini derivations (no remaining uses after the three guard blocks are migrated). Wire extendedHookEvents from the capability registry in bin/install.js call site. Add behavioral regression test (enh-1076-extended-hook-events-drive.test.cjs) confirming the drive is purely descriptor-based and runtime-name-agnostic. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1055): drive resolveRuntimeConfigIntent from the runtime descriptor; retire hand-kept REGISTRY - Rewrites src/runtime-config-adapter-registry.cts to require capability-registry.cjs and read installSurface / writesSharedSettings / permissionWriter from runtimes[id].runtime; deletes the hand-kept REGISTRY const (ADR-857 phase 5g drive 2). - ALLOWED_CONFIG_RUNTIMES is now derived from descriptor entries that have installSurface. - Fixes the configFormat parity gate in scripts/gen-capability-registry.cjs to read installSurface directly from capMap descriptor bodies, breaking the require cycle (adapter now requires the generated registry; gen-script must not require the adapter). - Adds golden-master test tests/enh-1055-config-intent-descriptor-drive.test.cjs (41 tests) pinning all 16 runtimes' return shapes and the TypeError-on-unknown contract. - Updates scripts/lint-test-file-count.allowlist.json (config module, +1 file). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1076): make hooksSurface descriptor load-bearing for the settings-json hook-skip - Adds hooksSurface?: string to ApplySettingsJsonHooksOpts and destructuring in applySettingsJsonHooks (src/runtime-hooks-surface.cts). - Replaces the hardcoded !isOpencode && !isKilo hook-skip guard with hooksSurface !== 'none'; removes the now-unused isOpencode/isKilo derivations (ADR-857 phase 5g drive 3). - Passes hooksSurface from the runtime descriptor at the applySettingsJsonHooks call site in bin/install.js using the established _capabilityRegistry?.runtimes?.[runtime]?.runtime?.hooksSurface idiom. - Extends tests/enh-1076-extended-hook-events-drive.test.cjs with two new suites proving: (a) hooksSurface:'none' writes no hooks regardless of runtime name; (b) hooksSurface:'settings-json' writes hooks even for 'opencode' (previously hardcoded to skip). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: record installSurface/writesSharedSettings/permissionWriter/extendedHookEvents descriptor axes in ADR-1016 Add Decision 7a documenting the four axes added in the 5f-completion pass, update axis counts from "six" to "twelve", note 5f-completion drives as done in Decision 8's ladder, update Out of scope to reflect #1055/#1076 are done and 5g (InstallPlan capstone) remains the only open phase. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1055): parity gate must fire on configFormat↔installSurface mismatch (read installSurface at the descriptor level) The test fixture makeRuntimeCapMap did not include installSurface in the runtime object, so the gate's typeof r.installSurface !== 'string' guard always skipped the entry and never threw. Added installSurface as an optional third parameter to makeRuntimeCapMap and passed the correct installSurface values ('settings-json' for claude, 'codex-toml' for codex) to the two THROWS tests. The gate implementation already reads r.installSurface correctly from the descriptor level. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1076): add installSurface↔hooksSurface + extendedHookEvents↔hookEvents consistency gates with rejection tests GATE A: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES map in validateRuntimeBody enforces that a runtime's hooksSurface is valid for its installSurface (e.g. profile-marker-only only allows none, codex-toml only allows codex-hooks-json). Derived from the 16 real runtime descriptors. GATE B: validateRuntimeBody checks that if extendedHookEvents contains Gemini agent-events (BeforeAgent/AfterAgent/BeforeModel), hookEvents must be 'gemini'; if it contains Claude-family events (SubagentStop/Stop/PreCompact/FileChanged), hookEvents must be 'claude'. Added 10 rejection tests in suite 27 covering each gate + each new field validator. All 16 real runtimes satisfy both gates (verified before coding). Exports: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES, VALID_INSTALL_SURFACES, VALID_EXTENDED_HOOK_EVENTS, VALID_PERMISSION_WRITERS, validateRuntimeBody. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1076): strengthen hooksSurface-drive assertions; defensive hooksSurface fallback; drop vacuous dup 1. bin/install.js: add explicit literal fallback for hooksSurface when the committed capability registry fails to load (opencode/kilo → 'none', all others → 'settings-json'). The descriptor is always the source of truth in normal operation. 2. enh-1076 Suite 7: change SessionStart assertion from key-presence (hasOwnProperty) to at least-one-command (hasHooksFor), so the test fails if hooks are initialized-but-empty. ensureHooksDist() in before() guarantees hook files exist. 3. enh-1055 Test 2: remove vacuous duplicate suite that re-asserted intent.runtime === row.runtime already fully covered by Test 1's deepStrictEqual over all four fields. 4. capability-registry.test.cjs: fix stale comments in the grok-skip test that claimed the parity gate uses the adapter registry; gate reads purely from the descriptor (installSurface absent → typeof r.installSurface !== 'string' → soft-skip). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1082): materialize the InstallPlan — collect install-level descriptor axes into resolveInstallPlan; route install()/finishInstall() through it (ADR-58/5g) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: record 5g InstallPlan materialization (ADR-58 Accepted, ADR-1016 phase-5 complete) ADR-1016 Decision 8 step 7 updated to DONE: resolveInstallPlan(runtime) in runtime-config-adapter-registry collects install-level descriptor axes into the typed InstallPlan consumed by install()/finishInstall(). Out-of-scope section updated: 5g capstone is complete, phase 5 fully materialized. ADR-1016 line ~20 updated: InstallPlan IS now materialized (both halves). ADR-58 Implementation note added (2026-06-11): realized in runtime-config-adapter-registry (co-located with adapter-selection). CONTEXT.md Runtime Config Adapter Registry entry extended to document resolveInstallPlan and both-halves realization. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1082): update install drift guard to the resolveInstallPlan seam (5g) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
fbd62cd84f |
feat(#1035): phase 5a — author 16 role:runtime capability descriptors (registry-only) (#1039)
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).
Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).
Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).
Closes #1035
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
b2222fa769 |
docs(#860): ADR-3660 addendum — Nth-runtime-via-existing-layout is an addendum, not a new ADR
Maintainer governance decision (waiving a standalone ADR for Qoder, PR #1021): registering a runtime that reuses the existing profile-marker-only install surface + a single skills kind is an enhancement governed by ADR-3660 via addendum. Records the addendum-vs-new-ADR qualifying criteria, the normative agent-frontmatter contract (name+description only — the sibling converter shape), and Qoder as the first runtime logged under this path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5f48a42514 |
docs(#1022): resolve step-vs-gate model question — steps additive, gates block, mode self-gates (#1025)
Record the resolution of #1022 (surfaced scoping the §5.6 ui-phase cutover): a step is purely additive and never halts the host; host-blocking preconditions are gates (blocking/onError:halt) — no hook-model change. Runtime/mode context (auto/chain vs manual) self-gates in the skill, not via when (config-only). §5.6 decomposes into the existing plan:pre step (ui-phase, self-gates on frontend + pipeline) + a new plan:pre gate (frontend-and-no-UI-SPEC → halt, when: workflow.ui_safety_gate) that blocks planning in manual mode — preserving the "run /gsd:ui-phase first" UX (maintainer call: pipelines-only auto-fire). The render-hooks dispatch template grows to handle gates, not just steps. Recorded in ADR-894 (clarification) + CONTEXT.md (RULESET.CAPABILITY.step-additive-gate-blocks). Unblocks the §5.6 cutover. Closes #1022 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
61b29d2af6 |
docs(#1016): ADR-1016 runtime capability descriptor (ADR-857 phase 5 design) (#1019)
Realize ADR-857 Branch 8 (host-CLI support as role:runtime Capabilities) + materialize ADR-58's InstallPlan. Closes the 6-axis descriptor vocabulary, absorbs the hard-case runtimes as data, and stages the install migration as a 5a→5g ladder (4/6 axes already modularized; InstallPlan is the 5g capstone, reachable by collection, not a big-bang rewrite). Amended after a plugin-side design grill (rubber-duck + grill-with-docs vs #956 MemPalace / #999 Impeccable): - New term Connected Capability (CONTEXT.md) — a Capability whose integration shape brings its own external process/service/state; orthogonal to authorship. Named, tracked gap (vehicle #956); current schema does not express it. - Narrowed the dogfood claim: the descriptor dogfoods the runtime interface only, not the feature-plugin/Connected path. - Structural "off means off" rule (CONTEXT.md RULESET.CAPABILITY.off-means-off): the host derives shared outputs from active hooks; a hook adds/is-counted, never mutates host source. Ratify in ADR-894; proven by spike #1018. - Hand-waves resolved by code: sandboxTier real-but-thin; model-catalog orthogonal (not a 7th axis); converters closed into a ConverterName enum + added the kimi-agents artifact kind; configHome is pure read-only. - Hook-firing path is unproven (render-hooks built, never consumed) → spike #1018 must prove render-hooks→live-workflow execution before phase-5 build. Design-only; no code. Status: Proposed. Closes #1016 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1c86368785 |
fix(#947): restore gsd- prefix on Hermes skills for canonical dispatch (#955)
* test(#947): add regression tests and update stale Hermes assertions - Add bug-947-hermes-gsd-prefix.test.cjs: 12 TDD tests covering fresh install canonical layout, bare-stem migration, manifest key format, and non-Hermes runtime isolation - Update hermes-skills-migration.test.cjs: bare-stem → gsd-prefixed path and name assertions (#947 canonical layout) - Update install-nested-layout.test.cjs: Hermes NEST matrix prefix '' → 'gsd-' - Update install-regressions.test.cjs: Defect #1 now seeds bare-stem dirs (help/, quick/) and asserts gsd-help/ canonical output; use real GSD stems so readGsdCommandNames() migration finds them - Update install-runtime-artifacts.test.cjs: Hermes nested layout and legacy migration assertions align with gsd- prefix - Update install.test.cjs: Hermes install test uses gsd- prefixed paths - Update runtime-artifact-layout.test.cjs: prefix '' → 'gsd-' Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#947): restore gsd- prefix on Hermes skills for canonical dispatch Hermes skills were installing under bare-stem paths (skills/gsd/<stem>/SKILL.md, name: <stem>) due to prefix: '' set in ADR-3660 / #3664. This broke /gsd-<stem> dispatch and forced users to invoke skills without the gsd- namespace prefix. - src/runtime-artifact-layout.cts: change Hermes skillsKind prefix from '' to 'gsd-'; skills now land at skills/gsd/gsd-<stem>/SKILL.md with name: gsd-<stem> - bin/install.js _runLegacyInstallMigrations: invert the #3664 migration — remove stale bare-stem dirs (using readGsdCommandNames() to distinguish GSD-owned stems from user content), keep gsd-* dirs which are now canonical - bin/install.js _runLegacyUninstallCleanup: also remove bare-stem dirs on uninstall for clean teardown - bin/install.js uninstallRuntimeArtifacts: post-cleanup removes DESCRIPTION.md and empty skills/gsd/ category dir on Hermes - bin/install.js: remove skillListPrefix Hermes exception (now uses shared 'gsd-' path) - docs/adr/3660-runtime-artifact-layout-module.md: document #947 reversal of the bare-stem sub-decision Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore: add changeset for #947 fix (#955) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#947): remove ALL pre-migration bare-stem Hermes skills on reinstall (adversarial review) Replace readGsdCommandNames()-based bare-stem cleanup (which missed skills not in the commands source tree, e.g. dev-preferences) with _removeHermesBareStemDirs(), called AFTER the install loop when the exact set of installed gsd-<stem>/ dirs is authoritative. For every gsd-<stem>/ written this run, the corresponding bare skills/gsd/<stem>/ is removed. User-owned bare dirs with no gsd-<stem> counterpart are preserved. Add two adversarial-review regression tests that FAIL on old code: - bare skills/gsd/dev-preferences/ removed when gsd-dev-preferences/ installed - user-owned bare dir with no gsd-<stem> counterpart is preserved (no over-deletion) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
34435736e4 |
docs(#959): ADR-959 capability command contribution (ADR-857 phase 4d design) (#960)
Design how a Capability contributes a gsd-tools CLI command family and how the hardcoded 73-case runCommand switch opens to registry-driven dispatch. Realizes ADR-857 decision 7's reserved commands/module field (deferred by ADR-894). Grilled to its leanest form: the registry DISCOVERS a standard route*Command (no rebuilt handler table, no new arg convention); dispatch sits in the default case (collision structurally impossible, no shadowing gate needed); graphify is the first real cutover (lowest blast radius, has skill+cluster+config gate, full-only so 4c stays no-op), proven equivalent and serving as the phase-6 template. Design-only; CONTEXT.md gains a "Capability Command Family [Planned]" entry. Closes #959 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
86845340dc |
chore(#930): remove self-masking next dist-tag repoint from release finalize (#931)
* chore(#930): remove self-masking next dist-tag repoint from release finalize The "Clean up next dist-tag" step silently failed under OIDC trusted publishing (which can't write dist-tags) while unconditionally reporting success via || true + an echo. It also violated the release model by trying to repoint @next→stable; @next is managed exclusively by the rc job's --tag next publish. Closes #930 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * docs: update ADR-660 to reflect removal of next dist-tag repoint The finalize job no longer runs `npm dist-tag add … next`; update the ADR-660 description of step 4 to match the new behavior — @next is managed exclusively by the rc job. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
197d6bffe2 |
docs(#894): ADR-894 Capability declaration format + registry generation (#895)
* docs(#894): ADR-894 Capability declaration format + registry generation ADR-857 rollout phase 3a (design-only). Resolve ADR-857's deferred open question — the on-disk Capability declaration format — as a reviewable design ADR before any generator code. Specifies: the capabilities/<id>/capability.json folder layout (migration-staged ownership — declarations reference existing stems until the phase-6 move); the capability.json schema for role:feature (skills/agents/hooks/federated config/ loopHooks) and role:runtime (the six closed projection-primitive axes); the 12 named Loop Extension Points; the gen-capability-registry.cjs generator design (validation + cross-capability invariants + --write/--check drift gate, mirroring gen-inventory-manifest); the generated capability-registry.cjs shape (by-id / by-skill / by-loop-point indexes + requires-closure); and a full worked example (the UI capability: ui-phase + ui-review + agents + config + two loop hooks). No code — design artifact only; the generator build, federated config loader (3b), and loop seam (3c) implement against this contract. Closes #894 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#894): amend ADR-894 with grilled capability declaration format Stress-tested the declaration format before merge; the format changed materially. Amendments: - loopHooks[] -> three typed arrays (steps/contributions/gates), each with its own shape (step: ref+produces/consumes; contribution: fragment+into agent-role; gate: check+blocking). - Add the Loop Host Contract (§3): each step publishes its points, agent roles, and core artifacts so the generator validates hooks against reality, not trusted strings. - requires = capability ids only (host implicit); add tier-monotone invariant; drop the requires:["plan"] error from the example. - Config federation = atomic move: a migrated key leaves the central schema in the same PR; presence in both is a collision (invariant stays). - One registry, role-partitioned indexes (feature indexes vs runtimes index). - Rework the UI worked example to the split-array shape (2 steps + 1 gate) + a contribution illustration. Adds a "Grilling amendments" section recording the six changes. * docs(#894): amend ADR-894 with round-2 grilling (operational reality) Second design-grill round, folded in before merge: - Loop Host Contract is GENERATED from structured workflow markers (<loop-point>/<agent-role>/<loop-artifact>) via gen-loop-host-contract.cjs — it can't drift from the real workflows. - Hook activation `when`: cheap deterministic config-level gating evaluated by loop.render-hooks; deeper phase-context applicability self-gates inside the dispatched skill (no phase-context vocabulary to keep honest). - `tier` is the source of install-profile + cluster membership; profiles and clusters are generated from tier + requires-closure (/gsd:surface operates on capabilities) — collapses ADR-857's multiple toggle systems. - Gate `check` = query | declarative-predicate | agentVerdict; agentVerdict is forced advisory; only deterministic checks may block. - byLoopPoint ordering is materialized in the registry; render-hooks filters the active set + renders. Same-capability hooks degrade gracefully when an entry step self-gates. - Rollout: registry-only until atomic per-feature cutover (no double-execution with still-inlined workflow features). Updates the Grilling amendments / Consequences / Alternatives / Open questions sections; reworks the UI example with `when`. --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
03b467f2e7 |
docs(#857): ADR-857 Capability system — five-step loop core, features as plug-ins (#858)
Record the final-state architecture: the five-step loop (Discuss → Plan → Execute → Verify → Ship) plus shared-infrastructure skills are the privileged host/core; every other feature is a Capability (plug-in) selectable at install and toggleable after restart, attaching through ~12 Loop Extension Points. Adds docs/adr/857-capability-system.md (Status: Proposed) capturing the eight resolved design decisions, the resolved design details (extension points, contribution merge, declaration shape, runtime descriptor, deferred trust gate), alternatives, consequences, and a six-phase rollout. Records the new domain terms in CONTEXT.md: Capability, Capability Registry, Loop Extension Point, Runtime Capability. Runtime/CLI support is itself a Capability (role: runtime) — Claude Code, Codex, Antigravity tier-1; the seam is declarative-over-primitives so third-party CLI support lands later as an additive loader, no rework. Closes #857 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f66c4a082c |
feat(#766): distribute gsd-core as a native Claude Code plugin (#797)
* feat(#766): distribute gsd-core as a native Claude Code plugin Add an additive .claude-plugin/plugin.json manifest plus hooks/hooks.json so gsd-core can be installed as a first-class Claude Code plugin (marketplace or zero-friction @skills-dir), with /gsd-core: namespaced commands and lifecycle management — alongside the unchanged npm/file-copy installer. - .claude-plugin/plugin.json: validated with 'claude plugin validate --strict' - hooks/hooks.json: mirrors the installer's always-on Claude hook wiring via ${CLAUDE_PLUGIN_ROOT} - package.json: ship .claude-plugin in the npm tarball - tests/issue-766-plugin-manifest.test.cjs: manifest + always-on-hook-contract drift guards - docs: install-on-your-runtime.md + FEATURES.md Closes #766 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#766): add ADR-766 + glossary entry for Claude Code Plugin Manifest Module Record the plugin manifest as the Seam projecting gsd-core's artifact surfaces onto the Claude Code plugin contract (sibling of the Runtime Artifact Layout Module, ADR-3660), with the defined kind->field mapping and the always-on hook projection rule. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3011b937ad |
docs(#58): add ADR for Runtime Install Policy Module boundary (#762)
Record the Runtime Install Policy Module decision and ownership boundary: install policy projects a pure, typed install plan by composing artifact placements (ADR-3660) and command text (ADR-0009) plus per-runtime config intentions, with no filesystem IO; runtime adapters consume the plan and execute concrete file mutations and format-specific config rendering. Explicitly records what stays outside the policy module (TOML/JSON/Markdown serialization, merge semantics, filesystem effects). Adds the ADR index row in docs/adr/README.md and a glossary entry in CONTEXT.md. Leads the installer-refactor chain (#58 -> #60 -> #56). Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
11afca2968 |
feat(#656): Research module — content-addressed cache + provider seam + registry-API legitimacy (#664)
* feat(#656): add Research Store module (content-addressed cache, TTL staleness) Content-addressed research cache behind a clock seam: researchKey (sha256, deterministic), putResearch/getResearch ({hit,stale}, never throws), ttlForSource (curated HIGH 30d / MED 7d / web LOW 1d), two-tier resolveStorePath (curated -> ~/.gsd/research-cache, web/synthesis -> project .planning/research/.cache). 28 behavioral + property tests; boundary coverage at ttl-1/ttl/ttl+1. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#656): add Research Provider module (waterfall + confidence + plan) Single source of truth for the Balanced provider waterfall (docs Context7->Ref->Jina, web Exa+Tavily, fallback Perplexity/Brave, Firecrawl scrape-only). classifyConfidence stamps HIGH|MEDIUM|LOW by provider (never throws). providerAvailability maps config flags to usable providers. planResearch checks the Research Store (injected seam) and returns cache-hits + a per-question fetch plan, falling through the waterfall to the always-available websearch terminal. 22 behavioral + property tests. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#656): add Package Legitimacy module (registry-API verdicts, slopcheck optional) Replaces the pip-install-or-degrade slopcheck prose gate with code: classifyPackage (pure, never throws) computes OK|SUS|SLOP from tunable thresholds (minAgeDays 30, minWeeklyDownloads 1000, requireRepo). checkPackages queries injectable npm/PyPI/crates registry adapters (real https with 5s timeout, degraded-not-thrown on failure); slopcheck is one optional adapter that can only escalate severity, never degrade to [ASSUMED]. 34 behavioral + property tests; boundary coverage on age and downloads (limit-1/limit/limit+1). Known follow-up: real npm adapter must add api.npmjs.org last-week downloads fetch (currently null -> unknown-downloads). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#656): detect Tavily/Ref/Perplexity/Jina provider keys; complete npm downloads adapter config: add tavily_search/ref_search/perplexity/jina availability flags (env var or ~/.gsd/<x>_api_key), mirroring brave_search/exa_search/firecrawl, so the Research Provider waterfall can gate them. package-legitimacy: real npm adapter now fetches api.npmjs.org last-week downloads (bounded, degraded-not-thrown) so weeklyDownloads is populated. +12 config tests; 34 legitimacy tests unchanged. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#656): expose Research seam via gsd-tools query (research-plan, research-store, package-legitimacy) Routes the L2-hybrid surface so agents reach it as CLI: 'query research-store get/put' (cache, HOME-sandboxable), 'query research-plan --input' (cache-hits + fetch plan from planResearch), 'query package-legitimacy check --ecosystem' (async registry verdicts). Commands skip .planning root resolution and appear in top-level usage. 5 behavioral runGsdTools tests; command-contract unchanged (335). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#656): document Research module (CONTEXT predicates, ADR-0656, architecture, changeset) Adds GSD-RESEARCH.* + DEFECT.RESEARCH-PROVIDER-PROSE-DRIFT predicates to CONTEXT.md, ADR-0656 recording the L2-hybrid seam decision, a docs/ARCHITECTURE.md Research Module subsection, and an Added changeset fragment (pr:0, backfill on PR). Notes the #657 deferrals (agent collapse + install.js MCP mapping). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#656): sync inventory for research modules Regenerate INVENTORY-MANIFEST.json and bump docs/INVENTORY.md CLI Modules count 82->85 with rows for research-store/research-provider/package-legitimacy (DEFECT.INVENTORY-DRIFT). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#656): eslint-ignore generated research .cjs artifacts (ADR-457) research-store/research-provider/package-legitimacy .cjs are tsc-generated from src/*.cts, so they belong in the ESLint ignore block (lint the .cts source, not the emitted .cjs). Fixes tests/551-eslint-bin-lib-coverage. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#656): backfill changeset pr number to #664 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#656): satisfy eslint lint-tests gate Fix 20 eslint errors in the new research files: use helpers.cleanup() instead of raw fs.rmSync() in tests (local/no-raw-rmsync-in-tests, Windows-EBUSY retry budget); drop redundant '| string' union members and unnecessary type assertions; deterministic object normalization in researchKey (no-base-to-string). Logic unchanged; 6180 tests still green. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): harden package legitimacy per review (W1/W2/I3/I4) W1: httpsGet now reads statusCode; npm/PyPI/crates map 404 -> exists:false -> SLOP (registry-existence is the #1 slopsquatting defense; previously only npm caught it). Transport made injectable (_setHttpGet) for hermetic 404 tests. W2: suspicious-postinstall is now terminal SLOP independent of the optional slopcheck adapter, and the regex drops the bare https?:// arm (over-fired on esbuild/sharp/node-gyp) for shell-exec/download-exec signatures only. I3: checkPackages now threads version to registry.lookup and adapters verify that specific version exists. I4: moreServerVerdict -> moreSevereVerdict. +11 regression tests (all RED-first); 45 total green. Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): research-store tier coherence + freshness + version TTL (W4/I1/I2/I4) I1: tier now derives from source (curated -> user ~/.gsd, else -> project .planning), not kind, so put-tier and get-tier can't diverge; kind is a key component only. W4: getResearch searches both tiers and returns the freshest (non-stale preferred), never letting a stale curated entry shadow a fresh web one; blank version caps TTL at 1 day (no 30d on version-blind keys). I2: atomic platformWriteSync instead of raw fs.writeFileSync on the shared global path. I4: dropped the dead ttlForSource arm. CLI get now searches both tiers. +5 RED-first regression tests; 38 green. Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): expose classifyConfidence as a CLI route, killing dead code (W3) Adds 'gsd-tools query classify-confidence --provider X [--verified]' so research agents get the confidence tier FROM CODE (provider waterfall + verification lever) instead of asserting it in prose. classifyConfidence previously had no runtime caller. HIGH means 'trusted provider'; --verified raises web results to MEDIUM (verification semantics documented in ADR-0656). +4 behavioral tests. Addresses review by @davesienkowski on #664 (W3). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): close Codex adversarial-review findings (path-traversal, version-age, malformed-cache) HIGH: research key must be 64-hex sha256 (isValidResearchKey) + resolved-path containment check in put/get + CLI validation -> blocks '../../x' arbitrary-file-write. HIGH: package legitimacy now derives publishedAt from the REQUESTED version (npm time[version], PyPI releases[version] upload_time, crates versions[].created_at) so a new malicious version of an old package can't inherit old age and evade 'too-new'. MEDIUM: getResearch validates entry shape (finite fetched_at + positive ttl + required fields) -> malformed cache entry is a miss, not fresh-forever. +regression tests (RED-first); 111 green. Codex adversarial review (required pre-PR gate). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): close code-review correctness findings (1) package-legitimacy CLI now rejects unknown --flags instead of silently consuming the following package as a flag value; only --ecosystem takes a value. (2) crates recent_downloads (90-day) normalized to a weekly figure before the minWeeklyDownloads threshold (was ~13x too lenient). (3) research-plan --input validates parsed JSON is an object with an Array questions before destructuring -> clean usage error instead of an uncaught TypeError on null/bad input. (4) research-store put rejects a flag value that is itself a --flag (no more storing '--source' as content). (5) planResearch skips questions whose text is not a non-empty string instead of emitting question:undefined. +13 RED-first regression tests; 143 green. Code-review gate. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(#657): extract researcher documentation_lookup to shared @-reference 6 researcher agents carried a near-duplicate <documentation_lookup> block; consolidate into gsd-core/references/research-documentation-lookup.md (@-included). Unifies the ctx7 CLI fallback to the safer 'command -v ctx7' guard (drops silent 'npx --yes ctx7@latest' execution in 5 agents). Behavior-preserving dedup; inventory 63->64 references. Phase A of the agent collapse. Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(#657): extract researcher philosophy + verification-protocol to shared @-references philosophy and the pitfalls+pre-submission-checklist common-core were near-duplicated in project/phase researchers; consolidate into gsd-core/references/research-{philosophy,verification-protocol}.md (@-included). phase-researcher keeps its 3 extra checklist items inline. Pre-submission domains checklist made agent-agnostic so project-researcher doesn't lose features/architecture coverage. Write-contract intentionally left inline (bug-214 tests assert it verbatim). Inventory 64->66 refs. Behavior-preserving. Phase A. Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#657): wire gsd-phase-researcher to the Research seam (Phase B / S1) The phase researcher now CALLS the code seam instead of carrying inline mechanics: provider waterfall -> 'gsd-tools query research-plan' (+ research-store put to cache digests); confidence-tier prose -> 'gsd-tools query classify-confidence'; slopcheck pip-install protocol -> 'gsd-tools query package-legitimacy check'. This makes the Research module a real runtime consumer (validates the seam end-to-end, addresses reviewer S1) and removes the duplicated waterfall/confidence/slopcheck prose. RESEARCH.md output contract, commit step, structured returns, and Phase-A @-includes unchanged. package-legitimacy-gate.test.cjs rewritten prose-grep -> behavioral (asserts the seam invocation). Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#657): wire gsd-project-researcher to the seam + add tavily/ref/jina MCP tools (Phase C.1) project-researcher now calls gsd-tools query research-plan / classify-confidence (+ research-store put) instead of the inline provider waterfall + confidence-tier prose (mirrors the phase-researcher rewire; no package-legitimacy — phase-only). Output contract (STACK/FEATURES/ARCHITECTURE/PITFALLS/SUMMARY.md + sections, no-commit, structured returns, Phase-A @-includes) unchanged. Adds mcp__tavily/ref/jina__* to the project/phase/ui researcher tools frontmatter (Balanced provider set) so install.js MCP mapping (C.2) has a consumer. Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#657): cover tavily/ref/jina MCP install handling + frontmatter parity guard (Phase C.2) Investigation: exa/firecrawl have no explicit per-runtime tool-mapping — every mcp__<server>__* except context7 rides the generic passthrough (Copilot lowercases; OpenCode/Cursor/Windsurf/Augment keep as-is; Gemini auto-discovers). tavily/ref/jina are handled identically, no install path broken. Added 12 copilot-install passthrough tests + a mcp-tool-inheritance parity guard (tavily co-declared with exa, jina with firecrawl, ref present across the 3 web researchers) so the MCP set can't drift. No io.github registry ids invented (none sourceable in-repo); documented as a follow-up. 488 tests green. Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#657): profiles as source of truth for researcher agents + drift-guard (Phase C.3) scripts/research-profiles.cjs declares each of the 7 researcher agents' identity + contract (name, description, color, tools, required @-includes, required gsd-tools seam calls, output-contract markers). scripts/gen-research-agents.cjs --check validates every committed agent against its profile; --write regenerates ONLY the frontmatter from profiles (body untouched) and is a verified no-op against the current agents (zero diff = fidelity). tests/research-agent-profiles.test.cjs is the DEFECT.GENERATIVE-FIX drift guard. Design note: profiles govern the generatable/contract surface rather than destructively regenerating the disparate operational prose bodies (those were deduped via @-includes in Phase A). scripts/ is not inventoried (no inventory change). Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#657): complete agent provider-dispatch + parity guard; align legitimacy field; validate profiles Adversarial-review findings: (HIGH) the seam-wired agents' Step-C dispatch only mapped 6 providers, so a planResearch result of jina/ref/perplexity/brave (reachable via the waterfall fallbacks) had no handling -> agent stall; completed both agents' dispatch to all 9 PROVIDER_WATERFALL ids + a catch-all, and added a parity test asserting agent dispatch stays in sync with research-provider PROVIDER_WATERFALL (DEFECT.GENERATIVE-FIX). (MEDIUM) phase-researcher package-legitimacy JSON example used 'package' but the module returns 'name' -> aligned. (LOW) gen-research-agents checkAgent now returns a clear failure for a malformed profile instead of throwing. +parity/validation tests (RED-first). Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): make classifyConfidence verification-evidence-driven (W3) Confidence conflated provider authority with claim verification — context7/ref stamped HIGH purely by provider identity, and the only verification lever was a self-set --verified flag. Split into two axes: provider authority (static) + verification evidence (code-computed). HIGH now requires ground-truth corroboration (legitimacyVerdict OK), independent of provider; authority alone caps at MEDIUM; SLOP caps at LOW; the self-reported --verified is demoted to a MEDIUM-only web lever. HIGH = corroborated-against-authoritative-source, not a correctness guarantee. Adds --legitimacy-verdict to the classify-confidence CLI; updates CONTEXT.md predicate + ADR-0656 (tier set unchanged, ADR-consistent). Addresses davesienkowski's W3 review on #664. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#656): bind classify-confidence verdict to code, closing CLI self-grading Adversarial review found the new --legitimacy-verdict flag was caller-supplied, so an agent could self-assert OK->HIGH without any real legitimacy check — reintroducing the exact self-grading hole W3 closes. Remove the free flag; the CLI now computes the verdict via checkPackages only when --package/--ecosystem is given (code-computed, not agent-asserted). Update the stale CLI test (context7 alone -> MEDIUM) and extend the property test to vary legitimacyVerdict. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0afed31904 |
docs(#660): add ADR for release-from-next-head release model (#661)
Replaces the persistent/frozen release branch + hand-moved tag with: release always cut from next's head, immutable tags minted once at finalize, next on a -dev stream, and @next dist-tag as the RC surface. Closes #660 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
463cffd894 |
chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/ Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary (`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers are unaffected. Mechanical (bulk, ~90% of the diff): - `git mv get-shit-done gsd-core` - Swept path/identifier references across the repo via `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead preserves the five legitimate slug variants that are NOT the directory: get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names). - Build/manifest wiring: package.json (bin, files, coverage globs), tsconfig.build.json (outDir), ~86 .gitignore build-output entries, stryker.config.mjs, scan-ignore files, install.js path strings. - Frozen (not rewritten): CHANGELOG.md history; translated docs (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/). New logic (review here): - src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper ADR-0008 installer migration. On upgrade it walks the legacy `~/.claude/get-shit-done/` tree, classifies each file via the prior install manifest, and emits remove-managed / backup-and-remove for managed files while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked root and symlinked entries; bounds-checks every path under configDir). The framework rolls back on install failure. Emptied dirs may remain (framework has no recursive dir-removal primitive) — documented. - scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare `get-shit-done` directory token (split token to avoid self-match; case- insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines). Wired into the lint-tests CI job. - Restored scripts/lint-package-identity-drift.cjs detection regexes (the mechanical sweep had wrongly rewritten the old-name patterns it exists to detect) and marked them as intentional legacy references. - TDD tests for the migration and the guard; do.md slash-command guard regex tightened so a `/gsd-core/bin` path segment is not mistaken for a command; changeset + docs/installer-migrations.md row added. Breaking: the installed runtime path moves `~/.claude/get-shit-done/` -> `~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed files (preserving user files) on upgrade. Users with custom hooks/configs hardcoding the old path must update them. Closes #604 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unsweep pending changesets + allowlist injection-example docs CI fixes for the rename PR: - Do not sweep pending .changeset/*.md (ephemeral release-note fragments, like CHANGELOG); reverted those body edits so 5 pre-existing malformed fragments (missing type/pr) no longer enter the PR diff and trip docs-lint. Allowlisted .changeset/ in the legacy-name guard accordingly. - Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in prompt-injection-scan.sh: they contain intentional injection examples / security-model prose; the path-reference rewrites are kept. CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR; none in the new migration/guard) and are out of scope for the rename. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): resolve CodeQL alerts surfaced on this PR The rename diff touched files carrying pre-existing CodeQL findings; per the no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving them off. All behavior-preserving: - scripts/ci-test-scope.cjs: build the config-path match from string .includes() instead of a RegExp over an arg-derived value (js/regex-injection). - src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName so the table-cell escape is complete (js/incomplete-sanitization). - tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization). - tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace, keep the meaningful POSIX-class conversion (js/identity-replacement). Verified: build:lib green; the touched test files + ci-test-scope + profile-output suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization) The prior commit's fixes for two alerts were ineffective: - ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file` reaching static regex `.test(file)` calls (not the config rule). Removed ALL regex over file/t — startsWith/includes/=== string checks + an isWindowsHint helper — so there is no regex sink for the tainted value. - js/incomplete-multi-character-sanitization (3 test files): a single `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint loop (replace until stable) plus a final bare-opener strip. Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL CodeQL flags the regex PATTERNS syntactically (regex-injection on the --files arg split; incomplete-multi-character-sanitization on the <!--...--> replace), so loop fixes do not satisfy it. Made these paths regex-free: - ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/). - 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)). Behavior preserved; ci-test-scope + the 3 suites pass; guard clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unblock security base64 scan on the large rename diff The security job hit its 10m timeout: base64-scan.sh choked on the binary test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/ non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings), and the ~800-file rename diff is slow to scan regardless. - scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they can't carry base64-obfuscated *text* and feeding NUL bytes through the per-line scanner is pathologically slow. collect_files already filtered binary *extensions*; this catches binary *content* in text extensions. - .github/workflows/security-scan.yml: raise the security job timeout 10m->30m to accommodate very large diffs (the scan itself is unchanged). Verified locally: scan skips the fixture, 0 "ignored null byte" warnings, 0 findings, exit 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): sweep get-shit-done refs introduced by merging next The branch was updated with next (#614/#384/#618 etc.), which reference the get-shit-done/ dir (still named that on next). Swept the stale references in the merged files to gsd-core so the rename stays consistent and lint:legacy-name passes: - commands/gsd/discuss-phase.md (runtime-launcher shim paths) - src/core.cts (getAgentsDir layout comments) - tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib) Verified: guard 0 violations; build green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant The #614 runtime-launcher shim added to discuss-phase.md references `${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it mis-read the directory path as a dangling `/gsd-core` command ref (same class as the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path segments are not treated as slash-command references. Verified locally on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22 image) full suite: 0 failures - bug-3683 + bug-2954 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI) CI intermittently failed state.test's gsd-tools subprocess with "findProjectRoot is not a function" (flip-flopping across legs; not reproducible on mac full suite, gsd-test linux full suite, test:unit, or state.test x8). findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs); binding it via destructure at module-load can be undefined under a load-ordering edge. Resolve it lazily at call time via a small wrapper so the lookup happens after core.cjs is fully initialized. Verified green on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22) full suite: 0 failures - state.test.cjs: 106/106; gsd-tools loads cleanly. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): allowlist verification-patterns.md placeholder examples in secret scan The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var examples (illustrative Stripe test-key / database-URL / API-key placeholders) — not real credentials. Added it to .secretscanignore with the strict annotation, mirroring the existing gsd-core/workflows/plan-phase.md exception. Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next exits 0 with 0 findings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3bb2f8f1c5 |
docs: rebrand to GSD Core and restructure docs with Diataxis (#605)
* chore: wire docs/agents config into AGENTS.md Agent skills section
Add the `## Agent skills` discovery block pointing the engineering
skills at the existing docs/agents/{issue-tracker,triage-labels,domain}.md
files (issue tracker, triage label mapping, single-context domain docs).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: rebrand to GSD Core and restructure docs with Diataxis
Reorganise the root README and docs/ around the Diataxis framework
(tutorials, how-to guides, reference, explanation), add new how-to
guides and schema references (STATE.md / CONTEXT.md / PLAN.md /
planning artifacts), and cross-link the whole set. Update the lone
legacy gsd-build reference to open-gsd; keep internal get-shit-done/
filesystem paths unchanged (directory rename tracked separately in
open-gsd/gsd-core#604). Regenerate the ja-JP, ko-KR, pt-BR and zh-CN
localised trees to mirror the new structure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: backfill changeset PR number (#605)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
5cd52eb151 |
enhancement(#537): pilot TS build-at-publish for bin/lib (semver-compare) (#541)
* docs(#457): rewrite ADR-457 to ground truth and accept build-at-publish The prior draft asserted a codebase state that never existed (13 tsc-generated files, src/ trees, a tests/cjs-ts-parity.test.cjs). Corrected to verified ground truth (84 bin/lib .cjs, 1 value-baked package-identity.cjs, no tsc pipeline), distinguished value-baking from transpilation so package-identity stops being miscited as precedent, made check-in-the-artifact vs build-at-publish the central decision, and flipped status to Accepted (build-at-publish). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * build(#537): pilot TS build-at-publish for bin/lib (semver-compare) First hand-written module collapsed to a TypeScript source of truth per ADR-457. src/semver-compare.cts compiles (tsc, strict, noEmitOnError) to a gitignored get-shit-done/bin/lib/semver-compare.cjs. build:lib is wired into build, pretest, pretest:coverage, and prepublishOnly so the artifact is built before test and shipped on publish. Type-aware ESLint on src/**/*.cts immediately caught the params were over-typed as `unknown` (no-base-to-string); narrowed to a honest VersionInput domain type. Behavior preserved: semver-compare.test.cjs (14) and bug-10 (4) pass against the generated output; runtime consumer changeset/cli.cjs unaffected. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): make build-at-publish robust across all CI paths (codex review) Adversarial review found the pilot's generated artifact would be missing on clean CI checkouts. `pretest`/`pretest:coverage` only fire for `npm test`, but CI runs `test:unit`/`test:integration`/`test:install` and `node run-tests.cjs` directly — none of which built the artifact, so any suite requiring semver-compare.cjs would hit module-not-found on a clean checkout, and install-smoke's `npm pack` could ship without it. - Add a `prepare` script (`npm run build:lib`). `npm ci` runs it automatically, so every CI test job and install-smoke's pack emit the artifact before use. This is the idiomatic npm mechanism for compiled-output-not-in-git and fixes both the test and pack paths in one place. - Add `src/` + `tsconfig.build.json` to ci-test-scope and the install-smoke / mutation path filters, so a source-only edit to a migrated module still triggers its tests and mutation coverage (prevents silent CI skips). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): map src/*.cts to built artifact in mutation changed-files detection Follow-up to the codex re-review. The prior commit added src/**/*.cts to the mutation workflow's path trigger but left its "compute changed core lib files" step diffing only get-shit-done/bin/lib/**/*.cjs — which are now gitignored and never appear in a diff. A source-only edit would trigger the workflow then early-exit ("no core lib files changed"), silently skipping mutation testing. Map each changed src/*.cts to its built get-shit-done/bin/lib/*.cjs path (the on-disk artifact Stryker mutates after prepare/build:lib), merge with the hand-written .cjs diff, and apply the test/excluded-module filters once. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): use 'src/' pathspec in mutation diff (git glob doesn't match top-level) Codex review caught that `git diff -- 'src/**/*.cts'` returns empty for a top-level file like src/semver-compare.cts — git's default pathspec glob does not match `**` across zero directories (verified on git 2.50.1). The prior commit's src-detection therefore never fired, so source-only changes still skipped mutation. Switch to the dir-scoped pathspec 'src/' + a `.cts` grep (robust for flat and nested layouts), and broaden the workflow path trigger to 'src/**' to match install-smoke. Verified end-to-end: a change to src/semver-compare.cts now resolves to get-shit-done/bin/lib/semver-compare.cjs in the --mutate list. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#537): add changeset fragment for build-at-publish pilot Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): replace prepare with prepack + build-if-missing; defer mutation wiring CI surfaced three real issues the local run and codex review missed: 1. lockfile-sync failed on every platform. Root cause: `npm ci --dry-run` (the repo's lockfile health check) RUNS the `prepare` script, but in dry-run the devDependencies aren't installed, so `tsc` is not found (exit 127) and the check reports a misleading "out of sync". `prepare` is the wrong hook for a build needing a devDep. Replace it with `prepack` (runs only on pack/publish, when node_modules exists) for the tarball path, and build the artifact inside scripts/run-tests.cjs (build-if-missing) for the test path — the universal chokepoint every CI test invocation funnels through, including the direct `node run-tests.cjs --files-from` step that bypasses npm lifecycle hooks. The guard is a no-op once built, so the run-tests harness test is unaffected. 2. The Stryker mutation gate ran only 1 test against semver-compare (~0% score, 71/71 mutants surviving) — a Stryker test-selection problem orthogonal to the build migration, and raising the score needs property tests (ADR-456). Revert the mutation.yml src wiring; mutation coverage for src-authored modules is a separate follow-up tracked in #537. (The deletion of the gitignored top-level .cjs does not match the workflow's `bin/lib/**/*.cjs` git pathspec, so the gate skips cleanly.) Verified: clean-room `npm ci --dry-run` exits 0; deleting the artifact then running a suite rebuilds it; run-tests harness 22/22 green; `npm pack` includes the built artifact via prepack. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d2ff4ac092 |
docs(#22): add ADR for plan-vs-codebase drift guard (defaults + resolver seam) (#484)
Consolidated decision record: source-grounding verification default-on (plan_review.source_grounding), intel.enabled stays opt-in, and the three-valued symbol-resolver seam with a climbable adapter ladder. Refs #22 Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ed1c20061b |
docs(#456): add testing standards, ADRs 452/456/457, and CONTEXT.md entries (#458)
- docs/adr/452-eslint-lint-harness.md (Accepted): adopt ESLint flat config with typescript-eslint, eslint-plugin-n, eslint-plugin-no-only-tests, and local AST-rule plugin; retire homegrown scripts/lint-*.cjs regex checkers; three test-rigor rules ship at warn, promoted to error after #453 cleanup - docs/adr/456-test-rigor-architecture.md (Accepted): deterministic-over-racing via injectable clock seam + node:test mock.timers; antagonistic tier with fast-check + Stryker at 80% threshold; typed-surface mandate; delete-bad-tests policy with no-permanent-quarantine - docs/adr/457-generated-cjs-single-source.md (Proposed): future direction to collapse ~59 hand-written bin/lib/*.cjs to TS-generated single source; eliminates tsconfig.lint.json stopgap; marked Proposed / not yet executed - TESTING-STANDARDS.md: orients to existing docs; codifies six test-rigor contracts; adds new policies (no-timing-assertion, clock-seam, property-based, mutation-score, delete-bad-tests); pairs each with exact ESLint rule names; markdownlint-clean (MD040 fences, MD056 table columns) - CONTEXT.md: adds six RULESET.TESTS.* predicates (no-timing-assertion, clock-seam, property-based-testing, mutation-score, delete-bad-tests, eslint-harness) and five glossary terms (clock seam, deterministic scheduler, property-based test, mutation testing/score, ESLint harness) - docs/adr/README.md: adds index rows for ADRs 452, 456, 457 Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5ca646f015 |
feat(#443): unified cross-provider effort controls + fast-mode-aware routing (#463)
* test(#443): RED unified effort + fast_mode + resolve-execution All 68 tests failing as expected — no implementation yet. Covers: effort cascade (tier defaults, overrides, invalid fallthrough), fast_mode cascade (boolean-only, tier defaults), resolveEffortForTier escalation, renderEffortForRuntime clamping, resolve-execution CLI, config schema new keys, QA hostile-input matrix. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(#443): unified cross-provider effort + fast_mode knobs and resolve-execution query Adds config-driven effort control (universal ladder: minimal<low<medium<high<xhigh<max) and fast_mode propagation knobs, with per-runtime rendering that clamps the unique tail values (max=Anthropic-only clamps to xhigh on Codex; minimal=Codex-only clamps to low on Claude). Key changes: - config-schema.manifest.json: add effort.default, fast_mode.enabled as validKeys; add 4 dynamicKeyPatterns for effort.routing_tier_defaults, effort.agent_overrides, fast_mode.routing_tier_defaults, fast_mode.agent_overrides; fix stale _comment - config-defaults.manifest.json: add effort and fast_mode blocks with tier defaults - model-catalog.cjs: add EFFORT_RENDERING map, renderEffortForRuntime(), RUNTIMES_WITH_FAST_MODE - model-profiles.cjs: re-export new catalog exports - core.cjs: add resolveEffortInternal, resolveFastModeInternal, resolveEffortForTier, VALID_EFFORTS, EFFORT_SET, nextEffort; pass effort/fast_mode through loadConfig - commands.cjs: replace reasoning_effort in cmdResolveModel with unified effort; add cmdResolveExecution (superset command with effort_rendered, effort_param, effort_propagation, fast_mode, fast_mode_supported) - gsd-tools.cjs: add resolve-execution case with --effort/--fast-mode/--attempt flags - tests/feat-443: 69 tests covering cascade, rendering, escalation, CLI, schema, QA matrix - tests/commands.test.cjs: convert 3 reasoning_effort assertions to unified effort - docs/CONFIGURATION.md: document effort + fast_mode + resolve-execution sections - settings-advanced.md: list new effort/fast_mode keys in confirmation table Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#443): remove dead catalog effort lane; unify codex effort through renderEffortForRuntime - Remove resolveReasoningEffortInternal (catalog-driven effort function) from core.cjs and its export; remove from commands.cjs destructure import - Convert tests/issue-2517-runtime-aware-profiles.test.cjs: all 11 effort assertions now use resolveEffortInternal + renderEffortForRuntime; Claude effort is first-class (output_config.effort); unknown runtimes assert param===null - Convert tests/feat-3023-model-phase-types.test.cjs: replace the entire resolveReasoningEffortInternal describe with unified effort assertions; effort derives from AGENT_DEFAULT_TIERS routing tier, not phase-type tier Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * docs(#443): ADR for unified cross-provider effort + fast-mode routing Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(#443): architecture-level QA invariants + test-strategy doc Add 48-test integration suite (feat-443-effort-fast-mode.integration.test.cjs) covering 8 architectural invariants: cross-provider validity (never emit a value the real API would 400 on), param/channel contract stability, resolve-execution JSON contract (all 8 keys + correct types), totality across the full 33-agent registry, fast-mode honesty (claude always fast_mode_supported=false), precedence first-valid-wins matrix for both effort and fast_mode cascades, dynamic-routing composition (effort escalation independent of model tier), and config-set round-trip for all new effort/* and fast_mode/* key namespaces. Append test-strategy section with invariant rationale and E2E gap documentation to docs/TESTING-SUITES.md. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#443): add failing install-wiring tests for effort per-runtime injection (RED) TDD RED: 10 failing tests covering: - Claude .md gets effort: injected per tier (planner=xhigh, mapper=low, executor=high) - Gemini .md does NOT get effort: (already passing — Gemini-safe) - Codex .toml gets model_reasoning_effort via unified resolver - Config-driven: effort.agent_overrides drives both Claude .md and Codex .toml - Source purity: agents/*.md have no effort: key (already passing) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(#443): wire effort per-runtime at install (Claude .md frontmatter + Codex .toml unified) - Import AGENT_DEFAULT_TIERS and renderEffortForRuntime from model-catalog.cjs - Add readGsdEffectiveEffortConfig(targetDir): reads merged effort config from .planning/config.json (per-project wins) + ~/.gsd/defaults.json (global fallback), same probe pattern as readGsdRuntimeProfileResolver - Add resolveInstallTimeEffort(effortCfg, agentName): pure function matching resolveEffortInternal() precedence (agent_overrides > routing_tier_defaults > default > 'high') without loadConfig side-effects (no sub-repo detection, no migration writes) - Claude agent copy loop: inject `effort: <value>` into frontmatter ONLY for runtime === 'claude'; all other .md runtimes (Gemini, Qwen, Hermes, etc.) stay effort-free (Gemini-safe source contract preserved in agents/*.md) - generateCodexAgentToml: add effortCfg param; emit model_reasoning_effort from unified resolver (replaces old catalog entry.reasoning_effort); Codex clamps max → xhigh via renderEffortForRuntime('codex', ...) - installCodexConfig: pass readGsdEffectiveEffortConfig(targetDir) to generateCodexAgentToml so per-project config wins for Codex .toml too - Update failing tests to GREEN: 12/12 pass; all 17 install tests pass; 2847/2848 unit tests pass (1 pre-existing failure: policy-shell-pinning) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#443): source install effort defaults from manifest (kill drift) + guard test Replace hardcoded _GSD_EFFORT_MANIFEST_TIER_DEFAULTS and the 'high' fallback in resolveInstallTimeEffort with values read from config-defaults.manifest.json at module init, using the same __dirname-relative path install.js already uses for all shared manifests. Add feat-443-effort-defaults-drift.test.cjs to assert equality between install.js's runtime constants and the manifest on every CI run. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#443): reconcile Codex TOML tests with unified effort design The #443 unified effort resolver makes generateCodexAgentToml always emit model_reasoning_effort (driven by resolveInstallTimeEffort, not model_profile_overrides). The test 'generated TOML omits reasoning_effort when runtime has none' had an obsolete premise — model_profile_overrides.reasoning_effort:'' no longer suppresses unified effort. Convert it to assert the new invariant: Codex TOML always carries a valid model_reasoning_effort from the agent's routing tier (xhigh for gsd-planner, a heavy-tier agent), while model_profile_overrides model override is still respected. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#443): make install.js effort resolution lazy (no load-time side effects breaking launcher-parity) Replace module-load-time IIFE + hard throw (config-defaults.manifest.json read) and top-level require of model-catalog.cjs with a lazy _getGsdEffortCatalog() getter that initialises on first call from resolveInstallTimeEffort / generateCodexAgentToml / Claude .md effort injection. Requiring install.js in unrelated test contexts (e.g. runtime-launcher-parity) no longer triggers manifest IO or throws, eliminating the load-time side effect that changed subprocess exit codes / stderr on the bench. Drift-guard exports (_GSD_EFFORT_MANIFEST_TIER_DEFAULTS / _GSD_EFFORT_MANIFEST_DEFAULT) preserved as lazy getter properties on module.exports so feat-443-effort-defaults-drift still validates them without forcing eager load. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#443): isolate install-wiring test HOME to stop \$HOME/.claude pollution breaking launcher-parity runGlobalInstall() now redirects HOME to a per-call isolated tmpdir in addition to the existing runtime-specific env-var redirects (CLAUDE_CONFIG_DIR, GEMINI_CONFIG_DIR, CODEX_HOME). This ensures install.js code that uses os.homedir() directly — including the ~/.cache/gsd update-check deletion, ~/.gsd/defaults.json reads, and any HOME-relative npm subprocess writes — never touches the real \$HOME during the test. Without the HOME isolation the install test (which is new to this branch and is now picked up by Docker's raw \`tests/*.test.cjs\` glob) could write or delete files under the real \$HOME, causing runtime-launcher-parity test (D) to fail: (D) asserts a loud non-zero exit when \$RUNTIME_DIR/gsd-tools.cjs is absent and gsd-tools is not on PATH, but the launcher's \$HOME/.claude fallback arm succeeds if \$HOME/.claude/get-shit-done/bin/gsd-tools.cjs exists. Also sets GSD_SKIP_STALE_SDK_CHECK=1 to suppress the \`npm ls -g\` subprocess that the global installer spawns — irrelevant to effort-wiring assertions, slow, and potentially writes to ~/.npm cache. All 12 feat-443 install-wiring assertions preserved. Drift-guard 5/5. Unit suite 2848/2850 (pre-existing policy-shell-pinning.test.cjs failure on next). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(#443): add changeset fragment for effort + fast-mode routing Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(#443): set GSD_TEST_MODE before requiring install.js in drift-guard test to prevent HOME leak Without GSD_TEST_MODE=1, require('bin/install.js') runs the module's main install block (guarded by !GSD_TEST_MODE), performing a real global Claude install into $HOME/.claude/. On CI ubuntu where node is on standard PATH, the launcher's $HOME/.claude fallback arm then finds gsd-tools.cjs, causing runtime-launcher-parity test (D) to exit zero when it must exit non-zero. Root cause: feat-443-effort-defaults-drift.test.cjs (unit suite) runs alphabetically before runtime-launcher-parity.test.cjs in the same node --test invocation. Each runs in a separate worker process but shares the same HOME. The drift test's install leaks gsd-tools.cjs into that HOME, then the launcher test's bash subprocess finds it via the $HOME/.claude arm. Fix: add process.env.GSD_TEST_MODE = '1' at the top of the drift-guard test, before the require(installPath) call. This matches the pattern used by feat-443-effort-fast-mode.test.cjs and feat-443-effort-install-wiring .install.test.cjs. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#443): deterministic resolve-execution arg parsing + validate install-time effort (Codex adversarial findings) Finding 1: resolve-execution --effort low gsd-planner misrouted 'low' as the agent. Replace find(non-dash) with a proper flag-consuming loop that collects a single positional; validate missing/extra positionals and malformed --attempt values. Finding 2: resolveInstallTimeEffort returned unvalidated effort strings (e.g. "ultra") verbatim. Each precedence layer now checks GSD_EFFORT_SET (imported once from core.cjs) before accepting a value, mirroring resolveEffortInternal exactly. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#443): newline-agnostic effort frontmatter injection (Windows CRLF) + CRLF-safe assertions Extracts injectEffortFrontmatter(content, effortValue) pure helper that detects EOL (LF vs CRLF) from the opening '---' line and inserts 'effort: <value>' before the closing '---' delimiter using the same EOL as the surrounding frontmatter. Regex now uses /^---\r?\n([\s\S]*?)^---\r?$/m instead of the LF-only /^(---\n[\s\S]*?)(---)(\n|$)/ that silently skipped CRLF files on Windows (git core.autocrlf=true checkout). Also adds 7 unit tests covering LF, CRLF, idempotency, no-frontmatter, and complex frontmatter cases. Exports injectEffortFrontmatter from module.exports. Fixes 6 CI failures in tests/feat-443-effort-install-wiring.install.test.cjs on windows-latest runners (lines 138, 145, 152, 261, 345, 356). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |