653f95e39f7338d08e04f7620c366fe49ce1dfd0
110 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ffd5370464 |
fix(#2903): use the command form that actually works in reader-facing docs (#3047)
* fix(#2903): use the command form that actually works in reader-facing docs Docs told readers to type the colon form, which no runtime registers -- 18 of 19 runtimes use slash-hyphen and the 19th uses shell-var -- so anyone copying an example got an unrecognized command. Swept 178 occurrences across 53 files, locale mirrors included so they do not re-diverge from English. The colon form is a source-authoring token, not a user-facing one: install-time converters key on it to produce the hyphen form runtimes actually register. So the sweep is scoped, and three things are deliberately left alone: - ADRs, which are a historical record; editing their prose falsifies what was written at the time. - The legacy release-notes archive, pending a maintainer decision on whether it follows the same historical carve-out. Excluding it keeps a later reversal additive rather than a revert. - Source artifacts under commands, workflows and agents, where the colon form is load-bearing. Rewriting those would break the installed-skill guarantee across every runtime -- the single largest hazard here. The plugin namespace form is a real, separate token and survives untouched. Adds a lint enforcing exactly that boundary, since the correct form genuinely differs by directory and nothing previously caught the drift. Also fixes a hardcoded colon form in the capability-matrix generator. The sweep alone would have left the generated matrix disagreeing with the template that produces it, so the fix is at the source and the output regenerated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2903): stop the sweep misquoting source frontmatter Adversarial review caught three lines where the sweep rewrote a citation of the literal YAML name: key from a source command file. That key genuinely is the colon form -- this change's own carve-out logic says source-authoring tokens keep it -- so the docs ended up misquoting the real files. One of the three is an acceptance-checklist assertion, which the sweep turned into a false statement. Restored the three citations to match their sources verbatim, surgically: where a line carried both a name: citation and a real reader-facing slash command, only the citation reverted and the command stayed corrected. The guard needed the same distinction, or it would have flagged the restoration and reddened the build: a gsd:<cmd> token preceded by name: is a citation of a source token and is now permitted. The exemption is deliberately narrow -- a bare gsd:<cmd> anywhere else still fails -- with a test pinning that narrowness. Also makes the detection case-insensitive. Review found /GSD:next slipped through silently; no such casing exists in the tree today, so this closes a latent gap rather than fixing a live one. Swept the whole tree for further corrupted citations: none beyond the three. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2903): retire the stale-next invariant and sweep next like every other command Maintainer decision on a genuine conflict between two contracts. Invariant #3054 banned the literal /gsd-next from user-facing docs because it named a retired workflow-advance command. But commands/gsd/next.md is a live command -- the state-aware smart-entry launcher -- and this issue requires docs to use the hyphen form every runtime actually registers. Both could not hold for this one command, so docs had been sidestepping the ban by keeping the colon form, which is exactly the defect this issue exists to remove. FEATURES.md already recorded the reassignment: the hyphen form "is not the retired workflow-advance command; it is reserved for the state-aware smart-entry launcher. Workflow advancement remains under /gsd-progress --next." With that reassignment the invariant's premise is obsolete and the guard now contradicts the documented command form, so it is retired with a comment recording why rather than deleted silently. next is now swept like every other command, and the earlier exemption added to the new guard is removed so nothing is special-cased. Four citations of the literal name: frontmatter key stay in colon form, because the source file really does carry name: gsd:next and a doc quoting it must reproduce it verbatim. Two of those lines were reworded to say which side is the frontmatter key and which is the slash command, since they previously conflated the two. Verified the retired scan would now genuinely fail against this tree -- the conflict was real and resolved, not dodged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2903): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
de78f2eef2 |
docs(#2775): align package-legitimacy docs to the ADR-0656 registry-API gate (#3010)
* docs(#2775): align package-legitimacy docs to the ADR-0656 registry-API gate security-model.md, USER-GUIDE.md, ARCHITECTURE.md, COMMANDS.md, FEATURES.md, and gsd-planner.md's STRIDE template (+ ja-JP mirrors) described the pre-ADR-0656 design: slopcheck as the install-or-degrade gate, with unavailability degrading every package to [ASSUMED]. ADR-0656 inverted this months ago — registry-API verdicts (npm/PyPI/ crates.io) are the gate; slopcheck is an optional escalate-only adapter that no shipped configuration wires. Verified every replacement claim against src/package-legitimacy.cts (checkPackages, classifyPackage, lookupNpm/lookupPypi/lookupCrates) via Memtrace before writing it, so the corrected prose matches the live implementation rather than restating the ADR from memory. Restored docs/explanation/security-model.md:79-84 (and its ja-JP mirror) to original wording after an orthogonal spec review caught that an earlier draft had edited the "Why WebSearch packages are always [ASSUMED]" paragraph — inside the range issue #2775 explicitly named as correct and to leave alone. The ja-JP mirror was missing the closing clause present in the corrected English original ("its absence leaves registry-API verdicts intact rather than downgrading everything to [ASSUMED]") — added for parity. This completes the ja-JP mirror the issue's acceptance criteria named explicitly. zh-CN/ko-KR/pt-BR (not named by #2775, but carrying the same stale design) get the mechanical portion of the same fix: command-string swaps, table headers, ARCHITECTURE.md diagram labels, and technical- term swaps that reuse a word already attested elsewhere in the same file (合法性/적법성/legitimidade for "legitimacy") — surrounding prose untouched. The remainder in those three locales — full-paragraph rewrites of the corrected degrade-path mechanism, deleted "External dependency" bullets, and "manually install slopcheck" code blocks — needs prose composed by a fluent speaker of each language and is filed as open-gsd/gsd-core#3002 with an exact file:line inventory. * test(#2775): acknowledge gsd-planner.md byte growth from the STRIDE-row fix agents/gsd-planner.md grew 14 bytes (49309 -> 49323) from the STRIDE supply-chain row correction (slopcheck -> package-legitimacy gate). Emitted agent/workflow files are byte-tracked; this fragment acknowledges the growth per tests/emitted-attribution.test.cjs's "differential attribution over the real tree" check. * docs(#2775): close ja-JP FEATURES.md gap; fix a ko-KR transliterated heading docs/ja-JP/FEATURES.md:2808 still read the katakana transliteration "スロップチェック verdict" in REQ-PKG-GATE-01 — invisible to a literal "slopcheck" grep, so it was missed when ja-JP parity was checked and declared complete. Corrected to "正当性判定" (legitimacy verdict), matching the term already established in ja-JP/explanation/ security-model.md and ja-JP/USER-GUIDE.md. This was the only remaining ja-JP gap; a full sweep for the transliterated form across docs/ja-JP/ now returns zero hits, and the ja-JP mirror is genuinely at parity. docs/ko-KR/USER-GUIDE.md:398's heading "슬롭체크 판정:" had the same transliteration problem. Fixed inline to "적법성 판정:", reusing the 적법성/legitimacy word already attested two lines below in the same table. A parallel sweep of zh-CN and pt-BR found no transliterated forms of "slopcheck" in either locale. The remaining transliterated occurrence in ko-KR (USER-GUIDE.md:406, the lead-in to the pip-install code block) needs prose composition like the rest of that block and is added to open-gsd/gsd-core#3002's inventory. * chore(#2775): backfill changeset PR number to 3010 --------- Co-authored-by: sim <sim@local> |
||
|
|
7372d99a26 |
enhance(#2800): derive reviewer flag lists and gate reviewer lane docs across locales (#2882)
* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales The reviewer lane roster was hand-enumerated across five documentation surfaces and three workflow files that had drifted apart: --kimi-code was missing from all four translated COMMANDS.md mirrors, --coderabbit from every workflow forwarding list, and --antigravity from FEATURES.md. Adds checkReviewerDocsParity, a second pure gate deliberately separate from checkReviewerLaneParity so a stale doc cannot make the runtime checker look red. Workflows now derive their flag lists from a new review-lane flags query instead of hand-enumerating them, which also retires the unanchored grep that matched --agy inside --antigravity. Documents the previously absent reviewer body and hostBehaviors field in the capability manifest reference. Closes #2800 Closes #2781 Closes #2272 * fix(#2800): key the docs parity table arm on first-cell position Review found the flag arm was file-scoped, so the forwarding row that lists every flag in its third cell satisfied it on its own. Deleting a lane's own reviewer-table row -- the #2781 regression this gate exists to prevent -- therefore passed undetected. Arm 4 keys on the FIRST table cell, which separates a lane row from the forwarding row structurally and in every locale. Regression test included. * fix(#2800): shape-filter the flags subcommand output All three consumers read review-lane flags through an unquoted command substitution so the output word-splits into loop items. Phase 2 admits third-party overlay lanes, so an overlay flag containing whitespace would inject a second loop item and one containing a glob would expand against the cwd. Emit only well-formed flags so neither reaches the shell. * fix(#2800): remove the regex length ceiling and count only prose mentions Review found two real defects in the docs parity gate. The never-throws contract was false: building a RegExp from a declared flag or section title throws SyntaxError past ~100k chars, and Phase 2 admits overlay lanes whose declared strings are untrusted in length. Every one of these matches is literal, so String.includes replaces the regex outright, which also deletes escapeLiteral and the llama.cpp escaping it existed for. Arm 1 was context-blind: a flag mentioned only inside a fenced example or a commented-out row counted as documented. Both are stripped before matching. Also advertises all 13 lane flags in the argument-hint and corrects a stale eleven-lane count in the slug grammar note. * test(#2800): repoint the convergence suite off deleted workflow text The derived flag loop deleted the literal per-flag grep lines four tests matched on. Two of those failed loudly. The behavioral and property tests failed SILENTLY instead: their end marker no longer resolved, so the parse block extracted empty and both passed vacuously, and the property test's gsd_run stub had a no-op default that hid it. All now share one extractor and execute the real deployed block through a gsd_run shim backed by the actual binary. The whitelist assertions become an anti-parity check: re-adding a hand-written flag list must fail. Also repairs two vacuous cases in the docs parity suite. The unreadable-doc test called its own mock rather than the reader, and the integration test bounded nothing, so a doc losing its marker would have been silently skipped and still passed green. * fix(#2800): run the derived flag loop after the launcher preamble The remote matrix caught a real runtime bug, not a test artifact. In autonomous.md and plan-review-convergence.md the launcher preamble that defines gsd_run lives in a separate, LATER bash fence than the derived loop. Each fence is its own shell, so gsd_run was undefined where the loop ran: the command substitution yielded nothing and zero reviewer flags would have been forwarded. Worse than the drift this epic fixes, and silent. The whole CONVERGENCE_ARGS construction moves as one unit, because the --max-cycles append sits between the loop and the preamble and would otherwise have run against an uninitialized variable and then been dropped by the relocated initializer. Also documents all 13 lane flags in help/modes/full.md, which the repo gates bidirectionally against each command's argument-hint. * test(#2800): repoint the two converge suites off deleted flag literals Both asserted workflow.includes('--codex') against the hand-enumerated list the derived loop removed. They now assert the derivation itself, keep --all and --text (convergence controls, still literal), and add an anti-parity guard so re-adding a hardcoded list fails. The lost pass-through proof is replaced with a real one: every flag the tests used to hardcode is asserted present in the actual roster emitted by the binary, which is the property the old assertion was protecting. * test(#2800): acknowledge the workflow byte growth from the derived flag loop * chore(#2800): backfill changeset pr number to 2882 * fix(#2800): strip HTML comments to a fixed point in the parity gate CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick construction smuggles a commented-out row past the gate and it counts as documented. Not an injection risk here since nothing is rendered, but it is the exact false pass this helper exists to prevent. Strips to a fixed point, then treats any surviving opener as unterminated so the multi-line branch closes it on a later line. Terminates because every pass strictly shortens the string. * test(#2800): pin the comment-smuggling regression with a real reproducer The obvious fixture for this class does not reproduce it: <!--<!---->--> leaves a dangling --> rather than a live <!--, and is caught either way, so it would have passed with and without the fix. The join-trick construction (<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely regresses on the single-pass strip and is what the test now uses. --------- Co-authored-by: Test <test@example.com> |
||
|
|
3f6b063fbb |
chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers iterate declared lanes instead of hand-authored per-CLI bash. Five additive descriptor amendments, each forced by a lane that ships today: - LaneHandler gains 'opencode' — the lane rebuilds its review from assistant text parts of a --format json stream; a plain stdout copy re-breaks #1936. - modelConfigKey — antigravity's key is review.models.agy, not .antigravity, so resolving by slug silently dropped a configured model. - defaultHost/fallbackModel — Phase 4 federated every *_host with a default of empty string; the real fallback only existed in the bash. - args becomes an argv template with a closed four-placeholder vocabulary. Positional splicing produced 'codex --model M -o F exec --ephemeral', which is not a valid invocation: codex injects in the middle, twice. - kimi-code lane, with the bounded command-capability probe (needle --output-format) that tells Kimi Code from the legacy python kimi-cli. Parity gate re-pointed: the workflow-text families it scanned are the text this phase deletes, so they are replaced by descriptor-to-registry parity plus an anti-parity check that no bespoke leg returns. jq, curl and external timeout/gtimeout all drop out of the review path. Refs #2782 * chore(#2799): add review-lane query surface and widen the manifest vocabulary Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow loops over, projects all twelve lanes into their capability manifests, and widens capability-validator for the amendments. opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's own admission rule: one lane, justified by a documented upstream defect data cannot express (#1936 — the agent can end its turn with zero output tokens and --format default then drops the assistant text entirely). Two bugs caught by an end-to-end stub run and fixed here: - loadConfigResolved returns a provenance wrapper, not the config; using it directly resolved every key to undefined, which reads as 'nothing configured' and silently dropped every model override. - hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced with a PATH scan that spawns nothing at all. Refs #2782 * chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved lanes, and renders REVIEWS.md sections from each lane's declared reviewsSection instead of thirteen hardcoded headings. review.md drops from 1104 lines to 507 (61KB to 28.7KB). Parity gate re-pointed, as agreed: the leg-marker and section-heading families scanned exactly the text this phase deletes, so they are replaced by descriptor-to-registry parity in both directions, plus an anti-parity check that fires if a bespoke leg is ever re-added. Enum, emitting sites and the Object.keys lock moved together. The budget-trim helper is hoisted out of the Ollama leg: it was always lane-agnostic, and any lane may now declare a promptBudgetKey. Refs #2782 * feat(#2799): bind the consented egress host and re-verify it at invocation Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3 but was not implemented: ConsentRecord had no host field and nothing in the tree bound one, so this phase's rule-4 comparison had no baseline. ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design: isValidConsentRecord does not require it, so every record already on disk stays valid and no re-consent storm fires (D4 rule 5). It is deliberately excluded from disclosureSignature — the loader has no config resolver, so folding a config-derived value in would make loader and lifecycle compute different signatures for the same manifest and re-prompt forever. Install resolves hostConfigKey (falling back to the lane's declared defaultHost, which is what the invocation path uses) and records it. Invocation re-resolves and blocks on mismatch rather than silently redirecting. Absence allows: no record, or a record predating the field, means nothing to compare — denying there would break every existing local-model user on upgrade. Refs #2782 * test(#2799): cover the resolver, runner and handlers; retarget the parity suites Adds the golden invocation-plan table (one row per shipped lane, derived from the bash legs rather than the descriptor types) plus runner coverage for the probe, empty-output policy, the three handlers and the egress check. Retargets the existing suites onto the new contract: descriptor-to-registry parity, the anti-parity check, the opencode handler, and the twelfth lane. Two corrections found by running them: - modelConfigKey was required; that breaks D4 rule 2, since a reviewer manifest authored before this phase would fail validation on upgrade. It is optional, read as null when absent. - the antigravity non-zero-exit test pre-seeded the transcript, which asserted that a STALE entry leaks through — the exact bug the watermark prevents. The spawn now appends, as the real tool does. Refs #2782 * fix(#2799): restore agy --add-dir and the self-report prompt in the handler Retargeting the three legacy reviewer suites off the deleted bash surfaced two real regressions in the port, both #2176: - --add-dir was dropped. Without it agy's permission context never receives the cwd repo, so the agent anchors on its own scratch dir and reviews the plan text in isolation — the exact failure the Review Instructions forbid. It is capability-probed, because an older agy rejects the unknown flag outright and a lane that fails to start is worse than one running on the prompt anchor. - the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS self-report, which is what makes a blind review distinguishable from a grounded one. antigravity now builds its own prompt variant. Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own log is the only evidence that anything failed. The three suites now assert against the plan and the handler instead of matching fence text, so they no longer need allow-test-rule exemptions. Refs #2782 * docs(#2799): document the declared lanes, the new flag, and dropped prerequisites COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph, which is now false: no lane requires jq, curl or an external timeout. Adds the changed-egress-destination behavior, since a blocked lane is something a user can hit. CONFIGURATION.md records that the model config key is declared per lane rather than derived from the flag — antigravity's is review.models.agy — and adds review.models.kimi-code. reviewer-instances.md now routes an instance through its lane's single invocation seam instead of a copied per-adapter bash block, which is what lets a cross-cutting fix reach instances for free. That required implementing the --model/--agent/--as flags it documents; --model re-resolves through the lane's argv template rather than splicing, so the flag lands where the lane declares it rather than ahead of a subcommand. CONTEXT.md glossary gains both new modules. Refs #2782 * chore(#2799): drop the stale emitted-drift acknowledgment The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md. That file now shrinks by ~32KB and every emitted hash that moved is attributable to this diff, so the ack no longer explains anything. Removing the last entry means removing the file: its presence is the alarm, and an empty one signals nothing. Verified by deleting it and re-running the attribution and provenance gates plus lint:ci — all green without it. Refs #2782 * docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782 Five additive amendments, each forced by a lane that ships today, plus two corrections the phase had to make rather than work around: - D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so this phase's rule-4 comparison had no baseline. Recorded because an ADR asserting a rule was delivered is exactly what stops a later phase checking. - The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families scanned the text this phase deletes. Also records that D7's 'skip the probe where no bounding mechanism exists' carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded on every stock macOS host, which ships neither timeout nor gtimeout. Refs #2782 * fix(#2799): close four defects found by adversarial review Two confirmed bugs, both reproduced before fixing: - resolveLanePlan was not total. An openai-http lane with a missing or non-object invoke dereferenced inv.hostConfigKey and threw, contradicting the module's own documented contract; the spawn branch guarded correctly and the http branch did not. The CLI seam resolves every selected lane in one map, so one malformed overlay manifest would have aborted the whole review rather than dropping its own lane. Guarded, plus a per-lane try/catch at the seam so a throw can never take down siblings. - A reviewer-instance model was silently dropped for any lane declaring modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates that cli is a known slug but never that the slug accepts a model, so a user could configure one, get a clean run, and never learn a different model reviewed their plan. Now warns explicitly. Two hardening fixes: - The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in the resolver rather than inherited from a validator that does not run on this path — the module documents itself as the overlay-manifest trust boundary, so it should not depend on someone else having checked. - normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses with an empty hostname, so it became 'localhost://11434' and was compared and requested as if real. An empty hostname now means not-a-URL. Also documents the one gap that cannot be closed here: the antigravity watermark is keyed by workspace, so two concurrent reviews of the same repo share a transcript. agy exposes no per-invocation id to filter on, so the handler now states which half of its never-stale guarantee actually holds. Refs #2782 * test(#2799): retarget the remaining eight review.md-asserting suites The remote runner found 37 failures the local sweep missed (it hit the shell's two-minute cap before reaching these). All eight extract per-CLI bash from review.md that this phase deletes; each protects a real invariant, so each is retargeted onto the plan, the runner or the handler rather than removed. Three real defects surfaced by doing so: - effort args never reached ANY lane. model-resolver.cjs exports no resolveExecution, so effortFor silently returned [] every time. Restored by calling the same bounded resolve-execution query the bash legs used — and NOT with --raw, which prints the resolved effort rather than the picked field, so claude got 'low' instead of '--effort low'. - the timeout guidance lost 'a silent empty output is a timeout kill, not a crash' — the operator note that exists because of the Codex 0xc0000142 misdiagnosis. Restored. - the opencode handler dropped EMPTY assistant text parts. The shipped jq was , and only substitutes for false/null — an empty string is truthy in jq and contributed a blank line. Found by a property test shrinking to ['', '']. The opencode property suite no longer spawns jq at all, which deletes the #2099 hang mechanism it was architected around rather than mitigating it. Refs #2782 * fix(#2799): register the two new generated modules, and untrack them The remote runner caught build output committed to git. Both new modules compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each bin/lib/*.cjs is linted xor ignored according to migration state" failed. Registered both in .gitignore and eslint.config.mjs alongside the Phase 1 module, and dropped them from the index. Nothing about the shipped behaviour changes; the artifacts are rebuilt by build:lib. This is the new-.cts-module registration ripple, and it is the one part of it I had not completed - the CONTEXT.md glossary and the inventory manifest were already done. Refs #2782 * chore(#2799): backfill changeset pr number to 2861 * chore(#2799): backfill changeset pr number to 2861 --------- Co-authored-by: Test <test@example.com> |
||
|
|
c87f6f358e |
enhance(#1854): offer restore for user-added files backed up on update (#2679)
* test(#1854): failing-first coverage for user-files-backup restore Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#1854): offer restore for user-added files backed up on update Adds a restore-custom-files gsd-tools verb and wires it into update.md as a restore_custom_files step: plan, compatibility-check against the newly installed release, then restore only on explicit opt-in. The backup is never deleted, a shipped path is never overwritten, and a single unwritable entry does not abort the rest. Also drops the jq pipe from update-context field extraction (#2589 class, missed by that sweep) and repairs a broken code fence in docs/CLI-TOOLS.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): reject symlinked restore destinations and backup roots Self-review of the restore path found two write-through holes: copyFileSync follows a symlinked destination, so a link planted at the restore target wrote outside the config dir with every ancestor still a real directory; and statSync on the backup root followed a link, letting the walk read arbitrary files and present them as the user's own backup. Both now lstat. Also marks the report's path/detail strings as untrusted data in update.md so the rendered step cannot carry instructions into the runtime model. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1854): move the update-context jq guard into the #2589 sweep update.md joins the AUDITED list rather than carrying a duplicate assertion in the backup-restore suite, and the guard gains a negative-proof companion so 'no jq pipe' cannot pass by the fields simply no longer being read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): validate manifest files map shape before trusting it Security review flagged that Object.keys on a non-plain-object files field yields numeric-index keys matching nothing, so the managed-path check dies silently while manifest_found still reports true. Shape, not just type (ADR-227): an array or scalar files map is now an unusable manifest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): size the restore prompt by eligible_count Spec review found the prompt was driven by entries.length, so a backup holding only blocked entries asked "Restore 1 file(s)?" when accepting would restore zero. The question now reads eligible_count, and an all-blocked backup reports its reasons instead of offering a choice that cannot be honored. The decline path names the resolved backup_dir rather than the bare directory name. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1854): use t.skip on hosts without symlink support A bare return in a node:test body registers as a PASS, so the four symlink guards silently reported green on unprivileged Windows instead of skipping. Adds the dangling-link destination case the security review called out, and moves outside-dir teardown to t.after so a failing assert cannot leak it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): unfence restore hint, regen goldens, widen install timeout Three gate failures from the c99d612a5 run, all root-caused: 1. capability-registry (3): update.md's decline message put an instructional 'gsd-tools ...' line in an UNTAGGED fence, and the guard treats untagged fences as shell blocks. Retagged both display blocks as text and switched the hint to the resolved 'node <config-dir>/.../gsd-tools.cjs' form users can actually paste. 2. golden-install-parity (19): update.md and gsd-tools.cjs ship, so every runtime fixture moved. Regenerated; the diff is exactly those two hashes per fixture, no other drift. 3. install.test.cjs (5): one real failure, four cascades. The Cursor suite's before hook died on 'spawnSync ETIMEDOUT' at the 60s cap while the node22 lane passed the SAME commit in 12.7s. A full install measures 13-30s idle, so 60s was under 2x headroom and shrinks with every file added to the payload. Raised to 120s, matching the heavy case already in this file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1854): backfill changeset pr number to 2679 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c5e0371775 |
feat(#1951): reversibility tagging — gate one-way-door decisions (#2471)
* test(#1951): add failing-first tests for reversibility tagging Red phase for issue #1951 (reversibility tagging: classify decisions by undo cost, gate one-way doors behind a checkpoint:decision). Tests assert, per the issue's acceptance criteria: - discuss-phase CONTEXT.md template records a **Reversibility:** field with a rationale on captured decisions, and states it is optional - gsd-planner @-references planner-reversibility.md and stays under the 49152-char agent cap (LARGE_CAP, tests/agent-size-budget.test.cjs) - a one-way rating inserts a checkpoint:decision before the dependent task; reversible inserts none; costly is flagged but never blocks - the taxonomy defaults to reversible when unsure (checkpoint-fatigue guard) and inserting a checkpoint implies autonomous: false - docs/reference/plan-md.md documents <reversibility> as optional with all three ratings - --no-reversibility-gates parses to REVERSIBILITY_GATES=false, is injected into the planner prompt, and is advertised in the command argument-hint and help full mode (argument-hint parity) - the override suppresses the gate but still persists the rating - cmdVerifyPlanStructure accepts every rating and the absent case (additive-validator guarantee, behavioral via runGsdTools) - parity: thinking-models-planning.md #4 adopts the canonical three-level taxonomy and the binary REVERSIBLE/IRREVERSIBLE vocabulary is gone - no content loss from the planner extraction made to fit under the cap Prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately — regression guards proving the validator already accepts unknown optional tags. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1951): reversibility tagging — gate one-way-door decisions Classify planning decisions by what undoing them would cost, and give a one-way door a human beat before the agent walks through it (issue #1951, The Pragmatic Programmer Topic 15 'Reversibility'; Bezos's one-way/two-way door framing). Acceptance criteria met: - discuss-phase records an optional reversibility rating with a rationale on <decisions> entries in the phase CONTEXT.md template. Unrated decisions are treated as reversible, so existing phases are unaffected. - a one-way rating makes gsd-planner insert a checkpoint:decision before the task that implements the decision, reusing the existing checkpoint mechanism -- no new checkpoint machinery. - reversible ratings trigger no checkpoint; costly ratings are flagged in the plan but never block. - the rating persists on the task as the optional <reversibility rating=> element. cmdVerifyPlanStructure accepts every rating and the absent case; the structural validator does not reject unknown optional tags. - --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses checkpoint insertion for intentionally-unattended runs while still recording ratings -- the override changes what stops the run, not what the plan remembers. Single taxonomy, not two: references/thinking-models-planning.md #4 already shipped a binary REVERSIBLE/IRREVERSIBLE classification and is loaded by both gsd-planner and gsd-plan-checker. It is rewritten onto the canonical three-level vocabulary and now points at planner-reversibility.md as the taxonomy owner, with a parity test that fails if the surfaces diverge (DEFECT.GENERATIVE-FIX-DIVERGENCE). agents/gsd-planner.md sat 47 chars under the 49152 LARGE_CAP, so the checkpoint DO/DON'T guidance was relocated verbatim into planner-antipatterns.md -- already @-referenced from the same section for the same topic, so the planner still loads it and nothing was dropped. A test guards the relocation against content loss. Files: gsd-core/references/planner-reversibility.md (NEW, canonical taxonomy + emission rules + anti-patterns), gsd-planner.md, plan-phase workflow/command/help (flag wiring + parity), plan-md.md schema, discuss-phase context template, CONTEXT.md glossary, INVENTORY + manifest, size baselines, install goldens, plugin skills regen, changeset. Closes #1951 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1951): address orthogonal review findings Two isolated reviewers (correctness + security), neither of which authored the change. Every finding fixed: Security — the rationale is untrusted input (ADR-1577). It originates in conversation and flows CONTEXT.md -> planner -> PLAN.md -> executor, each hop an LLM reading the previous hop's output, with no validation on the path. planner-reversibility.md and the discuss-phase template now state it is data and never instructions, and name the </reversibility> early-termination hazard explicitly -- a rationale that closes its own element injects sibling structure the executor reads as real tasks. Four tests guard it. Correctness 1 — nothing machine-enforced the feature's own promise: a task rated one-way with no preceding checkpoint:decision validated as fully clean, so a planner error silently reopened the gap this feature exists to close. cmdVerifyPlanStructure now warns on an ungated one-way rating. A warning, not an error: <reversibility> stays additive and the plan stays valid. Four tests cover ungated (warns), gated (silent), still-valid, and reversible/costly never flagged. Correctness 2 — pass-always test. The --no-reversibility-gates parse test substring-matched the whole workflow file, and plan-phase.md prose mentions both tokens in one sentence, so it passed with the bash conditional deleted: it was testing the documentation, not the parser. Now scoped to the fenced bash blocks and matched as one physical line, with a negative control confirming prose alone cannot satisfy it. Correctness 3 — costly had no itemized emission rule, only one-way did, so two agents could diverge on whether to tag costly at all. Correctness 4 — template convention break: the example ratings were bare while every sibling field uses [...] to signal substitution, inviting an LLM to copy one-way/costly forward as boilerplate. Now bracketed. Correctness 5 — latent false-green: .includes('reversible') also matches inside irreversible/irreversibility, which appear in anti-pattern prose, so a surface that dropped the real taxonomy entry would still pass. Now word-boundary matched. ADR-857 phase-6 ceiling — the first gsd-test run caught plan-phase.md 1216 bytes over its frozen 94519 ceiling (it had 49 bytes of headroom on next). The ceiling may only rise for privileged host machinery, and reversibility gating is optional-feature logic, so the wiring was slimmed to its minimum and the explanatory prose moved to the reference files the planner already loads. plan-phase.md is now 94400 bytes -- 119 under the ceiling and 70 bytes SMALLER than on next, so the host loop shrank while gaining the feature, which is what phase 6 ratchets toward. The tracer contract (tests/tracer-bullet.test.cjs) is unchanged. Lint — fixed an unnecessary non-null assertion in verify.cts and a CRLF-fragile bare \n regex in the new test (DEFECT.WINDOWS-CRLF-TEST- PORTABILITY, the #1658/#1668/#2206/#2449/#2450 class). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture must carry the common task elements The gated-one-way fixture built a checkpoint:decision task from the abbreviated skeleton in gsd-planner.md, which shows only the checkpoint-specific elements (<decision>/<context>/<resume-signal>). cmdVerifyPlanStructure requires <name> and <action> on EVERY task regardless of type, so the fixture failed validation for reasons that had nothing to do with reversibility: errors: ["Task missing <name> element", "Task 'unnamed' missing <action>"] Caught by gsd-test on 14d14a39 (2 failures, both this fixture). The canonical shape is in tests/verify.test.cjs:266 — a checkpoint task carries <name>/<files>/<action>/<verify> like any other. Fixture corrected to match. Verified behaviorally against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Not a product defect: the validator's every-task contract is intentional and pre-existing, and docs/reference/plan-md.md scopes its required-element list to type=auto/tracer only because those are the elements a planner must author, not because checkpoints are exempt from <name>. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1951): backfill changeset pr number to 2471 * fix(#1951): CodeQL incomplete-sanitization + prompt-injection scan collision Both CI failures were real defects in code this PR added, not false positives. CodeQL js/incomplete-sanitization (high), reversibility-tagging.test.cjs:46 — the namesRating helper built its regex with `rating.replace(/[-]/g, '\\-')`, which escapes the hyphen but not backslash, so the escape was incomplete. It was also unnecessary: `-` carries no special meaning outside a character class. Replaced with a complete metacharacter escape (backslash included). Word-boundary behavior verified unchanged across all three ratings — notably that "irreversible" prose still does not satisfy a "reversible" match, which is the false-green this helper exists to prevent. Prompt injection scan — the checkpoint fixture used the human-verification child element inside <verify>. That tag name is a fake-instruction-boundary pattern in scripts/prompt-injection-scan.sh, and the scan runs over changed files, so copying the shape from tests/verify.test.cjs (unflagged only because it is not in this diff) tripped the gate. Switched to the documented plain-prose <verify> form. The first attempt at that fix failed the same gate a second time: the comment explaining the collision quoted the offending tag literally. The comment now names it in prose instead — the scanner does not care whether a match is code or commentary, which is the whole point of the DEFECT.PROMPT-INJECTION-SCAN-COLLISION note in CLAUDE.md. Verified locally before push: scan reports 0 findings across 57 changed files, eslint clean, and both fixtures still validate as designed (gated one-way silent, ungated one-way warns, neither errors). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): record measured cost and halve gsd-tools spawns The Windows shard 1/3 job timeout was traced to the sharding layer, not to this PR's assertions — see #2472. Two contributing factors were this file's own, and are fixed here. 1. tests/test-timings.json had no entry for reversibility-tagging.test.cjs, so scripts/run-tests.cjs weighted it at the table's median fallback (~315ms) for LPT chunk packing. It actually measures 5595ms — an 18x under-weight. Recorded the measured value from the green gsd-test run (max across the node22/node24 lanes, per gen-test-timings.cjs's convention). Only this one entry: a full regen churns 634 entries of run-to-run drift, and the table is explicitly advisory and un-gated, so a 637-line diff does not belong in a feature PR. 2. Each verifyPlan() spawns gsd-tools, which dominates this file's cost. Spawns cut from 9 to 6 with no coverage lost: - the ungated-one-way warning and its stays-valid assertion now share one plan instead of building the same plan twice; - the reversible/costly never-flagged-as-ungated test was strictly subsumed by the additive suite, which already runs those two ratings ungated and asserts no /reversibilit/ warning at all — and the gate warning's text contains both "reversibility" and "one-way", so the broader assertion catches it. It only re-spawned gsd-tools twice to prove the same thing. Both are symptom fixes. The shard imbalance itself (19/11/10 minutes against a 20-minute cap, from a cost-blind round-robin partition that also reshuffles downstream files whenever one is inserted) is tracked in #2472. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture adopts the #2444 type-branched contract Surfaced by rebasing onto next, which gained #2444 (branch plan-structure validation on task type=checkpoint:*) while this PR was in review. cmdVerifyPlanStructure no longer applies one required-element set to every task. A checkpoint:decision now requires <name> + <resume-signal> + <decision> + <options>, and is exempt from the <action>/<verify>/<done>/ <files> set that auto and tracer tasks carry. The gated-one-way fixture predated that split and failed on the new requirement: errors: ["Task 'Task 0: Confirm the on-disk format' missing <options>"] Fixture rewritten to mirror the checkpoint:decision contract exactly — real <options> with two <option> children — rather than padding it with fields checkpoints no longer need. That also drops the plain-prose <verify> the earlier revision carried purely to dodge the prompt-injection scan; a checkpoint task has no <verify> requirement at all, so the workaround is moot. Verified against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d16a66479a |
feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship Adds a new capability (#1950) that operationalizes GSD's no-defer discipline as a tracked, enforced artifact: accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths across phases, and /gsd-ship blocks while any entry is open. Implementation: - src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR + I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed error assertions. Windows-safe atomic rename with retry on transient EPERM/EBUSY/EACCES. - gsd-tools.cjs: new subcommand (status | append | waive | fixed), wired via routeWindows + HOST_COMMAND_ROUTERS.windows. - capabilities/broken-windows/capability.json: one ship:pre gate with artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0. activationKey windows.enabled (default true) + sibling windows.enforce (default true, separate so tracking can precede enforcement). - gsd-core/workflows/ship.md: capId==broken-windows branch in preflight, sibling to security — reads gsd_run windows status --raw, fails closed on open_count > 0 or unreadable ledger. - agents/gsd-executor.md: extends the existing ## Known Stubs instruction to also append to WINDOWS.md via gsd_run windows append (best-effort, never blocks execution). - agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify items in WINDOWS.md. - gsd-core/workflows/progress.md: surfaces open + waived counts. - docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md: document the gate, waiver mechanism, and new module. - tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check roundtrip property; fail-closed on malformed ledger; security boundary on path traversal in --file. Backward-compatible: a project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly. Disable enforcement per-project with gsd config-set windows.enforce false (tracking continues, gate stays open). * chore(#1950): ratchet size baselines, defer verifier integration - Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632 (broken-windows preflight branch + open-windows surface). - Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also appends to WINDOWS.md). gsd-verifier.md unchanged. - LARGE_CAP (49152) preempted the planned verifier integration (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the documented 'real headroom'). Verifier integration deferred to a follow-up PR that extracts the VERIFICATION.md template (lines 739-859) to gsd-core/references/ — a pre-existing cap-tightness defect this PR exposed but does not expand scope to fix. Verifier integration is not in the issue's acceptance criteria (executor writes is; unmet-truths recording was an enhancement, not a gate). * fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44 cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all caps off → empty hooks'): - capability manifest: rename windows.enabled+windows.enforce (default true) → single federated key workflow.windows_enforce (default FALSE, opt-in). Matches security's workflow.security_enforce convention and makes the adr857 all-caps-off test pass without modification (the test's buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps the gate out of the registry's default ship:pre resolution so existing loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId 'security') stay valid; users opt in via gsd config-set workflow.windows_enforce true. - drop activationKey (security doesn't have one either; workflow.* key doubles as the activation toggle). - regenerate docs/reference/capability-matrix.md to include broken-windows (capability-matrix-sync test). - regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) — installer now emits the new capability + lib file. - update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md, agents/gsd-executor.md to use the new key name and /gsd:colon slash syntax (slash-command-namespace test). - restore accidentally-regressed /gsd:capture in progress.md. Tracking-only by default; enforcement is opt-in. Acceptance criterion '/gsd-ship fails while any ledger entry is open' is met when workflow.windows_enforce=true (test fixture enables it). * test(#1950): update ship:pre structural invariants for 2-gate registry - loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre (security + broken-windows), regardless of activation. Activation tests above still pin security-only or empty behavior via fixtures; these structural tests pin the REGISTRY shape, which has 2 gates as of #1950. - workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce rename added 17 bytes). * fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker Adversarial isolated review (Step 6.3) found 2 HIGH findings that block the PR and 3 mediums. All addressed: H1 (HIGH): description containing the markdown 3-backtick fence would terminate the ledger's JSON code block early inside JSON.stringify output (JSON doesn't escape backticks), corrupting the file and bricking the next parse. Fix: use a 4-backtick fence (json ... ) which JSON.stringify cannot produce on its own, AND validate that no entry text field contains a 4-backtick run (reject at append time with new WINDOWS_INVALID_TEXT reason code). Locked by a regression test. H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger', silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate would then pass on an unreadable ledger — the precise vector the workflow doc claims is impossible. Fix: only ENOENT returns null; every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the gate blocks and the operator sees a real diagnostic. Locked by a regression test that chmod 000s a ledger with open_count=1 and asserts the result is never a false-green 0. M1: writeLedgerAtomic left an orphaned .tmp file on rename failure. Wrapped renameWithRetry in try/catch with best-effort unlink. M2: validateLine silently coerced 'abc' → NaN → null, hiding type drift. Removed the line === 0 special case (was undocumented) and made the error message match the strict check. Now any non-positive- integer line value throws, including strings. M3: tests/broken-windows.test.cjs (with its fast-check property test) was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate src/broken-windows.cts but no test would catch the mutations, producing false surviving-mutant scores. Added to the list. L1 (dead throw e after error()), L7 (line boundary tests, H1/H2 regression tests, 4-backtick CLI test) also addressed. * docs(#1950): inline concurrency + busy-wait notes (review L2+L3) * fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test gsd-test v4 caught two issues: - goldens I regenerated earlier (commit 526682084) predated the L1 routeWindows catch-block cleanup (commit dd844d565). Regenerated via 'npm run gen:golden' against current HEAD so the install parity hash for gsd-tools.cjs matches. - 'append --line boundary' test expected --line 0 to succeed with null entry.line, but the M2 fix correctly rejects 0 (lines are 1-indexed; 0 is not a valid source line). Updated the boundary test to assert --line 0 fails alongside -1 and 'abc'. * chore(#1950): regen goldens after rebase onto next * chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase) * chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md * fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization) CodeQL flagged the markdown-table cell escaper: String(s ?? '').replace(/\|/g, '\\|') — it escapes pipe but not backslash first. A description containing '\|' would render as '\\|' which markdown parses as 'literal backslash' + 'cell separator', splitting the column. Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|). Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped pipe), which markdown renders as a single '\|' inside the cell. The JSON code block (the parse source-of-truth) was already correctly escaped via JSON.stringify; only the display-only table was affected. Locked by a regression test that: 1. Verifies the JSON block reparses with the description intact. 2. Walks the rendered table row counting unescaped pipes — must be exactly 11 (the row separators for 10 cells), proving no in-cell pipe added a split. |
||
|
|
c4237df8e6 |
docs(#2276): 1.7.0 release documentation — what's-new, EoS explanation, feature index (#2282)
Add a curated 1.7.0 release-highlights page (docs/whats-new-1.7.0.md) and a conceptual Embeddable Orchestration System (EoS) explanation (docs/explanation/embeddable-orchestration-system.md), extend docs/FEATURES.md with a v1.7.0 feature section, and wire both new docs into the docs index (docs/README.md) and the root README. Covers the release's marquee changes: the ADR-1239 Host-Integration Interface / EoS (Embeddable Orchestration System) runtime expansion, the Capability + EoS discoverability registries, the gsd-mcp-server companion, model-catalog advances (GPT-5.6, (1M) badge), statusline enhancements, the compact GSD-state format, plus a themed summary of the 100 fixes and 4 security hardenings. Also corrects a stale CONTEXT.md glossary entry: the Capability Registry Overlay now documents the #2009 fail-open behavior for a load-failed gate-declaring capability (previously described as fail-closed). American house style; no parity-gated reference docs hand-edited. Refs #2276, #1678 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8022c5c864 |
fix(#2198): remove dead scan exports, correct injection-scan docs
scanEntropyAnomalies + shannonEntropy were dead exports with zero production callers — the live hooks (gsd-prompt-guard.js, gsd-read-injection-scanner.js) inline their own pattern subsets for hook independence and never called these functions. Changes: - Remove scanEntropyAnomalies + shannonEntropy from src/security.cts - Remove scanEntropyAnomalies test block from tests/security.test.cjs - Correct REQ-SCAN-INJ-02/-03 in FEATURES.md (EN/zh-CN/ja-JP) to describe what actually runs live (injection patterns, invisible Unicode) vs CI-only (base64-decode, codebase scan) - Correct docs/security/baseline.md §2.4 to clarify live hooks inline patterns, not import from security.cts - Add regression test asserting the corrected contract - scanForInjection retained: it serves as the CI codebase-scanner engine |
||
|
|
60e3c4988a |
feat(#2182): scaffold community capability + EoS registry (tests + stubbed core)
Adds the discoverability-registry surface for issue #2182: JSON-sourced capability/eos catalogs, a pure schema/vocab module (registry-schema.cjs) with the ADR-857 loop points + ADR-1239 axes, thin validate/gen CLIs, the registry-entry PR template, README spec, and CONTEXT.md glossary terms. The three pure functions (isValidGsdRange/validateEntries/renderMarkdown) are stubbed here so the comprehensive test suite fails first (red), per the feature-implementation red-first directive; the next commit implements them. Refs #2182 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4483300253 |
fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.
Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
→ ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
→ REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
→ FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
reviewer's own override was ignored); init.quick now resolves `reviewer_model`
(gsd-code-reviewer) and the spawn threads it.
resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.
Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.
Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.
Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
2c878c966f | no-mistakes(document): Sync onboarding docs | ||
|
|
a3cca0704d | no-mistakes(document): Sync onboard documentation | ||
|
|
ed79902509 |
feat(#2007): implement mempalace memory_mode kg_backend and replace routing (#2010)
Wire the two forward-declared mempalace.memory_mode modes so they actually route recall/capture instead of silently behaving as `augment`: - kg_backend: the palace temporal KG is the primary knowledge-graph source; native .planning/graphs/ is the fallback. Non-KG drawer recall stays additive. - replace: recall resolves through the palace as the source of truth; native artifacts are the fallback. Every mode stays onError:skip and default-resilient — an unreachable palace degrades to native memory and GSD keeps writing .planning/graphs/, so no memory is lost. Cross-mode .planning/graphs/ migration remains a documented open question (PRD/ADR §17), out of scope here. Surfaces updated (instruction-only contract): recall/capture commands (+ generated skills), discuss/wave fragments, curator agent, capability.json schema. Docs: how-to Step 3, CONFIGURATION, FEATURES, CONTEXT glossary. Regenerated capability-registry, golden install-parity fixtures (mempalace hashes only), agent-size-baseline. Added a routing-contract + cross-surface parity test. Incidental (folded per no-defer rule): removed pre-existing unused imports (spawnSync in capability-registry.test.cjs; fs in issue-498-package-identity.test.cjs) that eslint flagged in/alongside the touched files. Closes #2007 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bd77b40107 |
feat(#1825): configurable graphify graph location (graphify.graph_path) (#2013)
* feat(#1825): configurable graphify graph location (graphify.graph_path) Add a graphify.graph_path config key (.planning/config.json) that overrides where /gsd-graphify query|status|diff read the knowledge graph, so one curated umbrella-level cross-repo graph can serve multiple sibling projects without N drifting ~5 MB mirror copies. Previously the graph location was hardcoded to <cwd>/.planning/graphs/. - src/graphify.cts: resolveGraphLocation(cwd, planningDir) honors the key (resolved relative to project root; absolute paths honored via path.resolve); falls back to the historical .planning/graphs/graph.json when unset/blank/ non-string (byte-identical). Wired into graphifyQuery, graphifyStatus, graphifyDiff (snapshot travels with the configured graph via dirname), and writeSnapshot. Configured-but-missing -> actionable error naming the path. Build stays project-scoped (skill hardcodes the cp dest); umbrella graph is built in the umbrella project, sub-projects only READ it. - config-schema.manifest.json: register graphify.graph_path in validKeys. - tests/graphify-graph-path.test.cjs: boundary matrix (unset byte-identical, set+present reads configured graph not default, set+missing actionable error, relative resolved vs project root, blank treated as unset, snapshot alongside configured graph, diff from configured dir, build project-scoped) + VALID_CONFIG_KEYS registration. - docs: CONFIGURATION.md row, FEATURES.md REQ-GRAPH-06, CONTEXT.md module note, .changeset (Added). Closes #1825 * docs(#1825): backfill changeset pr number 2013 |
||
|
|
8f2ebbe9bf |
feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap). --gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): backfill changeset PR number (#1996) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): drop Gemini CLI from issue templates (review nit) Removes the sunset Gemini CLI runtime from the two GitHub issue-template runtime lists that the removal PR missed, per @davesienkowski's review nit: - feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise request a feature for a runtime GSD no longer supports) - bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json retrieval-help line Leaves the post-removal templates fully consistent with the Antigravity redirect. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e5ef323b15 |
feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command
Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.
* feat: add /gsd-start smart-entry command
State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.
- src/smart-entry.cts: deterministic situation classifier (no-project,
paused, blocked, verify-failed, needs-first-phase, planning, executing,
verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
dispatcher presenting an AskUserQuestion menu (with --text fallback for
non-Claude runtimes) and dispatching to existing commands. Falls back
to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
(markdown-layer invariants + every emitted command resolves to a real
slash command).
Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.
* refactor: rename smart-entry command to /gsd:next
Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.
All affected tests (188) pass; lint:ci clean.
* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)
Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.
- detectSignals now reads total_phases + percent from nested progress{}
first, then scalar fm, then body; current_phase falls back to the
body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
body Phase field) covering verify-pending + executing situations.
Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.
* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)
Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.
Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
layer is down — read .planning/STATE.md directly with the Read tool
and synthesize a minimal situation + actions menu so /gsd:next stays
useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
/gsd:progress as before.
Matches the direct-read resilience the live agent already did by hand.
* docs: add gsd-next skill surface
* chore: trigger no-mistakes validation
* no-mistakes(review): Fix smart-entry phase ordering
* no-mistakes(review): Fix decimal smart-entry phase ordering
* no-mistakes(test): Fix smart-entry next test contracts
* no-mistakes(document): Docs synced for smart entry
* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: shorten next.md description and update golden install parity fixtures
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: trigger no-mistakes validation
* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files
Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.
* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: regenerate golden install parity fixtures for /gsd:next
Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.
* Fix smart-entry verify-failed phase scoping and empty resolve shim step
Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.
* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: recapture all 16 golden fixtures with updated smart-entry.md hash
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: regenerate fixtures + inventory manifest after rebase onto next
Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next
Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).
Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)
The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs
Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo
Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
recommended action, 1-4 unique-id /gsd:* actions (previously the
one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout
Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:
- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
consolidation file doing dozens of real installs — 41s even on a fast Mac
(much worse on the slow Windows I/O path), plus an install-heavy cluster.
Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.
Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
6d072435d0 |
test(#1975): consolidate 51 CLI + scripts-tooling regression tests into module suites
Fold 51 issue-named CLI black-box + scripts-tooling regression files into their canonical module suites (runtime-launcher-parity, worktree-safety, install-*, managed-hooks, read-guard, capability-registry, etc.), plus a NEW slash-command-namespace.test.cjs grouping the 4 slash/colon-namespace-leak invariant suites that had no canonical owner. Verbatim block-scoped describe wrappers; 427 subtests conserved 1:1. Host-env pre-check (per B2): no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value. One folded suite (bug-3668 runtime resolver) creates an extension-less PATH gsd-tools stub + bash -c; co-locating it with the host's chmodSync tripped local/no-unguarded-nonportable-exec, so it's now Windows-guarded (skip on win32) matching the host suite's own bash -c guard. Regenerates regression-name allowlist (222->182), ratchets file-count allowlist (graphify 7->6, docs entry removed), makes 26 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes stale ids). Repoints 13 tests/ references across CONTEXT.md, COMMANDS.md/FEATURES.md (EN + ja/ko/pt/zh) and ADR-0002. lint:ci green. Part of epic #1969. Closes #1975. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
18995380ce |
feat(#1154): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1738)
* feat(verify-phase): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1154) Carry the edge-probe's existing `backstop` (non-inferable) tier through the plan-phase projection as a structured flat-scalar marker instead of a prose parenthetical, and make verify-phase abstain -> human_needed (never silent-pass) on a backstop truth it cannot confirm with explicit evidence. Truth-axis mirror of #644's prohibition judgment-tier (ADR-550 D4). Engine (deterministic, CI-tested per ADR-550 D5 — never the LLM verdict): - src/probe-core.cts: truthStatement/truthVerification normalizers, projectTruths (conservative serializer), dispositionForUnverifiableTruth (backstop+no-evidence -> unverified/flagged/insufficient_spec; backstop+evidence -> green; inferable -> green, the over-abstention guard). - src/roadmap.cts: coerceTruthToString now reads `statement` first so an object-form backstop truth is surfaced, not dropped (Hyrum backward-compat for truth-readers). Workflow/agent/docs: plan-phase emits the structured marker (flat scalar, ADR-550 #1278); verify-phase + gsd-verifier add the abstain arm; new references/honest-verifier.md; FEATURES/COMMANDS document insufficient_spec; ADR-550 amended (truth-axis D4 mirror). Decisions adopted (trek-e review): insufficient_spec feeds existing human_needed with a distinguishable reason (no new VERIFIER_STATUS); changeset Changed; round-trip parity test; abstain-on-unconfirmed-backstop regression test red-first. Implementation notes (deviations from the issue's proposed file list, verified live): - frontmatter.cts needs no change — its flat parser already round-trips object-form truths. - verify.cts needs no change — it grades artifacts/key_links structurally; truths are LLM-graded at the workflow layer, so consumption lives there + the deterministic helper. - No CJS<->SDK hand-sync — the SDK seam was retired (ADR-0174); src/*.cts is sole source. Regenerated artifacts: golden-install-parity fixtures, INVENTORY-MANIFEST, size baselines. * chore(#1154): add changeset (Changed) for honest verifier User-facing changelog fragment for #1738. Typed `Changed` (not `Added`) per trek-e review condition 3 — the verify behavior shifts for backstop-bearing specs (a confident silent `passed` becomes `human_needed`), which is user-visible even though the schema marker is additive. * docs(#1154): score-formula also excludes abstained insufficient_spec truths (review nit-1) trek-e review nit: the verify-phase score sentence said PRESENT_BEHAVIOR_UNVERIFIED truths were "the only ones excluded" from verified_truths. Post-#1154 an abstained `insufficient_spec` backstop truth is also excluded (it is not ✓ VERIFIED and routes to human_needed). Behavior was already correct; this tightens the wording. Regenerated golden-install-parity fixtures + workflow-size baseline for the touched verify-phase.md. (Nit-2 — a dedicated insufficient_spec_items frontmatter list — is intentionally not taken: the current design is ADR-550-D4-conformant, the abstain cause rides as a distinguishable report reason, and adding it would exceed the approved scope.) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
e12a2abfd8 |
feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed), enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way to browse or audit parked seeds on demand. This adds a read-only listing, following the established --list → workflow pattern (per the approved scope on - gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the seeds dir, returns { count, seeds[], summary } JSON with each seed's id, slug, status, scope, trigger_when, planted, title. Optional case-insensitive status filter. User-controlled content is sanitized (sanitizeForDisplay) and every path validated (requireSafePath); read-only. Independent of audit.scanSeeds, which only returns unimplemented seeds for the milestone surface. - /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that renders the seed table. Closes #441 * chore(#441): point changeset fragment at PR #722 * test(#441): allowlist list-seeds test in prompt-injection scan The test asserts that list-seeds neutralizes injection payloads (<system>, [INST]) embedded in seed content, so the fixtures legitimately contain those patterns — same as the sibling security tests already on the allowlist. * fix(#441): use canonical /gsd:capture colon form in list-seeds workflow Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new list-seeds workflow used the hyphen form. * docs(#441): sync help full.md + INVENTORY for --list-seeds Adds the --list-seeds entry to the help reference (help/modes/full.md, per bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json. * docs(#441): add --list-seeds how-to + drop phantom statuses Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers): - USER-GUIDE.md Seeds section (how-to): extend the task to cover auditing parked seeds on demand via --list-seeds, including the status filter — kept task-oriented per Diataxis how-to mode. - CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected from the list-seeds filter vocabulary; the system only produces dormant|active|triggered (src/audit.cts scanSeeds). Reference must be factually accurate and complete. * fix(#441): guard non-scalar status frontmatter in cmdListSeeds A seed with a bare `status:` line (extractFrontmatter yields {}) or a `status: [a, b]` value (yields an array) crashed the whole audit list: `(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string. Coerce every frontmatter read through a `fmStr` helper (mirrors the existing `typeof fm.id === 'string'` guard), so a non-scalar status falls back to dormant and non-scalar scope/trigger_when/title can no longer leak a raw array/object into the JSON contract. Title is now capped symmetrically. Adds regression coverage for empty and array `status:` and non-scalar fields. Refs #441 * docs(#441): align list-seeds workflow status vocabulary The load_seeds step listed `implemented` as an example status filter, but the real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds); `implemented` has no producer. Matches the earlier CLI-TOOLS.md correction. Refs #441 * refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested in-process (review minor #1). No behavior change. Filter comparison now matches the raw lowercased status (both sides already normalized) instead of sanitizeForDisplay(status); sanitization is for output, not matching (review nit #3). * test(#441): add fast-check property coverage and count=1 boundary for list-seeds Adds tests/list-seeds.property.test.cjs with four fast-check properties over deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing (review minor #1). Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2). * chore(#441): sync runtime launcher snippet into list-seeds workflow Propagate the current _runtime-launcher.snippet.sh (with non-Claude runtime home probes) into the new list-seeds.md workflow via scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation. * test(#441): record list-seeds.md in workflow size baseline (#1074) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7c93d9e222 |
feat(#1463): add capability outdated (per-source update check); drop phantom slash-command docs
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
34bc096ec2 |
feat(#1451): wire gsd capability install/update/remove/list/disable/enable management CLI (#1457)
* feat(#1451): wire gsd capability install/update/remove/list/disable/enable CLI ADR-1244 D5/D6: the management command was built as a library (capability-lifecycle.cjs install/upgrade/remove + capability-ledger.cjs) across Phases 3-5 but never wired to a user-facing command — gsd-tools.cjs 'capability' only handled state/set. This adds the six subcommands, dispatching to the existing lifecycle/ledger: - install <spec> [--integrity] [--scope global|project] [--yes] [--shared-file <rel>]… - update [<id>|--all] [--scope] [--yes] [--shared-file] (re-resolves recorded source) - remove <id> [--purge-data] [--scope] (first-party rejected) - list [--json] (first-party + overlay, both scopes, JSON array) - disable|enable <id> (activation-state alias of capability set --off/--on) Scope→runtimeDir mapping matches capability-loader exactly (global=$GSD_HOME||home, project=project root; caps at <root>/.gsd/capabilities/<id>, ledger at <root>/.gsd-capabilities.json). Consent is non-interactive: --yes grants; without it an executable install aborts after printing the disclosure and writes nothing. Best-effort reconcile before each mutation. Tests: tests/capability-cli.test.cjs (20 behavioral, real resolver via local specs, GSD_HOME-sandboxed) — install consent/block/usage matrix, list, update round-trip, remove round-trip + first-party guard, disable/enable, unknown subcommand. Docs: docs/reference/gsd-capability-command.md reconciled to the real surface (ledger paths, --shared-file, consent model, disable mechanism, outdated marked planned); docs/COMMANDS.md gains the gsd capability entry. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1451): resolve adversarial-review findings + root-cause the --raw silent-output bug Adversarial-review (Codex) fixes: - capReadStrict passes a malformed strict_known_registries value THROUGH so the trust gate fail-closes on it (was silently downgrading to permissive) - installCapability/upgradeCapability gain an expectedId guard + first-party-id rejection (capability-lifecycle.cts): an overlay can't shadow a first-party id, and 'update <id>' can't act on a different id if the recorded source was retargeted - capability update: prints the consent disclosure, exits non-zero on --all partial failure, no longer masks the resolved id - capability remove: ledger-first ordering so an overlay is removable even if it shadows a first-party name; first-party guard only fires for ids not in the ledger - gsd-capability-command.md: disable/enable doc corrected (registry-known ids; overlay toggle not yet wired through this path) Silent-output bug (root cause, not waved off as pre-existing): - captureStdoutSyncWrites buffered fd-1 output and DISCARDED it on the throw path — any --raw command that emitted a result/error envelope then threw (to set a non-zero exit) lost ALL of stdout. Now it flushes the captured buffer before re-throwing (exit code preserved). - cmdCapabilitySet threw via process.exit() (bypassing the capture wrapper entirely); now throws ExitError so the wrapper flushes — matches the repo's no-process-exit architecture. - Regression test: capability disable <unknown> --raw must emit the JSON error envelope on stdout. Verified: capability suite 165/165, @file/json-errors/phase 183/183, lint clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1451): address adversarial-review R2 — shared-file confinement, MCP no-clobber, config fail-closed - confinedSharedFile(): realpath-confine every shared-config write/strip to the scope root (mirrors safeRmUnder), so a --shared-file whose parent is a symlink escaping the scope can't write outside it. - mcpServers shared edits: never overwrite an UNOWNED entry — a name collision with the user's (or another capability's) server is skipped, so install/remove can't silently clobber user MCP config (hooks already append; the map-keyed mcpServers path was the gap). - capReadStrict: a PRESENT-but-unparseable .planning/config.json now fails CLOSED (lockdown) instead of silently downgrading the strict_known_registries policy to permissive. - Tests: symlink-escape shared-file writes nothing outside scope; colliding user mcpServers entry preserved; unparseable config blocks an external install. capability suite 83/83, lint clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1451): address code-review — aborted-status robustness + coverage + project-scoped strict doc - install/update: handle an 'aborted' result independently of the requiresConsent flag so it can never fall through to the generic 'blocked: unknown reason' arm (aborted always means consent-needed per the lifecycle contract; latent today, hardened for future status additions). - Clarify capResolveScope comment (project scope === already-resolved cwd) and document that strict_known_registries is a PROJECT-scoped policy (read regardless of --scope; no machine-wide allowlist) in gsd-capability-command.md. - Tests: update --all over an empty ledger returns an empty result set (exit 0); a flag value that looks like another flag (--integrity --scope) is rejected, not swallowed. CLI suite 33/33, lint clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1451): FEATURES.md entry #147 + Added/Fixed changesets for the capability CLI Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1451): backfill changeset PR number → #1457 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f3c06f59df |
fix(#1326): stop emitting Codex agents/openai.yaml sidecars; clean up stale ones (#1360)
* fix(#1326): stop emitting Codex agents/openai.yaml sidecars; clean up stale ones Codex installs wrote an agents/openai.yaml sidecar under every managed gsd-* skill dir. Recent Codex builds index both SKILL.md and the sidecar, so each GSD skill appeared twice in autocomplete (canonical gsd-* name + humanized display_name). - Replace writeCodexSkillMetadataFiles / generateCodexSkillMetadataYaml with cleanupCodexSkillMetadataSidecars: Codex-only (if isCodex), removes stale managed gsd-*/agents/openai.yaml and prunes the now-empty agents/ dir. - Preserve user-owned dirs (gsd-dev-preferences), non-empty agents/ dirs, and non-gsd dirs; lstat-guard against symlinked agents/ so a delete can never escape the skills tree; fail-open per directory. - Codex relies on SKILL.md alone for /skills discovery. - Update USER-GUIDE/FEATURES docs and rewrite the #774 emission tests into cleanup tests. Scope: the sidecar duplicate only. The separate multi-root (~/.agents/skills shared-skills) duplicate facet is a distinct concern, not addressed here. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1326): add changeset for Codex sidecar cleanup Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
49bef1927e |
Merge remote-tracking branch 'origin/next' into feat/1279-fail-first-prover
# Conflicts: # docs/adr/550-spec-phase-probe-contract.md # gsd-core/workflows/verify-phase.md # tests/workflow-size-baseline.json |
||
|
|
e50ead7ad2 |
enhance(verify-phase): deterministic auto-locate of the prohibition check descriptor (#1278) (#1301)
* test(1278): RED-first descriptor parity + fail-closed guards + CHK-07 byte-stability (wave 1) - CHK-03 (RED): extend PROB-14 parity in prohibition-probe.schema.test.cjs to carry the flat check_kind/check_target/check_rule scalars through project->write->parseMustHavesBlock; the non-droppable check_kind-presence assertion is the load-bearing RED trigger (fails because projectProhibitions strips check_* on the current build). - CHK-07 (GREEN forward-guard): probe-core.test.cjs pins descriptor-less byte-stability + dispositionForProhibition fail-closed policy, with a t.todo marker forward-locking plan 01-02. - CHK-06 (RED): prohibition-enforcement.test.cjs asserts descriptorFromProjection export + fail-closed on absent/partial/unknown descriptors via the projection adapter (RED until 01-03). - No src/*.cts or .cjs edits; no new test files; lint-test-file-count clean. * feat(1278): add optional flat-scalar check descriptor fields to Prohibition interface (wave 2) - check_kind?/check_target?/check_rule? mirror CheckDescriptor.kind/target/rule (minus caller-attested failFirst, #1279) - optional so existing Prohibition consumers compile unchanged * feat(1278): project check descriptor as flat scalars in projectProhibitions (wave 2) - emit check_kind/check_target (+ check_rule only for lint-rule with a rule) when descriptor well-formed - under-specified/descriptor-less items project byte-identically (CHK-07); flat scalars ride existing parseMustHavesBlock continuation-KV path (no parser rewrite) - add CHK-02 probe-core unit cases pinning the projection - turns CHK-03 parity test GREEN; dispositionForProhibition untouched * feat(1278): descriptorFromProjection read-back adapter feeds fail-closed locate (wave 3) - Add descriptorFromProjection(projected) -> CheckDescriptor | null to src/prohibition-enforcement.cts: renames the projected flat scalars check_kind/check_target/check_rule -> {kind,target,rule?}, or null when the descriptor is absent/non-object (no check_kind key). - failFirst is NEVER sourced from the projection (stays caller-attested; #1279). - rule is set only when check_rule is a non-empty string; the adapter does NOT re-validate kind/target/rule — an under-specified descriptor reconstructs to one the EXISTING runProhibitionEnforcement LOCATE guard rejects (located:false, never green). The merged #1259 guard stays the single source of fail-closed truth. - Turns the RED CHK-06 fail-closed tests (plan 01-01) GREEN end-to-end; CHK-03 / CHK-07 stay green. CheckDescriptor type, locate guard, dispositionForProhibition, and parseMustHavesBlock are unchanged (additive +36/-0). * feat(1278): verify-phase locates prohibition check from projected descriptor (wave 3) - request.check kind/target/rule sourced from projected check_kind/check_target/check_rule via descriptorFromProjection, not verifier invention (CHK-05) - replaces the #1278 author-supplied / tracked-follow-up note with the delivered deterministic-locate behavior - preserves fail-closed routing: absent/partial descriptor -> never green, hard-gate in both modes - failFirst stays a verify-time caller attestation; #1279 bounds the remaining fail-first proof * feat(1278): spec-phase captures wired-check descriptor on test-tier resolution (wave 3) - Step 5.6 'Keep it' / verification: test path captures check_kind/check_target/check_rule, projected onto must_haves.prohibitions for verify-phase deterministic locate (CHK-04) - SOFT capture: a test-tier prohibition without a descriptor is still allowed (no hard authoring block); stays fail-closed/flagged downstream - --auto captures only an unambiguous descriptor, never fabricates a check path - failFirst NOT captured at spec-phase (verify-time attestation; #1279) - PROB-06 soft-gate + text-mode (PROB-09) behavior unchanged * chore(1278): re-baseline workflow size for grown verify-phase + spec-phase prose (wave 3) - spec-phase.md 28438 -> 30343 (+1905), verify-phase.md 35362 -> 36498 (+1136) - regenerated via npm run size:baseline (no hand-picked numbers); growth is the #1278 deterministic-locate + descriptor-capture prose - workflow-size-budget guard green (122/122) * docs(1278): ratify optional check descriptor in dated ADR-550 addendum + type:Changed changeset - Append dated 2026-06-15 ADR-550 addendum ratifying the D3 prohibition-item shape extension (optional flat-scalar check_kind/check_target/check_rule) - Document flat-scalar rationale, deterministic projection/read-back, fail-closed on partial/invalid/absent, #1279/policy out-of-scope - Add .changeset/1278-prohibition-check-descriptor.md (type: Changed) * docs(1278): document optional check descriptor in prohibition-probe reference + FEATURES - Add 'Optional wired-check descriptor (deterministic locate, #1278)' section to the prohibition-probe reference (flat-scalar keys, projection/read-back, fail-closed + backward-compat, failFirst stays attested) - Add deterministic prohibition-check descriptor source entry to FEATURES.md - No CONTEXT.md glossary change: descriptor reuses existing wired-check / verification:test vocabulary, no new glossary term introduced * fix(1278): pass packaging gates — changeset pr field + retired slash-form fix - Add required pr: 1278 to changeset (lint:changeset MISSING_PR hard requirement; plan's 'omit if unknown' was inaccurate — issue number per #1259 convention, updated to real PR number when opened) [Rule 3 - blocking] - Fix retired /gsd-spec-phase -> /gsd:spec-phase at verify-phase.md:83 (wave-3 prose; caught by slash-namespace invariant #3443/bug-2543, blocked CHK-09 full-suite-green) [Rule 1 - bug] - size:baseline + INVENTORY manifest verified in-sync post-build (no diff) * docs(1278): add check descriptor + descriptorFromProjection to CONTEXT.md prohibition glossary * fix(1278): harden descriptorFromProjection round-trip (numeric-coercion + stray-rule) per review - MD-01/LW-01: narrow projected scalars to primitives + String()-coerce, so a numeric-looking check_target (parseMustHavesBlock coerces ^\d+$ to number) reconstructs as a string and locates instead of silently un-locating; no as-string type-lie, satisfies no-base-to-string. - LW-02: attach rule only for the lint-rule kind (drop a stray node-test rule). - LW-03: document the optional check_* keys in the reference Output schema. RED->GREEN tests added in prohibition-enforcement.test.cjs. * chore(1278): set changeset pr to 1301 * test(1278): add fast-check property for the check-descriptor round-trip + fail-closed (trek-e review) RULESET.TESTS.property-based-testing: the projectProhibitions -> render -> parseMustHavesBlock -> descriptorFromProjection chain is a bijective/transformation contract. Adds 2 fc properties to tests/probe-core.property.test.cjs (no new file; ratchet stays at 2 for probe-core): - well-formed descriptors survive the round-trip across the full string domain incl. the numeric-coercion case (target/rule reconstruct as strings); - under-specified/invalid descriptors (absent / target-less / rule-less / unknown-kind) are always fail-closed (never green, flagged, unlocated). Stability is asserted at the descriptorFromProjection layer (the raw parse step is intentionally lossy for numeric scalars; the shared parser is unchanged). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
8cf4716c99 |
docs(#1279): ratify ADR-550 addendum + FEATURES/reference/verify-phase + changeset
- ADR-550 dated 2026-06-15 addendum: machine-proven fail-first REALIZED (D5d CLOSED); ratify violationFixture, GSD_PROHIB_SUBJECT, failFirst demotion - FEATURES REQ-PROHIB-07 reads shipped (machine-proven, not caller-attested) - prohibition-probe reference + verify-phase descriptor shape updated with violationFixture + convention - tag GSD_PROHIB_SUBJECT + violationFixture PROPOSED/renamable (zero live consumers) as a PR-review flag - add .changeset/1279-machine-proven-fail-first.md (type: Changed) |
||
|
|
6a9e6cb0ae |
fix(1259-01): address trek-e review — B1 (fatal/suppressed), B2 (bounded subprocess), M1/M2, minors
Maintainer CHANGES_REQUESTED (reviewed |
||
|
|
31b822b312 |
fix(1259-01): close adversarial review findings — genuine enforcement, honest fail-first scope
Adversarial pre-submission review found the injected-runCheck tests masked a non-functional real runner. Fixes: - BL-01 (false green on vacuous test): the node-test runner now parses the TAP summary and requires a NON-VACUOUS pass (>=1 test, >=1 pass, 0 fail) AND a reported test named distinctly from the file — node --test counts an empty file as one passing test, so counts alone could not catch it. - SF-01 (lint anchor never greened): the lint-rule runner now runs the project eslint as --format json and filters by ruleId, so plugin rules (local/*) load via the flat config — bare --rule cannot load a plugin. local/no-source-grep now genuinely greens (covered by a real, non-injected test). - BL-02 (tautological fail-first): the runner no longer echoes the caller's failFirst as if confirmed. failFirst is documented as caller-ATTESTED; the producer requires attestation + a genuine non-vacuous pass. Machine-proven fail-first (needs a violation fixture) is flagged as a tracked follow-up in ADR-550, the changeset, FEATURES, the reference doc, and verify-phase. - SF-02: added real-runner end-to-end tests (no injected runCheck) + pure, exported parse/filter helpers (parseNodeTestSummary, tapTestNames, eslintJsonHasRule, eslintFileResultCount) so the shipping branches are mutation-pinned. - NIT-01/02: LOCATE guard rejects empty-string rule and unknown kinds. - Hardening: spawn checks with NODE_TEST_CONTEXT/NODE_OPTIONS scrubbed so an ambient test-runner context cannot corrupt a verify-time result. - Docs reconciled to the shipped behavior (no 'confirms fail-first' overclaim). |
||
|
|
ce01e1376b |
docs(1259-01): wire verify-phase consumer + ADR-550/FEATURES/reference + changeset
- verify-phase.md: replace test-tier 'fail-closed/deferred' bullet with the check prohibition-enforcement enforcement step (locate -> fail-first -> run -> evidence -> green-or-hard-gate); update determine_status tree - ADR-550 addendum: mark D5d enforcement half LANDED (#1259); cite ADR-857 open-question §147 + D6 (core verify rail) - FEATURES §146: enforcement wording + add REQ-PROHIB-07; keep REQ-PROHIB-06 intact - references/prohibition-probe.md: test-tier enforced + hard-gates via check prohibition-enforcement - changeset (type: Changed) with the D5 '2 no-source-grep invalid cases, not 96' correction |
||
|
|
3556450b0d |
feat(spec-phase): surface zero-classification edge-probe requirements as unclassified candidates (#1110) (#1117)
Edge-probe now surfaces a zero-classification requirement (non-empty prose, no shape cue matched, no `shapes` override) as a single soft `unclassified — review manually` candidate instead of silently dropping it — the exact blind spot the probe exists to catch. Dismissible like any edge; the `shapes: []` opt-out stays silent; `TAXONOMY` (the closed 8 categories) is unchanged. Under `--auto` the candidate is left `unresolved`, never auto-`backstop` (a missing shape is not evidence an edge exists). Closes #1110 |
||
|
|
395fb519e7 |
feat(spec-phase): prohibition probe — surface "must-NOT" constraints (#644) (#1149)
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes #644. |
||
|
|
a375c4b354 |
feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) (#1201)
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) Adds an opt-in, default-resilient ADR-857 feature capability that wires MemPalace (local-first memory: MCP server + CLI) into the GSD loop: deliberate recall before discuss/plan and verbatim + temporal-KG capture at phase boundaries. Three memory modes (augment default; kg_backend and replace forward-declared). Master gate mempalace.enabled (default off); every hook onError:skip, zero gates; absent/disabled MemPalace => loop unchanged. Transport is rendered-markdown only — MemPalace runs out-of-process, no third-party code in gsd-core (ADR-857 §7). Capability: capabilities/mempalace/ (manifest + 2 fragments), skills commands/gsd/mempalace-{recall,capture}.md, agent agents/gsd-mempalace-curator.md. Registration: ns-context router, utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot install list, size baselines; regenerated capability-registry + inventory manifest. ship:post wired into ship.md (wire-on-demand). HELD on #1196: this capability also declares hooks at discuss:pre and discuss:post, which are structurally un-wireable until the host-loop conformance model covers the discuss phase (discuss-phase.md is not in HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails on exactly those two orphaned points by design — see #1196. Once #1196 lands, rebase onto next and the gate goes green with no further change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#956): backfill changeset PR number (#1201) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
827011b865 |
fix(#1098): guard generate-claude-md against clobbering hand-crafted files; redirect default to .claude/CLAUDE.md (#1118)
/gsd-new-project wrote a repo-root CLAUDE.md full of broad project docs, overwriting/diluting a hand-crafted instruction file. --force was parsed but silently dropped, and nothing guarded an existing non-GSD file. - Guard: an existing instruction file with no `<!-- GSD:<section>-start` markers (hand-crafted) is left untouched; report action:"skipped". --force (now wired through CmdGenerateClaudeMdOptions) overwrites intentionally. The marker check uses /<!-- GSD:[a-z]+-start/ so a file merely documenting GSD syntax is safe. - Redirect: the Claude-family default output is now ./.claude/CLAUDE.md (a valid auto-loaded project-memory location) instead of repo-root ./CLAUDE.md. Aligned across the handler default, config-defaults.manifest.json, buildNewProjectConfig, the config template, new-project.md, and cmdGenerateClaudeProfile; advisory read-CLAUDE.md hints in plan-phase/quick/profile-user updated. Codex still writes AGENTS.md. Closes #1098 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3e836fef0d |
feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/
Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.
Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.
Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
source-checkout-gated build fallback) instead of LLM re-derivation; the engine
capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.
Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.
* test(#550): RED — status×verification re-cut + probe-core engine specs
Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):
status: resolved | dismissed | unresolved (lifecycle, shared)
verification: explicit | backstop | null (only when resolved)
- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
to be extracted — validateResolution(r, validators), validateRequirement,
analyzeCoverage(items, resolutions?, validators), byVerification rollup,
runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
{resolved,backstop}; coverage gains byVerification.{explicit,backstop};
proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
COUNT preserved on every fixture (closed set = resolved+dismissed; doc
line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
rewritten to the two-axis model.
Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).
* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)
Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.
probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
items[] (core never assumes propose is deterministic — edge resolves via
LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
{categories, verification, requiredFieldsByVerification} (ADR-550 #5)
edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.
* chore(#550): register probe-core.cjs artifact in ledgers + inventory
New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:
- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).
probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.
* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]
trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.
Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.
* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard
Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:
- validateResolution now enforces the 'verification is null unless resolved'
invariant for EVERY status (not just resolved): a dismissed/unresolved
resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
(was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
contract; a new test locks that an all-dismissed run is NOT affirmatively covered
(byVerification is the honest gate).
Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.
* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics
Re-review #5 (trek-e) clarity edits:
- Decision 5: annotate that only contract item (a) ships on #584 (the edge
adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
parenthetical — the blessed/implemented semantics are count-preserved = the
CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
carrying the per-tier resolved-status breakdown. The old parenthetical
contradicted the shipped count.
* test(#550): cover runProbeCli structural-guard numeric-count branch
Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.
* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)
The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.
* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)
A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.
* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)
templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.
* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)
The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.
* docs(#550): add how-to for resolving edge-coverage findings (B1)
Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.
* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)
trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.
* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)
trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.
* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)
trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.
* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe
Rebased onto next (
|
||
|
|
9223f2f4c8 |
feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results (#1063)
* feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results Wire the already-reserved `phase.uat-passed` alias (subcommand `uat-passed`, mutation:false) into the phase command router with a new markdown-aware predicate that evaluates HUMAN-UAT results and reports pass only when every required check passes. Post-SDK-retirement (ADR-0174/#174) successor to the SDK-framed #70, with no SDK-specific API surface. New pure module src/uat-predicate.cts: - stripFalsePositiveContexts: frontmatter -> HTML-comment -> CommonMark-style fenced-block state machine (tracks delimiter char+length) -> blockquote, each a small composable step, so a `result: passed` inside frontmatter, a fenced/~~~ block (incl. ~~~ nested in a ``` fence), a comment, or a blockquote is never counted. - parseUatResultItems: heading-block parser, column-0-anchored same-line result; a heading with no result -> `missing` (fail-closed). - analyzeMarkdown: unterminated fence/comment detection (malformed -> blocker). - evaluateUatPassed: allowlist pass/verification semantics; passed = no blockers && >=1 check && all passing; no_uat_artifacts discriminator (no vacuous pass); optional requireVerification policy hook. Thin cmdPhaseUatPassed handler in phase.cts; router closure rejects unknown flags via makeInvalidArgs. Hardened across two Codex adversarial passes (vacuous pass, dropped failing tests, permissive verification status, nested-fence escape, cross-line result value, masked unterminated comment) — all fixed fail-closed. New unit + CLI-integration suites incl. a fast-check property test; docs, CONTEXT glossary, inventory, and changeset updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#247): backfill changeset PR number (#1063) * fix(#247): indexOf paired-scan for unterminated-comment detection CodeQL js/incomplete-multi-character-sanitization (high) flagged the `raw.replace(/<!--[\s\S]*?-->/g,'')`-then-`.includes('<!--')` detection in analyzeMarkdown as incomplete sanitization (a single regex pass can leave a residual `<!--`). Replace it with a paired left-to-right indexOf scan that contains no `.replace()` of the comment token — CodeQL-clean and strictly more correct (a closed earlier comment can never mask a later unterminated one). Behaviour unchanged; 98 predicate tests + scoped docker run green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
cd5db1f8db |
test(suites): seed security/slow/integration suites via measured retags
Renames (git mv) with all references updated (ci-test-scope RULES, windows-parity allowlist, test-file-count allowlist, docs in 6 locales): - 5 scanner tests -> *.security.test.cjs — the 'Run security tests' CI step ran zero files since the suite taxonomy landed; it is now honest. - graphify-auto-update -> *.slow.test.cjs (36s, slowest file in the suite; e2e gsd-tools spawns) — runs on full-matrix lanes and push to next. - installer-migration-install-integration -> *.integration.test.cjs (13s; an integration test by its own name). Coverage gate measured after retags: 88.55% lines (gate 70%). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3697e6768f |
fix(#853): gate manager/autonomous background dispatch by runtime (#863)
* fix(#853): gate manager/autonomous bg dispatch by runtime /gsd-manager and /gsd-autonomous --interactive dispatched Plan/Execute via Agent(run_in_background=true). On Claude Code a backgrounded agent has no Agent/Task tool, so it cannot spawn the nested subagents those pipelines need — per-plan worktree-isolated executors, the plan-checker, and the verifier. The phases reported complete but isolation and independent verification silently never ran, even with use_worktrees / plan_check / verifier enabled. Both workflows now resolve the runtime (config-get runtime, default claude) before dispatching: run plan/execute INLINE on Claude Code so the nested pipeline runs, and background-dispatch only on runtimes where a backgrounded agent can still nest. Mirrors execute-phase.md's existing Codex fail-closed precedent. Reconciles the stale unconditional background/overlap/lean-context claims elsewhere in both workflows and in the docs. Adds a content regression test pinning the gate. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#853): add changeset for runtime-gated bg dispatch Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
29c0a2f5a1 |
docs(#849): capture 1.4.0 release features across the docs base (#850)
Diataxis review of the 1.4.0 content (52 changesets, multi-runtime maturation plus native packaging and new flags) against the existing docs base found most per-feature docs already landed with their PRs. Fill the four remaining gaps, each in its Diataxis quadrant: - Reference: FEATURES.md Feature #36 (Multi-Runtime Support) updated in place with 1.4.0 additions — native skills emission (Cline/Kilo/OpenCode), new slash-command surfaces (CodeBuddy/Augment/Cursor), cross-runtime lifecycle hooks for context-headroom tracking, and the Gemini CLI extension package. - Reference: CONFIGURATION.md gains a dedicated worktree.baseRef entry (values, .claude/settings.local.json location, auto-set-on-install behaviour). - How-to: plan-a-phase.md gains an 'override planning granularity for one phase' section for the --granularity flag. - Explanation: context-engineering.md gains a 'Lifecycle hooks and context headroom' section (the why of lifecycle hooks + forked context), cross-linked from multi-agent-orchestration.md. Docs-only; documents already-shipped features, so no changeset required (docs/ is not in the changeset-lint user-facing prefixes). Closes #849 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
40d48c0508 |
feat(#815): add /gsd-update --next to install the @next RC channel (#839)
Adds an opt-in --next (alias --rc) flag to /gsd-update targeting the @next RC dist-tag (ADR #660), with a {latest,next} allowlist enforced at three layers, channel-aware version check + banner, and byte-for-byte unchanged default @latest behavior. Closes #815 |
||
|
|
cb284962bf |
feat(#789): elevate CodeBuddy — slash commands (#830)
* feat(#789): elevate CodeBuddy — emit slash commands (+ document subagent/MCP scope) Emit a CodeBuddy slash-command surface so GSD workflows appear in the '/' menu, reaching parity with other elevated runtimes. - Add convertClaudeCommandToCodebuddyCommand and register a commands/ artifact kind for the codebuddy runtime (commands/gsd-<name>.md), consistent with the Cursor (#785) and Augment (#790) commands surfaces. - Mark emitted skills user-invocable:false so the commands surface is the sole '/' entry point (no duplicate /gsd-* entries); skills stay model-invocable. CodeBuddy's SKILL.md supports this field. - Normalize $HOME/.codebuddy (bare + slash) path forms in runtime rewrites so --config-dir/local installs don't leak the default home. - Report installed commands/ count on install; uninstall prunes gsd-* commands while preserving user-owned commands. Scope: subagents (~/.codebuddy/agents/) are already emitted by the generic agents block (unchanged); no mcp.json is written (gsd ships no MCP server, and CodeBuddy's mcp.json registers only external servers). Closes #789 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#789): set changeset pr number to 830 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
74a818308e |
feat(#785): write .cursor/commands/ Cursor 1.6 slash-command surface (#805)
* feat(#785): write .cursor/commands/ as Cursor 1.6 slash-command surface Cursor 1.6 (released 2025-09-12) introduced plain-markdown slash commands in `.cursor/commands/<name>.md` — no frontmatter, invocable via `/` in the Agent input. GSD previously emitted only `~/.cursor/skills/` for Cursor. This PR wires a second artifact kind for `cursor` in `runtime-artifact-layout.cts`: `convertedCommandsKind('commands', 'gsd-', 'convertClaudeCommandToCursorCommand', configDir)`. The new kind applies the same `convertClaudeToCursorMarkdown` transforms (tool renames, brand substitution, slash-command normalisation) and then strips YAML frontmatter so the output is plain prose. Skills output is unchanged. `stageCommandsForRuntimeFlat` in `install-profiles.cts` stages each source `.md` as a flat `<stem>.md` in a temp dir; the existing `_copyStaged` commands path then prefixes and copies to `<configDir>/commands/`. `.cursor/mcp.json` is explicitly OUT OF SCOPE: GSD ships no MCP server; the `mcpServers` schema cannot be usefully populated by the installer. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#785): address review nit --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
f66c4a082c |
feat(#766): distribute gsd-core as a native Claude Code plugin (#797)
* feat(#766): distribute gsd-core as a native Claude Code plugin Add an additive .claude-plugin/plugin.json manifest plus hooks/hooks.json so gsd-core can be installed as a first-class Claude Code plugin (marketplace or zero-friction @skills-dir), with /gsd-core: namespaced commands and lifecycle management — alongside the unchanged npm/file-copy installer. - .claude-plugin/plugin.json: validated with 'claude plugin validate --strict' - hooks/hooks.json: mirrors the installer's always-on Claude hook wiring via ${CLAUDE_PLUGIN_ROOT} - package.json: ship .claude-plugin in the npm tarball - tests/issue-766-plugin-manifest.test.cjs: manifest + always-on-hook-contract drift guards - docs: install-on-your-runtime.md + FEATURES.md Closes #766 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#766): add ADR-766 + glossary entry for Claude Code Plugin Manifest Module Record the plugin manifest as the Seam projecting gsd-core's artifact surfaces onto the Claude Code plugin contract (sibling of the Runtime Artifact Layout Module, ADR-3660), with the defined kind->field mapping and the always-on hook projection rule. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7e76f1a736 |
feat(#703): add --granularity override flag to /gsd:plan-phase (#750)
* feat(#703): add --granularity override flag to /gsd:plan-phase Add a `--granularity <coarse|standard|fine>` flag to /gsd:plan-phase that overrides the configured planning granularity for a single invocation. The override is a new highest-priority tier above the existing precedence chain (granularities[phaseType] -> granularity -> planning.granularity -> 'standard') in resolveGranularityInternal; when the flag is absent, resolution is byte-for-byte unchanged. cmdInitPlanPhase now resolves with phaseType 'planning' so granularities.planning participates, and emits the resolved value in the init JSON, which the plan-phase workflow forwards to the planner prompt. Invalid values are rejected at the CLI boundary via a shared assertValidGranularityOverride helper. Closes #703 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#703): set changeset pr to 750 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
14d0238cf5 |
docs: roll up legacy release notes into a single archive (#658) (#659)
Consolidate all pre-rename release notes (the retired get-shit-done-cc / get-shit-done-redux lineage, 1.0.0 -> 1.50.0-canary.1) into one condensed, read-only archive so the legacy 1.x version numbers no longer collide with the current @opengsd/gsd-core line. - Add docs/RELEASE-NOTES-LEGACY.md: rename banner, master version-index table, condensed per-version sections (1.42.3 -> 1.0.0), and a separate pre-release & canary builds section. Stale install commands stripped. - Trim CHANGELOG.md to the current @opengsd/gsd-core line only; replace the Legacy Release History block with a pointer to the archive and drop the orphaned numbered legacy reference-link definitions. - Remove the 10 standalone docs/RELEASE-v*.md files. - Repoint docs/CANARY.md and docs/FEATURES.md links to the new archive. Closes #658 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
463cffd894 |
chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/ Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary (`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers are unaffected. Mechanical (bulk, ~90% of the diff): - `git mv get-shit-done gsd-core` - Swept path/identifier references across the repo via `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead preserves the five legitimate slug variants that are NOT the directory: get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names). - Build/manifest wiring: package.json (bin, files, coverage globs), tsconfig.build.json (outDir), ~86 .gitignore build-output entries, stryker.config.mjs, scan-ignore files, install.js path strings. - Frozen (not rewritten): CHANGELOG.md history; translated docs (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/). New logic (review here): - src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper ADR-0008 installer migration. On upgrade it walks the legacy `~/.claude/get-shit-done/` tree, classifies each file via the prior install manifest, and emits remove-managed / backup-and-remove for managed files while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked root and symlinked entries; bounds-checks every path under configDir). The framework rolls back on install failure. Emptied dirs may remain (framework has no recursive dir-removal primitive) — documented. - scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare `get-shit-done` directory token (split token to avoid self-match; case- insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines). Wired into the lint-tests CI job. - Restored scripts/lint-package-identity-drift.cjs detection regexes (the mechanical sweep had wrongly rewritten the old-name patterns it exists to detect) and marked them as intentional legacy references. - TDD tests for the migration and the guard; do.md slash-command guard regex tightened so a `/gsd-core/bin` path segment is not mistaken for a command; changeset + docs/installer-migrations.md row added. Breaking: the installed runtime path moves `~/.claude/get-shit-done/` -> `~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed files (preserving user files) on upgrade. Users with custom hooks/configs hardcoding the old path must update them. Closes #604 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unsweep pending changesets + allowlist injection-example docs CI fixes for the rename PR: - Do not sweep pending .changeset/*.md (ephemeral release-note fragments, like CHANGELOG); reverted those body edits so 5 pre-existing malformed fragments (missing type/pr) no longer enter the PR diff and trip docs-lint. Allowlisted .changeset/ in the legacy-name guard accordingly. - Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in prompt-injection-scan.sh: they contain intentional injection examples / security-model prose; the path-reference rewrites are kept. CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR; none in the new migration/guard) and are out of scope for the rename. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): resolve CodeQL alerts surfaced on this PR The rename diff touched files carrying pre-existing CodeQL findings; per the no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving them off. All behavior-preserving: - scripts/ci-test-scope.cjs: build the config-path match from string .includes() instead of a RegExp over an arg-derived value (js/regex-injection). - src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName so the table-cell escape is complete (js/incomplete-sanitization). - tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization). - tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace, keep the meaningful POSIX-class conversion (js/identity-replacement). Verified: build:lib green; the touched test files + ci-test-scope + profile-output suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization) The prior commit's fixes for two alerts were ineffective: - ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file` reaching static regex `.test(file)` calls (not the config rule). Removed ALL regex over file/t — startsWith/includes/=== string checks + an isWindowsHint helper — so there is no regex sink for the tainted value. - js/incomplete-multi-character-sanitization (3 test files): a single `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint loop (replace until stable) plus a final bare-opener strip. Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL CodeQL flags the regex PATTERNS syntactically (regex-injection on the --files arg split; incomplete-multi-character-sanitization on the <!--...--> replace), so loop fixes do not satisfy it. Made these paths regex-free: - ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/). - 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)). Behavior preserved; ci-test-scope + the 3 suites pass; guard clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unblock security base64 scan on the large rename diff The security job hit its 10m timeout: base64-scan.sh choked on the binary test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/ non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings), and the ~800-file rename diff is slow to scan regardless. - scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they can't carry base64-obfuscated *text* and feeding NUL bytes through the per-line scanner is pathologically slow. collect_files already filtered binary *extensions*; this catches binary *content* in text extensions. - .github/workflows/security-scan.yml: raise the security job timeout 10m->30m to accommodate very large diffs (the scan itself is unchanged). Verified locally: scan skips the fixture, 0 "ignored null byte" warnings, 0 findings, exit 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): sweep get-shit-done refs introduced by merging next The branch was updated with next (#614/#384/#618 etc.), which reference the get-shit-done/ dir (still named that on next). Swept the stale references in the merged files to gsd-core so the rename stays consistent and lint:legacy-name passes: - commands/gsd/discuss-phase.md (runtime-launcher shim paths) - src/core.cts (getAgentsDir layout comments) - tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib) Verified: guard 0 violations; build green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant The #614 runtime-launcher shim added to discuss-phase.md references `${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it mis-read the directory path as a dangling `/gsd-core` command ref (same class as the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path segments are not treated as slash-command references. Verified locally on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22 image) full suite: 0 failures - bug-3683 + bug-2954 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI) CI intermittently failed state.test's gsd-tools subprocess with "findProjectRoot is not a function" (flip-flopping across legs; not reproducible on mac full suite, gsd-test linux full suite, test:unit, or state.test x8). findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs); binding it via destructure at module-load can be undefined under a load-ordering edge. Resolve it lazily at call time via a small wrapper so the lookup happens after core.cjs is fully initialized. Verified green on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22) full suite: 0 failures - state.test.cjs: 106/106; gsd-tools loads cleanly. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): allowlist verification-patterns.md placeholder examples in secret scan The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var examples (illustrative Stripe test-key / database-URL / API-key placeholders) — not real credentials. Added it to .secretscanignore with the strict annotation, mirroring the existing gsd-core/workflows/plan-phase.md exception. Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next exits 0 with 0 findings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3bb2f8f1c5 |
docs: rebrand to GSD Core and restructure docs with Diataxis (#605)
* chore: wire docs/agents config into AGENTS.md Agent skills section
Add the `## Agent skills` discovery block pointing the engineering
skills at the existing docs/agents/{issue-tracker,triage-labels,domain}.md
files (issue tracker, triage label mapping, single-context domain docs).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: rebrand to GSD Core and restructure docs with Diataxis
Reorganise the root README and docs/ around the Diataxis framework
(tutorials, how-to guides, reference, explanation), add new how-to
guides and schema references (STATE.md / CONTEXT.md / PLAN.md /
planning artifacts), and cross-link the whole set. Update the lone
legacy gsd-build reference to open-gsd; keep internal get-shit-done/
filesystem paths unchanged (directory rename tracked separately in
open-gsd/gsd-core#604). Regenerate the ja-JP, ko-KR, pt-BR and zh-CN
localised trees to mirror the new structure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: backfill changeset PR number (#605)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
a11ba2dfcb |
feat(#68): per-phase granularity overrides (granularities.<phaseType>) (#595)
Closes #68. Per-phase-type granularity overrides via granularities.<phaseType>, mirroring models.<phaseType>. Includes maintainer-authorized sdk-seam reference cleanup. |
||
|
|
9ffe45a7c3 |
feat(#163): tighten gsd-roadmapper granularity defaults to reduce thin-phase fragmentation (#591)
* feat(#163): tighten gsd-roadmapper granularity defaults to reduce thin-phase fragmentation Tighten the Granularity Calibration buckets in gsd-roadmapper (Coarse 3-5->2-4, Standard 5-8->4-6, Fine 8-12->6-10) and append inline Key guidance naming the thin-phase failure pattern (single requirement / internal-quality goal / task-shaped success criteria) with instruction to fold into a neighbor rather than create a standalone phase. Implements the maintainer-approved proposal verbatim. Update the canonical English docs that hardcoded the old phase-count numbers: docs/CONFIGURATION.md and docs/FEATURES.md. Translated docs are community-maintained and are not updated per-PR (CONTRIBUTING.md language policy). Prompt/doc text only; no code, format, or downstream-consumer changes. Agent size-budget and skills-awareness tests pass; full suite green. Closes #163 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#163): add Changed changeset for roadmapper granularity tightening Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#163): lock tightened gsd-roadmapper granularity buckets source-text-is-the-product test asserting the Granularity Calibration table holds the tightened ranges (Coarse 2-4, Standard 4-6, Fine 6-10), that no row maps to an old bucket, and that the Key paragraph carries the thin-phase folding guidance. Would fail if the values regress. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
de2f73d21a |
enhancement(#34): add Antigravity CLI (agy) as a peer reviewer in /gsd-review
Closes #34 Squash-merged via admin override — all CI green (28/28 checks), branch protection review gate bypassed with maintainer authorization. |
||
|
|
79002a00cb |
chore(#518): rename npm package + bin to @opengsd/gsd-core (#519)
* chore: rename npm package + bin to @opengsd/gsd-core (functional) - package.json: name @opengsd/get-shit-done-redux → @opengsd/gsd-core, bin key get-shit-done-redux → gsd-core, repository/homepage/bugs URLs - package-lock.json: regenerated (npm install --package-lock-only) - tests/**, scripts/**, bin/**, .github/**, agents/**, commands/**, get-shit-done/bin/**, get-shit-done/workflows/**: applied the 4-rule replacement (scoped npm ref, GitHub repo path, bin/clone invocations) per #505 single-source refactor Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: sweep live references to @opengsd/gsd-core Update all live documentation (README.md + translations, docs/**, CONTRIBUTING.md, VERSIONING.md, SECURITY.md, CONTEXT.md, docs/CANARY.md) to reflect the renamed package and repository. Rules applied: - @opengsd/get-shit-done-redux → @opengsd/gsd-core (scoped npm name) - open-gsd/get-shit-done-redux → open-gsd/gsd-core (GitHub repo) - GSD-redux/get-shit-done-redux → open-gsd/gsd-core (stale badge org) - bare bin/clone refs → gsd-core CHANGELOG.md, docs/adr/**, docs/RELEASE-*.md, docs/research/**, and .changeset/** are preserved byte-identical. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix: add negative lookbehind to slash-command regex in bug-2954 test The extractSlashReferences regex matched /gsd-core inside npm package URLs (@opengsd/gsd-core), producing a false /gsd:core command reference. Adding a negative lookbehind (?<![a-z]) excludes matches preceded by a letter, so only standalone /gsd-<cmd> and /gsd:<cmd> tokens are found. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#518): add changeset for package rename Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#518): update package-identity expectations to the renamed coordinates The rebase regenerated the seam to @opengsd/gsd-core (bin gsd-core, repo open-gsd/gsd-core). The #498 seam tests assert deriveIdentity against the REAL package.json, so their expected literals must follow the rename. The drift-lint unit test is left as-is — its SEAM is a self-consistent fixture and its stale-literal detection cases would shift if altered; the live-repo scan in it already passes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |