b9f51836e644e71df07930ac36c12b8ea340f49d
433 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ad3b9ec486 |
chore(#1671): fragmentize plan-phase.md and repair flag forwarding to the init bundle — Phase 6.2 (#3019)
* chore(#2993): fragmentize plan-phase.md onto the fragment model Epic #1671 Phase 6.2. plan-phase.md is the largest workflow in the repo and carried zero markers; it was deferred out of the Phase 3 pilot for two reasons, both now dead. The 36-byte PRE_PHASE6 headroom was never the blocker it looked like — fragmentizing is net-negative on host source, so the trim is what creates the room. The --mvp interleaving was resolved by measurement in #2992 and no sub-line mechanism is built. - widen WHEN_VOCABULARY 14 -> 19 via a second coordinated ADR-1671 amendment: flag:--ingest, flag:--prd, flag:--research-phase, flag:--reviews, state:chunked-mode - state:chunked-mode is `--chunked` OR config workflow.plan_chunked, and that disjunction is resolved in the FACT, never in the grammar, so a compound condition never becomes an operator - parse the new flags on the plan-phase route; extract six gated bodies to gsd-core/workflows/plan-phase/steps/ behind manifest-gated stubs - prd-express-path.md was already extracted but read unconditionally; its wrapper is now gated, so the existing extraction finally pays off plan-phase.md 94,483 -> 87,575 bytes (cap 94,519): headroom goes from 36 bytes to 6,944. Also closes a surfaced docs gap: five real plan-phase flags (--chunked, --skip-ui, --bounce, --skip-bounce, --granularity) were documented in neither the argument-hint nor help. Making --chunked load-bearing without fixing its siblings would leave the defect class half-open. Refs #2993 * fix(#2993): forward flags to the init bundle so section gating actually fires Blocker found by the correctness review, confirmed directly, and missed by both the isolated reviewer and every test in this branch. Neither workflow forwarded its flags to the init CLI: plan-phase.md:71 INIT=$(gsd_run query init.plan-phase "$PHASE" $GRAN_PARAM) execute-phase.md:84 INIT=$(gsd_run query init.execute-phase "${PHASE_ARG}") So every flag: atom was permanently false in production and its section permanently excluded. For plan-phase that made the PRD express path UNREACHABLE — a regression, since it was an unconditional read before. For execute-phase this is PRE-EXISTING: #2932 shipped `flag:--wave` gating that has never once been true, so `--wave` silently dropped its own wave-filtering guidance. Fixed here under the no-defer rule. Why every test missed it: they drive the init CLI directly with flags, which works. Production goes through the workflow's bash line, which did not pass them — the exact "assert against the shape production uses" trap this branch's own test matrix warns about. - parse and forward --prd/--ingest/--research-phase/--reviews/--chunked (plan-phase) and --wave (execute-phase), using the anchored regex idiom the neighbouring GRAN_PARAM line already uses - add a regression guard DERIVED FROM THE MANIFEST: for every flag:--X section, the owning workflow's init line must forward --X. It fails against the pre-fix files and covers any future atom, rather than spot-checking today's six. Verified through the workflow shape, not the CLI shape: `3 --prd spec.md` now yields ["prd-express-gate"] (was []), `2 --wave 2` yields ["partial-wave"] (was []). Refs #2993 * test(#2993): acknowledge the execute-phase ripple and regenerate install-tree fixtures Remote matrix was red with 46 unique failures, identical on both lanes. Both causes are mechanical consequences of changing shipped workflow content, and neither is visible to any local gate. - emitted-attribution: execute-phase.md grew 163 bytes from the WAVE_PARAM forwarding fix and was unacknowledged, while the ack fragment named plan-phase.md, which SHRANK and therefore needed no ack at all — a stale entry is itself a failure. The reason now names the real ripple. The entry had to merge into the existing 2930 fragment: the ack linter does unconditional cross-fragment duplicate-key detection with no spent/live exception, so a second fragment declaring execute-phase.md collides even when the first is already merged and inert. Resolved per the linter's own guidance and that file's precedent of appending successive ripple reasons to one entry. - golden-install-tree: tests/fixtures/install-tree/*.json are committed and deliberately excluded from the ADR-2719 attribution cutover, so they must be regenerated when shipped tree content changes. Regenerated after build:lib per the ordering landmine. 19 runtimes each gained exactly the six new plan-phase step files; zero paths removed, which is the absolute failure shape those fixtures exist to catch. Refs #2993 * fix(#2993): restore the launcher preamble in an extracted step and follow moved content in its drift guards Second red run: 26 unique failures, identical on both lanes, in two classes. RUNTIME BUG (runtime-launcher-parity, 7 failures) — chunked-planning-mode.md calls gsd_run but carried no canonical launcher preamble, which is what DEFINES gsd_run(). On any non-Claude runtime that step would fail outright. The preamble is now copied verbatim from the canonical source of truth, gsd-core/workflows/_runtime-launcher.snippet.sh, and the fence dedented to column 0 to match the prd-express-path.md sibling (a list-continuation indent breaks the byte-equal preamble match). prd-express-path.md already had a correct one. This is the same defect #2932 hit when it extracted steps; the parity test caught a real bug, not a stale assertion. DRIFT GUARDS (plan-phase-drift-guard, issue-2762-plan-reviews-chunked, skill-frontmatter-contract) — these assert plan-phase.md contains content this branch moved into step files. Retargeted at where the content now lives, with the asserted property unchanged; the ALL-RUNTIMES label COUNT test now reads host + every step file so the count is preserved across the split rather than reduced. Each retargeted guard was verified to still fail when its step file is stripped, so none was weakened into vacuity. No emitted-drift ack was needed: currentSizes() enumerates gsd-core/workflows/*.md non-recursively, so files under plan-phase/steps/ are never in the size ratchet's scope. Refs #2993 * chore(#2993): backfill changeset pr number to 3019 --------- Co-authored-by: sim <sim@local> |
||
|
|
f1af47766a |
chore(#1671): widen the when= grammar and key the section manifest per workflow — Phase 6.1 (#3013)
* chore(#2992): widen the when= grammar and key the section manifest per workflow Epic #1671 Phase 6.1. Two blockers stopped the fragment model reaching any file beyond execute-phase.md: the when= vocabulary was frozen at 4 atoms (3 execute-phase-specific), and the section manifest was single-workflow by construction with 'execute-phase' hardcoded into buildSectionManifestField. - widen WHEN_VOCABULARY 4 -> 14 via a coordinated ADR-1671 amendment; the grammar stays CLOSED (one atom, no operators, negation or nesting) and WHEN_PREDICATES stays a hand-written literal map, never deriving a predicate from its atom string - InvocationFacts gains flags: ReadonlySet<string> plus three computed state booleans; add the missing reverse vocabulary/predicate parity guard - key the manifest artifact per workflow; a stale flat {sections:[...]} artifact now fails shape validation instead of being misattributed - wire the field into six init entry points and parse the flags each needs An atom ships only with both a real consuming section and a fact the init seam actually computes. Six surveyed atoms are withheld because their workflows have no dedicated init entry point; an atom without a computed fact evaluates false forever and silently disables its own section. Fixes a defect found while wiring: parseNamedArgs always materializes a boolean flag key, so folding its false into the absent sentinel is required or every flag reads as present and gating is silently always-on. Also resolves ADR-1671:194 by measurement: --mvp stays unmarkable, because its interleaved sites are always-run flag resolution and a ~340 byte block that already delegates lazily. Refs #2992 * fix(#2992): treat any falsy option value as an absent flag and reject unsafe manifest read paths Findings from two orthogonal reviews (Claude /code-review + an isolated adversarial pass); both independently reproduced the first one. - MAJOR: the flags-builder treated only `undefined` as absent, but parseNamedArgs yields `null` for an absent value-flag and `false` for an absent boolean-flag, so `--granularity` read as present on every plan-phase invocation. Fixed at the root: a flag is present iff its option value is truthy. The six per-handler `|| undefined` folds are now redundant and removed, which also closes the duplicate-translation and missed-onboard-handler findings. - MAJOR: state:needs-codebase-map had zero coverage. Added unit, property and real-CLI integration tests. - MINOR: reject absolute, UNC/drive and `..`-traversing `read` paths in the manifest, degrading the whole load to null like every other shape violation. Verified: `/etc/passwd` previously reached section_manifest.read. - MINOR: corrected a stale "4 to 20" doc comment; the vocabulary is 14. Refs #2992 * test(#2992): update the generator suite for the per-workflow manifest shape The remote matrix went red with 5 unique failures, identical on linux-node22 and linux-node24, all in tests/gen-section-manifest.test.cjs. Re-keying the artifact to {workflows:{...}} left this suite asserting the old flat {sections:[...]} shape; nothing else in the tree still does. - three tests read manifest.sections.length, now undefined; retargeted at workflows.<name> with their original intent preserved (a fenced or loop-host marker still asserts NO section is produced, not merely a changed count) - the stale-manifest test wrote its fixture in the OLD shape, so it tripped shape validation and stopped exercising staleness at all. Its fixture is now valid-but-mismatched so FAIL_STALE is genuinely reached again. - added the coverage that exposed: a pre-6.1 flat artifact must report FAIL_MANIFEST_MALFORMED_SHAPE. That is the real upgrade path for an installed tree and nothing covered it. Refs #2992 * chore(#2992): backfill changeset pr number to 3013 --------- Co-authored-by: sim <sim@local> |
||
|
|
4df6d884b3 |
fix(#2641): treat absent capture_artifacts as enabled (schema default) (#2982)
* test(#2641): add regression for mempalace-capture gate default inversion * fix(#2641): treat absent capture_artifacts as enabled (schema default) The gate used `capture_artifacts !== true` which treated absent (undefined) as disabled — inverted from the capability registry's declared default of true. Changed to `capture_artifacts === false` (disabled only on explicit false), matching the sibling gsd-mempalace-recall skill's correct pattern. * chore(#2641): add changeset fragment * chore(#2641): backfill changeset PR number 2982 --------- Co-authored-by: sim <sim@local> |
||
|
|
7372d99a26 |
enhance(#2800): derive reviewer flag lists and gate reviewer lane docs across locales (#2882)
* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales The reviewer lane roster was hand-enumerated across five documentation surfaces and three workflow files that had drifted apart: --kimi-code was missing from all four translated COMMANDS.md mirrors, --coderabbit from every workflow forwarding list, and --antigravity from FEATURES.md. Adds checkReviewerDocsParity, a second pure gate deliberately separate from checkReviewerLaneParity so a stale doc cannot make the runtime checker look red. Workflows now derive their flag lists from a new review-lane flags query instead of hand-enumerating them, which also retires the unanchored grep that matched --agy inside --antigravity. Documents the previously absent reviewer body and hostBehaviors field in the capability manifest reference. Closes #2800 Closes #2781 Closes #2272 * fix(#2800): key the docs parity table arm on first-cell position Review found the flag arm was file-scoped, so the forwarding row that lists every flag in its third cell satisfied it on its own. Deleting a lane's own reviewer-table row -- the #2781 regression this gate exists to prevent -- therefore passed undetected. Arm 4 keys on the FIRST table cell, which separates a lane row from the forwarding row structurally and in every locale. Regression test included. * fix(#2800): shape-filter the flags subcommand output All three consumers read review-lane flags through an unquoted command substitution so the output word-splits into loop items. Phase 2 admits third-party overlay lanes, so an overlay flag containing whitespace would inject a second loop item and one containing a glob would expand against the cwd. Emit only well-formed flags so neither reaches the shell. * fix(#2800): remove the regex length ceiling and count only prose mentions Review found two real defects in the docs parity gate. The never-throws contract was false: building a RegExp from a declared flag or section title throws SyntaxError past ~100k chars, and Phase 2 admits overlay lanes whose declared strings are untrusted in length. Every one of these matches is literal, so String.includes replaces the regex outright, which also deletes escapeLiteral and the llama.cpp escaping it existed for. Arm 1 was context-blind: a flag mentioned only inside a fenced example or a commented-out row counted as documented. Both are stripped before matching. Also advertises all 13 lane flags in the argument-hint and corrects a stale eleven-lane count in the slug grammar note. * test(#2800): repoint the convergence suite off deleted workflow text The derived flag loop deleted the literal per-flag grep lines four tests matched on. Two of those failed loudly. The behavioral and property tests failed SILENTLY instead: their end marker no longer resolved, so the parse block extracted empty and both passed vacuously, and the property test's gsd_run stub had a no-op default that hid it. All now share one extractor and execute the real deployed block through a gsd_run shim backed by the actual binary. The whitelist assertions become an anti-parity check: re-adding a hand-written flag list must fail. Also repairs two vacuous cases in the docs parity suite. The unreadable-doc test called its own mock rather than the reader, and the integration test bounded nothing, so a doc losing its marker would have been silently skipped and still passed green. * fix(#2800): run the derived flag loop after the launcher preamble The remote matrix caught a real runtime bug, not a test artifact. In autonomous.md and plan-review-convergence.md the launcher preamble that defines gsd_run lives in a separate, LATER bash fence than the derived loop. Each fence is its own shell, so gsd_run was undefined where the loop ran: the command substitution yielded nothing and zero reviewer flags would have been forwarded. Worse than the drift this epic fixes, and silent. The whole CONVERGENCE_ARGS construction moves as one unit, because the --max-cycles append sits between the loop and the preamble and would otherwise have run against an uninitialized variable and then been dropped by the relocated initializer. Also documents all 13 lane flags in help/modes/full.md, which the repo gates bidirectionally against each command's argument-hint. * test(#2800): repoint the two converge suites off deleted flag literals Both asserted workflow.includes('--codex') against the hand-enumerated list the derived loop removed. They now assert the derivation itself, keep --all and --text (convergence controls, still literal), and add an anti-parity guard so re-adding a hardcoded list fails. The lost pass-through proof is replaced with a real one: every flag the tests used to hardcode is asserted present in the actual roster emitted by the binary, which is the property the old assertion was protecting. * test(#2800): acknowledge the workflow byte growth from the derived flag loop * chore(#2800): backfill changeset pr number to 2882 * fix(#2800): strip HTML comments to a fixed point in the parity gate CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick construction smuggles a commented-out row past the gate and it counts as documented. Not an injection risk here since nothing is rendered, but it is the exact false pass this helper exists to prevent. Strips to a fixed point, then treats any surviving opener as unterminated so the multi-line branch closes it on a later line. Terminates because every pass strictly shortens the string. * test(#2800): pin the comment-smuggling regression with a real reproducer The obvious fixture for this class does not reproduce it: <!--<!---->--> leaves a dangling --> rather than a live <!--, and is caught either way, so it would have passed with and without the fix. The join-trick construction (<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely regresses on the single-pass strip and is what the test now uses. --------- Co-authored-by: Test <test@example.com> |
||
|
|
c5e0371775 |
feat(#1951): reversibility tagging — gate one-way-door decisions (#2471)
* test(#1951): add failing-first tests for reversibility tagging Red phase for issue #1951 (reversibility tagging: classify decisions by undo cost, gate one-way doors behind a checkpoint:decision). Tests assert, per the issue's acceptance criteria: - discuss-phase CONTEXT.md template records a **Reversibility:** field with a rationale on captured decisions, and states it is optional - gsd-planner @-references planner-reversibility.md and stays under the 49152-char agent cap (LARGE_CAP, tests/agent-size-budget.test.cjs) - a one-way rating inserts a checkpoint:decision before the dependent task; reversible inserts none; costly is flagged but never blocks - the taxonomy defaults to reversible when unsure (checkpoint-fatigue guard) and inserting a checkpoint implies autonomous: false - docs/reference/plan-md.md documents <reversibility> as optional with all three ratings - --no-reversibility-gates parses to REVERSIBILITY_GATES=false, is injected into the planner prompt, and is advertised in the command argument-hint and help full mode (argument-hint parity) - the override suppresses the gate but still persists the rating - cmdVerifyPlanStructure accepts every rating and the absent case (additive-validator guarantee, behavioral via runGsdTools) - parity: thinking-models-planning.md #4 adopts the canonical three-level taxonomy and the binary REVERSIBLE/IRREVERSIBLE vocabulary is gone - no content loss from the planner extraction made to fit under the cap Prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately — regression guards proving the validator already accepts unknown optional tags. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1951): reversibility tagging — gate one-way-door decisions Classify planning decisions by what undoing them would cost, and give a one-way door a human beat before the agent walks through it (issue #1951, The Pragmatic Programmer Topic 15 'Reversibility'; Bezos's one-way/two-way door framing). Acceptance criteria met: - discuss-phase records an optional reversibility rating with a rationale on <decisions> entries in the phase CONTEXT.md template. Unrated decisions are treated as reversible, so existing phases are unaffected. - a one-way rating makes gsd-planner insert a checkpoint:decision before the task that implements the decision, reusing the existing checkpoint mechanism -- no new checkpoint machinery. - reversible ratings trigger no checkpoint; costly ratings are flagged in the plan but never block. - the rating persists on the task as the optional <reversibility rating=> element. cmdVerifyPlanStructure accepts every rating and the absent case; the structural validator does not reject unknown optional tags. - --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses checkpoint insertion for intentionally-unattended runs while still recording ratings -- the override changes what stops the run, not what the plan remembers. Single taxonomy, not two: references/thinking-models-planning.md #4 already shipped a binary REVERSIBLE/IRREVERSIBLE classification and is loaded by both gsd-planner and gsd-plan-checker. It is rewritten onto the canonical three-level vocabulary and now points at planner-reversibility.md as the taxonomy owner, with a parity test that fails if the surfaces diverge (DEFECT.GENERATIVE-FIX-DIVERGENCE). agents/gsd-planner.md sat 47 chars under the 49152 LARGE_CAP, so the checkpoint DO/DON'T guidance was relocated verbatim into planner-antipatterns.md -- already @-referenced from the same section for the same topic, so the planner still loads it and nothing was dropped. A test guards the relocation against content loss. Files: gsd-core/references/planner-reversibility.md (NEW, canonical taxonomy + emission rules + anti-patterns), gsd-planner.md, plan-phase workflow/command/help (flag wiring + parity), plan-md.md schema, discuss-phase context template, CONTEXT.md glossary, INVENTORY + manifest, size baselines, install goldens, plugin skills regen, changeset. Closes #1951 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1951): address orthogonal review findings Two isolated reviewers (correctness + security), neither of which authored the change. Every finding fixed: Security — the rationale is untrusted input (ADR-1577). It originates in conversation and flows CONTEXT.md -> planner -> PLAN.md -> executor, each hop an LLM reading the previous hop's output, with no validation on the path. planner-reversibility.md and the discuss-phase template now state it is data and never instructions, and name the </reversibility> early-termination hazard explicitly -- a rationale that closes its own element injects sibling structure the executor reads as real tasks. Four tests guard it. Correctness 1 — nothing machine-enforced the feature's own promise: a task rated one-way with no preceding checkpoint:decision validated as fully clean, so a planner error silently reopened the gap this feature exists to close. cmdVerifyPlanStructure now warns on an ungated one-way rating. A warning, not an error: <reversibility> stays additive and the plan stays valid. Four tests cover ungated (warns), gated (silent), still-valid, and reversible/costly never flagged. Correctness 2 — pass-always test. The --no-reversibility-gates parse test substring-matched the whole workflow file, and plan-phase.md prose mentions both tokens in one sentence, so it passed with the bash conditional deleted: it was testing the documentation, not the parser. Now scoped to the fenced bash blocks and matched as one physical line, with a negative control confirming prose alone cannot satisfy it. Correctness 3 — costly had no itemized emission rule, only one-way did, so two agents could diverge on whether to tag costly at all. Correctness 4 — template convention break: the example ratings were bare while every sibling field uses [...] to signal substitution, inviting an LLM to copy one-way/costly forward as boilerplate. Now bracketed. Correctness 5 — latent false-green: .includes('reversible') also matches inside irreversible/irreversibility, which appear in anti-pattern prose, so a surface that dropped the real taxonomy entry would still pass. Now word-boundary matched. ADR-857 phase-6 ceiling — the first gsd-test run caught plan-phase.md 1216 bytes over its frozen 94519 ceiling (it had 49 bytes of headroom on next). The ceiling may only rise for privileged host machinery, and reversibility gating is optional-feature logic, so the wiring was slimmed to its minimum and the explanatory prose moved to the reference files the planner already loads. plan-phase.md is now 94400 bytes -- 119 under the ceiling and 70 bytes SMALLER than on next, so the host loop shrank while gaining the feature, which is what phase 6 ratchets toward. The tracer contract (tests/tracer-bullet.test.cjs) is unchanged. Lint — fixed an unnecessary non-null assertion in verify.cts and a CRLF-fragile bare \n regex in the new test (DEFECT.WINDOWS-CRLF-TEST- PORTABILITY, the #1658/#1668/#2206/#2449/#2450 class). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture must carry the common task elements The gated-one-way fixture built a checkpoint:decision task from the abbreviated skeleton in gsd-planner.md, which shows only the checkpoint-specific elements (<decision>/<context>/<resume-signal>). cmdVerifyPlanStructure requires <name> and <action> on EVERY task regardless of type, so the fixture failed validation for reasons that had nothing to do with reversibility: errors: ["Task missing <name> element", "Task 'unnamed' missing <action>"] Caught by gsd-test on 14d14a39 (2 failures, both this fixture). The canonical shape is in tests/verify.test.cjs:266 — a checkpoint task carries <name>/<files>/<action>/<verify> like any other. Fixture corrected to match. Verified behaviorally against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Not a product defect: the validator's every-task contract is intentional and pre-existing, and docs/reference/plan-md.md scopes its required-element list to type=auto/tracer only because those are the elements a planner must author, not because checkpoints are exempt from <name>. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1951): backfill changeset pr number to 2471 * fix(#1951): CodeQL incomplete-sanitization + prompt-injection scan collision Both CI failures were real defects in code this PR added, not false positives. CodeQL js/incomplete-sanitization (high), reversibility-tagging.test.cjs:46 — the namesRating helper built its regex with `rating.replace(/[-]/g, '\\-')`, which escapes the hyphen but not backslash, so the escape was incomplete. It was also unnecessary: `-` carries no special meaning outside a character class. Replaced with a complete metacharacter escape (backslash included). Word-boundary behavior verified unchanged across all three ratings — notably that "irreversible" prose still does not satisfy a "reversible" match, which is the false-green this helper exists to prevent. Prompt injection scan — the checkpoint fixture used the human-verification child element inside <verify>. That tag name is a fake-instruction-boundary pattern in scripts/prompt-injection-scan.sh, and the scan runs over changed files, so copying the shape from tests/verify.test.cjs (unflagged only because it is not in this diff) tripped the gate. Switched to the documented plain-prose <verify> form. The first attempt at that fix failed the same gate a second time: the comment explaining the collision quoted the offending tag literally. The comment now names it in prose instead — the scanner does not care whether a match is code or commentary, which is the whole point of the DEFECT.PROMPT-INJECTION-SCAN-COLLISION note in CLAUDE.md. Verified locally before push: scan reports 0 findings across 57 changed files, eslint clean, and both fixtures still validate as designed (gated one-way silent, ungated one-way warns, neither errors). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): record measured cost and halve gsd-tools spawns The Windows shard 1/3 job timeout was traced to the sharding layer, not to this PR's assertions — see #2472. Two contributing factors were this file's own, and are fixed here. 1. tests/test-timings.json had no entry for reversibility-tagging.test.cjs, so scripts/run-tests.cjs weighted it at the table's median fallback (~315ms) for LPT chunk packing. It actually measures 5595ms — an 18x under-weight. Recorded the measured value from the green gsd-test run (max across the node22/node24 lanes, per gen-test-timings.cjs's convention). Only this one entry: a full regen churns 634 entries of run-to-run drift, and the table is explicitly advisory and un-gated, so a 637-line diff does not belong in a feature PR. 2. Each verifyPlan() spawns gsd-tools, which dominates this file's cost. Spawns cut from 9 to 6 with no coverage lost: - the ungated-one-way warning and its stays-valid assertion now share one plan instead of building the same plan twice; - the reversible/costly never-flagged-as-ungated test was strictly subsumed by the additive suite, which already runs those two ratings ungated and asserts no /reversibilit/ warning at all — and the gate warning's text contains both "reversibility" and "one-way", so the broader assertion catches it. It only re-spawned gsd-tools twice to prove the same thing. Both are symptom fixes. The shard imbalance itself (19/11/10 minutes against a 20-minute cap, from a cost-blind round-robin partition that also reshuffles downstream files whenever one is inserted) is tracked in #2472. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture adopts the #2444 type-branched contract Surfaced by rebasing onto next, which gained #2444 (branch plan-structure validation on task type=checkpoint:*) while this PR was in review. cmdVerifyPlanStructure no longer applies one required-element set to every task. A checkpoint:decision now requires <name> + <resume-signal> + <decision> + <options>, and is exempt from the <action>/<verify>/<done>/ <files> set that auto and tracer tasks carry. The gated-one-way fixture predated that split and failed on the new requirement: errors: ["Task 'Task 0: Confirm the on-disk format' missing <options>"] Fixture rewritten to mirror the checkpoint:decision contract exactly — real <options> with two <option> children — rather than padding it with fields checkpoints no longer need. That also drops the plain-prose <verify> the earlier revision carried purely to dodge the prompt-injection scan; a checkpoint task has no <verify> requirement at all, so the workaround is moot. Verified against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
903182fed3 |
fix(#2414): mempalace-capture rooms example must be dicts with name key (#2464)
* fix(#2414): mempalace-capture rooms example must be dicts with name key
The skill's Step 3 'Add the drawer (verbatim)' example wrote a flat list
of bare strings under rooms: in the embedded mempalace.yaml. mempalace's
miner (detect_room + _mine_impl) indexes room["name"] — a bare-string
list crashes the first 'mempalace mine' invocation with
TypeError: string indices must be integers, not 'str'
Following the documented example verbatim and running the capture crashed
every time, before any file was routed. The bug shipped in #2220's fix
(commit
|
||
|
|
352876ff0c |
fix(#2315): respect review.default_reviewers in bare convergence invocation (#2451)
* test(#2315): regression test for review.default_reviewers precedence A bare /gsd-plan-review-convergence invocation (no reviewer flags) is supposed to let users configure a persistent reviewer lineup via review.default_reviewers and just run the loop. Instead, the orchestrator's argument parser silently discards that configuration and forces --codex on every no-flag invocation — with no warning that the configured reviewers were ignored. This commit adds a regression test that fails against the pre-fix workflow (the buggy unconditional --codex fallback is still present at this commit) and passes after the fix lands: - Structural: the workflow must NOT contain an unconditional 'if [ -z "$REVIEWER_FLAGS" ]; then REVIEWER_FLAGS="--codex"; fi' line before the workflow.plan_review_convergence config gate. - Structural: the workflow must query review.default_reviewers AFTER the config gate and document that empty REVIEWER_FLAGS lets gsd-review apply the default. - Behavioral: matrix across {configured, unset, empty-array, explicit-flag} invoking the actual deployed parse + resolution blocks with a stubbed gsd_run. Also updates two existing tests whose assertions the fix makes stale: - #2293 behavioral: endMarker was the buggy unconditional fallback line; the bare invocation assertion was 'run("5") === "--codex"'. Both flip post-fix (endMarker is now the last --all grep line; bare invocation returns empty from the parse block, default applied later in step 1.5). - command-default-claim: the pre-fix command documented '--codex (default if no reviewer specified)' which was the user-facing mirror of the bug. The assertion now requires the command to document the review.default_reviewers precedence. * fix(#2315): respect review.default_reviewers in bare convergence invocation Root cause: plan-review-convergence.md step 1 (Parse and Normalize Arguments) contained an unconditional fallback that set REVIEWER_FLAGS=\"--codex\" whenever no explicit reviewer flag was supplied. This value was then interpolated verbatim into the gsd-review args, so gsd-review saw --codex as an explicit flag (precedence rule 1) and never reached rule 3 (review.default_reviewers). The same path silently dropped any configured review.reviewer_instances (instances participate ONLY via review.default_reviewers per ADR-1517). Fix: - Remove the unconditional --codex fallback from step 1. - Add a config-gated resolution in step 1.5 (after CONVERGENCE_ENABLED check) that queries review.default_reviewers and either leaves REVIEWER_FLAGS empty (letting gsd-review apply its own rule-3 default) or falls back to --codex when no default is configured — preserving the pre-fix default for unconfigured users (AC3). - Replace the banner {REVIEWER_FLAGS} token with {REVIEWER_DISPLAY} so the startup banner reflects what will actually run (AC4), not a hardcoded value. The fix upholds the documented precedence contract (ADR-0011, ADR-0015) that the bug was actively violating. Explicit-flag invocations (--gemini, --all, etc.) are unaffected (AC5). References: #2315; ADR-0011 (review.default_reviewers precedence); ADR-0015 (autonomous cross-AI convergence); ADR-1517 (reviewer instances). * chore(#2315): bump plan-review-convergence.md baseline + changeset - Bump plan-review-convergence.md size baseline 23713 → 25536 (the new step-1.5 default-resolution block). - Add .changeset/plucky-yaks-roar.md documenting the user-visible change. * fix(#2315): restore /gsd: colon syntax + skip behavioral test when jq missing Two follow-ups to the #2315 fix discovered by gsd-test: 1. The fix commit accidentally regressed the slash-command namespace in the disabled-feature exit message: /gsd:plan-review-convergence (correct, from PR #3452) became /gsd-plan-review-convergence (retired dash syntax). The slash-command-namespace invariant test caught this. Restored the colon form. 2. The behavioral test exercises the deployed reviewer-resolution block, which pipes through jq. jq is a documented production dependency (review.md:244 "install jq if missing") and is present in every production deployment, but is NOT on PATH in the gsd-test linux-node{22,24} containers (same constraint as tests/opencode-review-reconstruction. property.test.cjs). Without jq, the printf|jq pipeline fails silently, the ||echo 0 fallback yields DEFAULT_REVIEWERS_COUNT=0, and the resolution falls through to the --codex branch — producing a false negative. Added a jqAvailable guard at module load (matching the existing pattern) and skip the behavioral test when jq is absent. The structural tests (no bash execution) still run and validate the fix. * chore(#2315): regenerate golden-install-parity fixtures Source changes to plan-review-convergence.md (workflow + skill mirror + command doc) changed the install-tree hashes. Regenerated via 'npm run gen:golden' after rebuilding gsd-core/bin/lib/install-engine.cjs ('npm run build:lib') — the on-disk lib was stale relative to src/install-engine.cts (isSymlinkedDestOptIn) and blocked fixture gen. * test(#2315): address review findings — strengthen structural tests + property test Code-review + security-review (isolated subagent passes) surfaced Low/Nit findings; this commit addresses the test-side findings: - Structural test 1 ("unconditional --codex one-liner") now asserts the buggy line is absent EVERYWHERE, not just before the config gate. The earlier assertion allowed a maintainer to re-add the line after the gate (passing the structural test) while the bug would still bite at runtime before the gate runs. - Structural test 4 ("banner uses REVIEWER_DISPLAY") now asserts the LITERAL banner placeholder "Reviewers: {REVIEWER_DISPLAY}" and forbids "Reviewers: {REVIEWER_FLAGS}". The earlier workflow.includes( "REVIEWER_DISPLAY") was satisfied by a comment mention. - New structural test for command/skill content parity: both files must document the review.default_reviewers precedence on the --codex flag (catches a manual edit to one that the gen:plugin-skills mirror misses). - Behavioral test stub now passes default_reviewers via env var ($GSD_TEST_DEFAULT_REVIEWERS) instead of inline-interpolating into a bash single-quoted string. Removes the (currently-safe) fragility where a future test input containing a single quote would close the bash quote and execute as bash under execFileSync. - New property test (fast-check, numRuns=25) for the JSON-classification contract: non-empty arrays of slugs -> empty REVIEWER_FLAGS; empty array / scalar JSON / malformed JSON -> --codex fallback. Locks the parser contract per CLAUDE.md mandate. * fix(#2315): address review findings — defensive jq-missing warning + banner cleanup Code-review surfaced two Low-severity workflow-side findings: - Defensive jq-missing warning: if jq is not on PATH in production (it is a documented dependency per review.md:244, but the dependency can be absent in degraded environments), the printf|jq pipeline fails silently to "0" and a user with review.default_reviewers configured gets --codex with no indication their configured default was unreadable. Added a command -v jq guard at the top of the resolution that falls back to --codex AND emits a stderr warning explaining the reason. This makes the failure diagnosable instead of silently reproducing the #2315 override. - Banner leading-space cleanup: REVIEWER_FLAGS accumulates with a leading space ("$REVIEWER_FLAGS --gemini" from ""), so the explicit-flag branch of REVIEWER_DISPLAY="$REVIEWER_FLAGS" rendered "Reviewers: --gemini" (double space). Pre-existing but worth fixing alongside the AC4 banner work. Strip one leading space with the ${VAR# } parameter expansion in the explicit-flag branch only (the configured-default and --codex branches already produce clean strings). * chore(#2315): bump plan-review-convergence.md baseline + regen golden fixtures The defensive jq-missing warning grew plan-review-convergence.md (25536 -> 26285 bytes). Bumps the per-file baseline snapshot and regenerates the golden-install-parity fixtures for the resulting install-tree hash changes. * chore(#2315): backfill pr:2451 in .changeset/plucky-yaks-roar.md |
||
|
|
a7d83dc234 |
fix(#2390): warn on goal-shaped phase.add titles, correct auto-detect docs (#2425)
* fix(#2390): phase.add title warning + auto-detect doc fix phase.add now returns a `warning` field when a description reads as goal-shaped (>80 chars and/or multi-sentence) rather than title-shaped, instead of silently writing the whole paragraph verbatim as the `### Phase N:` header. The CLI still creates the phase as-is (the strict two-layer slash-vs-CLI interface is unchanged); the warning just surfaces the gap. Also clarifies six doc sites (command argument hints, workflow detection steps, and how-to/reference docs) that described the phase-number argument as "auto-detecting" the next unplanned phase -- that detection is an orchestrating-workflow/LLM step reading ROADMAP.md (concretely: `query roadmap.analyze`'s `next_phase` field), not a `gsd-tools.cjs` CLI feature. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2390): regenerate fixtures + lint gate-prep * fix(#2390): repair failing tests after gate verification * chore(#2390): add changeset (#2425) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ada79bee97 |
fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md (#2338)
* fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md Step 4 rewrote the `## Current Milestone` heading in the shared root PROJECT.md unconditionally. references/workstream-flag.md marks PROJECT.md `# Shared`, and per-workstream milestone state already lives in the workstream's own STATE.md / ROADMAP.md / REQUIREMENTS.md. With parallel milestones — the sanctioned design — whichever workstream ran new-milestone last silently won the shared heading. Step 4 is now skipped when a workstream is active; step 6 no longer stages PROJECT.md in that mode (cmdCommit returns nothing_to_commit rather than failing when a staged path is unchanged). Also fixes a second defect found while diagnosing this, same root cause (the workflow was workstream-unaware): step 1 parsed only --reset-phase-numbers and the milestone name, so GSD_WS was never set — yet ${GSD_WS} was interpolated at the routing lines. It always expanded to empty, so `/gsd:new-milestone --ws x` suggested `/gsd:discuss-phase [N]` with the workstream scope silently dropped, violating the routing-propagation contract. Step 1 now parses --ws using the established idiom from verify-work.md. Guard is keyed on GSD_WS, not $GSD_WORKSTREAM: the runtime launcher does not export the latter and it is only priority 2 of 5 in resolution, so it would miss the --ws flag case that is the actual repro. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2308): regenerate install goldens for the new-milestone workflow change gsd-core/workflows/ ships as an installed artifact, so new-milestone.md's content hash is pinned in all 18 runtime golden fixtures. Only that hash changed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2308): address review — inert step-6 guard, dropped Evolution repair, tautological tests Independent review found the first pass was partly cosmetic: 1. The step-6 `if [ -n "$GSD_WS" ]` branch was INERT. GSD_WS is assigned in step 1's shell and each step's bash block runs in its own shell — this file already proves it, since step 5 round-trips OUTGOING_MILESTONE through a file for exactly that reason (#2288). The guard read an unset variable, always took the flat branch, and staged PROJECT.md anyway. Rather than re-deriving GSD_WS in step 6, the branch is removed entirely: step 4 Part A's guard is what protects the shared heading, so post-guard the only content PROJECT.md can carry is Part B's idempotent Evolution backfill — which must be staged, not stranded. A regression test now asserts no cross-step GSD_WS branch returns. 2. Skipping ALL of step 4 also dropped the `## Evolution` structural repair — a shared, idempotent backfill that is not workstream state. A pre-Evolution project running only `--ws` would never get the section that transition and complete-milestone expect. Step 4 is now split: Part A (milestone-state write) is workstream-guarded; Part B (Evolution) always runs. 3. The tests were tautological prose-pinning — including one asserting a comment mentions "#2308". The step-6 test asserted the guard's TEXT was present, so it passed on the inert guard it existed to catch. Replaced with executable tests that extract the step-1 and step-6 fences and run them under bash with stubbed gsd_run, asserting real parse and --files behavior. 4. --ws is now stripped from the milestone name (step 1 previously left "--ws search" in the remaining text), and documented in argument-hint, help/modes/full.md, and docs/COMMANDS.md. 5. Changeset no longer overstates: --ws reaches the prose guard and routing hints only, not the SDK calls (state.milestone-switch/phases.clear/init.new-milestone still take no ${GSD_WS} — out of scope here). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * chore(#2308): regenerate SKILL.md, goldens, and size baseline for the argument-hint change skills/gsd-new-milestone/SKILL.md is generated from commands/gsd/new-milestone.md, so documenting --ws in the argument-hint made it stale (caught by lint:ci's gen-plugin-skills --check). Regenerated it plus the install goldens and workflow size baseline, since commands/, skills/, and gsd-core/workflows/ all ship. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2308): backfill PR number 2338 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1bb724048a |
fix(#2293): recognize --agy/--antigravity in plan-review-convergence whitelist (#2325)
* fix(#2293): recognize --agy/--antigravity in plan-review-convergence whitelist The convergence reviewer-flag whitelist predated the 1.7.0 Antigravity CLI adapter and silently dropped --agy/--antigravity, so convergence fell back to --codex only and the working adapter was unreachable (worse after Gemini CLI's upstream shutdown). Add both flags to the workflow grep whitelist, the command argument-hint + flag docs, and the regenerated SKILL.md; they pass through to /gsd-review unchanged. --gemini behavior is untouched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2293): backfill PR number 2325 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
315d94f6d4 |
feat(#1945): tracer-first planning default + executor feedback gate (#2294)
* feat(#1945): tracer-first planning default + executor feedback gate Make "thin end-to-end slice first, verify, then expand" the default planning + execution discipline instead of the opt-in --mvp mode. - gsd-planner: first-class `type="tracer"` task; every plan LEADS with one production-quality end-to-end tracer slice by default; --no-tracer restores horizontal layers; --mvp/--tdd compose on top. - gsd-executor + execute-plan: post-tracer feedback gate — autonomous runs halt-on-fail before expansion, interactive runs emit checkpoint:human-verify after the tracer. - --no-tracer flag wired through plan-phase workflow/command/help/skill. - CONTEXT.md glossary defines tracer bullet vs prototype; docs + references reconciled. - tests/tracer-bullet.test.cjs: prose-contract + behavioral (verify plan-structure accepts tracer) coverage. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1945): backfill changeset PR number to 2294 --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d8af61be44 |
fix(#2220): replace invalid mempalace mine --room with detect_room() staging (#2260)
* fix(#2220): replace invalid 'mine --room' with detect_room() staging approach mempalace mine has no --room flag (only search does) — verified against MemPalace 3.5.0 official docs (mempalaceofficial.com/reference/cli.html). The headless capture path used --room, causing every headless/no-MCP run to fail with 'unrecognized arguments: --room' and silently skip capture. Fix: replace the flag with a staging-based approach that uses detect_room()'s documented folder-path match — stage the artifact under a room-named subfolder with a mempalace.yaml room taxonomy, then run 'mempalace mine <stage> --wing'. Docs sources cited in-file: - CLI reference: https://mempalaceofficial.com/reference/cli.html - Mining guide: https://mempalaceofficial.com/guide/mining.html - Config guide: https://mempalaceofficial.com/guide/configuration.html Changes: - skills/gsd-mempalace-capture/SKILL.md: headless staging instructions - commands/gsd/mempalace-capture.md: same - capabilities/mempalace/fragments/capture-problems.md: reference staging - .gitignore: exclude .planning/.mempalace-stage/ - tests/mempalace-capture-headless-invocation.test.cjs: regression test - Golden install parity fixtures + workflow-size baseline regenerated * docs: backfill changeset PR number (#2260) * fix: regenerate golden fixtures after next merge |
||
|
|
64c4a105f6 |
fix(#2116): use resolvable paths in surface.md require + correct package name
Four require() examples in commands/gsd/surface.md used bare 'gsd-core/...' specifiers that Node cannot resolve (wrong package name + runtime-mirror layout off module path). Now derives the path from runtimeConfigDir. Also fixes reinstall hint from 'npm i -g gsd-core' to 'npm i -g @opengsd/gsd-core'. |
||
|
|
90ebe4ea73 | no-mistakes(review): Include manager in onboard installs | ||
|
|
2c878c966f | no-mistakes(document): Sync onboarding docs | ||
|
|
1797207280 |
fix(onboard): route onboard under ns-project and drop from core profile
Integrate the brownfield /gsd:onboard skill into the skill subsystems so the full CI suite passes: - Route onboard under commands/gsd/ns-project.md (requires + routing row) so it nests as gsd-ns-project/skills/onboard on nested-layout runtimes instead of leaking as a 7th top-level skill dir (fixes install-nested-layout + issue-69). - Remove onboard from PROFILES.core (src/install-profiles.cts) so the frozen main-loop core stays at 8 skills; onboard remains in standard/full. - Add the TEXT_MODE plain-text fallback note to gsd-core/workflows/onboard.md for non-Claude runtimes (#2012). - Allowlist onboard.md as a user-invocable skill (enh-2790 ratchet). - Regenerate docs/INVENTORY-MANIFEST.json, golden-install-parity fixtures, and the workflow size baseline to match. Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> |
||
|
|
53f81ab5a9 | no-mistakes(lint): Fix command dependency lint | ||
|
|
a3cca0704d | no-mistakes(document): Sync onboard documentation | ||
|
|
896c2740d3 | feat(#1990): add onboard command for brownfield setup | ||
|
|
ed79902509 |
feat(#2007): implement mempalace memory_mode kg_backend and replace routing (#2010)
Wire the two forward-declared mempalace.memory_mode modes so they actually route recall/capture instead of silently behaving as `augment`: - kg_backend: the palace temporal KG is the primary knowledge-graph source; native .planning/graphs/ is the fallback. Non-KG drawer recall stays additive. - replace: recall resolves through the palace as the source of truth; native artifacts are the fallback. Every mode stays onError:skip and default-resilient — an unreachable palace degrades to native memory and GSD keeps writing .planning/graphs/, so no memory is lost. Cross-mode .planning/graphs/ migration remains a documented open question (PRD/ADR §17), out of scope here. Surfaces updated (instruction-only contract): recall/capture commands (+ generated skills), discuss/wave fragments, curator agent, capability.json schema. Docs: how-to Step 3, CONFIGURATION, FEATURES, CONTEXT glossary. Regenerated capability-registry, golden install-parity fixtures (mempalace hashes only), agent-size-baseline. Added a routing-contract + cross-surface parity test. Incidental (folded per no-defer rule): removed pre-existing unused imports (spawnSync in capability-registry.test.cjs; fs in issue-498-package-identity.test.cjs) that eslint flagged in/alongside the touched files. Closes #2007 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e5ef323b15 |
feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command
Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.
* feat: add /gsd-start smart-entry command
State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.
- src/smart-entry.cts: deterministic situation classifier (no-project,
paused, blocked, verify-failed, needs-first-phase, planning, executing,
verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
dispatcher presenting an AskUserQuestion menu (with --text fallback for
non-Claude runtimes) and dispatching to existing commands. Falls back
to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
(markdown-layer invariants + every emitted command resolves to a real
slash command).
Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.
* refactor: rename smart-entry command to /gsd:next
Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.
All affected tests (188) pass; lint:ci clean.
* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)
Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.
- detectSignals now reads total_phases + percent from nested progress{}
first, then scalar fm, then body; current_phase falls back to the
body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
body Phase field) covering verify-pending + executing situations.
Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.
* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)
Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.
Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
layer is down — read .planning/STATE.md directly with the Read tool
and synthesize a minimal situation + actions menu so /gsd:next stays
useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
/gsd:progress as before.
Matches the direct-read resilience the live agent already did by hand.
* docs: add gsd-next skill surface
* chore: trigger no-mistakes validation
* no-mistakes(review): Fix smart-entry phase ordering
* no-mistakes(review): Fix decimal smart-entry phase ordering
* no-mistakes(test): Fix smart-entry next test contracts
* no-mistakes(document): Docs synced for smart entry
* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: shorten next.md description and update golden install parity fixtures
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: trigger no-mistakes validation
* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files
Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.
* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: regenerate golden install parity fixtures for /gsd:next
Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.
* Fix smart-entry verify-failed phase scoping and empty resolve shim step
Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.
* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: recapture all 16 golden fixtures with updated smart-entry.md hash
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: regenerate fixtures + inventory manifest after rebase onto next
Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next
Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).
Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)
The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs
Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo
Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
recommended action, 1-4 unique-id /gsd:* actions (previously the
one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout
Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:
- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
consolidation file doing dozens of real installs — 41s even on a fast Mac
(much worse on the slow Windows I/O path), plus an install-heavy cluster.
Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.
Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
b0d5ca3379 |
feat(#1517): support custom reviewer instances for /gsd:review (#1766)
* feat(#1517): support custom reviewer instances for /gsd:review Add a bounded review.reviewer_instances config surface so one model-capable adapter (e.g. opencode) can run as several independent reviewer identities in a single /gsd:review pass. Instances participate only via review.default_reviewers, expand before built-in slugs, are available iff their cli is detected, and a non-matching entry is a hard error (typo must be loud). >=2 same-cli instances emit a shared-adapter caveat in REVIEWS.md. Default path with no instances is byte-for-byte unchanged. Single-source instance->cli resolution lives in resolveReviewerSelection / normalizeReviewerInstances (parity-locked in tests/review-reviewer-instances.test.cjs). cli validated against KNOWN_REVIEWER_SLUGS only (never arbitrary shell); model/agent opaque, never shell-interpolated. Closes #1517 * chore(#1517): backfill changeset pr:1766 --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
e12a2abfd8 |
feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed), enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way to browse or audit parked seeds on demand. This adds a read-only listing, following the established --list → workflow pattern (per the approved scope on - gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the seeds dir, returns { count, seeds[], summary } JSON with each seed's id, slug, status, scope, trigger_when, planted, title. Optional case-insensitive status filter. User-controlled content is sanitized (sanitizeForDisplay) and every path validated (requireSafePath); read-only. Independent of audit.scanSeeds, which only returns unimplemented seeds for the milestone surface. - /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that renders the seed table. Closes #441 * chore(#441): point changeset fragment at PR #722 * test(#441): allowlist list-seeds test in prompt-injection scan The test asserts that list-seeds neutralizes injection payloads (<system>, [INST]) embedded in seed content, so the fixtures legitimately contain those patterns — same as the sibling security tests already on the allowlist. * fix(#441): use canonical /gsd:capture colon form in list-seeds workflow Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new list-seeds workflow used the hyphen form. * docs(#441): sync help full.md + INVENTORY for --list-seeds Adds the --list-seeds entry to the help reference (help/modes/full.md, per bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json. * docs(#441): add --list-seeds how-to + drop phantom statuses Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers): - USER-GUIDE.md Seeds section (how-to): extend the task to cover auditing parked seeds on demand via --list-seeds, including the status filter — kept task-oriented per Diataxis how-to mode. - CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected from the list-seeds filter vocabulary; the system only produces dormant|active|triggered (src/audit.cts scanSeeds). Reference must be factually accurate and complete. * fix(#441): guard non-scalar status frontmatter in cmdListSeeds A seed with a bare `status:` line (extractFrontmatter yields {}) or a `status: [a, b]` value (yields an array) crashed the whole audit list: `(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string. Coerce every frontmatter read through a `fmStr` helper (mirrors the existing `typeof fm.id === 'string'` guard), so a non-scalar status falls back to dormant and non-scalar scope/trigger_when/title can no longer leak a raw array/object into the JSON contract. Title is now capped symmetrically. Adds regression coverage for empty and array `status:` and non-scalar fields. Refs #441 * docs(#441): align list-seeds workflow status vocabulary The load_seeds step listed `implemented` as an example status filter, but the real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds); `implemented` has no producer. Matches the earlier CLI-TOOLS.md correction. Refs #441 * refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested in-process (review minor #1). No behavior change. Filter comparison now matches the raw lowercased status (both sides already normalized) instead of sanitizeForDisplay(status); sanitization is for output, not matching (review nit #3). * test(#441): add fast-check property coverage and count=1 boundary for list-seeds Adds tests/list-seeds.property.test.cjs with four fast-check properties over deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing (review minor #1). Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2). * chore(#441): sync runtime launcher snippet into list-seeds workflow Propagate the current _runtime-launcher.snippet.sh (with non-Claude runtime home probes) into the new list-seeds.md workflow via scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation. * test(#441): record list-seeds.md in workflow size baseline (#1074) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
a0dbf8bbdf | fix(#1319): use portable Claude skill effort (#1352) | ||
|
|
0b3a2e5f9c |
feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15) (#1237)
* feat(#1190): wire --converge primary surface into /gsd:progress --next (ADR-15) ADR-15 designates /gsd-progress --next --auto --converge as the PRIMARY plan-convergence surface, but only the secondary surface (autonomous.md) was wired. next.md now parses --converge/--cross-ai into a plan strategy, gates on workflow.plan_review_convergence, forwards reviewer flags + --max-cycles, and routes Route-3 planning through /gsd:plan-review-convergence (mirroring autonomous.md); --auto chaining preserves converge mode. Adds argument-hint + help/full.md + COMMANDS.md + how-to parity and a structural regression test. Closes #1190 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1190): add changeset for progress --converge surface (#1237) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1190): keep progress --converge docs skill-dep-clean + regen workflow size baseline CI surfaced two ripples from the ADR-15 workflow edits: (1) lint-skill-deps + profile-closure flagged /gsd:plan-phase and /gsd:plan-review-convergence SlashCommand tokens in progress.md's flag docs as undeclared deps — reworded to plain prose since progress.md only advertises the flag (the real invocation lives in next.md); (2) the per-file workflow size baseline needed regenerating after the next.md/help edits. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a375c4b354 |
feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) (#1201)
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in) Adds an opt-in, default-resilient ADR-857 feature capability that wires MemPalace (local-first memory: MCP server + CLI) into the GSD loop: deliberate recall before discuss/plan and verbatim + temporal-KG capture at phase boundaries. Three memory modes (augment default; kg_backend and replace forward-declared). Master gate mempalace.enabled (default off); every hook onError:skip, zero gates; absent/disabled MemPalace => loop unchanged. Transport is rendered-markdown only — MemPalace runs out-of-process, no third-party code in gsd-core (ADR-857 §7). Capability: capabilities/mempalace/ (manifest + 2 fragments), skills commands/gsd/mempalace-{recall,capture}.md, agent agents/gsd-mempalace-curator.md. Registration: ns-context router, utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot install list, size baselines; regenerated capability-registry + inventory manifest. ship:post wired into ship.md (wire-on-demand). HELD on #1196: this capability also declares hooks at discuss:pre and discuss:post, which are structurally un-wireable until the host-loop conformance model covers the discuss phase (discuss-phase.md is not in HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails on exactly those two orphaned points by design — see #1196. Once #1196 lands, rebase onto next and the gate goes green with no further change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#956): backfill changeset PR number (#1201) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
092340d18a |
fix(#711): wire autonomous convergence flag (#729)
* fix(#711): wire autonomous convergence flag * Update wise-ibex-tumble.md --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
f61b97276e |
fix(#724): block convergence on actionable review findings (#728)
* fix(#724): block convergence on actionable review findings * merge: integrate clean next (#936 inline) onto author tip + re-apply cursor fixes (Mode field, REVIEWS.md extraction) and review hardening (#724) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
dc139b38e0 |
feat(#949): install/surface consume derived profiles/clusters (ADR-857 phase 4c) (#954)
Make install + surface read the registry's derived profileMembership/ capabilityClusters so a capability's tier drives what installs + surfaces. resolveProfile (when given the registry) unions capability skills for the profiles its tier implies before the requires: closure; resolveSurface merges capabilityClusters into the cluster map. bin/install.js, /gsd:surface, and the capability-state resolver all thread the registry. Shipped as a proven no-op: the UI capability is reconciled to tier:full (its skills were full-only in the hand-authored profiles), so it contributes only to the full profile (already the '*' sentinel) and core/standard are unchanged. Equivalence tests prove resolveProfile/resolveSurface/listSurface/staging/ capability-state are identical with vs without the registry; the core-alias staging path is verified equivalent (empty manifest → raw PROFILES.core). Closes #949 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e4e8a9fcb8 |
fix(#936): run plan-phase inline in convergence; guard against nested spawner wraps (#939)
Both sites in plan-review-convergence.md that wrapped gsd-plan-phase in Agent() (initial planning + replan loop) are now bare Skill() calls at depth 0. On Claude Code, a depth-1 Agent has no Agent tool so wrapped plan-phase could never spawn gsd-planner/gsd-plan-checker — the replan loop silently produced no revised plan when HIGHs were found. Running plan-phase inline from the depth-0 orchestrator (which retains the Agent tool) restores the full sub-agent chain. A full audit of all workflow files confirmed these two sites were the only instances of the anti-pattern (no other workflow wraps a spawner orchestrator in Agent() without a RUNTIME carve-out). Added structural guard test bug-936-no-nested-spawner-wrap.test.cjs that dynamically derives the spawner set (workflows containing subagent_type=) and asserts no workflow wraps a spawner inside Agent() without a RUNTIME != claude carve-out — prevents silent regression. Test passes on fixed code, would fail on pre-fix code at the two de-wrapped sites. Also applied two low-severity prose nits flagged in review: - commands/gsd/plan-review-convergence.md: orchestrator role updated to describe inline plan-phase + Agent for review (was generic "spawn Agents") - gsd-core/workflows/plan-review-convergence.md success_criteria: narrowed "Each Agent fully completes" to the review Agent (plan-phase is inline now) Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
b866b95296 |
fix(#921,#922): orchestrators must not fork; plan-phase Agent gate is attempt-based (#926)
`context: fork` strips the `Agent` tool from a subagent's environment. Spawning orchestrators (`/gsd-autonomous`, `/gsd-execute-phase`, `/gsd-plan-phase`) depend on `Agent` to dispatch sub-agents; running them forked silently disables the core capability they exist to provide (#921). Remove `context: fork` from all three command frontmatter files. `effort: xhigh` (introduced by #769) is preserved. The `<runtime_compatibility>` Agent-availability guard added by #913 was checking whether `Agent` was present *before* attempting the call. On runtimes where the tool list is dynamically resolved this produced false-negative aborts in sessions that have the tool (#922). Replace the introspection-based pattern with an attempt-based gate: always attempt the `Agent()` call; stop only if a real tool-unavailable error is returned. This preserves #853's backgrounded-session close-off and #913's intent of preventing inline role-collapse, while eliminating false negatives. Tests updated: enh-769-context-fork-effort.install.test.cjs asserts the three orchestrators lack `context: fork` and that the converter still passes the field through for non-orchestrator commands; plan-phase-drift- guard.test.cjs adds four assertions for the attempt-based gate language; workflow-size-budget unchanged (budgets not exceeded). Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a90c654745 |
fix(#891): probe non-Claude runtime homes in gsd-tools launcher shim detection (#911)
- Updated `gsd-core/workflows/_runtime-launcher.snippet.sh` with 15 new `elif` arms covering Hermes, Cursor, Codex, Gemini, Copilot, Windsurf, Augment, Trae, Qwen, CodeBuddy, Cline, Grok, Antigravity, OpenCode, and Kilo (respecting each runtime's env-var override with a `$HOME`-relative default). - Re-ran `scripts/sync-runtime-launcher.cjs` to propagate the expanded snippet into all `gsd-core/workflows/*.md` files (~70 files). - Manually applied the same snippet update to `commands/gsd/import.md` (1 occurrence) and `commands/gsd/graphify.md` (5 occurrences) — these are not covered by the sync script. - Updated `tests/workflow-size-budget.test.cjs` budgets (XL/LARGE/DEFAULT + discuss-phase target) to account for the ~3 KB snippet expansion. - Added regression test `tests/bug-891-non-claude-runtime-home-fallback.test.cjs` (6 tests: structural probe presence, ordering, behavioral HERMES_HOME env-var + default-path stubs, resolution order, and workflow propagation). - Added `.changeset/891-launcher-non-claude-runtime-homes.md` (Fixed). Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0a11d361ca |
feat(#69): nest concrete skills under namespace routers at install (#883)
Emit the 6 gsd-ns-* routers as the only top-level skill bundles and nest the ~61 concrete skills under <router>/skills/<name>/SKILL.md on runtimes with confirmed non-recursive skill loaders (claude global, cline, qwen, hermes, augment, trae, antigravity). Router bodies rewrite their routing tables from Skill-tool dispatch to a Read skills/<name>/SKILL.md pattern. Recursive/unconfirmed loaders (cursor, codex, copilot, windsurf, codebuddy, opencode, kilo) keep the flat layout. Completes the v1.40 namespace architecture (#2792) so the eager skill listing drops to ~6 entries. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
40d48c0508 |
feat(#815): add /gsd-update --next to install the @next RC channel (#839)
Adds an opt-in --next (alias --rc) flag to /gsd-update targeting the @next RC dist-tag (ADR #660), with a {latest,next} allowlist enforced at three layers, channel-aware version check + banner, and byte-for-byte unchanged default @latest behavior. Closes #815 |
||
|
|
41f91b2e88 |
feat(#769): adopt context:fork + effort on heavy workflow skills (#820)
* feat(#769): emit context:fork + effort: frontmatter on heavy workflow skills Add `context: fork` and `effort: xhigh` to the three heaviest workflow commands (plan-phase, execute-phase, autonomous) and `effort: low` to the two quick-status commands (progress, stats). On Claude Code, `context: fork` runs the skill in an isolated subagent context window so the main session's context budget is protected. `effort: xhigh` / `effort: low` signal the appropriate token-budget tier to the runtime. Both fields are silently ignored by runtimes that do not recognise them (Gemini, Codex, Cursor, etc.) — no behaviour change outside Claude Code. Update convertClaudeCommandToClaudeSkill in bin/install.js to preserve `context:` and `effort:` when rewriting source command files to SKILL.md for a Claude global install. Add install-suite tests to assert the fields are present in both source commands and the installed SKILL.md output. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#769): tighten regex assertions + add execute/plan-phase effort coverage Fix low-severity adversarial finding: tighten test regex patterns from `\s*` to `[ \t]*` so they cannot match across newlines (CRLF parity). Add missing effort: xhigh assertions for gsd-execute-phase and gsd-plan-phase SKILL.md install output to complete the black-box coverage gap. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
42b74100f1 |
feat(#159): auto-use existing RESEARCH.md in /gsd:plan-phase --research-phase (#718)
* feat(#159): auto-use existing RESEARCH.md in /gsd:plan-phase --research-phase When RESEARCH.md already exists in research-only mode and neither --research nor --view is passed, emit a one-line notice and exit cleanly instead of prompting update/view/skip. This matches the promptless auto-use of standard /gsd:plan-phase <N> (§5.1) and removes the §5.0/§5.1 inconsistency, making AI-agent and CLI invocations non-interactive in the common case. The two explicit-flag escape hatches (--research to refresh, --view to print) cover any deviation. Closes #159 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#159): point changeset fragment at PR #718 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#159): tighten research-phase reference register (Diataxis) Make the 'no modifier' research-phase entries descriptive rather than imperative and drop the trailing 'pass --research/--view' clauses, which duplicated the adjacent --research/--view documentation. Reference docs describe; the recovery flags are documented in their own entries. The emitted runtime notice in the workflow keeps naming the flags (in-band recovery), unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0e259a589c |
fix(#705): route hardcoded $HOME gsd-tools invocations in agents/commands through gsd_run (#707)
* fix(#705): route hardcoded $HOME gsd-tools invocations in agents/commands through gsd_run The hardcoded `node "$HOME/.claude/gsd-core/bin/gsd-tools.cjs" <cmd>` form (fixed for workflows in #621/#637) survived in agent/command surfaces and misresolves on global/shim-only installs. Route every agent-executed invocation through the resolved `gsd_run` launcher in gsd-phase-researcher, gsd-planner (load_graph_context extracted to a shared reference to stay under the planner size budget), import, and graphify. Add a regression guard over agents/ + commands/ + gsd-core/references/ bash blocks. User-facing display messages and docs are intentionally left untouched. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#705): use repo changeset fragment format (type: Fixed, pr: 707) The hand-written fragment used the standard changesets package format (package: bump) which lacks the type:/pr: frontmatter the repo's docs-required lint consumes (fail_malformed_fragment / missing_type). Regenerated via scripts/changeset/new.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
de7d6add15 |
fix(#622): make graph.html copy optional in /gsd-graphify build chain (#634)
* fix(#622): make graph.html copy optional in /gsd-graphify build chain The Step 3 shell chain in commands/gsd/graphify.md copied graphify-out/graph.html with an unconditional `cp` linked by `&&`. When a graph exceeds graphify's HTML viz node limit (default 5000), `graphify update .` deliberately omits graph.html, so the `cp` failed with "cannot stat" and aborted the chain — skipping the GRAPH_REPORT.md copy, the diff-snapshot write, and the status report, and reporting BUILD FAILED even though the graph data was rebuilt successfully. Guard the graph.html copy with `{ [ -f graphify-out/graph.html ] && cp ... || true; }`, mirroring the already-correct tolerant copy in hooks/lib/gsd-graphify-rebuild.sh. A skipped optional HTML artifact no longer aborts the chain. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#622): add changeset for graph.html optional-copy fix Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#622): CRLF-tolerant fence regex + allowlist the new graphify test Two CI guards flagged the new regression test: - windows-test-parity (fenceRegexLiteralNewline): the block-extraction regex matched ```bash with a literal \n, which breaks on Windows CRLF checkouts. Use ```bash\r?\n per the guard's sanctioned fix. - lint-test-file-count: the new file is a 5th test in the grapify bucket (cap 2, grandfathered at 4). Add it to the graphify allowlist files array and give the entry a real tracking issue (TBD -> 622). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
463cffd894 |
chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/ Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary (`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers are unaffected. Mechanical (bulk, ~90% of the diff): - `git mv get-shit-done gsd-core` - Swept path/identifier references across the repo via `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead preserves the five legitimate slug variants that are NOT the directory: get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names). - Build/manifest wiring: package.json (bin, files, coverage globs), tsconfig.build.json (outDir), ~86 .gitignore build-output entries, stryker.config.mjs, scan-ignore files, install.js path strings. - Frozen (not rewritten): CHANGELOG.md history; translated docs (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/). New logic (review here): - src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper ADR-0008 installer migration. On upgrade it walks the legacy `~/.claude/get-shit-done/` tree, classifies each file via the prior install manifest, and emits remove-managed / backup-and-remove for managed files while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked root and symlinked entries; bounds-checks every path under configDir). The framework rolls back on install failure. Emptied dirs may remain (framework has no recursive dir-removal primitive) — documented. - scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare `get-shit-done` directory token (split token to avoid self-match; case- insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines). Wired into the lint-tests CI job. - Restored scripts/lint-package-identity-drift.cjs detection regexes (the mechanical sweep had wrongly rewritten the old-name patterns it exists to detect) and marked them as intentional legacy references. - TDD tests for the migration and the guard; do.md slash-command guard regex tightened so a `/gsd-core/bin` path segment is not mistaken for a command; changeset + docs/installer-migrations.md row added. Breaking: the installed runtime path moves `~/.claude/get-shit-done/` -> `~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed files (preserving user files) on upgrade. Users with custom hooks/configs hardcoding the old path must update them. Closes #604 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unsweep pending changesets + allowlist injection-example docs CI fixes for the rename PR: - Do not sweep pending .changeset/*.md (ephemeral release-note fragments, like CHANGELOG); reverted those body edits so 5 pre-existing malformed fragments (missing type/pr) no longer enter the PR diff and trip docs-lint. Allowlisted .changeset/ in the legacy-name guard accordingly. - Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in prompt-injection-scan.sh: they contain intentional injection examples / security-model prose; the path-reference rewrites are kept. CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR; none in the new migration/guard) and are out of scope for the rename. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): resolve CodeQL alerts surfaced on this PR The rename diff touched files carrying pre-existing CodeQL findings; per the no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving them off. All behavior-preserving: - scripts/ci-test-scope.cjs: build the config-path match from string .includes() instead of a RegExp over an arg-derived value (js/regex-injection). - src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName so the table-cell escape is complete (js/incomplete-sanitization). - tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization). - tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace, keep the meaningful POSIX-class conversion (js/identity-replacement). Verified: build:lib green; the touched test files + ci-test-scope + profile-output suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization) The prior commit's fixes for two alerts were ineffective: - ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file` reaching static regex `.test(file)` calls (not the config rule). Removed ALL regex over file/t — startsWith/includes/=== string checks + an isWindowsHint helper — so there is no regex sink for the tainted value. - js/incomplete-multi-character-sanitization (3 test files): a single `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint loop (replace until stable) plus a final bare-opener strip. Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL CodeQL flags the regex PATTERNS syntactically (regex-injection on the --files arg split; incomplete-multi-character-sanitization on the <!--...--> replace), so loop fixes do not satisfy it. Made these paths regex-free: - ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/). - 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)). Behavior preserved; ci-test-scope + the 3 suites pass; guard clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unblock security base64 scan on the large rename diff The security job hit its 10m timeout: base64-scan.sh choked on the binary test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/ non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings), and the ~800-file rename diff is slow to scan regardless. - scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they can't carry base64-obfuscated *text* and feeding NUL bytes through the per-line scanner is pathologically slow. collect_files already filtered binary *extensions*; this catches binary *content* in text extensions. - .github/workflows/security-scan.yml: raise the security job timeout 10m->30m to accommodate very large diffs (the scan itself is unchanged). Verified locally: scan skips the fixture, 0 "ignored null byte" warnings, 0 findings, exit 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): sweep get-shit-done refs introduced by merging next The branch was updated with next (#614/#384/#618 etc.), which reference the get-shit-done/ dir (still named that on next). Swept the stale references in the merged files to gsd-core so the rename stays consistent and lint:legacy-name passes: - commands/gsd/discuss-phase.md (runtime-launcher shim paths) - src/core.cts (getAgentsDir layout comments) - tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib) Verified: guard 0 violations; build green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant The #614 runtime-launcher shim added to discuss-phase.md references `${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it mis-read the directory path as a dangling `/gsd-core` command ref (same class as the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path segments are not treated as slash-command references. Verified locally on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22 image) full suite: 0 failures - bug-3683 + bug-2954 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI) CI intermittently failed state.test's gsd-tools subprocess with "findProjectRoot is not a function" (flip-flopping across legs; not reproducible on mac full suite, gsd-test linux full suite, test:unit, or state.test x8). findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs); binding it via destructure at module-load can be undefined under a load-ordering edge. Resolve it lazily at call time via a small wrapper so the lookup happens after core.cjs is fully initialized. Verified green on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22) full suite: 0 failures - state.test.cjs: 106/106; gsd-tools loads cleanly. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): allowlist verification-patterns.md placeholder examples in secret scan The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var examples (illustrative Stripe test-key / database-URL / API-key placeholders) — not real credentials. Added it to .secretscanignore with the strict annotation, mirroring the existing gsd-core/workflows/plan-phase.md exception. Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next exits 0 with 0 findings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
b177c1704f |
fix(#614): resolve gsd-tools via runtime shim in discuss-phase mode routing (#618)
* fix(#614): resolve gsd-tools via runtime shim in discuss-phase mode routing The discuss-phase mode-routing snippet and the codebase-drift gate called the bare `gsd-tools` binary. On a shim-only install (gsd-tools.cjs present but `gsd-tools` not on PATH) the call exits 127, `2>/dev/null` hides it, and `|| echo` silently substitutes a default — so `workflow.discuss_mode: assumptions` was ignored and routing always fell back to standard discuss mode. Both sites now resolve the binary through the canonical `_GSD_SHIM_NAME` probe and call `gsd_run`. Discuss-phase fails loudly on a genuinely missing shim (interactive — wrong mode is worse than an error); the non-blocking drift gate uses a soft `return 127` fallback so it still skips gracefully when nothing is resolvable. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#614): set changeset pr to 618 * fix(#614): scope to discuss-phase mode routing; revert drift-gate change The runtime-launcher-parity invariant requires exactly one canonical (byte-equal) gsd_run preamble per workflow .md before the first gsd_run call. codebase-drift-gate.md has two independent gsd_run bash blocks; hardening its first block cleanly conflicts with that invariant and risks the auto-remap block's separate execution scope. Descope the drift-gate hardening to a follow-up and keep this PR focused on the titled bug: the discuss-phase mode-routing snippet now resolves gsd-tools via the runtime shim (gsd_run) instead of the bare PATH command, so shim-only installs no longer silently fall back to standard discuss mode. Drift-gate file reverted to its next state; its test assertion removed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
de2f73d21a |
enhancement(#34): add Antigravity CLI (agy) as a peer reviewer in /gsd-review
Closes #34 Squash-merged via admin override — all CI green (28/28 checks), branch protection review gate bypassed with maintainer authorization. |
||
|
|
79002a00cb |
chore(#518): rename npm package + bin to @opengsd/gsd-core (#519)
* chore: rename npm package + bin to @opengsd/gsd-core (functional) - package.json: name @opengsd/get-shit-done-redux → @opengsd/gsd-core, bin key get-shit-done-redux → gsd-core, repository/homepage/bugs URLs - package-lock.json: regenerated (npm install --package-lock-only) - tests/**, scripts/**, bin/**, .github/**, agents/**, commands/**, get-shit-done/bin/**, get-shit-done/workflows/**: applied the 4-rule replacement (scoped npm ref, GitHub repo path, bin/clone invocations) per #505 single-source refactor Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: sweep live references to @opengsd/gsd-core Update all live documentation (README.md + translations, docs/**, CONTRIBUTING.md, VERSIONING.md, SECURITY.md, CONTEXT.md, docs/CANARY.md) to reflect the renamed package and repository. Rules applied: - @opengsd/get-shit-done-redux → @opengsd/gsd-core (scoped npm name) - open-gsd/get-shit-done-redux → open-gsd/gsd-core (GitHub repo) - GSD-redux/get-shit-done-redux → open-gsd/gsd-core (stale badge org) - bare bin/clone refs → gsd-core CHANGELOG.md, docs/adr/**, docs/RELEASE-*.md, docs/research/**, and .changeset/** are preserved byte-identical. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix: add negative lookbehind to slash-command regex in bug-2954 test The extractSlashReferences regex matched /gsd-core inside npm package URLs (@opengsd/gsd-core), producing a false /gsd:core command reference. Adding a negative lookbehind (?<![a-z]) excludes matches preceded by a letter, so only standalone /gsd-<cmd> and /gsd:<cmd> tokens are found. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#518): add changeset for package rename Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#518): update package-identity expectations to the renamed coordinates The rebase regenerated the seam to @opengsd/gsd-core (bin gsd-core, repo open-gsd/gsd-core). The #498 seam tests assert deriveIdentity against the REAL package.json, so their expected literals must follow the rename. The drift-lint unit test is left as-is — its SEAM is a self-consistent fixture and its stale-literal detection cases would shift if altered; the live-repo scan in it already passes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7a3822fce1 |
fix: replace removed gsd-sdk prompt references (#355)
* fix: replace removed gsd-sdk prompt references * chore: add changeset |
||
|
|
d8b432da1e |
fix(#14): wire --auto flag through progress→next handoff (#148)
* fix(#14): document --auto in progress.md and wire chaining logic in next.md The --auto flag was accepted by /gsd:progress --next --auto but silently ignored: it was not documented in the <flags> section of progress.md and had no handling in the next.md show_and_execute step, so it was dropped at the handoff boundary and never produced step chaining. - Add --auto and --next --auto entries to progress.md <flags> - Update progress.md <process> to explicitly list --auto as a passthrough arg - Add --auto chaining logic to next.md show_and_execute: after each step completes, re-invoke /gsd:progress --next --auto until milestone complete or a blocking decision is required - Add regression test (4 assertions) covering all three fix points Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#14): bump lint-test-file-count progress ceiling for bug-14 test bug-14-progress-auto-flag-dropped.test.cjs resolves to the "progress" effective prefix and legitimately grows the cluster to 6. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * chore(#14): add changeset fragment Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(#14): address review feedback — differentiate duplicate tests, scope assertions to specific blocks Test 2 now extracts the <process> block and asserts --auto within it, distinguishing it from test 1's <flags>-level check. Remaining assertions use semantic token matches (--auto, --next --auto) that are robust to benign reformatting. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
99b52a302a |
fix(3659): applySurface prunes skill dirs on cluster disable (#3766)
* test(3659): add regression tests for applySurface skill-dir pruning on cluster disable Tests that applySurface with claude global scope correctly prunes ~/.claude/skills/gsd-STEM/ dirs for disabled clusters, preserves gsd-STEM dirs in enabled clusters, leaves non-gsd user dirs untouched, and is idempotent across two consecutive calls. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(3659): applySurface now prunes ~/.claude/skills/gsd-STEM/ on cluster disable Root cause: surface.md directed the AI to use RUNTIME_CONFIG_DIR=~/.claude/skills (the skills sub-directory) instead of the base Claude config dir (~/.claude). When runtimeConfigDir=~/.claude/skills and scope=global, the layout computes dest=~/.claude/skills/skills — the wrong target — so pruning never reached the actual gsd-STEM dirs in ~/.claude/skills/. Fix: - surface.md: correct RUNTIME_CONFIG_DIR to use the base config dir (~/.claude), add explicit SCOPE=global, and update all path references in execution_context. Surface state file moves from ~/.claude/skills/.gsd-surface.json to ~/.claude/.gsd-surface.json, matching install/uninstall conventions. - surface.cjs: extract pruneSkillDirs() as a shared helper (single point of truth for gsd-STEM dir removal). _syncGsdDir now delegates to it instead of having the ownership/prune logic inline. Export pruneSkillDirs for callers that need stand-alone pruning without a full applySurface pass. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(3659): update changeset to reference PR #3766 * fix(3659): manifest-membership gate on pruneSkillDirs prevents user gsd-* dir data loss Finding 1 (CRITICAL): the prefixed branch previously deleted any on-disk dir that matched the 'gsd-' prefix and was not in retainedNames. A user-created gsd-mything/ would be silently destroyed. Fix: deletion now requires BOTH prefix match AND manifest membership (stem present in manifest). Dirs that match the prefix but are not manifest-known are preserved with a process.stderr warning so the user knows the dir was kept. Finding 2 (type guard): the Hermes (empty-prefix) branch passed manifest directly to new Set([...manifest.keys()]) without verifying it is actually a Map. A truthy non-Map would throw. Fix: safeManifest = (manifest instanceof Map) ? manifest : null, used in both branches. Non-Map manifest triggers the same conservative no-deletions path already used when manifest is absent. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(3659): counter-test for all-clusters-disabled + user gsd-* dir preservation Finding 3: add test (e) that disables every cluster (Object.keys(CLUSTERS)) and asserts three things: 1. All GSD-owned skill dirs (gsd-explore/, gsd-help/) are removed. 2. Non-gsd user dir (my-custom-skill/) is preserved. 3. User-created gsd-mything/ (prefix match, not in manifest) is preserved — this is the critical regression guard for the Finding 1 data-loss fix. Also update the existing _syncGsdDir skills-kind test in surface-apply.test.cjs to pass a manifest that declares old-skill as GSD-owned. Without a manifest the new conservative path correctly preserves all unknown gsd-* dirs, which broke the pre-existing no-manifest assertion; supplying the manifest restores the expected pruning behavior and documents the required calling contract. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(3659): address pr-review-toolkit + codex review findings - Collapse redundant if/else in _syncGsdDir - Collapse duplicate canonicalStems branches in pruneSkillDirs - Update stale module-header comment (config-dir root) - Clarify dead isGsdOwned guard comment - Log rmSync failures to stderr - Add pruneSkillDirs to module-header Exports JSDoc - Remove unused imports in bug-3659 test file Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
09ab16f9e5 |
fix(3683): normalize /gsd:<cmd> → /gsd-<cmd> in command, workflow, and reference bodies (#3685)
* fix(3683): normalize /gsd:<cmd> → /gsd-<cmd> in command, workflow, and reference bodies Extends #3677's agent-body normalizer to all body text staged through copyWithPathReplacement (commands, workflows, references). The initial isCommand guard was structurally redundant — normalizeAgentBodyForRuntime already self-gates on shouldNormalizeHyphenNamespaceInAgentBody(runtime), so dropping it covers all hyphen-name runtimes (Claude / Qwen / Hermes) without affecting colon-canonical runtimes (Gemini). Addresses the user-visible symptom in #3683: workflows like get-shit-done/workflows/discuss-phase.md (7 colon refs) leaked /gsd:<cmd> markers to the model context, which the model echoed at the end of /gsd-discuss-phase runs. Source-prose drift caught by the new cross-reference invariant test: - commands/gsd/plan-phase.md: removed a slash-form mention of the deleted /gsd-research-phase command (#3042) - commands/gsd/profile-user.md: replaced a slash-form artifact reference with a backticked bare name (the referenced item is a skill config, not a user-callable slash command) Tests: - tests/bug-3683-command-colon-namespace-leak.test.cjs — runtime-form regression for commands/gsd/*.md staging - tests/bug-3683-command-cross-reference-invariant.test.cjs — locks cross-reference coherence: every /gsd-X / /gsd:X reference in a command body must resolve to commands/gsd/X.md (so a future rename forces every cross-reference to update) - tests/bug-3683-workflow-colon-namespace-leak.test.cjs — runtime-form regression for workflows + references; negative test for gemini asserting the colon form is preserved Fixes #3683 Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(changeset): remove undefined cycle reference --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
b8fa89b5b6 | fix(3663): address CodeRabbit surface/layout follow-ups | ||
|
|
0531313789 |
docs(3663): update surface.md runbook to layout-passing applySurface signature
Replace bare applySurface(runtimeConfigDir, commandsDir, agentsDir, ...) calls in profile/disable/enable sections with resolveRuntimeArtifactLayout + applySurface(runtimeConfigDir, layout, ...) pair. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
6a2bf05de7 |
feat(#3039): tier /gsd-help output (--brief, default, --full, <topic>) (#3040)
Replace the single 747-line /gsd-help reference with a progressive-disclosure dispatcher (#2551 pattern). Newcomers get a one-page tour; returning users get a 10-line refresher with --brief; the complete reference stays available behind --full; /gsd-help <topic> emits one section; and /gsd-help --brief <topic> is a compact scoped lookup (signature + one-line summary). - workflows/help.md becomes a small dispatcher routing on $ARGUMENTS - workflows/help/modes/{brief,default,full,topic}.md hold the tier bodies - commands/gsd/help.md passes $ARGUMENTS through, advertises composable form - docs/COMMANDS.md documents the new flags and topic form - existing tests that read help.md repointed at help/modes/full.md (bug-2836, bug-2950, bug-2954, cursor-reviewer, execute-phase-wave) - new feat-3039-help-tiered test enforces structure, size budgets, dispatcher routing, shim arg passthrough, topic→section coverage, orphan-heading detection, conflict-resolution rules, routing preamble, and compact-scope rule Trek-e review fixes (PR #3040): - topic.md output rules split into 5a/5b/5c — explicit handling for single sections, multi-section "plus" joins, and bold-line sub-block anchors; each rule also takes scope (full vs compact) into account - explicit resolved-routing preamble line emitted by topic.md before content ("**Topic:** `<alias>` → `<heading>` *(scope: full | compact)*") so the user sees which alias matched at which scope (review finding #3) - composable `--brief <topic>` invokes topic.md in compact scope: heading + first `**/gsd:*`** signature line + one-line summary. Dispatcher and topic.md cooperate via $ARGUMENTS pass-through (review finding #4) - full.md capped at LARGE-tier budget (FULL_BUDGET = 1500); the non-recursive workflow-size-budget test does not reach modes/ subdirs - structural <progressive_disclosure> table parse (5-row assertion) replaces substring-soup regex matching — 5 rows = 4 base tiers + composable scope - forward /gsd:* sub-block token coverage + reverse orphan-heading allowlist catch alias-table drift in both directions - four conflict-resolution tests guard dispatcher promises (--brief+--full without topic → --full; --brief <topic> → compact; --full <topic> or bare → full; dispatcher retains --brief when delegating to topic.md) - hardcoded topic lists removed from docs/COMMANDS.md and full.md (drift surfaces reduced from 5 to 2) - topic.md alias bloat trimmed (~75 → ~25 rows); cleanup/update split into distinct sub-block rows under ### Utility Commands - comment-rot ("~750 lines") removed from default.md and full.md - dispatcher size guard tightened from < 100 to <= 40 lines - commands/gsd/help.md <process> block trimmed to one line - MD040 fence languages added to all plain code blocks across mode files Main-merge conflict resolution: - workflows/help.md kept as dispatcher (body lives in help/modes/full.md) - /gsd-<cmd> → /gsd:<cmd> rename from #3452 reapplied to the mode files (full.md, default.md, brief.md, topic.md) — the six namespace routers (/gsd-context, /gsd-ideate, /gsd-manage, /gsd-project, /gsd-quality, /gsd-workflow) and wildcards (/gsd-*) preserved in hyphen form per main's convention - bug-2950 test combines branch's path repointing with main's namespaced replacement strings |
||
|
|
da21edfb59 |
feat(workflow): add git.create_tag config to disable milestone tagging
Adds boolean config key `git.create_tag` (default: true, fully backcompat) so projects with their own release flow can disable GSD's automatic `git tag -a v[X.Y]` on milestone completion. Also adds tag-collision pre-check to prevent silent failure on re-run. Closes #3086 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |