6cfa0c55d205eaeaf005f4d21f7b5e5cefa4448c
23 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6cfa0c55d2 |
refactor: drop 12 runtimes, keep Claude, Codex, OpenCode, Cursor, ZCode, Antigravity
Removes kilo, kimi, kimi-code, copilot, windsurf, augment, trae, qwen, hermes, cline, codebuddy and pi end to end: capability descriptors, installer branches and converters (bin/install.js 14.9k -> 11.2k lines), TypeScript converters, hook surfaces and runtime homes, review lanes qwen/kimi-code, the two pi migrations, Kimi payload normalization in the hook guards, dead hostBehaviors vocabulary, launcher home probes, fixtures, runtime-specific tests and the prose that presented them as supported. Installer output for the six kept runtimes is byte-identical to before the prune. The Kimi tool-vocabulary tests in workflow-guard, read-guard and read-injection-scanner are left in place pending a decision. |
||
|
|
a9a7a328e6 |
refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD across contents and paths, upstream package/repo coordinates -> @golem15/msd-core and golem15com/msd-core. Deep links into upstream history, sibling upstream packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is. Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line, package/plugin identity, regenerated lockfile, install-tree fixtures, derived registries and benchmark baseline; migration checksum baseline re-locked (MSD keeps its own install state, so no install had applied the old sums); sort-order and regex-escaped expectations in tests adjusted. |
||
|
|
9a41a95212 |
fix(#4717): consult the per-install runtime marker at both identity seams (#4861)
* test(#4717): add failing-first coverage for the two runtime-identity marker seams * fix(#4717): consult the per-install runtime marker at both identity seams resolveReportedRuntime (agent_runtime) and loadConfigResolved (config.runtime) both ignored the per-install .gsd-runtime marker that resolveRuntime and the model-resolver gate already read. On a multi-runtime machine (e.g. a globally exported CODEX_HOME), host sniffing misreported every Claude Code session as codex, and a shared defaults.json stamped by the first non-Claude install leaked its runtime to every other one. Seam 1: the reported-runtime ladder becomes explicit > install marker > host detection > claude. Seam 2: loadConfigResolved fills an empty config.runtime from GSD_RUNTIME then the marker, copy-on-write (the builtin-defaults branch returns a shared object). Explicit runtimes and marker-less trees are unchanged. * fix(#4717): a marker-detected runtime opts into its tier map (decision a) * fix(#4717): stamped-defaults leg, marker fail-safe, docs, review fold-ins * chore(#4717): backfill changeset PR number (4861) --------- Co-authored-by: sim <sim@local> |
||
|
|
26e650f312 |
fix(#4462): normalize whitespace workstream environment values (#4509)
* test(#4462): expose whitespace workstream scope * fix(#4462): normalize the workstream environment scope * docs(#4462): add changeset for #4509 * fix(#4462): normalize sibling planning scope readers * test(#4462): cover workstream normalization properties * fix(#4462): share normalized workstream resolution --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
925a363879 |
enhance(#4032): apply configured agent tool grants (#4238)
* test(4032): add failing installed-agent grants contract Cover global and project agent_tools precedence at the real Claude installer seam before adding implementation. * feat(4032): apply configured agent tool grants during staging Resolve selector-level global and project config once per staging call, then append validated grants before runtime conversion. * test(4032): cover host grant and quoted MCP contracts Exercise installed host artifacts and prove ZCode must treat quoted MCP scalars like plain MCP grants. * feat(4032): apply configured agent tool grants across runtimes Move augmentation and scalar identity into the converter seam so every staged artifact preserves host policy. * fix(4032): register agent tool grants in configuration Accept documented agent_tools config without unknown-key warnings.\n\nKeep installer fixtures on the shared temporary-directory helper. * fix(4032): translate configured MCP grants for Kilo Reuse the converter-owned scalar decoder so quoted canonical grants reach Kilo's native permission keys without altering other host policies. * fix(4032): decode YAML-escaped tool grants * fix(4032): emit valid inline agent tool grants * fix(4032): reject invalid trailing-colon grants * test(#4032): cover cross-review remediation gaps * fix(#4032): close cross-runtime grant gaps * test(#4032): expose Kimi global project context * fix(#4032): preserve Kimi project config context * chore(#4032): add release note * test(#4032): expose fork review regressions * fix(#4032): address fork review findings * test(#4032): make byte-stability assertion portable Compare repeat installs at one root so platform-specific path rendering cannot masquerade as an agent_tools behavior change. * chore(#4032): bind changeset to upstream PR 4238 * fix(#4032): address trek-e review findings (2,3,4,5,6,7,8) Fixes fail-closed decode-failure handling in ZCode's mcp__ stripper, a comment-only `tools:` header mis-parse that silently dropped configured grants, and a naive comma-split that could tear a quoted scalar containing a literal comma. Documents Kilo's inherent `{server}_{tool}` MCP-permission-key collision (external, fixed format — not ours to widen) and locks the existing first-seen-wins resolution in with a regression test. Opts kimi/kimi-code out of the ADR-1235 pre-converter path-rewrite step: routing Kimi through that pipeline (needed so project-scoped agent_tools selectors reach it) was short-circuiting Kimi's own neutralizeKimiAgentPrompt, which expects the original ~/.claude/gsd-core text rather than a pre-rewritten Kimi path. Extends the fast-check token pool and per-runtime install coverage with the missing comment/comma/broad-runtime cases the prior review flagged as untested. * docs(#4032): add CONTEXT.md glossary entries for agent_tools resolver + pre-converter step Documents readGsdEffectiveAgentTools (Install Model Override Resolver Module) and the appendAgentTools pre-converter pipeline step (Runtime Artifact Conversion Module), per contributor-standards.md's new-seam glossary requirement (finding 1). * fix(#4032): address agy adversarial review findings An agy (gemini-3.8-flash-high) adversarial pass over the prior review-fix commit found the fixes for findings 3, 4, 6 and 8 had unfixed sibling gaps, plus a genuine new regression and two CONTEXT.md inaccuracies: - ZCode's comment-only `tools: # note` header matched the inline-value branch instead of falling through to the block-list scan, so a following mcp__* item leaked through unstripped — the exact defect finding 4 fixed in appendAgentTools, unfixed in this sibling function. - Reverted capabilities/kimi-code/capability.json's noPathRewrite: true. kimi-code uses the standard 'agents' kind with converter: null (not kimi-agents — confirmed by reading the descriptor, not its prose description), so it never went through the pipeline change finding 5 fixed, and disabling its path rewrite broke every ~/.claude/ embed in its shipped agents instead. - decodeToolScalar never stripped a trailing ` # comment` from a bare (unquoted) scalar, so a comment after a block-list item, or after an appended grant on an inline line, became part of the "tool name" — fixed at the source (one call site fixes every consumer). - appendAgentTools's comment-index scan wasn't quote-aware, so a `#` inside a quoted scalar (`"mcp__server #1"`) was mistaken for a comment start and corrupted the quote. - parseFrontmatterTools (Kimi/Qwen's tool-list reader, downstream of appendAgentTools's own output) had the same naive comma-split and comment-only-header gaps as findings 4 and 6, unpatched. - The all-runtime smoke test's presence assertion was built on a guessed omit-list; empirically only 7 of 17 runtimes keep an arbitrary mcp__ grant recognizable, replaced with a verified allowlist. - CONTEXT.md claimed a `project:<agent>` selector prefix that does not exist (project override is a same-key merge across two config files) and mislabeled stageAgentsForRuntimeWithConverter's module. * fix(#4032): address full-PR review (Opus critical/ponytail + agy) A whole-PR pass (critical-code-reviewer + ponytail-review on Opus, plus a second agy full-source adversarial pass) surfaced defects the earlier finding-scoped passes couldn't reach: - appendAgentTools corrupted a `tools:` line whose ENTIRE value is a leading quoted scalar (`tools: "Read"` -> `tools: "Read", Write`, invalid YAML) — there is no safe line-surgical rewrite here, so it now refuses to touch that shape instead of emitting broken frontmatter. - decodeToolScalar's malformed-trailing-quote check ran BEFORE comment stripping, so a bare tool name with a quote inside its own trailing comment (`Bash # note: "internal"`) was wrongly rejected. Reordered. - findUnquotedCommentIndex (added in the prior remediation commit) was built on a wrong model of YAML: a `#` after whitespace starts a real comment in a plain scalar regardless of nearby quote characters — verified against the actual parser. The one case that DOES need protection (a leading quoted scalar) is now refused outright above, so the quote-tracking scan was dead weight solving a problem that no longer reaches it. Removed; reverted to the plain `[ \t]#` scan. - Kilo has a SEPARATE agent-frontmatter parser (convertClaudeToKiloFrontmatter, distinct from the buildKiloAgentPermissionBlock fixed earlier) with the same comment-only-header and naive-comma-split gaps as findings 4 and 6 — unfixed in both its src/ and bin/install.js copies. Fixed in both, exporting splitToolScalars for bin/install.js to reuse rather than reimplementing it. - Pipeline docstring in stageAgentsForRuntimeWithConverter still listed 5 steps, omitting appendAgentTools (now step 3 of 6). - docs/CONFIGURATION.md didn't state that a --global install still discovers agent_tools from the cwd's .planning/config.json (confirmed intentional and already covered by a dedicated test, not a bug). - Removed install-engine.cts's deps.cwd injection seam: zero callers or tests ever populated it. Two claims from this round were verified and rejected, not fixed: prototype pollution via a `__proto__` selector key (empirically confirmed `Object.prototype` is never touched — only reassigns the resolver's own local object's prototype, with no observable effect), and a `*` grant value crashing YAML parsing as an alias reference (empirically confirmed it parses as plain scalar text, no crash). A pre-existing, unrelated defect (extractFrontmatterField returns null for block-list `tools:` on Copilot/Antigravity/Cursor/Codex/Qwen, affecting two shipped agents today) was filed as a follow-up rather than fixed here — it predates #4032 and isn't caused or worsened by this PR. * fix(#4032): update stale slug-derivation-drift-guard fixture line normalizeKimiSkillName's real closing brace moved from line 616 to 635 as a side effect of this PR's edits to runtime-artifact-conversion.cts; the MAJOR-1 fixture's hardcoded realEndLine had gone stale. * fix(#4032): address CodeRabbit findings on projectDir threading and flow-sequence tools bin/install.js's installAgentsKindStandalone call site omitted the projectDir argument the function already supports, so a global install through this legacy branch silently fell back to the runtime config dir instead of process.cwd() when resolving project-scoped agent_tools grants — inconsistent with the sibling installOpencodeFamilyArtifacts call site, which already threads it correctly. appendAgentTools' leading-quoted-scalar bailout did not cover a YAML flow sequence (`tools: [Bash, Read]`): splitToolScalars tore it apart on the in-sequence commas and appended past its closing bracket, producing invalid frontmatter. Extended the bailout regex to also refuse a value starting with `[`, matching the same "whole node, nothing may follow" reasoning already applied to quoted scalars. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
400db94e02 |
fix(#3894): quick path honors workflow.research_before_questions; key resolves from global defaults (#4047)
* test(#3894): research_before_questions must resolve globally and order quick.md (failing first) * fix(#3894): quick path honors workflow.research_before_questions; key resolves from global defaults Two layers, one key. The quick workflow ran its discussion phase before its research phase unconditionally — neither quick.md nor its steps ever read workflow.research_before_questions, though the key is documented, schema-registered, /gsd-settings-writable, and honored by /gsd-discuss-phase and /gsd-new-project. A gray-area answer given without research is then written to <quick_id>-CONTEXT.md as a locked decision downstream agents are told not to revisit — an evidence-free choice made unfalsifiable (the reporter's #3714 misresolution). - quick.md Step 4 now carries the same research-before-questions check the two honoring paths make: when enabled, research-phase executes before discussion-phase; false/unset keeps the written order. Both sections stay section-manifest gated. - src/config-loader.cts forwarded workflow.post_planning_gaps from ~/.gsd/defaults.json but silently dropped this key — same file, same nesting, one resolved and one didn't. Now forwarded with the same flat + nested-alias fallback shape, added to the resolution-keys lockstep canary and the #3532 shadowed-warning set (nested alias reporting generalized over both keys). Emitted-Drift-Ack-Growth: quick.md — #3894: +Step 4 ordering rule (the research-before-questions check the discuss-phase and new-project paths already make); a real behavioral gate, not incidental bloat. * fix(#3894): review fold-ins — gate the CONTEXT.md reference, colon slash-forms - quick/steps/research-phase.md directed the researcher subagent to read <quick_id>-CONTEXT.md under DISCUSS_MODE with no existence hedge — but under the new ordering (research BEFORE discussion) the file cannot exist yet when the researcher is dispatched. The reference now says read-only-if-present with the #3894 reason; the alignment purpose still applies on the default ordering. - quick.md's new rule used the hyphen slash forms (/gsd-discuss-phase, /gsd-new-project); source artifacts under gsd-core/workflows must author the colon form the install-time converters key on — the same file already uses /gsd:new-project and /gsd:quick elsewhere. * docs(#3894): planning-config row names the flat CONFIG_DEFAULTS alias config-field-docs requires every CONFIG_DEFAULTS key to appear in the doc; the row documented the canonical namespaced form only. Adds the same alias sentence post_planning_gaps's row carries, plus the #3894 quick-path note. * chore(#3894): changeset fragment (pr number backfilled after PR creation) * chore(#3894): backfill changeset PR number (4047) --------- Co-authored-by: sim <sim@local> |
||
|
|
3fd03bec4c |
fix(#3760): refuse a legacy-key migration into a non-object config section (#3767)
* test(#3760): failing-first regression for non-object config section Locks the contract from the issue's Expected section before any fix exists: a legacy-key section holding a string, number, boolean or array must be preserved verbatim, reported, and never persisted in an expanded form. Covers both blocks the issue names (branching_strategy -> git.*, sub_repos -> planning.*), the migrateOnDisk multiRepo branch that shares the shape, and the loader write paths that are what actually reach the user's config.json. Includes the negative-space cases that must keep hoisting ({} , null, absent section, canonical-nested-wins) and two fast-check properties. Refs #3760 * fix(#3760): refuse a legacy-key migration into a non-object config section normalizeLegacyKeys hoisted a legacy top-level key into its canonical nested section by spreading `result[section] ?? {}`. `??` guards only null and undefined, so a section holding a string was enumerated by index — `{...'main'}` is `{0:'m',1:'a',2:'i',3:'n'}` — while a number or boolean spread to `{}` and the value vanished. Because a fired block always pushed a Normalization, and every caller treats a non-empty normalizations array as 'config is dirty', that shape was written back to .planning/config.json and the original value became unrecoverable. Both blocks the issue names are fixed via one shared hoistLegacyKey helper, plus the two further sites that share the shape and are reachable from the same input: migrateOnDisk's multiRepo branch, and the loader's two `if (!planning) planning = {}` guards, where a non-empty string is truthy and the following assignment threw a strict-mode TypeError that the enclosing catch swallowed — discarding the user's entire config. A present non-object section now blocks its own migration. The section, the legacy key, and the file are left byte-identical; no Normalization is pushed, so nothing marks the config dirty; the refusal is reported in-band as `skipped[]` and out-of-band through the ADR-1411 warnUnusableInput seam (new frozen reason config_section_not_object). null and undefined keep their long-standing 'absent' meaning and still create the section. This is the nested-section analog of the ADR-227 shape check _readConfigFile already performs on the top-level document: valid JSON is not a config object. isConfigSection is exported and shared by both modules rather than copied. Fixes #3760 * fix(#3760): keep the multiRepo marker when planning cannot receive it Follow-up from the isolated adversarial review, and the same defect class as the two blocks the issue names — in the block it did not name. normalizeLegacyKeys block 3 deleted `multiRepo` and pushed a Normalization before anything consulted the planning section, deferring 'can this section receive sub_repos?' to the caller that runs filesystem detection. By then the marker was already gone and the config was already dirty, so with {"multiRepo":true,"planning":"docs"} the loader wrote the file back with multiRepo removed, the sub_repos injection silently no-opped against the string, and no diagnostic was emitted at all. migrateOnDisk warned for the same input; the ~30-caller loadConfig path did not. Section validity is knowable from the parsed config alone — detection is only needed for the VALUE, not for whether the destination can hold it. The refusal moves into block 3: the marker is kept, no Normalization is pushed, and a skipped entry is recorded, so all three callers inherit the preservation and the diagnostic together. The caller-side guards drop to pure narrowing. Also from review: skipped[] now reports sectionType ('string' | 'number' | 'boolean' | 'array') instead of sectionValue. migrateOnDisk's report is printed verbatim by `migrate-config`, and this module already masks config values on the set/unset output path; the type is the whole diagnostic and the value is still in the file. And `migrate-config --raw` no longer answers a refused migration with 'No legacy keys found — config is already canonical.' Legacy keys WERE found and declined, and the decline is the one thing only the user can fix by hand. Refs #3760 * fix(#3760): keep configuration.cjs dependency-free; emit from its callers The remote matrix caught a regression my own change introduced: adding `require('./unusable-input.cjs')` to configuration.cts broke the #3571 install-layout contract. `configuration.cjs` must load from a layout holding only itself plus bin/shared/*.manifest.json — the installer does not co-locate arbitrary siblings — so the new require failed at load time: Cannot find module './unusable-input.cjs' Require stack: - /tmp/gsd-3571-.../.codex/gsd-core/bin/lib/configuration.cjs pinned by 'co-located bin/shared manifests let configuration.cjs load without sdk/shared' in tests/install.test.cjs (3 failures). The contract is deliberate and the test is right, so the module goes back to zero sibling requires and the out-of-band diagnostic moves to the callers that already carry a dependency budget and hold the resolved path: cmdMigrateConfig (config.cts) and loadConfigResolved (config-loader.cts). normalizeLegacyKeys keeps reporting refusals in-band via skipped[], which is what lets it be pure and dependency-free at the same time. The emission-count and dedup assertions move to tests/config-loader.test.cjs, where the diagnostic now originates. A new assertion pins the inverse for the module itself — migrateOnDisk must emit ZERO diagnostics while still reporting skipped[] — so regrowing a sibling require fails a unit test instead of only the install suite. CONTEXT.md records why the emitter is the caller. Refs #3760 * chore(#3760): backfill changeset pr number to 3767 --------- Co-authored-by: sim <sim@local> |
||
|
|
a280054040 | fix(#3532): hermetic child GSD_HOME, typed-IR canaries, nested alias, list parity | ||
|
|
4fe072283d | test(#3532): failing-first suite for shadowed global-defaults diagnostic | ||
|
|
bc5c362610 |
fix(#2665): replace the hand-synced TEST_ENV_BASE copies with the canonical import
Eight declarations were kept in sync by hand with no parity assertion. The drift was already in the tree: api-coverage-gate-e2e and representative-corpus declared TERM_SESSION, but the real variable is TERM_SESSION_ID, so both scrubbed nothing for that slot and leaked TERM_SESSION_ID into every child. Deleting the copies removes the dead key with them -- there is no longer a second place to get wrong. run-tests-harness keeps its local session-identity literal: that helper mirrors the production runGsdTools to prove its CONTRACT, so it must not re-import what it is testing. The config-LOCATION keys are a safety scrub rather than part of that contract, so it spreads the canonical derived set and keeps the rest local. Addresses review findings: Blocker 2, Blocker 3. |
||
|
|
08021b02c0 |
fix(#2665): scrub config-location env vars in every TEST_ENV_BASE declaration
TEST_ENV_BASE blanks session-identity variables but none of the three that
decide WHERE a child process writes: CLAUDE_CONFIG_DIR, GSD_RUNTIME and
CODEX_HOME. The config-home resolver is env-first (runtime-homes.cts, the
dot-home case consults the env var before the home-derived fallback), so an
ambient CLAUDE_CONFIG_DIR in the developer's shell beats a call site that
sandboxes only HOME. The suite then writes into the developer's real config
directory -- including a registered skill under <configDir>/skills/ whose
body carries behavioural directives that load into later sessions.
Blank all three alongside the session-identity vars. `...env` still spreads
last, so the five call sites that already constrain these locally keep
winning with their explicit values.
TEST_ENV_BASE is re-declared in nine files, so the three lines are added
nine times rather than once. Consolidating the nine into a single exported
constant -- and fixing the TERM_SESSION / TERM_SESSION_ID drift between the
copies -- is deliberately left out of this change; see the PR body.
One call site needed adjusting. capability-state.test.cjs's
`capability state --runtime claude` CLI test passed no env at all and
compared the CHILD's resolved config dir against the PARENT process's
getGlobalConfigDir('claude'). That agreed only because the child inherited
the developer's ambient CLAUDE_CONFIG_DIR -- i.e. it passed *because of*
the leak. It now redirects both runtime homes into the sandbox and asserts
against values the test controls, so it is hermetic with the variable set
or unset.
Regression case folded into the owning module's test file rather than a new
bug-NNNN file, per scripts/lint-regression-test-names.cjs. It sets the
variable on the PARENT process, which is the actual vector; setting it in
the per-call env argument would exercise a path that was never broken.
|
||
|
|
9faacc0c15 |
test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist Migrates the final 170 unbounded sync spawn sites across 49 files, then removes the allowlist entirely. local/no-unbounded-spawn now runs with no exemption surface across tests/**: there is no file to add a name to. drift-detection's throw-native git() helper routes to gitOrThrow -- bare runGit would have taken 16 call sites quiet on failure. commands.test.cjs has two independently-scoped runGsdTools/runCli helpers, one already bounded and one not; they are kept distinct rather than unified, the same trap as the two same-named git() helpers in Wave 1. runNpm's bound was erasable. Its options spread callerOptions after the defaults, so an explicit timeout:undefined silently dropped the 180000ms bound -- the rule flagged it and was right; it was not a false positive. Fixed by destructuring with a default, with a test that fails when the default is removed. Two sites stay on a raw spawn with an explicit timeout because the seam cannot express them: one needs shell:true for npm.cmd on Windows, one redirects stdout to a real fd. Both are the rule's own documented second option, not an escape from it. Closure verified rather than asserted: the derivation scan reports 0 unbounded spawn helpers and 0 unbounded direct git call sites, and a temporary file carrying an unbounded spawn still errors with the allowlist gone. Closes #3064. * test(#3148): close a hole in the guard's own eslint-disable ban The ban listed only the top level of tests/, so it was blind to 37 .cjs files under tests/helpers, qa, observability, fixtures and dispatch. With the allowlist deleted this test is the sole remaining way to detect someone silencing the rule inline, so the gap was load-bearing: a nested file could carry an unbounded spawn plus an eslint-disable and pass everything. Proven before and after. A probe planted under tests/helpers with both was invisible to the guard and clean under eslint; after making the listing recursive the guard fails on it. The scanned set goes from 771 files to 808. Pre-existing since the guard shipped, but this wave is what promoted it to sole defense, so it is fixed here rather than filed. Also converts the last hand-rolled throw check to throwIfFailed and the last re-derived legacy shape to compose toLegacyResult, which makes the epic's none-remain claim true rather than nearly true. toLegacyResult itself is not widened -- eight callers depend on its shape and one consumer does not justify changing a shared contract. * fix(#3148): correct seam incoherence at the bound and a slow review-lane error path Two real failures from the remote runner, both fixed at the cause. The seam could return outcome TIMED_OUT together with exitCode 0. At the exact bound spawnSync reports ETIMEDOUT while the child has already exited with a real status, and toSeamResult classified on the error code while passing status straight through -- an incoherent pair its own boundary test was written to catch, and did. A status that is not null is direct evidence the child exited on its own, so it now decides the outcome before the error-code branches run. process-seam.cjs was deliberately untouched by every earlier wave; this is a defect in the module itself, kept surgical, with a unit test that fails against the old logic. review-lane with an unknown subcommand fell through to its usage error only after loading the capability registry and building a per-lane plan, which spawns one child process per lane -- up to twelve. The error path took ~1288ms instead of ~119ms, and under bench load it outran a caller's spawn timeout and was killed before writing anything, which is the empty stdout and stderr CI saw. It now fails fast before any of that work begins. This is the epic's first production change. It is user-facing, so it carries a changeset rather than a no-changelog label. * test(#3148): replace a real-race timeout test with a deterministic one E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a warm container git finishes first, spawnSync returns status 0 with no error at all, the seam correctly classifies EXITED, and gitOrThrow correctly does not throw -- so the test failed on both lanes. A probe confirms a genuine timeout always carries status null, so this was never the seam misbehaving. Raising the bound would only lengthen the odds, which is the same defect with better luck. The test now drives gitOrThrow against a stubbed runGit that returns a synthetic TIMED_OUT result, so it asserts exactly what it always meant to -- that a timeout propagates as a throw -- with no timing dependence. Five consecutive runs are identical where the old one varied. I wrote this test in Wave 0; it is a real-race test by construction and CLAUDE.md says to replace those rather than re-run them. * chore(#3148): backfill changeset PR number 3192 --------- Co-authored-by: sim <sim@local> |
||
|
|
c7da62b682 |
Merge pull request #3094 from open-gsd/test/3057-wave2-liveness
chore(#3057): remove tests that report coverage they do not have — Wave 2 |
||
|
|
aa7697fe97 |
fix(#2997): include phase_id_convention in resolved config (#3098)
* fix(#2997): include phase_id_convention in resolved config _baseConfig in config-loader.cts is an explicit allowlist of keys copied from parsed into the resolved config. phase_id_convention was in VALID_CONFIG_KEYS (manifest line 83) and survived the unknown-key filter, but was never copied into _baseConfig — silently dropped on a clean read. The milestone-prefix validation check could only be activated via the ROADMAP frontmatter fallback, not the documented project-config surface. Added phase_id_convention: get('phase_id_convention') ?? null to _baseConfig. 3 tests: survives resolution, null round-trips, absent resolves to null. * chore(#2997): backfill changeset PR number 3098 --------- Co-authored-by: sim <sim@local> |
||
|
|
7dd9e59f6b |
test(#3090): stop exempting violations under categories that do not fit
An allow-test-rule annotation citing a category that does not apply is worse than no annotation, because it reads as reviewed. Eight were confirmed by reading the assertions each one covered, and auditing the rest found five more plus one refutation — a converter test whose wording described the wrong mechanism while the covered assertion genuinely was deployed-text. The instructive one used the CANONICAL string for the same mistake: STATE.md command output labelled as a deployed artifact. A canonical string is not evidence the category fits, which is why normalising strings alone would have laundered the problem rather than fixed it. Every mapping the audit had inferred rather than code-verified was spot-checked before rewriting, and the ones that turned out not to fit were re-annotated rather than relabelled. Fourteen STATE.md assertions had a typed extractor available all along and now use it; their annotations came out because nothing needs exempting. Eight assertions genuinely need a production change first — CLI stdout and stderr with no structured mode — and are tagged pending-migration-to-typed-ir citing #3090, which is what that category is for. It had zero real uses before this, while one file carried a real citation to migration issue #2974 under a non-canonical tag. Six annotations covered assertions that do no text matching at all. An exemption for a violation that does not exist is noise that makes the real ones harder to audit; those are removed. atomic-write-coverage gains the annotation it always warranted — its own docstring describes a structural-regression-guard while the file carried none. Fifty-nine non-canonical strings across roughly thirty files are normalised, and the allow-test-rule allowlist is regenerated to match. 472 annotations became 463: every one now uses a canonical category, and the two remaining non-canonical strings are ESLint RuleTester fixtures, not annotations. Refs #3057 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3eb1cede26 |
fix(#1880): distinguish a corrupt config from an absent one (epic #1879 Phase 1) (#2688)
* test(#1880): prove corrupt config is indistinguishable from absent Failing-first. Encodes the issue's runtime repro: a trailing comma in .planning/config.json currently yields source:builtin-defaults with degraded:false - byte-identical to the file not existing - and the user's entire configuration is silently discarded. Asserts on the typed surface (CONFIG_REASON, _warnedUnusableConfig) rather than diagnostic prose, per the ADR-1411 amendment's test-methodology clause and CONTRIBUTING.md's raw-text-matching rule. IO failure is injected by monkeypatching fs.readFileSync and restoring in t.after(), never chmod 0o000 (root bypasses mode bits). Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): distinguish a corrupt config from an absent one loadConfigResolved wrapped the read, the JSON.parse and the entire config build in one try with one catch, so ENOENT, EACCES and SyntaxError all fell through to the same defaults and the branches returned degraded:false - actively asserting health over discarded configuration. A single trailing comma in .planning/config.json silently replaced the user's whole config, reporting source:builtin-defaults degraded:false, byte-identical to having no config file at all. ConfigResolution now carries a machine-readable reason. Genuine absence keeps degraded:false / not_configured; a file that exists but cannot be used sets degraded:true with config_unparseable or config_unreadable. The same split applies to the root config and to ~/.gsd/defaults.json. Control flow is deliberately unchanged. preflight_check reports cyclomatic 141 / cognitive 196 and 93 dependents on this function, with the guidance that small edits beat one big one, so faults are CAPTURED at the existing read sites and stamped onto the returns rather than the try/catch being restructured. Also carries the ADR-1411 amendment's wiring clause: loadConfig returns .config alone to ~51 call sites and would never see the new field, so an unusable file emits a deduplicated stderr diagnostic keyed on resolved path plus errno. Without it the reason would be an unreachable field and the user whose config was discarded would still get no signal - the actual defect. Registers the config-loader seam in lint-resolution-provenance, which until now guarded only agent-skills. Caller audit: ConfigResolution.degraded has exactly one consumer outside this module, cmdAgentSkills (src/init.cts:2259), which destructures {config, source, degraded} - adding a field does not break it. Its --json IR now reports degraded:true for a corrupt config, which is the intended fix and the one observable behavior change. Closes #1880 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): degrade when any config on the path is unusable, not just the last Two defects found by isolated adversarial review of the first cut. BLOCKER: the success-path return did not consult configFault. A corrupt ROOT config whose workstream override happened to parse returned degraded:false / reason:resolved - the root's settings silently dropped, which is the exact failure this issue closes, reappearing for any project using workstreams. The stderr diagnostic fired, so the out-of-band half worked while the in-band half reported a clean resolve; a --json consumer saw health. MAJOR: reason was derived from Object.keys(parsed) - the root+workstream MERGE - so an empty workstream file inheriting a non-empty root reported resolved despite carrying no settings. Emptiness is now judged on the file actually read, snapshotted before normalizeLegacyKeys mutates it. Also: corrects the ConfigResolution JSDoc, which still described the pre-#1880 degraded contract; adds a fast-check property asserting a PRESENT file is never reported not_configured whatever its bytes (CONTRIBUTING.md parser rule); and asserts the literal enum values so the provenance lint's configured_empty/not_configured markers check real assertions rather than incidental prose in test titles. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): reject valid JSON that is not a config object at the read seam The fast-check property added in the previous commit failed on both node lanes: a config.json containing 0, "str", [], null or true is valid JSON, so it parsed "ok", then threw downstream in normalizeLegacyKeys, and the outer catch reported not_configured - a PRESENT file reported as absent, which is precisely the collapse this issue exists to close. The property asserts a present file is never not_configured, and it caught it. _readConfigFile now validates shape, not just parseability (ADR-227: check the semantic shape at a trust boundary, not merely the type). A non-object JSON document is an unusable config, reported config_unparseable. Adds named regression cases for each non-object form alongside the property, so the class is documented and not only randomly sampled. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1880): backfill changeset pr number (pr:0 -> 2688) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2452): record a fetch-time shallow failure instead of crashing This guard failed CI on ubuntu-24 while passing on ubuntu-22 and windows-24 for the same commit, and passed on other PRs. Not a flake and not caused by the change under test - a real fragility in the test. runnerDiff ran the base fetch OUTSIDE its try and only guarded the diff, so it assumed the failure mode is always 'fetch succeeds, diff reports no merge base'. At a shallow boundary that lands short of the merge base, git can instead fail during the FETCH ('unable to parse commit' - the boundary commit's parent is not available). Which stage git fails at is version and transport dependent, so on some runners the error escaped runnerDiff and crashed the test rather than being recorded as the ok:false the assertions expect. Both stages mean the same thing for what this guard protects: a shallow base ref cannot resolve the three-dot diff. Also drops two assert.match calls against git's stderr prose. 'no merge base' and 'unable to parse commit' are the same condition reported at different stages, and CONTRIBUTING prohibits raw text matching on subprocess output. The typed outcome (ok === false) is the contract; the tests now assert that plus the presence of a cause. Found while investigating the red lane on #2688; fixed here per the no-defer rule rather than filed. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c3958018dd |
docs(#2674): amend ADR-1411 — corrupt is not absent (epic #1879 Phase 0) (#2678)
* docs(#2674): amend adr-1411 with the corrupt-is-not-absent house pattern ADR-1411 reasons only about a resolution miss. It is silent on input that is present but not usable, which is how five engine read paths (#1879) could fold an unusable input into the value meaning 'genuinely absent' without contradicting an Accepted ADR. Read together, ADR-1411 and ADR-227 converge and do not license throwing as the cluster's answer: ADR-227 requires malformed input to be coerced rather than propagated and carves out only genuinely-fatal fields, while ADR-1411 already permits a fallback provided it is 'a visible value, not a silent substitution'. The defect in these five sites is therefore not that they fall back but that they fall back invisibly. Records the pattern that follows: every current return value is preserved, and the cause is made visible in-band where the result already carries a provenance envelope, or out-of-band via a deduplicated stderr diagnostic where it returns a bare value it cannot extend. Throwing stays confined to ADR-227's genuinely-fatal carve-out, decided per call. Also names the per-applier caller audit and the lint-resolution-provenance registry gap. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2674): prove the warning-state reset misses the unknown-key dedup set The two existing cases in this suite only pass because each picks a key name no other case reuses, so neither can observe whether the reset the beforeEach calls actually runs. Failing-first: asserts the exported _warnedUnknownConfigKeys is empty after _resetRuntimeWarningCacheForTests(). It is not - the helper clears only _warnedConfigKeys despite documenting itself as resetting per-process warning state. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2674): reset the unknown-key dedup set with the runtime warning cache _resetRuntimeWarningCacheForTests documents itself as resetting per-process warning state but cleared only _warnedConfigKeys, leaving _warnedUnknownConfigKeys populated across cases. The suite that exists to test that set - 'loadConfig - unknown-key warning dedup' - calls the helper in beforeEach expecting exactly this, so the reset was a silent no-op for it; both cases passed only because each picked a key name the other never reused. Any later case reusing a key would have had its warning suppressed by leaked state. Found while amending ADR-1411, which names this dedup guard as the pattern five downstream PRs (#1880-#1884) will adopt - shipping the ADR without the fix would have propagated the footgun to each of them. Folded in here per CLAUDE.md's no-defer rule rather than filed. RED verified on 3c4895841 (test only, no fix): linux-node22 reported 'FAIL tests/config-loader.test.cjs - the documented per-process warning-state reset must clear the unknown-key dedup set too'. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2674): document src/ in the changeset-lint trigger list CONTRIBUTING.md presented the Changeset Required trigger list as bin/, gsd-core/, agents/, commands/, hooks/, sdk/src/ - omitting src/, which scripts/changeset/lint.cjs has in USER_FACING_PREFIXES. src/ is the TypeScript source of truth compiled into gsd-core/bin/lib/*.cjs, so it is the most-edited user-facing path in the repo and the omission sends any contributor who touches it into a CI failure the doc says cannot happen. Also documents that the lint reads GITHUB_BASE_REF, which only CI sets, so running it bare locally reports success without evaluating the branch. This PR hit exactly that: a local run said ok_fragment_present and CI failed fail_missing_fragment on the same diff. Found while opening this PR; folded in per the no-defer rule. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2674): add Fixed changeset for the src/ trigger-list and reset fixes Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2674): restore the round-2 review corrections to the amendment These edits were made in response to the second isolated review pass but never staged: later commits used targeted `git add <file>` for the test and the source fix, so the two markdown files stayed dirty and shipped nothing. The branch carried the round-1 text, including the ADR-227 misquote the reviewer raised as a blocker. Restores: the unconditional-diagnostic clause (ADR-227's GSD_DEBUG opt-in was never implemented, so citing it as the precedent was wrong), the dedup key, #1882 folded into the out-of-band mechanism instead of a fourth mechanism-less category, the narrowed caller-audit rationale, and the test-methodology clause. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6d072435d0 |
test(#1975): consolidate 51 CLI + scripts-tooling regression tests into module suites
Fold 51 issue-named CLI black-box + scripts-tooling regression files into their canonical module suites (runtime-launcher-parity, worktree-safety, install-*, managed-hooks, read-guard, capability-registry, etc.), plus a NEW slash-command-namespace.test.cjs grouping the 4 slash/colon-namespace-leak invariant suites that had no canonical owner. Verbatim block-scoped describe wrappers; 427 subtests conserved 1:1. Host-env pre-check (per B2): no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value. One folded suite (bug-3668 runtime resolver) creates an extension-less PATH gsd-tools stub + bash -c; co-locating it with the host's chmodSync tripped local/no-unguarded-nonportable-exec, so it's now Windows-guarded (skip on win32) matching the host suite's own bash -c guard. Regenerates regression-name allowlist (222->182), ratchets file-count allowlist (graphify 7->6, docs entry removed), makes 26 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes stale ids). Repoints 13 tests/ references across CONTEXT.md, COMMANDS.md/FEATURES.md (EN + ja/ko/pt/zh) and ADR-0002. lint:ci green. Part of epic #1969. Closes #1975. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
85ed50cc4f |
test(#1972): consolidate 94 command/module regression tests into subject suites
Fold 94 issue-named command/module regression files into the canonical test file that owns each subject-under-test, across 52 existing suites (state, config, frontmatter, roadmap-parser, capability-registry, shell-command-projection-dispatch, plan-phase-drift-guard, health-validation, runtime-converters, commands, etc.). Verbatim block-scoped describe wrappers; 881 subtests conserved 1:1. No new test files. Host-env pre-check (per B2): the only GSD_WORKSTREAM/GSD_PROJECT-touching destinations (intel, planning-workspace) clear those vars hermetically, so folded CLI tests are safe. Regenerates regression-name allowlist (222->162), ratchets file-count allowlist across 8 buckets (validate entry removed after dropping <=2), makes 34 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes 34 stale ids). Repoints CONTEXT.md + ADR-0002/443/1235/3524 test-file references. lint:ci green. Part of epic #1969. Closes #1972. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
75552f7ea0 |
fix(#1533): prototype-pollution guard in _deepMergeConfig (#1534)
* fix: prototype-pollution guard in _deepMergeConfig (audit M4)
The root↔workstream config merge iterated Object.keys(overlay) with no
__proto__/constructor/prototype guard, while four sibling paths in the same
file (lines ~315/319/331/341/549) guard them. A workstream/root config.json
with {"__proto__": {...}} could pollute the merged object's prototype chain
and spoof unset config flags (per-object, not global Object.prototype).
Adds the same three-key continue guard at the top of the overlay loop plus a
regression test for __proto__/constructor/prototype overlay keys.
Closes a gap missed by the closed config proto-pollution hardening
(#751/#1406/#663).
Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
* chore(changeset): Fixed fragment for #1534 (config proto-pollution guard)
Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
|
||
|
|
484a5b7b86 |
fix(#1415): loadConfig provenance + agent-skills diagnostic (Resolution Provenance P2) (#1424)
Implements ADR-1411 P2 / #1415, closing #1366. Part 1 — config-loader.cts: - Adds `loadConfigResolved(cwd, options) → ConfigResolution { config, source, degraded }` with six tagged branch return paths: A1 ws+wsconfig → source:'workstream', degraded:false A2 no-ws+config → source:'root', degraded:false B ws requested, wsconfig absent → source:'root', degraded:true (intercepts recursive call) C .planning/ exists, no config → source:'builtin-defaults', degraded:false D no .planning/, global defaults readable → source:'global-defaults', degraded:false E no .planning/, no global → source:'builtin-defaults', degraded:false - `loadConfigResolved` calls `findProjectRoot` at entry to anchor resolution to the nearest .planning/ ancestor (cwd-drift fix, heuristic 4 from P1) - `loadConfig` becomes a one-line delegation: `return loadConfigResolved(cwd, options).config` - Exports `loadConfigResolved` in the `export =` block Part 2 — init.cts cmdAgentSkills: - Imports `findProjectRoot` from `./project-root.cjs` - Anchors to project root before loading config (fixes #1366 cwd-drift) - Uses `loadConfigResolved` for provenance; passes projectRoot to buildAgentSkillsBlock - Computes `configured` + `reason` (AgentSkillsReason enum): 'resolved' | 'not_configured' | 'configured_empty' | 'configured_unresolved' - configured_empty and configured_unresolved emit stderr WARNING; not_configured is silent - --json IR gains: configured, reason, source, degraded (in addition to existing agent_type, block, skills_count, warnings) Tests (TDD): - tests/config-loader.test.cjs: 8 new provenance tests (RED before impl, GREEN after) - tests/agent-skills.test.cjs: 7 new diagnostic tests (RED before impl, GREEN after) - All 306 tests across 4 suites pass (config-loader:33, agent-skills:72, init:105, workstream:96) Docs: - CONTEXT.md: Config Loader Module entry updated with loadConfigResolved interface; Resolution Provenance entry notes P2 is now implemented - docs/CLI-TOOLS.md: --json field reference table added for agent-skills Closes #1366 Part of #1411 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c76827afbc |
refactor(#1291): T6 — migrate test files off the core spine ahead of deletion (#1293)
The convergence lint only scanned src/ + gsd-core/bin, so ~35 test files still imported core.cjs. Repoint all 33 behaviour importers to the leaf modules directly (same symbol->leaf map as the src migration; leaves are the objects core re-exported by reference), delete the now-meaningless shim-identity describe blocks in the 8 leaf tests, and delete tests/core.test.cjs (forwarded-behaviour coverage now lives at the leaves; resolveWorktreeRoot test relocated to worktree-safety in T0) and tests/lint-core-spine-imports.test.cjs (the lint is removed in T-final). Dropped the stale core.test.cjs entries from the allow-test-rule-refs allowlist; eslint-rules RuleTester fixture path pointed at io.cjs. After T6: ZERO test imports core.cjs. core.cts still builds (now fully unused); T-final deletes it. No behaviour change. Closes #1291 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a74e71b049 |
refactor(#885): extract loadConfig cluster into config-loader.cts (#886)
ADR-857 rollout phase 2e — the largest core.cts extraction. Move the configuration-loading subsystem (loadConfig + _getConfigDefault/ _getNestedConfigDefault/CONFIG_DEFAULTS/_deepMergeConfig, isGitIgnored + _gitIgnoredCache, _warnUnknownProfileOverrides + RUNTIME_OVERRIDE_TIERS + the dedup Sets, _resetRuntimeWarningCacheForTests) out of core.cts into a new leaf module src/config-loader.cts. core.cts re-exports the public surface (loadConfig, isGitIgnored, CONFIG_DEFAULTS, RUNTIME_OVERRIDE_TIERS, _resetRuntimeWarningCacheForTests); 12+ callers unchanged. Cycle-free: config-loader imports only leaves (configuration, config-schema, planning-workspace, shell-command-projection, core-utils, model-catalog). All core-internal helpers loadConfig touches moved with it to avoid a cycle. core keeps CANONICAL_CONFIG_DEFAULTS for its model-resolver functions, which now resolve loadConfig via the binding — this unblocks the final model-resolver extraction (2f). core.cts: 1275 -> 792 lines. Repointed tests/config-field-docs.test.cjs (a docs-parity source check) to read the CONFIG_DEFAULTS literal from its new home (config-loader.cjs). New-CLI-module checklist done (.gitignore, eslint, INVENTORY 95->96 + row, manifest, ARCHITECTURE, CONTEXT.md "Config Loader Module"). Adds tests/config-loader.test.cjs (27 tests: behavioral + shim-identity + adversarial config fixtures). Gates: lint, code-review, security-review (prototype-pollution guard confirmed intact), codex adversarial-review (0 findings; byte-identical move). Mac 4115 pass; clean-build docker: full-suite hit the local mirror's known incremental-tsc non-determinism on an unrelated re-exported symbol (findPhaseInternal, from already-merged 2d), but a clean targeted rebuild of the affected file passed 163/0 — CI's clean full matrix is the authoritative gate. Closes #885 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |