ae40529d31d7ccf1921077aba7848e57d34e7e7c
993 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0a0905705a |
fix(#4256): resolve todos from the root via todosDir everywhere (#4479)
* test(#4256): pin todos as root-scoped under workstreams (RED) * fix(#4256): resolve todos from the root via todosDir everywhere * chore(#4256): changeset fragment (pr number to backfill) * chore(#4256): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
6ebe6372ce |
fix(#4243): anchor stateReplaceProgressPercent bold form to line start (#4474)
* fix(#4243): anchor stateReplaceProgressPercent bold form to line start The bold branch of stateReplaceProgressPercent carried no ^ and no /m flag, so a bold percent-ish label quoted MID-SENTENCE inside prose — an Accumulated Context bullet mentioning **Progress:** — captured the machine-segment rewrite and destroyed the rest of its line, silently, while the real Progress line stayed stale (and the frontmatter moved on without it, breaking the #4213 surfaces-agree contract). Every caller (cmdStateUpdateProgress, syncCore's percent arm, applyPostSyncPreservation) feeds the whole document, so all three were exposed. Anchored to ^([ \t]*\*\*Progress:\*\*[ \t]*)([^\r\n]*)$ with /im — the exact idiom #4453 applied to stateReplaceField's bold branch (same-line confinement per #4010: the leading class is [ \t]*, deliberately not \s*, which can consume the newlines before the label into the match; $ is explicit-and-inert and documents end-of-line). #2177's recorded requirements all stand: frontmatter is stripped before matching, the suffix-preserving machine-segment swap is untouched, and bold-beats-plain priority now governs line-start forms, so an earlier free-text plain Progress: line still cannot capture the rewrite ahead of the real bold status line. Per the maintainer ruling (2026-09-07), #2177's incidental bold-anywhere matching was not load-bearing. * test(#4243): scope the C4 region check with splitLines, not a bare \n split lint:ci (local/no-crlf-fragile-split) flagged the free-text-plain-line row's content.split(/\n## /)[0] — a bare \n split on readFileSync content is CRLF-fragile under Windows autocrlf. Same scoping via splitLines() (src/text-lines.cts), which splits on \r?\n. * chore(#4243): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
c4b6dbd486 |
fix(#4247): refuse update-plan-progress on a roadmap with no writable phase entry (#4468)
* test(#4247): failing-first regressions for checklist-form update-plan-progress * fix(#4247): refuse update-plan-progress when the roadmap has no writable phase entry * fix(#4247): single local source for the phase-heading anchor grammar * docs(#4247): note the missing_phase_details refusal in cli-tools reference * docs(#4247): backfill pr number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
33e393ba4c |
fix(#4243): anchor stateReplaceField bold form; pin frontmatter round-trip (#4453)
* test(#4243): failing-first regressions for bold-field anchoring and frontmatter round-trip * fix(#4243): anchor stateReplaceField bold form to line start The bold branch of stateReplaceField carried no ^ and no /m flag, so a bold label quoted mid-sentence inside prose — the issue's **Status:** inside an Accumulated Context bullet — captured the rewrite and destroyed the rest of its line, silently, whenever a whole-body caller fed the function every section (beginPhaseCore's tryField, advancePlanCore's Status/Current Plan writes). The plain branch was always line-anchored; only the bold branch lagged. Anchored to ^([ \t]*\*\*Field:\*\*[ \t]*) with /im, reusing #4010's same-line confinement idiom for the leading class (deliberately not the issue's suggested ^\s* — it can consume the newlines before the label into the match) and #4186's recognition-by-anchoring discipline. Frontmatter half of the issue (unknown-key drops, invented milestone defaults) is already fixed on next by #2202/#3216/#4129; pinned here with the issue's requested regression fixtures. * test(#4243): pin survival contract, not derived percent, in frontmatter rows Bench RED run caught two assertion defects in the pin rows: the unknown progress subkey re-parses as a quoted scalar ('77' vs 77), and percent is a declared derived subkey - omitted under the #3573 no-roadmap withhold, recomputed when measured (#4129) - so pinning its value over-pins derived semantics. The rows now pin what the issue demands: unknown/custom keys survive, stored counters are kept under the withhold, milestone identity is never reset to invented defaults. * chore(#4243): changeset for the anchored bold-field fix * chore(#4243): backfill PR number in changeset |
||
|
|
8c8eda46b0 |
fix(#4225): scope the sibling-worktree phase-number horizon to the active workstream (#4450)
* test(#4225): failing-first matrix for phase.add --ws workstream-scoped numbering Nine rows driven through the real CLI: the issue's verbatim topology (root roadmap @39 committed, workstream @2, sibling git worktree carrying the root roadmap), same-workstream sibling boundary, empty-workstream first phase, coincidental root-maximum, no---ws control (the #3849 global horizon, byte-for-byte), cross-workstream isolation, sibling lacking the workstream (fail open), add-batch parity, and a next-decimal control. Rows 1/2/3/6/7/8 are RED on next @38e4ce5f62 (numbering computed from the sibling ROOT roadmaps: 40 instead of 3). * fix(#4225): scope the #3849 sibling-worktree widening horizon to the active workstream collectSiblingWorktreePhaseNums scanned each sibling git worktree's ROOT .planning/ (phases/ dirs + ROADMAP.md headers) unconditionally. Under --ws (GSD_WORKSTREAM), every local number source flows through planningDir(cwd) and lands in the workstream scope, but the widening horizon still merged the siblings' ROOT-roadmap numbers into it — so phase.add --ws in a workstream at Phase 2 inside a project whose root roadmap sits at Phase 39 minted Phase 40 (directory 40-<slug>, and a Depends on: Phase 39 that does not exist in the workstream's numbering universe). The horizon now resolves each sibling's planning dir through the SAME canonical resolver, planningDir(wt, ws), with the env workstream read once via planningDir's own discriminator: a workstream-scoped allocation scans the sibling's copy of the SAME workstream (a number taken by that workstream on another branch is still taken — the #3849 widening survives, scoped), and never the sibling's root roadmap or another workstream's. No workstream active: ws is null and the root-scope horizon is byte-for-byte the #3849 behavior. A sibling lacking the workstream directory contributes nothing (fail open, unchanged). phase.add and phase.add-batch share the helper; both scopes of both verbs are covered by the matrix in the previous commit. Output shape and the publishStateContract boundary are untouched — only the number changes. * fix(#4225): rename siblingPlanning -> siblingPlanningDir (review nit) * chore(#4225): changeset fragment (pr number to backfill) * chore(#4225): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
e54d3aa159 |
enhance(#4401): register workflow.compact_content as a validated config key (#4441)
* feat(#4401): register workflow.compact_content as a validated config key - Add compact_content: false to the nested workflow object in gsd-core/bin/shared/config-defaults.manifest.json - Add 'workflow.compact_content': false to SCHEMA_DEFAULTS in src/config.cts so an absent key resolves to false via config-get --raw - validKeys entry in config-schema.manifest.json already present Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(#4401): behavioral and boundary tests for workflow.compact_content - 19 behavioral tests covering config-set/config-get round trip, invalid-shape rejection (banana, 42, empty string), the corrected null-unset semantics (#2046), absent-key resolution against config-defaults.manifest.json, config-new-project wiring, and doc-row shape assertions - Drops the install-tree fixture-parity block (and its docstring item) that asserted gsd-core/references/compact-content-gate.md and gsd-core/workflows/compact/map-codebase.md fixture entries — those paths belong to #4402 and do not exist on this filtered branch Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(#4401): document workflow.compact_content in both config references - One 4-cell row in docs/CONFIGURATION.md (workflow.* run) - One 5-cell row under Workflow Fields in gsd-core/references/planning-config.md - Both cross-reference ADR-4139 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(#4401): add changeset - Added-type fragment, pr: 4401 (issue number; backfill to the real PR number is a required follow-up once the PR is opened, per D-08 and CHANGESET-PR- FIELD-DRIFT) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(#4401): backfill changeset pr field to #4441 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(#4401): derive workflow.compact_content default from CONFIG_DEFAULTS SCHEMA_DEFAULTS['workflow.compact_content'] hardcoded the literal false instead of deriving it from CONFIG_DEFAULTS the way 3 of its 8 sibling entries do (smart_zone_tokens, pr_strict, inline_plan_threshold), leaving a single-source-of-truth drift risk: a future manifest-only edit to the default could silently diverge from this literal, only caught later by the D-03 test if it ever happened to manifest. Adds compact_content to CONFIG_DEFAULTS in src/config-loader.cts and derives SCHEMA_DEFAULTS from it in src/config.cts, matching the majority sibling pattern. Found during maintainer review (review-open-prs) of this PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4401): map compact_content in config-field-docs NAMESPACE_MAP The previous commit added compact_content to CONFIG_DEFAULTS in src/config-loader.cts but missed the matching entry in tests/config-field-docs.test.cjs's NAMESPACE_MAP, which maps flat CONFIG_DEFAULTS keys to their namespaced doc form before checking gsd-core/references/planning-config.md for a match. Without it, the test looked for a bare `compact_content` doc reference instead of the actual `workflow.compact_content` row, and failed: "CONFIG_DEFAULTS keys missing from planning-config.md: compact_content". Found by actually running gsd-test against the branch rather than trusting the plausible-looking fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4401): register compact-content-4139 test in the docs-guard lane tests/compact-content-4139.test.cjs's D-06 tests read docs/CONFIGURATION.md directly (fs.readFileSync) to assert the workflow.compact_content doc row's shape, which makes it a doc-reading test file under the #3753 docs-guard lane. It was never added to scripts/docs-guard-registry.cjs's DOCS_GUARD_TESTS map and carries no docs-guard-exempt marker, so tests/ci-docs-guard-registry.test.cjs's registration lint correctly failed: "compact-content-4139.test.cjs reads a docs/ path but is not registered in the docs-guard lane and carries no docs-guard-exempt marker". Registers it with ['docs/CONFIGURATION.md'] (the only real docs/-prefixed path it reads; gsd-core/references/planning-config.md is outside this registry's docs/ scope, matching the sibling config-field-docs.test.cjs entry's existing convention). Found by actually running gsd-test against the branch — this gap predates the maintainer's config-loader.cts fix and was already present in the original PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> Co-authored-by: sim <sim@local> |
||
|
|
38e4ce5f62 |
fix(#4186): anchored status vocabulary, record-session arg guard, recount pin (#4381)
* fix(#4186): anchored status vocabulary, record-session arg guard, recount pin Three defects from #4186: 1. normalizeStateStatus ran a first-match-wins SUBSTRING chain over the free-prose body Status field, so prose merely mentioning a status word was silently rewritten to a credible wrong token (a .planning/ path in Italian prose -> status: planning; verifica -> verifying; completezza -> completed). Recognition is now an ANCHORED whole-field match against a declared vocabulary (STATUS_EXACT_TOKENS + STATUS_ANCHORED_PATTERNS, state-document.cts) — case/whitespace-tolerant, branch-order artifacts preserved (Planning complete -> planning; Phase complete — ready for verification -> verifying). The recorded lenient fallback (#3873 row 26) stands: unrecognized prose passes through verbatim. Read-side consumers (W011, statusline) ride the same function. 2. The progress recount skew (stray *-SUMMARY.md inflating completed_plans) is already dead on next via #1988/PR #2016 (countMatchedSummaries pairs summaries to plans) — verified live and pinned with regression rows composed against the #4129/#4359 ratchet. 3. state record-session with no args executed and wrote STATE.md; it now errors like state update (stopped-at or resume-file required), handler- side so SDK callers are covered too. Four tests pinning the bare-call write are updated to the new contract. * fix(#4186): update status pins to the anchored vocabulary contract Bench round 1 follow-ups: - Legacy bare 'Milestone complete' kept as reader-side vocabulary (ADR-2207 removed the writers, not recognition of legacy files). - state.test pins updated: 'Paused at Plan 3' and round-trip 'Executing Plan 5' were pins of the substring guessing itself — the round-trip now uses the real handler form 'Executing Phase 5'. - record-session no-op/no-fields tests repurposed to the usage-error contract (CLI + SDK-level ExitError), byte-unchanged assertions kept. - statusline tests repinned: vocabulary values collapse to keywords; narratives render the documented first-word fallback instead of a guessed token. Hook doc comment updated to match. - docs-guard exempt baseline: state.test.cjs now cites docs/CLI-TOOLS.md. - docs/CLI-TOOLS.md: record-session signature notes the required flag. * fix(#4186): repair a dangling sentence in the schema docstring * test(#4186): bound the completed_plans scan regex (#2128 class) * chore(#4186): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
f09e7ed08c |
fix(#4137): existsSync-guard the Homebrew Cellar rewrite in normalizeNodePath (#4375)
* test(#4137): keg-only Homebrew Cellar path falls back to raw execPath Regression tests for the Homebrew branch of normalizeNodePath: the rewrite to <prefix>/bin/node must be existsSync-guarded like the mise/volta branches, falling through to the raw execPath when the keg-only formula was never linked into <prefix>/bin. Also makes the existing #3181/#2185 Cellar assertions hermetic by injecting existsSync stubs (granting existence to exactly the one candidate each asserts) so they no longer depend on the runner machine's real /usr/local/bin/node. * fix(#4137): existsSync-guard the Homebrew Cellar rewrite in normalizeNodePath The Homebrew branch of normalizeNodePath returned <prefix>/bin/node unconditionally — the only one of five runtime branches that never probed its rewrite candidate. On a keg-only or versioned Homebrew install (node@24 never brew-linked) that path does not exist, so every managed hook command baked by resolveNodeRunner/buildBakedNodeToken/ buildNodeRunnerChainToken failed at invocation with exit 127, /bin/sh: <prefix>/bin/node: No such file or directory. Guard the rewrite with the already-injected existsSync exactly like the mise and volta branches: when <prefix>/bin/node exists (linked formula) the rewrite is byte-identical to today; when it does not, fall through to the raw execPath — a working keg path instead of an immediately broken one. Also drops two now-unused constants from the regression tests. * test(#4137): make the #977 non-fnm Cellar assertions hermetic too The Bug #977 folded block's two 'still maps to stable symlink' assertions called normalizeNodePath without an existsSync stub, silently depending on the runner machine's real /usr/local/bin/node (present on the Linux bench image, absent for /opt/homebrew). With the #4137 guard these become environment-dependent; grant each exactly the one candidate it asserts. * chore(#4137): add changeset fragment * chore(#4137): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
54085516c1 |
fix(#4211): materialize Kimi's agent tree recursively during surface apply (#4371)
* fix(#4211): materialize Kimi's agent tree recursively during surface apply kimiAgentsKind stages `gsd.yaml` + `gsd.md` + `subagents/gsd-*.{yaml,md}`, and install copies that tree recursively (_copyStaged). Surface apply fell through to _syncGsdDir's flat command/agent branch, which reads only top-level `*.md`: the YAML half and the whole subagents/ subtree were ignored, and `gsd.md` was written as `gsdgsd.md` because the flat branch re-applies kind.prefix to a name that already carries it. A surface change could therefore corrupt Kimi's installed artifacts while still reporting success. Three divergences from the install path, all in src/surface.cts: - _syncGsdDir gains a kimi-agents branch: recursive copy, then a prune scoped to exactly what install's _removeGsdEntries owns for this kind (the two root files, and gsd-*.{yaml,md} under subagents/). Everything else is user-owned and preserved. - applySurface stages kimi-agents WITH agentCtx and the `skills: '*'` rule for an unmodified full profile, as it already does for the agents kind and as createRuntimeArtifactInstallPlan does for every kind — without it Kimi's generated subagents lost their path-prefix rewrites and attribution trailer, and an unmodified full profile staged only the skill-referenced subset. - applySurface runs rewriteStagedSkillBodies for kimi-agents, which the install plan routes through it alongside skills. * chore: add changeset for #4211 --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
acb3cc974b |
fix(#4197): dedup the update-context fast path against the selected global dir (#4413)
* fix(#4197): dedup the update-context fast path against the selected global candidate The preferredConfigDir fast path derived scope from a cwd-relative match alone, so a global install reported LOCAL whenever the shell sat in $HOME — and run_update then drove the installer through its --local arm (settings.local.json + the #338 relocation) against a global install. Extract resolveGlobalCandidate (env candidates first, then $HOME-relative, first hasInstall hit wins) and use it in BOTH paths: the fast path now answers LOCAL only for a cwd-relative match that is not the selected global dir, which is the same dedup the cascade applies at its isLocal check. A preferred dir that is also the env-directed global now answers GLOBAL on both paths (the cascade's answer), pinned by a parity test. The discriminator is the selected global candidate, not the $HOME pathname: with CLAUDE_CONFIG_DIR directing the global elsewhere, $HOME/.claude probed from cwd === $HOME is a genuine local install, and a pathname check would re-break parity (regression-pinned). * chore(#4197): add changeset * chore(#4197): backfill PR number in changeset --------- Co-authored-by: agent-4197 <agent-4197@gsd.local> |
||
|
|
b7917882bb |
fix(#4398): render the pending-todo bullet link repo-relative (#4416)
* test(#4384): failing-first regression rows for the macOS long-base todo-cap failure The 240-char pending-todo bullet cap must be deterministic w.r.t. where the repo is checked out. Deterministic long-base-path fixtures (a single 110-char segment, no real macOS dependency) reproduce next's own macos shard 3/3 failure (run 34038716700) on every OS: with an absolute link the bullet exceeds the cap and the documented needs-first truncation drops the 'Needs <solution>' clause. Rows cover the determinism property (byte-identical bullets under short and long bases), the CLI surface, relative-path stability, legacy no-projectRoot behavior, drop-order preservation, and adversarial edges (outside-root, path===root, non-string path). * fix(#4384): render the pending-todo bullet link repo-relative renderPendingTodosMarkdown gains an optional projectRoot; when given and the todo's path is absolute, the bullet's [todo file](…) target becomes toPosixPath(path.relative(projectRoot, path)) — the idiom already used for project_exists. cmdInitTodos passes cwd. The JSON todos[].path field stays absolute (#2376). Only the rendered display link changes: embedding the machine-variable absolute base let macOS's /private/var/folders/… temp paths consume the 240-char budget and drop the 'Needs' clause on long-path machines only — next's own macos-latest shard 3/3 went red on exactly this (run 34038716700), Linux's short /tmp passed. The 240-char whole-bullet cap and the needs→title→area drop order are unchanged; this matches PR #4384's own canonical example, docs, and unit tests, which all show repo-relative links. Docs updated at all three surfaces that describe the bullet (COMMANDS.md, templates/state.md, reference/state-md.md — the last was still pre-#4384 'count and reference' prose). Fixes the macOS regression introduced by #4384; next is red on its own CI. * test(#4384): fix substring false positive in the outside-root regression row The ../-form relative link legitimately contains the absolute path as a substring, so !line.includes(absolutePath) fired on correct output (caught by the first remote verify run, linux-node24 44018/44019). Assert the property itself instead: extract the link target and require it to be non-absolute and not equal to the absolute path. * chore(#4398): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
708d9a0b82 |
fix(#4187): status reader resolves a bare VERIFICATION.md like resolve-file (#4388)
* test(#4187): bare VERIFICATION.md regression matrix for the status surface Both query verbs must agree on every row: bare file, suffixed variants, missing file, other-dir placement, and the staleness seam. Row 1 is the failing-first regression from the issue repro. * fix(#4187): status reader resolves a bare VERIFICATION.md like resolve-file readVerificationStatus and its internal staleness check (findStaleVerificationSummary) called the shared resolver without allowBare, so a phase whose only report was a bare VERIFICATION.md read as missing and was told to re-run execute-phase while verification.resolve-file, determinePhaseStatus, and both init verification_path projectors all resolved the same file. Both call sites now pass allowBare: true, matching the other five; tier order (dashed > bare) is unchanged, so only bare-only directories change behavior. * fix(#4187): correct call-site counts in allowBare docblocks Adversarial review caught the comments claiming five of six call sites opted in; the current tree has six call sites with four previously passing allowBare — the two module-internal status-path sites were both holdouts, not one. * chore(#4187): changeset for the bare VERIFICATION.md status fix * chore(#4187): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
fd4aac5670 |
fix(#4192): honor explicit model pins on the claude runtime (#4396)
* fix(#4192): honor explicit model pins on the claude runtime Two documented model-configuration contracts did not hold on the claude runtime (confirmed-bug scope from the issue triage): Finding 1 — model_profile_overrides.claude.<tier> was inert. Step 3 of resolveModelInternal gated runtime-aware tier resolution on configRuntime !== 'claude', so the key's only reader was never consulted, while workflows/settings-advanced.md writes it for claude-runtime users. A new step 4.5 resolves ONLY the user's override entry (never the builtin claude tier map, so unpinned installs keep resolving aliases). An override value that maps to a current tier alias collapses to that alias (byte-equivalent, the #2041 protection); anything else — a pinned older generation, a bare alias repoint, a non-Anthropic id — resolves verbatim. It sits after the resolve_model_ids:'omit' gate so an explicit project omit still wins (#2297) and before the alias return so resolve_model_ids:true cannot re-materialize the pin to the latest id. Finding 2 — fully-qualified claude-* ids in model_overrides were warn-dropped to tier resolution (mapClaudeOverrideForRuntime unmappable branch, #2041), while the docs promise any fully-qualified model id is valid. The unmappable branch now passes the pin through verbatim with a warn-once breadcrumb (text describes the pass-through). Dropping it silently unpinned the operator's explicit choice — the exact 'profile can misrepresent what actually runs' defect of #4192. Mappable ids and non-claude values behave exactly as before; resolveModelForTier shares the mapping; the tier honesty signal is unchanged (raw ids still report 'unknown'); the model_policy path is untouched. Docs updated to the agreed contract (CONFIGURATION.md false 'Claude example' corrected; how-to + shipped reference document the pin semantics, the fable alias, and the tier-override composition). * test(#4192): pin explicit model pin resolution on the claude runtime 28 failing-first rows across the resolver seam and the resolve-model CLI: pinned-generation fidelity (tier override + per-agent verbatim pins, object form, explicit runtime), unpinned controls byte-stable (no override, other runtime/tier, inherit, project omit, precedence), adversarial rows (prototype-chain keys, malformed values, warn-once dedupe, 64-char stderr cap), and behavioral AC1/AC2 rows through runGsdTools. The stale #2041 fall-through assertions now pin the pass-through contract; mappable-id collapse assertions unchanged. * chore(#4192): add changeset fragment * chore(#4192): backfill PR number in changeset fragment --------- Co-authored-by: ZCode <zcode@localhost> |
||
|
|
b7406b293f | enhance(#2618): render pending todos as one bounded bullet per todo (#4384) | ||
|
|
66e4034fe4 |
fix(#4138): begin-phase without --phase exits non-zero and writes nothing (#4380)
* test(#4138): failing-first regression — begin-phase without --phase must fail closed * fix(#4138): begin-phase without --phase exits non-zero and writes nothing * chore(#4138): changeset fragment for begin-phase arg validation * chore(#4138): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
03738824de | enhance(#2586): stop installing Codex context-monitor hooks without metrics (#4367) | ||
|
|
0aa4202f6a |
fix(#4135): headline baseline coverage, opt-in strict gate, git-history widening (#4376)
* test(#4135): regression rows for pristine regen coverage collapse RED skeleton: src/pristine-baseline.cts exports findPristineInGit as a null-returning stub (wired into verifyFile after the #4145 orphan tier, behavior-neutral) so the git-history rows fail behaviorally, not at require time. Failing-first rows: baseline_covered aggregate on a 1-of-13 multi-version fixture, coverageHeadline typed renderer, the opt-in --min-baseline-coverage gate (exit 3, >= threshold semantics, vacuous-pass and malformed-value boundaries), git-history baseline recovery (dropped-line catch + surviving-line verify + older-commit hop), findPristineInGit unit, Step 5a workflow headline contract, and the installer-side describeBaselineCoverage honest N-of-M summary with the collapse disk-state pinned. Negative-space rows pin today: non-git ok_no_baseline posture, no-match-no-adoption, #3657 drift never rescued, canonical precedence, and no git tier without --pristine-dir. * fix(#4135): headline baseline coverage, opt-in strict gate, git-history widening The #3407 promotion rule regenerates gsd-pristine/ baselines from the INCOMING release source and keeps only candidates byte-identical with the OUTGOING recorded hash — correct in isolation, but on a multi-version jump the surviving set is precisely the files upstream did NOT change. The verifier then reports ok_no_baseline (advisory, exit 0) for everything else, and no surface distinguishes a 12-of-13-unverified green run from a fully-verified one: the human summary printed Checked/Failures only, the JSON had no coverage aggregate, and the installer's update output gave per-bucket counts without N-of-M framing. All three issue directions, none exclusive: - Report coverage prominently: --json gains an additive baseline_covered aggregate; the human summary leads with 'Baseline coverage: N of M file(s)...' on every run plus an advisory section naming each skipped file and reason; the installer prints an honest covered-of-modified line via the exported describeBaselineCoverage helper (typed return, exact contract); workflow Step 5a computes and prints the headline before any pass/fail framing. - Fail louder on low coverage: opt-in --min-baseline-coverage <0..1> exits with new documented code 3 when coverage falls below the threshold (>= semantics; empty run vacuously passes; content failure exit 1 outranks it; malformed values are usage errors, exit 2). Default posture unchanged — no_baseline stays advisory per #934. - Widen the promotion rule (its only trustworthy form): when no baseline resolves under gsd-pristine/ and a hash is recorded, the verifier now recovers the baseline from the config dir's own git history — the workflow's documented Option A — anchored by the same authority every tier trusts, exact pristine_hashes sha-256 equality. Read-only (git log/git show, windowsHide per #685), bounded (100 commits/file, 10s/subprocess), null-on-any-failure so ok_no_baseline remains the universal fallback. Tier order: canonical join -> #4145 orphan scan -> git history -> OK_NO_BASELINE; #3657 drift and canonical precedence untouched. Hash validation in saveLocalPatches is NOT relaxed — the collapse is legitimate conservatism; hiding it was the bug. Measured on the issue's shape (13 files, 12 changed upstream, 1.10->1.12): non-git installs report baseline_covered 1/13 with the headline and can gate at exit 3; a git-managed config dir with the outgoing bytes in history verifies 13/13. Review fixes folded in: workflow headline derives the unverified count from checked - baseline_covered (not the drift+no_baseline sum), and the new site-scoped allow-test-rule annotation carries its ADR-456 see-ref on the marker line. Emitted-Drift-Ack-Growth: reapply-patches.md — #4135 — +20 lines / ~1.5 KB, prose and bash only: two additive parse lines (BASELINE_COVERED, CHECKED_COUNT), a Step 5a coverage-headline block printed BEFORE any pass/fail statement (documents the opt-in --min-baseline-coverage exit-3 gate), and one Option B sentence noting the verifier's read-only git-history fallback. No step ordering, gate, tool-invocation, or dispatch shape changed; 5a's fail/drift/advisory handling is unchanged, the headline only precedes it. * chore(#4135): backfill PR number into changeset fragment --------- Co-authored-by: agent-4135 <agent-4135@gsd.local> |
||
|
|
7bb366e836 |
fix(#4130): --context flag for check decision-coverage-plan + parseDecisions quadratic-backtracking hardening (#4374)
* test(#4130): failing-first regressions for --context flag + parseDecisions hardening Block A (flag): check decision-coverage-plan --context <path> must route identically to the positional form; flag wins over positional context; valueless --context falls through to the #2770 fail-closed caller error; verify keeps its positional surface (flag is plan-only). RED on base: the flag token lands in the args[2] phase slot (false uncovered) or the args[3] context slot (silent CONTEXT.md-missing skip). Block B (hardening): regex-lattice asserts pin the atomic-ID wrapper (?=(X))\1 and the em-dash first-separator narrowing [^*—–]*[—–] plus the no-adjacent-overlap property; a differential property compares the module against a frozen copy of the pre-hardening grammars (reference validated against the base build: 60k generated lines, 0 mismatches); 40k cliff shapes assert correct outcomes with no wall-time asserts (repo rule). A12: partitionPredicateArgs keeps one parser behind parsePredicateFlags. * fix(#4130): --context flag for check decision-coverage-plan + quadratic-backtracking hardening in parseDecisions (A) check decision-coverage-plan --context <path> — sibling convention (check predicate, #2008): --flag value pairs parsed by the new shared partitionPredicateArgs (parsePredicateFlags reimplemented as its flags half — one parser, cannot diverge), the flag winning over a same-purpose positional, positionals kept (no sibling deprecates them; the plan-phase workflow caller passes positionals), valueless --context falls through to the #2770 fail-closed caller error. Repair of the routing accident where --context landed in the args[2] phase slot (false uncovered) or the literal token in the args[3] context slot (silent green skip). (B) parseDecisions regex seam hardened, byte-identical on all legal inputs: the three bullet grammars consume the ID atomically via the (?=(X))\1 lookahead emulation (kills the tail/[^:*]* O(n^2) re-split, ~1.1s @ 40k), and the em-dash first separator narrows [^*]*[—–] to [^*—–]*[—–] (kills the dash-position O(n^2) retry, ~1.7s @ 40k). Group indices unchanged (handlers untouched). Pinned by regex-lattice tests, a differential fast-check property vs the frozen pre-hardening grammars, and 40k cliff/legal-shape outcome tests (no wall-time asserts per repo rule — no deterministic engine step counter exists in Node). * docs+test(#4130): document --context invocation; harden lattice test tooling - docs/CONFIGURATION.md Decision Coverage Gates: new 'Invoking the plan gate directly' block documenting both the positional and --context forms, flag precedence, and the valueless-flag fail-closed semantics (same place the gate's behavior is documented; sibling check predicate documents its flags the same way). - Two changeset fragments per the maintainer brief (Added: flag; Fixed: hardening), PR numbers to be backfilled. - tests/decisions.test.cjs review fixes: readRegExpTemplate template escaping (bare ')' SyntaxError), range-aware lattice checker with backreference skip and template unescape, honest A1 contract, lint escape warning. * fix(#4130): valueless --context fails closed per #2770; A8 isolates flag-vs-positional context Suite-caught fixes from the first verify run: - cmdDecisionCoveragePlan now refuses a flag-shaped token as the positional context path: a bare valueless --context stays a positional (sibling parser semantics, unchanged) but reading it as a PATH would turn a caller mistake into a silent 'CONTEXT.md missing' green skip — exactly what #2770's fail-closed law forbids. Now falls through to the missing-context-argument error, as documented. - A8 test compares decoy-positional+flag against flag-with-phase (phase held constant) so the row isolates WHICH context was read; the old form compared against a no-phase invocation that could never match. * chore(#4130): backfill PR number in changeset fragments (PR #4374) --------- Co-authored-by: sim <sim@local> |
||
|
|
6adf3098ac |
fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans (#4364)
* test(#4145): regression rows for hash-matching prefix-less pristine baselines RED skeleton: src/pristine-baseline.cts exports findPristineByHash as a null-returning stub so the new rows fail behaviorally, not at require time. Failing-first rows: verifier resolution (no_baseline must drop to 0 when an exact-hash orphan exists), findPristineByHash unit row, and the two saveLocalPatches relocation rows. Negative-space rows pin today's behavior: missing baselines still report ok_no_baseline, mismatching orphans are never adopted or deleted, canonical precedence and the #3657 drift posture are untouched. * fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans Both pristine readers joined the manifest-keyed path strictly, so a snapshot stored without the gsd-core/ prefix (an earlier release's writer) was reported as ok_no_baseline by the verifier and pushed into regeneration by saveLocalPatches — where incoming-release candidates can never satisfy the recorded outgoing hash, leaving the correct baseline permanently unconsumed. - src/pristine-baseline.cts (new, ADR-457): shared findPristineByHash — deterministic sorted scan of gsd-pristine/, exact sha-256 equality with the recorded pristine_hashes entry (the same authority the #3657 drift guard trusts), symlink-skipping, canonical path excluded via skipRel. - verify-reapply-patches.cjs verifyFile(): on canonical miss with a recorded hash, adopt byte-identical content found anywhere under gsd-pristine/ before reporting OK_NO_BASELINE. Drift posture (#3657), canonical precedence, and the frozen REASON/report shapes are untouched; the verifier stays read-only. - install.js saveLocalPatches(): preserve-check rescue — relocate a hash-matching orphan to the canonical path (copy, hash-verify, then remove the orphan) so the state self-heals on the next update instead of repeating forever. Honest accounting: new non-overlapping rescued counter. - Workflow doc: one-sentence note on hash-based snapshot resolution. - Derived ripples: INVENTORY-MANIFEST.json regen, eslint ignore + .gitignore entries for the compiled artifact, seedFixture mkdir fix in the new rows. Emitted-Drift-Ack-Growth: reapply-patches.md — one-sentence note on hash-based pristine snapshot resolution (#4145) * fix(#4145): review follow-up — orphan scan never consumes a canonical path Adversarial review finding: with two modified files sharing byte-identical outgoing content, recoverOrphanedPristine could adopt the OTHER file's canonical pristine as its rescue source — relocating it (copy + delete at its home path) and ping-ponging the single baseline between the two files across updates. findPristineByHash's skip parameter now accepts a Set, and saveLocalPatches passes the normalized manifest keys so every canonical path is excluded; only genuine non-canonical orphans are eligible for removal (no strict-join reader ever consults those). Adds the canonical-theft regression row, a Set-skip unit assertion, and tightens the workflow doc sentence the same pass flagged as overstated. * fix(#4145): INVENTORY roster row + symlink-fixture correction Two leftovers from the ab17b7a1e5 bench run, both root-caused: - docs/INVENTORY.md roster row for cli_modules/pristine-baseline.cjs (#3762 gate: every manifest entry carries a row). - The findPristineByHash symlink unit fixture placed its symlink target INSIDE the scanned root, so the walk legitimately matched the real target file. The implementation skips the symlink itself; the fixture now keeps the target outside the scanned tree so the assertion tests what it claims. * changeset(#4145): fixed fragment for pristine baseline hash resolution --------- Co-authored-by: gsd-agent <agent@gsd.local> |
||
|
|
c3e2da153b |
fix(#4134): refuse punctuation-only milestone heading names (#4358)
* test(#4134): fail-first regression — refuse punctuation-fragment milestone names A first-milestone ROADMAP.md H1 that puts the version after the name (# Roadmap: Project — Name (v1.13)) leaves exactly ')' after the heading's own version token, which the ADR-3180 §7.2 pinned name rule returns as a COMPLETE-scope milestone name. Failing-first coverage: - getMilestoneInfo: name-then-version H1 (STATE-anchored + ROADMAP-only fallback) must yield TRUNCATED {version, name: null}, never ')' - the refusal is level-agnostic (H2/H3) - punctuation-family remainders (')', '()', '**', '.,;:', ']}', emoji-only) - listMilestoneHeadings enumerates the heading with name: null - init manager CLI reports milestone_name: null and no lone ')' anywhere - property (seed 20260905, 300 runs): a word-char remainder is always a name, a punctuation-only remainder never is - negative space: canonical delimiter forms, parenthetical names (#3171), trailing markers, digit-only names, CRLF headings, version-last-no-parens control * fix(#4134): refuse punctuation-only milestone heading names extractMilestoneHeadingName returns everything after the heading's own version token as the name (ADR-3180 §7.2 pinned rule), which assumes version-then-name. A name-then-version heading — the H1 a first-ever ROADMAP.md drifts into ('# Roadmap: Project — Name (v1.13)') — leaves exactly ')' after the token, and that fragment was returned as a COMPLETE-scope milestone name, propagating into init.* JSON output and buildStateFrontmatter's STATE.md writes. A remainder with no letter or digit anywhere (any script) is heading structure, not a curated name: refuse it as name: null so callers report the honest §7.2 rule-6 answer (version kept, TRUNCATED scope). Names that merely contain punctuation are unaffected — '(' stays an ordinary name character (#3171) — and digit-only names qualify. Also closes the template gap that lets the shape occur: the roadmapper agent's output_formats now templates the version-free canonical H1 ('# Roadmap: [Project Name]', per templates/roadmap.md) instead of leaving a first milestone's title line to invention. The new section shifts the file's existing bare-gsd-tools prose mention from line 647 to 660, so its line-keyed PROSE_ALLOWLIST entry moves with it. Emitted-Drift-Ack-Growth: gsd-roadmapper.md — deliberate +498 bytes: new '### 0. Top-Level Title (H1)' output_formats section templating the canonical version-free H1, closing the first-milestone template gap that lets an H1 drift into 'Name (vX.Y)' and corrupt milestone_name extraction (#4134) * chore(#4134): add changeset * chore(#4134): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
e6d047decc |
fix(#4129): derive completed_phases from the ROADMAP authority; honor the progress-ratchet on every state write (#4359)
* test(#4129): failing-first regressions — completed_phases clobber on resyncing writes and phase-complete failure to increment * fix(#4129): completed_phases derives from the ROADMAP authority and the write path honors the progress-ratchet Three coordinated prongs (diagnosis in .gsd/bug/fix-4129-completed-phases-recompute/): P1 — buildStateFrontmatter's disk scan floors the completed-phases numerator at the milestone-scoped ROADMAP Complete-row count (deriveProgressFromRoadmap, the one owner), gated inside the same safeToUseRoadmapCount / not-withheld branch that owns the denominator. A completed phase whose verification routes stale (#2348 clean-commit-time drift) or is missing no longer under-counts forever. P2 — applyPreserveAlways's resync arm merges instead of wholesale-replacing on a measured scan: totals derived both directions (#2440), completed counters up-only (#2969 — the schema-declared progress-ratchet, now enforced on the write path like the read path always has), percent recomputed from the merged counters. The #3756 unmeasured guard and the #3242 explicit-progress contract are unchanged. P3 — phase complete's atomic 3-file commit passes the post-completion ROADMAP-derived counters through the #2736 authoritativeFm seam (new object direction for the progress key; completedOnlyRaise at the post-preservation re-assert), because the transaction's disk scan reads the pre-completion ROADMAP and failed to increment on the completing phase's own write. * fix(#4129): adversarial-review hardening — intent is a floor at BOTH authoritativeFm sites The pre-preservation merge could lower a correctly-higher disk-derived counter (a verification-passed phase whose ROADMAP table row drifted behind the disk signal). completedOnlyRaise now governs both application sites: the intent and the derivation agree on direction (up), never on subtraction. * fix(#4129): the ratchet merge keeps derived values verbatim when numerically equal The re-parsed derived block carries string scalars ("2") while the curated snapshot carries numbers (2); substituting the curated spelling over an equal derived one was a no-op in substance but a shape churn the ADR-3473 §8.7 reporting loop surfaced as a phantom preserved-over-disagreeing-derived warning on phase complete (ADR-3408 §8.5 Matrix B). Only a strictly-greater curated counter replaces the derived value now; percent gets the same verbatim rule. * changeset(#4129): backfill PR 4359 --------- Co-authored-by: sim <sim@local> |
||
|
|
06eba5fdb0 |
fix(#4130): parse phase-prefixed decision IDs (D4-01) (#4357)
* test(#4130): failing-first regression for phase-prefixed decision IDs Add the #4130 matrix: D4-01/D12-01 across all three bullet forms, tags, discretion, wrapped lead-ins, gate-level plan/verify end-to-end rows, and parity properties (well-formed digit-prefixed ids parse to their exact id; a non-digit injected into the prefix fails loud). Update the #2347 non-D-prefix fixture from D5-NN (now a legal grammar) to DEC-NN, and graduate the representative d5-prefix corpus fixture from could-not-parse to parsed-but-uncovered. All new rows are RED against origin/next; they go green with the parser fix in the next commit. * fix(#4130): parse phase-prefixed decision IDs (D4-01) The three declaration grammars, the parse-miss guard, the #3939 join regexes, and the token evidence all anchored on the literal 'D-' (or '**D-'), so an ID carrying a digit-run phase prefix between the leading letter and the hyphen matched nothing — while the #2347 shape detector correctly called those bullets decision-shaped, collapsing the whole CONTEXT.md to could-not-parse with 0 extracted instead of a coverage verdict. Derive the extractor ID grammar from one shared DECISION_ID_SOURCE ('D[0-9]*-' + the existing alnum tail, full id captured), widen the guard/join anchors to ID_ATTEMPT_SOURCE (bare 'D-' or a digit-initial prefix run, so a typo'd 'D4x-01' fails loud while letter-initial prose like 'Deferred-until' stays none-present), and align the bare-token evidence. Both gates and the gap-checker share the parser, so all three surfaces read phase-prefixed decisions now; the gate messages name the accepted forms including the phase-prefixed one. * docs(#4130): document the phase-prefixed decision identifier form The canonical CONTEXT.md reference said decisions carry 'a sequential D-NN identifier' with no mention of the optional phase-number prefix the parser now accepts (D4-01) or the alphanumeric tail it always accepted (D-INFRA-01). Name both in the Decision identifier format section, EN and ja-JP. * chore(#4130): changeset * chore(#4130): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
0be5bf865a |
enhance(#3783): audit-uat summary segments current-milestone vs archived debt (#4336)
* test(#3783): add failing coverage for audit-uat summary segmentation Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3783): segment audit-uat summary into current_milestone and archived buckets Additive: current_milestone/archived are new; total_items, total_files, parse_gap_files, by_phase, and by_category are unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3783): add changeset fragment for audit-uat summary segmentation Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3783): allowlist the new audit-uat-summary-segmentation test file lint-test-file-count.cjs baselines the "audit" module (keyed off bin/lib/audit.cjs) at 6 pre-existing files; this adds the new dedicated suite as a 7th, matching the module's existing one-file-per-feature-slice precedent. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#3783): fix phase/file number mismatch in the mixed-milestone fixture The active phase fixture used dir "02-current" with file "01-UAT.md" — a cross-phase stray per phase-id.cts's isPhaseArtifact/scopeToPhase (#3511), so the file was silently excluded from the scan and current_milestone read {files:0, items:0} instead of {files:1, items:1}. Confirmed by direct CLI run against a hand-built fixture before recommitting. Renamed the file to 02-UAT.md to match its directory's phase number, matching every other fixture in this suite. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3783): backfill changeset PR number to 4336 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
c20675cc4d |
fix(#3819): widen executor's pre-commit guard beyond worktree mode (#4343)
* fix(#3819): widen executor's pre-commit guard beyond worktree mode The pre-commit protected-branch assertion in the executor agent (#2924) only fired inside a Claude Code worktree and matched a hardcoded five-name branch list. It never ran in an ordinary checkout and never covered this repo's own default branch ("next"), so gsd-executor could commit planning-repo documents directly onto a shared checkout's default branch with no PR ever created. Widen the guard to run in every isolation mode, and resolve the protected branch via the repository's actual default branch (with the existing five-name list retained as a fallback when the resolver itself cannot be invoked) plus any configured git.protected_branches. Add a git.allow_default_branch_commits escape hatch for projects that intentionally execute on their default branch. Also point the separate <final_commit> commit helper back at the same guard, so it cannot be sidestepped by that path. Emitted-Drift-Ack-Growth: gsd-executor.md — widened pre-commit protected-branch guard (#3819); tightened comments to stay under the size cap. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3819): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
4c60879b5d |
fix(#4132): verify durable runtime surface sources (#4182)
* fix(#4132): verify durable runtime surface sources * chore(#4132): record PR number in changeset * test(#4132): cover rejected commands source alias * fix(#4132): reject aliased package fallback * test(#4132): cover rejected agents source alias * test(#4132): cover partially aliased marker provider * fix(#4132): reject partially aliased source providers * test(#4132): cover routed source identity probes * fix(#4132): route installed source identity probes * refactor(#4132): tighten installer source metadata * test(#4132): cover corpus trust boundary attacks * fix(#4132): close installed corpus trust gaps * refactor(#4132): keep installer authority private * fix(#4132): preserve private installer fallback * test(#4132): preserve fixture source authority * fix(#4132): reject overlapping source fallback * fix(#4132): avoid redundant installed corpus reads * refactor(#4132): simplify provider resolution * test(#4132): sync install tree fixtures after rebase --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
86b745b48b |
fix(#4270): forward Codex spawn model routing (#4281)
Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7ff196c505 |
fix(#4096): honor --dry-run in todo complete and write completion keys inside the frontmatter fence (#4325)
* fix(#4096): honor --dry-run in todo complete and upsert completion keys inside the frontmatter fence * review(#4096): tighten todo complete flag rejection to any dash-prefixed token * chore(#4096): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
3d03ae65e6 |
fix(#4094): withhold all four STATE.md progress counters under the milestone-unbounded guard (#4322)
* test(#4094): failing-first matrix for withholding all four progress counters * fix(#4094): withhold all four progress counters under the milestone-unbounded guard completed_phases/total_plans/completed_plans are accumulated from the same phaseDirs walk as total_phases, so the #3354/#3573 withhold condition makes them equally untrustworthy — yet only total_phases was withheld, and every resyncing state.* write silently clobbered the three stored siblings with the under-scoped disk numbers. Extend the withhold-then-fall-back-to-stored pattern to all three siblings: null sentinels in the disk-scan cache value, three new stored-counter readers threaded through all three buildStateFrontmatter call sites, and the same cached-else-stored consumer fallback. Milestone-bounded projects are untouched (gate-conditional). * fix(#4094): scope-requires for the new test block, keep the (#3573) warning token, and update two #3578 rows to the withheld-counter contract - the #4094 describe sat after the closing brace of the section that owned the module-level beforeEach destructure, so it needs its own local requires (mirroring the #3642 block); - the #3573 warning keeps its literal '(#3573)' tag (asserted by an existing test) with '#4094' appended as a separate token; - two #3578 status-guard rows in tests/state.test.cjs asserted the pre-#4094 unconditional disk-scan assignment of completed_phases under the roadmap-absent withhold — exactly the silent clobber #4094 removes; the status-guard conclusion (must not fire) is unchanged, the counter-value assertions now pin the withheld contract. * test(#4094): lint conformance — splitLines for the persisted-progress parser, local seeder, scoped rmSync disable * changeset(#4094) * changeset(#4094): backfill PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
2e1ede6d99 |
fix(#4093): give advance-plan's zero-labeled-fields failure a disk-derived recovery decline (#4318)
* test(#4093): regression matrix for advance-plan zero-labeled-fields decline * fix(#4093): give advance-plan's zero-labeled-fields failure a disk-derived recovery decline * refactor(#4093): collapse IIFE to a plain block (review finding) * docs(#4093): document the advance-plan recovery decline + changeset * chore(#4093): backfill PR number in changeset * fix(#4093): budget lint-compiled-artifact-sync's tsc compile as a compile, not a probe --------- Co-authored-by: sim <sim@local> |
||
|
|
70f22e4643 |
fix(#4213): keep STATE.md progress surfaces synchronized (#4231)
* fix(#4213): keep STATE.md progress surfaces synchronized * fix(#4213): clamp the shared progress bar and keep bold-first priority, changeset + property tests - formatProgressMachineSegment clamps through clampPercentFromFraction (ADR-3180 Decision 7 kernel) with a 0 floor, so a hand-edited out-of-range persisted percent renders a clamped bar instead of throwing RangeError on repeat() inside the write seam - stateReplaceProgressPercent restores the #2177 bold-first priority: **Progress:** anywhere in the body wins; a plain ^Progress: line is the fallback, so free text starting with Progress: cannot capture the rewrite ahead of the real status line - cross-reference comment names the three consumers and the cmdStateSync sanctioned exception (ADR-3408 §8.3) - CONTEXT.md: applyPostSyncPreservation reconciliation documented in the STATE.md Transition Module entry - property tests (never-throws/well-formed, idempotency, round-trip, bold-first) + two regression rows through the CLI --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
294ec29857 |
fix(#4053): quote decimal-shaped frontmatter scalars for spec YAML readers (#4165)
* fix(frontmatter): quote decimal-shaped scalars so a spec YAML reader preserves them A decimal phase identifier written to STATE.md frontmatter (e.g. `current_phase: 22.10`) was emitted BARE, because `scalarNeedsDoubleQuoting` only asks whether a value can OPEN a plain scalar — which `22.10` can. A YAML-spec reader (js-yaml, the statusline, any external tool) then reloads bare `22.10` as the float 22.1, colliding with `22.1` and dropping the trailing zero. gsd's own tolerant line-scanner (`extractFrontmatter`) round-trips the raw text and so hid the defect; a spec reader does not. Fix: `reconstructFrontmatter`'s general scalar path now also quotes numeric- looking strings that are not plain all-digit integers (decimals, exponents, sexagesimal, hex/oct/bin) via `generalScalarNeedsNumericQuoting`, reusing the existing `YAML_NUMERIC_RE`. Every all-digit string — integer counts, phase numbers, and leading-zero fixtures like `02` — stays bare, so the state-rebuild idempotency baseline and the rest of the state corpus are unchanged. This also quotes `gsd_state_version: 1.0` on write, which matches the authoritative STATE.md template (`src/state.cts` already emits it quoted). Regression test drives the real write path and asserts, via js-yaml, that `22.1` and `22.10` no longer collide and read back string-typed; guards that integers and free-text stay unquoted. Fixes #4053 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YBicDMJyh3AH56ZFbUsyxC * chore(changeset): add Fixed fragment for #4053 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YBicDMJyh3AH56ZFbUsyxC * docs(frontmatter): trim the generalScalarNeedsNumericQuoting comment Cut the over-long doc block down to the essential why and drop the inline comment that repeated it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011ZKeSj55VakqQajtoBgCTC * docs(test): drop the #4053 explanatory comments from the touched tests The assertions speak for themselves; remove the added narrative comments. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011ZKeSj55VakqQajtoBgCTC * fix(#4053): correct the trade-off comment, changeset PR number, and cover every claimed numeric form Review follow-ups (trek-e): - The doc comment claimed a plain integer round-trips harmlessly. That is false for leading-zero values (`02` -> 2, `017` -> 17 under js-yaml). Rewrite it to state the real, deliberate trade-off: all-digit strings stay bare because zero-padded ids (`plan: 01`, `phase: 02`) are the pervasive GSD convention and quoting them all is the blanket quoting #4053 asked to avoid; the loss is padding not identity (`02` and `2` normalize to the same phase, `22.1` and `22.10` do not). - Changeset carried the auto-closed draft's number (4151); correct to 4165. - Test exponent, hex, octal, binary and sexagesimal forms through js-yaml, and pin the leading-zero trade-off so the documented behaviour is asserted. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qc7VN4zTpTSDTS9JXM2cFB --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
5869febb16 |
enhance(#4155): invalidate verification results when covered inputs change (#4290)
* enhance(#4155): invalidate verification results when covered inputs change readVerificationStatus() now recomputes a deterministic sha256 fingerprint over a VERIFICATION.md's declared covered_files (phase PLAN/SUMMARY, requirements, implementation files in the verified change set) and returns stale on any mismatch, fail-closed when a covered file is missing, unreadable, or escapes the project root. Legacy reports with no fingerprint metadata keep the prior SUMMARY-mtime staleness check unchanged. The verifier computes covered_digest via the new verification.fingerprint CLI command rather than by hand, since a digest is deterministic math, not an LLM-estimated value. * chore(#4155): backfill fork PR number in changeset * fix(#4155): trim gsd-verifier.md fingerprint instructions to fit LARGE tier byte cap * fix(#4155): address CodeRabbit findings on fingerprint fail-closed behavior Partial fingerprint metadata (one of covered_files/covered_digest present, the other missing or malformed) now fails closed to stale instead of silently downgrading to the legacy mtime-only check. computeCoveredDigest also canonicalizes with realpathSync before re-confining, so an in-root symlink whose target escapes the project root can no longer produce a matching digest. gsd-verifier.md restores the completeness requirement and checklist item trimmed by the earlier size-budget fix, within the LARGE tier byte cap. * chore(#4155): acknowledge gsd-verifier.md growth for the #4155 fingerprint instructions Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the covered-input fingerprint instructions and frontmatter fields the #4155 verification staleness mechanism requires; trimmed to stay within the LARGE tier byte cap * fix(#4155): address gemini adversarial review findings computeCoveredDigest now threads the caller-supplied opts.fs seam through its confinement and read paths instead of always using raw node:fs — a caller like planning-inspect.cts's containmentEnforcingVerificationFs (GAP 2, #2790 follow-up) was silently bypassed for covered-input reads. The project-root anchor itself still canonicalizes through real fs (it is a trusted value the caller derived, not attacker-influenced covered-input data); only per-file candidate reads go through the injected seam. Covered-file paths are now canonicalized (./ prefixes, redundant slashes, internal .. segments) before becoming dedup/sort/hash keys or confinement subjects — closes both a spurious-stale false positive (two spellings of the same file hashing differently) and a confinement gap (an internal .. segment that doesn't start the string). gsd-verifier.md now states covered-file paths are project-root-relative, not phaseDir-relative, closing an ambiguity that would have made a real verifier agent's first fingerprint invocation fail closed. defaultFsImpl's methods now late-bind through fs.<method> rather than capturing function references at module load — the earlier direct-capture form was invisible to existing tests' t.mock.method(fs, 'statSync', ...) seams, a real regression caught by the full suite (not the reviewer). * fix(#4155): catch a plan/summary added to the phase dir after verification but never declared The content digest only recomputes hashes for paths the verifier actually declared in covered_files — it had no way to notice a plan or summary added to the phase directory after verification if that new file was never declared, silently regressing behind the legacy mtime check it replaces (which scans the live directory, not a declared list). findUncoveredCurrentArtifact re-scans the live phase directory for every current *-PLAN.md/*-SUMMARY.md and requires each to be represented in covered_files, closing that gap; a directory scan failure fails closed to stale rather than silently skipping the check. CONTEXT.md's Verification Module entry corrected to describe the fingerprint path's stricter fail-closed FS-error contract (routes to stale) instead of the module's original degrade-to-safe one (missing / not-stale), which only the legacy path still keeps. * refactor(#4155): extract canonicalizeCoveredFiles, add real nested-project e2e test computeCoveredDigest and cmdVerificationFingerprint each normalized/deduped/ sorted covered_files independently — one shared helper now backs both (gemini review's ponytail-lens finding). Adds one CLI-to-readVerificationStatus test against a genuine .planning/phases/NN-x/ project with an implementation file outside .planning/ entirely, closing the review finding that prior #4155 unit fixtures put phaseDir directly under an ownerless tmpdir (findProjectRoot falls back to phaseDir itself there) and never exercised real multi-level path resolution. * fix(#4155): route computeCoveredDigest through real fs, fail closed on unreadable plans/ Two independent review rounds (opus critical-reviewer + opus ponytail + agy, run twice) found two instances of the same fail-open class: - computeCoveredDigest's per-file reads routed through the caller's injected fsImpl. planning-inspect.cts passes a `.planning/`-confined containment fs into readVerificationStatus's opts.fs, so any covered implementation file outside `.planning/` (mandatory per the issue) made the confinement wrapper throw, which was caught and turned into a stale digest -- reporting every fingerprinted phase permanently stale via `planning.inspect`, regardless of actual drift. Per-file reads now always use real node:fs, matching the pre-existing treatment of root canonicalization; the realRel-vs-realRoot check is the real confinement boundary for this data and needs no seam. - allCurrentArtifactsCovered's try/catch never fired (scanPhasePlans reports readdir failures via a `scope` field, it never throws), so an unreadable nested plans/ dir was silently treated as "zero artifacts, all covered" instead of failing closed. Now branches on scope !== SCOPE.COMPLETE. Also, per ponytail's second-round findings: reverted an unwarranted FINGERPRINT_VERSION bump and digest length-prefix from the first fix (no v1 digest has ever existed -- the feature is unreleased -- and the prefix closed a collision that grants no capability beyond what a writer of covered_files already has more cheaply); removed a verifier-facing escape-hatch instruction whose own example was a case that should trigger staleness, not bypass it; corrected CONTEXT.md references to the renamed allCurrentArtifactsCovered and a stale "unconditional" rescan claim; simplified the isStale derivation, removed dead FsLike members, and tightened test coverage. Regression tests for both fail-open bugs are included and were each confirmed to fail against the pre-fix code before the fix landed. full test suite: 2558/2560 pass, 2 skipped, 0 fail * fix(#4155): trim gsd-verifier.md under the LARGE size cap Fork CI caught what my local runs missed: the superseded/nested-plans instruction added earlier pushed gsd-verifier.md to 49299 bytes, 147 over the LARGE tier's 49152-byte hard cap (tests/agent-size-budget.test.cjs). Tightened the #4155 instruction's wording and dropped a redundant inline comment tag; no content lost. * chore(#4155): point changeset at the upstream PR number pr: 19 was the fork PR opened for internal review-lane CI; now that open-gsd/gsd-core#4290 exists, the changeset field must match it per CONTRIBUTING.md's release-notes convention. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
5ad9a36f35 |
fix(#4255): resolve reviewer-lane effort from the lane, not from gsd-plan-checker (#4275)
`review-lane plan` resolved every cross-AI reviewer lane's reasoning effort by spawning `query resolve-execution gsd-plan-checker --host <slug>`. The agent id was a hardcoded literal, so `--host` chose only the argv RENDERING while the LEVEL always came from the installed plan-checker's frontmatter — `low` under every shipped model profile. Every prompt-fed lane therefore ran at a fast structural verifier's effort, and because the rendered argument is a CLI config override it silently beat the effort the operator had configured for that CLI. At `low` a large source-grounded prompt makes a model end its turn with no final message, so the lane came back empty and its stub read as a crash. Effort is a property of the review, so the lane declares it. Two new fields on ReviewerLane — `effortConfigKey` (`review.effort.<slug>`) and `defaultEffort` — carried through each capability manifest and the generated registry, set on the three lanes with an argv effort channel and null on the other nine. A new pure `resolveLaneEffort()` resolves config key -> lane default -> nothing, where "nothing" emits no effort argument at all and the reviewer CLI's own configuration decides; `inherit` selects that path explicitly and an unrecognized level falls back to the lane default rather than being forwarded to a CLI that would reject it. The host's negotiated effortSurface still gates the rendering, so ADR-1239/#2481's trust boundary holds on this path too. Resolving in-process also removes up to twelve subprocess spawns per review. The empty-output stub now names the effort the lane ran at and distinguishes a clean exit from a timeout kill, a non-zero exit, and a process that never ran — `status` is null for both a timeout and a signal, so those were indistinguishable before. The hint is hedged: a clean empty exit is most often a model stopping short, but it is also consistent with a CLI writing its output elsewhere. Also: the capability validator now knows both fields, rejects a malformed key or an out-of-vocabulary default, and rejects a default declared without a config key (a level the operator could never override). An existing end-to-end row in tests/effort-surface-axis.test.cjs asserted the old coupling; it now configures the lane's own key and pins the decoupling in the same real spawn, with the agent execution tier set to a level that must not appear. Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort; leaving the new key undocumented there is the same invisibility that made the plan-checker coupling survive this long. Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort, so leaving the new key undocumented there is the same invisibility that let the plan-checker coupling survive. Claude-Session: https://claude.ai/code/session_01CRMEuzNMWn3gs5uUW2ghcF Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
925a363879 |
enhance(#4032): apply configured agent tool grants (#4238)
* test(4032): add failing installed-agent grants contract Cover global and project agent_tools precedence at the real Claude installer seam before adding implementation. * feat(4032): apply configured agent tool grants during staging Resolve selector-level global and project config once per staging call, then append validated grants before runtime conversion. * test(4032): cover host grant and quoted MCP contracts Exercise installed host artifacts and prove ZCode must treat quoted MCP scalars like plain MCP grants. * feat(4032): apply configured agent tool grants across runtimes Move augmentation and scalar identity into the converter seam so every staged artifact preserves host policy. * fix(4032): register agent tool grants in configuration Accept documented agent_tools config without unknown-key warnings.\n\nKeep installer fixtures on the shared temporary-directory helper. * fix(4032): translate configured MCP grants for Kilo Reuse the converter-owned scalar decoder so quoted canonical grants reach Kilo's native permission keys without altering other host policies. * fix(4032): decode YAML-escaped tool grants * fix(4032): emit valid inline agent tool grants * fix(4032): reject invalid trailing-colon grants * test(#4032): cover cross-review remediation gaps * fix(#4032): close cross-runtime grant gaps * test(#4032): expose Kimi global project context * fix(#4032): preserve Kimi project config context * chore(#4032): add release note * test(#4032): expose fork review regressions * fix(#4032): address fork review findings * test(#4032): make byte-stability assertion portable Compare repeat installs at one root so platform-specific path rendering cannot masquerade as an agent_tools behavior change. * chore(#4032): bind changeset to upstream PR 4238 * fix(#4032): address trek-e review findings (2,3,4,5,6,7,8) Fixes fail-closed decode-failure handling in ZCode's mcp__ stripper, a comment-only `tools:` header mis-parse that silently dropped configured grants, and a naive comma-split that could tear a quoted scalar containing a literal comma. Documents Kilo's inherent `{server}_{tool}` MCP-permission-key collision (external, fixed format — not ours to widen) and locks the existing first-seen-wins resolution in with a regression test. Opts kimi/kimi-code out of the ADR-1235 pre-converter path-rewrite step: routing Kimi through that pipeline (needed so project-scoped agent_tools selectors reach it) was short-circuiting Kimi's own neutralizeKimiAgentPrompt, which expects the original ~/.claude/gsd-core text rather than a pre-rewritten Kimi path. Extends the fast-check token pool and per-runtime install coverage with the missing comment/comma/broad-runtime cases the prior review flagged as untested. * docs(#4032): add CONTEXT.md glossary entries for agent_tools resolver + pre-converter step Documents readGsdEffectiveAgentTools (Install Model Override Resolver Module) and the appendAgentTools pre-converter pipeline step (Runtime Artifact Conversion Module), per contributor-standards.md's new-seam glossary requirement (finding 1). * fix(#4032): address agy adversarial review findings An agy (gemini-3.8-flash-high) adversarial pass over the prior review-fix commit found the fixes for findings 3, 4, 6 and 8 had unfixed sibling gaps, plus a genuine new regression and two CONTEXT.md inaccuracies: - ZCode's comment-only `tools: # note` header matched the inline-value branch instead of falling through to the block-list scan, so a following mcp__* item leaked through unstripped — the exact defect finding 4 fixed in appendAgentTools, unfixed in this sibling function. - Reverted capabilities/kimi-code/capability.json's noPathRewrite: true. kimi-code uses the standard 'agents' kind with converter: null (not kimi-agents — confirmed by reading the descriptor, not its prose description), so it never went through the pipeline change finding 5 fixed, and disabling its path rewrite broke every ~/.claude/ embed in its shipped agents instead. - decodeToolScalar never stripped a trailing ` # comment` from a bare (unquoted) scalar, so a comment after a block-list item, or after an appended grant on an inline line, became part of the "tool name" — fixed at the source (one call site fixes every consumer). - appendAgentTools's comment-index scan wasn't quote-aware, so a `#` inside a quoted scalar (`"mcp__server #1"`) was mistaken for a comment start and corrupted the quote. - parseFrontmatterTools (Kimi/Qwen's tool-list reader, downstream of appendAgentTools's own output) had the same naive comma-split and comment-only-header gaps as findings 4 and 6, unpatched. - The all-runtime smoke test's presence assertion was built on a guessed omit-list; empirically only 7 of 17 runtimes keep an arbitrary mcp__ grant recognizable, replaced with a verified allowlist. - CONTEXT.md claimed a `project:<agent>` selector prefix that does not exist (project override is a same-key merge across two config files) and mislabeled stageAgentsForRuntimeWithConverter's module. * fix(#4032): address full-PR review (Opus critical/ponytail + agy) A whole-PR pass (critical-code-reviewer + ponytail-review on Opus, plus a second agy full-source adversarial pass) surfaced defects the earlier finding-scoped passes couldn't reach: - appendAgentTools corrupted a `tools:` line whose ENTIRE value is a leading quoted scalar (`tools: "Read"` -> `tools: "Read", Write`, invalid YAML) — there is no safe line-surgical rewrite here, so it now refuses to touch that shape instead of emitting broken frontmatter. - decodeToolScalar's malformed-trailing-quote check ran BEFORE comment stripping, so a bare tool name with a quote inside its own trailing comment (`Bash # note: "internal"`) was wrongly rejected. Reordered. - findUnquotedCommentIndex (added in the prior remediation commit) was built on a wrong model of YAML: a `#` after whitespace starts a real comment in a plain scalar regardless of nearby quote characters — verified against the actual parser. The one case that DOES need protection (a leading quoted scalar) is now refused outright above, so the quote-tracking scan was dead weight solving a problem that no longer reaches it. Removed; reverted to the plain `[ \t]#` scan. - Kilo has a SEPARATE agent-frontmatter parser (convertClaudeToKiloFrontmatter, distinct from the buildKiloAgentPermissionBlock fixed earlier) with the same comment-only-header and naive-comma-split gaps as findings 4 and 6 — unfixed in both its src/ and bin/install.js copies. Fixed in both, exporting splitToolScalars for bin/install.js to reuse rather than reimplementing it. - Pipeline docstring in stageAgentsForRuntimeWithConverter still listed 5 steps, omitting appendAgentTools (now step 3 of 6). - docs/CONFIGURATION.md didn't state that a --global install still discovers agent_tools from the cwd's .planning/config.json (confirmed intentional and already covered by a dedicated test, not a bug). - Removed install-engine.cts's deps.cwd injection seam: zero callers or tests ever populated it. Two claims from this round were verified and rejected, not fixed: prototype pollution via a `__proto__` selector key (empirically confirmed `Object.prototype` is never touched — only reassigns the resolver's own local object's prototype, with no observable effect), and a `*` grant value crashing YAML parsing as an alias reference (empirically confirmed it parses as plain scalar text, no crash). A pre-existing, unrelated defect (extractFrontmatterField returns null for block-list `tools:` on Copilot/Antigravity/Cursor/Codex/Qwen, affecting two shipped agents today) was filed as a follow-up rather than fixed here — it predates #4032 and isn't caused or worsened by this PR. * fix(#4032): update stale slug-derivation-drift-guard fixture line normalizeKimiSkillName's real closing brace moved from line 616 to 635 as a side effect of this PR's edits to runtime-artifact-conversion.cts; the MAJOR-1 fixture's hardcoded realEndLine had gone stale. * fix(#4032): address CodeRabbit findings on projectDir threading and flow-sequence tools bin/install.js's installAgentsKindStandalone call site omitted the projectDir argument the function already supports, so a global install through this legacy branch silently fell back to the runtime config dir instead of process.cwd() when resolving project-scoped agent_tools grants — inconsistent with the sibling installOpencodeFamilyArtifacts call site, which already threads it correctly. appendAgentTools' leading-quoted-scalar bailout did not cover a YAML flow sequence (`tools: [Bash, Read]`): splitToolScalars tore it apart on the in-sequence commas and appended past its closing bracket, producing invalid frontmatter. Extended the bailout regex to also refuse a value starting with `[`, matching the same "whole node, nothing may follow" reasoning already applied to quoted scalars. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
e8800287d5 |
enhance(#4153): fail closed unresolved update targets (#4237)
* test(#4153): cover unresolved update target * fix(#4153): fail closed unresolved update target * test(#4153): require a concrete recovery installer * fix(#4153): use concrete unresolved recovery command * chore(#4153): bind changeset to fork PR * test(#4153): cover portable update diagnostics * fix(#4153): keep update diagnostics portable * fix(#4153): harden update version diagnostics * test(#4153): reject jq in update version checks * test(#4153): expose step-local parser gap * fix(#4153): keep JSON parsing step-local * docs(#4153): align update target guidance * test(#4153): expose workflow runtime fallback * test(#4153): expose resolver runtime fallback * fix(#4153): leave unknown workflow runtime empty * fix(#4153): stop inferring Claude for unknown targets * test(#4153): preserve Claude workflow targeting * test(#4153): preserve known runtime directory identity * fix(#4153): recognize Claude workflow paths * fix(#4153): reuse known runtime directory identities * chore(#4153): acknowledge emitted workflow growth The fail-closed diagnostic and known-runtime preservation deliberately add 48 emitted bytes. Emitted-Drift-Ack-Growth: update.md — explicit unresolved-target diagnostics and known-runtime preservation * test(#4153): expose missing Windsurf workflow contract * docs(#4153): document Windsurf update targets * chore(#4153): bind changeset to upstream PR * fix(#4153): gate unresolved-target exit before the VERSION-missing fallback The VERSION-missing bullet in get_installed_version sat before the UPDATE_TARGET_UNRESOLVED exit and shared its trigger condition (version 0.0.0). An LLM agent reading the workflow top-to-bottom could satisfy "proceed to install" without ever reaching the fail-closed exit this PR adds, reopening the ill-defined mutating path #4153 closes. Reorder so the unresolved-target gate runs first and scope the VERSION-missing bullet to require an already-resolved target. Also drop two vacuous mutationSpies entries: they checked '--sync'/ '--reapply' (commands/gsd/update.md content) against `step`, a slice of workflows/update.md — always -1 regardless of correctness. Those routes bypass get_installed_version entirely and are already covered by install.test.cjs, reapply-patches.test.cjs, and skill-frontmatter-contract.test.cjs. * chore(#4153): point changeset pr field at fork PR #10 for fork CI * test(#4153): guard RUNTIME_DIRS/update.md table parity, confirm narrowing intent Nit 1: update.md's PREFERRED_RUNTIME prose and RUNTIME_DIRS (src/update-context.cts) are two independently maintained copies of the same runtime->dir mapping with no parity check; add one so a future edit to either surface without the other fails loudly instead of silently drifting. Nit 2: call out in the changeset that a custom --config-dir matching no known runtime, marker file, or env var now resolves unresolved instead of silently defaulting to claude -- this narrowing is intentional, it's the fail-closed behavior #4153 asks for. * fix(#4153): drop dead $UC fallback in check_latest_version's uc_field, cover unresolved-runtime fast path agy (gemini-3.8-flash-high) adversarial review of the full PR: 1. check_latest_version's uc_field() copy-pasted get_installed_version's `${2:-$UC}` fallback, but every call site here passes $2 explicitly and $UC does not exist in this step's scope -- dead, misleading reference. Use $2 directly. 2. No unit test covered resolveUpdateContext's preferredConfigDir fast path returning runtime: '' for a custom --config-dir matching no RUNTIME_DIRS suffix, marker file, or env var (the exact fail-closed case #4153 adds). Added. A third finding (update.md:90 using /gsd:update vs docs using /gsd-update) was investigated and rejected: /gsd:update is the actual registered Claude Code command name (commands/gsd/update.md name: gsd:update) and is locked by this PR's own test (tests/update-workflow.test.cjs); /gsd-update is a separate, pre-existing, intentional prose convention used in audience-facing docs (README/INVENTORY/FEATURES). Not a defect. * chore(#4153): backfill changeset pr field to upstream PR #4237 --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
77e2472ca0 |
enhance(#4221): replace installer Read() deny rules with a managed secret-read guard hook (#4236)
* feat(#4221): gsd-secret-read-guard PreToolUse hook + registration Add hooks/gsd-secret-read-guard.js, a blocking PreToolUse guard on Read|Grep|Bash that denies reads of .env, .env.<suffix> and .secrets (the .env.example/.sample/.template/.dist templates stay readable). Read checks file_path; Grep checks an explicit path and judges the glob per brace alternative; Bash runs a two-pass token scan (quotes, comments, redirects with fd digits, separators, $( )/backtick/<( ) recursion, heredoc bodies never scanned as commands, nested bash -c/eval rescans, git <ref>:<path> shapes) with a closed non-reading exemption set for existence checks. Fail-open crash policy; 1 MiB commands are denied as command-too-large; more than 64 glob alternatives as glob-too-complex. Why: Claude Code 2.1.259 makes every `cd DIR && grep …` compound prompt for approval whenever any Read() deny rule exists, even in auto mode. A hook denial is not a permission rule and never arms that check. The installer-written deny rules are retired in the follow-up commit. Registration: hooks.json (Read|Grep|Bash, timeout 5), build-hooks HOOKS_TO_COPY, managed-hooks-registry, runtime-hooks-surface (blocking guard with BLOCKING_GUARD_TIMEOUT_S; Kimi ReadFile|Grep|Shell), shell-command-projection managed sets, installer-migration-report, OpenCode/Kilo plugin (grep tool mapping, include -> glob, dispatch), docs tables in five locales, ADR-766 always-on list, regen:derived fixtures, and a new table-driven unit suite. * test(#4221): pin the secret-read guard in existing hook gates Register gsd-secret-read-guard.js in every existing hook gate: the hooks-crash-policy table (deny row; 6 -> 7 deny cases), plugin-manifest REQUIRED_HOOKS and its Read|Grep|Bash group, docs-hooks-table-parity EXPECTED_SURFACE_HOOKS, install.test MANAGED_JS_HOOKS, install-minimal- hooks JS_HOOKS/BLOCKING_GUARDS, portable-node-runner GUARD_HOOKS, kilo-upgrades PLUGIN_GUARD_HOOKS, the Kimi normalization-parity and typed-payload floors, the OpenCode adapter (grep mapping, include -> glob, three dispatch tests) and a Kimi TOML matcher assertion. * fix(#4221): retire installer Read() deny rules (legacy filter) Rename GSD_CLAUDE_DENY_PERMISSIONS to GSD_CLAUDE_LEGACY_DENY_PERMISSIONS and stop adding the three Read(.env) / Read(.env.*) / Read(.secrets) strings. mergeClaudePermissions now only filters them out of an existing permissions.deny: an absent deny key stays absent, a malformed one is still repaired to [], and an array emptied by the filter is deleted so no `"deny": []` residue is left. Uninstall filters the same legacy list and, symmetric with the Antigravity branch, drops an emptied allow or deny key and an emptied permissions object. Unlike the #2278 allow-side migration there is no surviving current deny list, so the constant is renamed rather than mirrored. Removal is byte-exact: a hand-written identical rule is indistinguishable from the installer's and is removed too (the manifest never recorded permission strings). USER-GUIDE and CONTEXT.md updated. * test(#4221): flip install-regressions deny-rule assertions to the retired shape The fresh-merge, non-destructive merge, idempotency, end-to-end install, reinstall and uninstall assertions now expect no Read(.env*) deny rules and no permissions.deny key on a fresh install; the deny:null repair case is kept. A new describe block covers the legacy filter: retired strings removed with a user entry kept, partial sets, near-miss strings untouched, idempotency, GSD-only deny array deleted, a pre-existing empty deny preserved, and uninstall symmetry for allow/deny/permissions. * chore(#4221): add changeset fragment for PR #4236 * fix(#4221): case-fold names; scan shell stdin and xargs pipes Review round 1 (trek-e): - Blocker: secret-name matching is now case-insensitive in the Read, Grep (path and glob) and Bash paths, so `.ENV` / `.Secrets` on a case-insensitive filesystem are recognized as the same secret file. - Major: a shell interpreter's script is now scanned wherever it comes from. The tokenizer keeps heredoc bodies as per-segment tokens and records separator operators; pass 2 groups by segment id and resolves bash/sh/zsh/dash/ksh/su invocation mode: `-c` (including combined `-lc`) scans the script operand, a file operand is checked as a file (a `<( )` operand's echo/printf output is reconstructed), otherwise stdin is the script and heredocs, here-strings and a piped echo/printf source are scanned. `eval` joins all its operands; `source`/`.` handle process substitution. Data heredocs (`cat <<EOF`, the commit-message shape) stay unscanned. - Major: `… | xargs <cmd>` checks the upstream segment's operands as file names when the sub-command reads (`echo .env | xargs cat`, `find . -name .env | xargs cat`); `-a`/`--arg-file` suppresses the inference; a shell sub-command's `-c` script is scanned. Header, USER-GUIDE bullet and changeset updated; documented gaps now include piped scripts from non-echo sources and `exec`/`timeout` wrappers. 60 new suite cases pin the block and allow shapes. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
c6efe2905c |
fix(#4087): stage the hook helpers the Codex bundle's hooks require (#4117)
* fix(#4087): stage the hook helpers the Codex bundle's hooks require CODEX_HOOKS_TO_COPY is a flat, hand-maintained filename allowlist that never recursed, and Codex is excluded from installSharedHooksBundle() — the path that stages hooks/lib/ for full-bundle runtimes — by an !isCodex gate. Excluding hooks/lib/ was a correct scoped decision for #3579 until #3911 ( |
||
|
|
0ea012c519 |
fix(#3939): parse decision bullets with a wrapped bold lead-in (#3953)
* fix(#3939): parse decision bullets with a wrapped bold lead-in parseDecisionLines matched every PHYSICAL line against the three decision-bullet grammars, and all three require the closing `**` in the same string as the `- **D-` anchor. A declaration whose bold lead-in wraps across a line break — the shape discuss-phase itself writes whenever a decision title runs past the wrap column — matched none of them and fell to the #1365 parse-miss guard, which forces `could-not-parse` and hard-blocks check.decision-coverage-plan on a well-formed CONTEXT.md. Fold physical lines into logical bullets before matching: a declaration whose bold lead-in is still open at end-of-line absorbs following lines until that run closes. The three grammars are untouched, so every single-line form parses exactly as before. Joining is bounded and preserves the fail-loud contract. A blank or whitespace-only line, any block-level construct (a list marker of any family, an ATX heading, a blockquote, a table row), or the end of the block stops it, and a lead-in that never closes is emitted unchanged — so a genuinely malformed bullet still reaches the parse-miss guard and still fails loud (#1365), and cannot be "closed" by an inline `**` belonging to the block below it. The joined line keeps the first physical line's indent, so the nested cross-reference signal (#3169) is unchanged. Absorbed lines are scanned once each rather than re-searching the accumulated candidate, keeping a pathological unterminated run linear on the plan gate's hot path. Regression coverage lands in tests/decisions.test.cjs (the owning module's file, per the regression-test placement policy): all three grammars wrapped, a three-line wrap, tags/category/continuation preservation, one-line parity (including inline bold and emphasis inside a wrapped title), CRLF, the markdown-header path, plus negative proof that every join terminator still yields could-not-parse and that the FIX-B and #3169 fixtures are unchanged. Fixes #3939 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3939): add changeset fragment for PR #3953 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3939): property-test the wrap-position invariant Review follow-up: RULESET.TESTS.property-based-testing requires a parsing / transformation contract to carry at least one fast-check property asserting a domain invariant, and the join added by the fix is exactly such a transformation. The example-based tests pinned four hand-picked wrap points; these generalize over the whole dimension. Three properties, on the shared tests/helpers/fast-check-setup.cjs config (numRuns 200, seeded): - round-trip: for every grammar (colon-immediate, titled-colon, em-dash), every id shape, every tag, with and without a category heading, wrapping the bold lead-in at ANY interior space is deepStrictEqual to not wrapping it — where a line happens to break carries no information; - domain invariant: a well-formed wrapped declaration never reaches the parse-miss guard (outcome `parsed`) and keeps its declared id; - fail-loud preservation: an unterminated bold run followed by 0-12 prose lines still yields `could-not-parse` with no decision manufactured, however many lines the join would have to absorb before giving up. The corpus is deliberately free of markdown metacharacters: `:` and `*` select a different grammar (#1639's `[^:*]*` discipline) and a block-construct token legitimately terminates the join. Both are separate behaviours, example-tested above; these properties isolate the wrap-position dimension. Rebuilding the module from `next` with these in place fails 14 (was 12); the two new failures are the round-trip and never-a-parse-miss properties. The fail-loud property passes before and after, which is the point of it. Refs #3939 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3939): fail loud when a wrap splices a decision tag token Addresses review rounds 2 and 3 on PR #3953. Folding a soft line break to a single space is markdown's own rule and is invisible everywhere in a decision bullet except inside the id-adjacent `[tags]` bracket, which the three grammars turn into `tags` and therefore into `trackable`. There a spliced space splits one tag token into two (`[defer` + `red]` -> `defer red`), which does not fail: it parses to a DIFFERENT tag, silently flipping whether check.decision-coverage-plan demands coverage for that decision. The join now stops at such a splice, so the bullet reaches the #1365 parse-miss guard and fails loud instead of guessing. The check is delimiter-aware, so wraps that land next to `[`, `,` or `]` still join and still parse identically to the one-line bullet -- a comma-separated tag list may wrap at any of its separators, across any number of lines. A bracket further along the title is ordinary text and does not restrict the join. Also in this round: - blockConstructRe's doc comment claimed parity with the sectionizer seam's `iterateBullets`, which recognises only the `N. ` ordered form while this set also stops at `N) `. The widening is deliberate and one-directional (a terminator set may recognise more block openers than a bullet iterator; a spare terminator can only make a malformed bullet fail loud, never manufacture a decision). Comment corrected to say so, both marker forms now tested, and a drift guard asserts the seam still does not yield `N)` so the divergence cannot widen silently. - Documented that the table-row alternative deliberately has no trailing whitespace requirement (CommonMark tables may open flush), and that over-termination on prose opening `10.` or `|` is accepted fail-loud behaviour -- now pinned by a test. - Coverage the review asked for: a WRAPPED bold lead-in nested under an already-open decision (#3169, the existing guard used a single-line nested bullet), and title/body whitespace fidelity across every wrap position around a double space. - A fourth fast-check property: wherever a wrap lands inside a `[tags]` bracket, the parse either matches the one-line bullet exactly or fails loud with nothing extracted -- never a decision whose tags differ. - Property helpers render through `renderBullet`, which asserts the form exists instead of letting an unchecked map lookup yield undefined. Fail-first: tests/decisions.test.cjs run against origin/next's decisions.cts fails 18 of 131; against the previous PR head it fails the 2 new tag-splice guards. All 131 pass with this change. Real-world CONTEXT.md from the report is unchanged at 37/44 parsed. Refs #3939 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3939): arm the tag-splice guard on any wrapped line, not just the first The #3953 round-3 guard read the id-adjacent `[tags]` bracket only from a bullet's FIRST physical line, via a regex anchored to the bullet start. A lead-in that wraps twice can open that bracket on a LATER absorbed segment, where the guard was never armed and `wouldSpliceTagToken` became a no-op: - **D-01 [inform ational]: A title.** body text here. folded to the tag `inform ational` and `trackable: true`, where the one-line form gives `informational` and `trackable: false` — a silently wrong answer to the coverage gate, with no thrown error and no parse-miss to signal it. Exactly the re-classification the round-3 guard exists to prevent, for the case it did not cover. `tagBracketOpenAtEolRe` becomes `tagRegionRe`, which asks whether the id-adjacent bracket REGION is still unsettled rather than whether it opened on one specific line: group 1 present means the bracket is open, group 1 absent means the id is read but a `[` may still follow. `joinWrappedBoldLeadIns` keeps the assembled text in `tagRegion` only while the bracket has yet to open, so a bracket opening on any segment arms `tagTail`; once armed, the pre-existing O(1) tail update takes over and `tagRegion` is dropped. A non-empty segment that is not a bracket-open settles the region immediately, so this bounds the string to a single extra join and leaves the 5000-line unterminated run linear. The id class widens to admit an empty id, so a bare `- **D-` still counts as unsettled. This regex only answers "may an id-adjacent bracket still open here?", where matching MORE shapes is the conservative direction: an over-broad match can only make a malformed bullet fail loud, a missed one re-classifies silently. The existing property test wraps at exactly one point, and only at spaces — which round-trip exactly, since the join re-inserts the space it replaced — so neither the bracket-opens-later state nor an observable splice was reachable from it. `wrapBoldLeadInMulti` breaks at two or more arbitrary positions after the id and asserts the same disjunction: parse identically to the one-line bullet, or fail loud with nothing extracted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BCPSU591zVS9vLPd3gKnqn --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
eb336e9f77 |
fix(#4081): decode git C-quoted paths in codebase-drift --name-status parser (#4307)
* test(#4081): failing-first regression for quotepath C-quoted paths in codebase-drift * fix(#4081): decode git C-quoted paths in codebase-drift --name-status parse * test(#4081): set drift_threshold 1 so decoded-path test triggers action_required * chore(#4081): add changeset fragment * chore(#4081): fix changeset fragment formatting * chore(#4081): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
02ad0b91f3 |
fix(#4078): phase.complete next-phase cascade reads dash-grammar checkbox rows (#4301)
* test(#4078): phase.complete mixed-grammar roadmap picks lowest outstanding phase, not positional-last * fix(#4078): accept dash-grammar checkbox rows in phase.complete next-phase cascade Stage 2 (roadmap identity scan) and stage 3 (#2028 lowest-outstanding override) required a colon separator after the phase number, while the canonical phase lookup has accepted the bullet-house dash grammar (- [ ] **Phase N — Name**, #2199) for years. On a mixed-grammar roadmap the only parseable row above N was a later phase.add-ingested colon-form phase - positionally last - and it won the numeric-minimum vote it should never have been alone in: completing Phase 1 of 18 selected Phase 18 and skipped phases 2-17 (#4078). The checkbox branches now accept the #2199 separator class (em/en-dash, hyphen, colon); heading branches stay colon-only, mirroring findRoadmapPhaseInContent exactly. * test(#4078): align regression fixtures with slug name + checked-box semantics * fix(#4078): drop unnecessary type assertion flagged by eslint * chore(#4078): add changeset fragment * chore(#4078): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
04ac8723b9 |
fix(#4024): flag quantitative-criteria trap shapes in verify plan-structure (#4288)
* test(#4024): pin quantitative-criteria trap shapes for verify plan-structure Rows 1-3 and 20 of the #4024 test matrix reproduce the issue's shapes (exact grep -c counts, bulk all-N observed-failing claims) and are expected to FAIL against unmodified next: nothing judges these shapes today. Corrected-arm rows pin that each rule is silent on its own fix. * fix(#4024): flag quantitative-criteria trap shapes in verify plan-structure Add scanQuantitativeCriteria, the third plan-discipline scanner in the cmdVerifyPlanStructure family (#429, #968). It judges criteria text in <acceptance_criteria>/<automated>/<verify> blocks against a six-rule ban list of shapes proven to be traps at HEAD: exact grep -c counts (R1), bulk all-N observed-failing claims (R2), unquoted $VAR in command position (R3), fallible git swallowed by a non-final pipeline stage (R4, warn), wc output compared by string equality (R5), and relative HEAD~N git anchors (R6; bare git diff warns). Legitimate exit: <!-- plan-criteria-allow: R# - reason -->. Pure text scan, fail open. * test(#4024): bind node:test before hook locally below the fold-point * fix(#4024): R3 command-position anchor tolerates list bullets and inline-code backticks * test(#4024): bind VERIFY_CJS locally in the unit block instead of relying on fold scope * fix(#4024): R6 argument span ends at inline-code backtick or redirection * fix(#4024): satisfy no-adhoc-markdown lint on the R4 stage-boundary regex * chore(#4024): add changeset fragment * chore(#4024): backfill PR number in changeset fragment * fix(#4024): escape backticks in regex literals so drift-lint tokenizers keep function attribution --------- Co-authored-by: sim <sim@local> |
||
|
|
2e056488d9 |
fix(#4067): derive advance-plan phase-complete from disk, not the plan counter (#4292)
* test(#4067): pin advance-plan phase-complete guard matrix (RED) Five-case matrix: decline on unsummarized plans (regression), fire on fully-summarized phase, fail-open on unresolvable phase dir, idempotent decline, normal advance untouched. * fix(#4067): derive advance-plan phase-complete from disk, not the plan counter The phase-complete branch of state.advance-plan was decided purely by STATE.md's scalar plan counter (currentPlan >= totalPlans). A stale counter carried into a newly planned phase, or a counter raced by wave-parallel executors, let 'Phase complete — ready for verification' land while sibling plans were still executing. cmdStateAdvancePlan now re-decides that branch from disk before the write: every plan in the Current Position phase's directory must have a SUMMARY.md (scanPhasePlans single owner, the same source state.update-progress recalculates from). Outstanding plans decline the entire write byte-identically (idempotent, concurrency-safe, counter stays display-only); an unavailable disk answer fails open to the counter-derived decision. * fix(#4067): review round 1 — route phase-dir lookup through listMilestonePhaseDirs #3185 drift guard: no hand-rolled phases-dir readdirSync. Windowed (current-milestone) lookup first so an archived milestone's stale dir cannot shadow the live one; unscoped retry when the window cannot answer. Also restore the transform's undefined-data error semantics and extract scanOutstanding. * chore(#4067): add changeset fragment * chore(#4067): backfill PR number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
8249ebcf6e |
fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do not exist yet; every row fails on require. Per #3770 only an intentional target-test failure may authorize GREEN; zero-test discovery, fixture crashes, unrelated failures, and unexpected green are INVALID_RED. * fix(3770): require intentional RED evidence before GREEN Only an intentional failure of the TARGET test (distinctly named, TAP-reported assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test discovery, fixture/load crashes (file-named failures), nonzero exits without a failing test, unrelated failures, unexpected greens, and malformed/missing records are INVALID_RED and block GREEN. - src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses the prohibition-enforcement TAP primitives; fail-closed, never throws) - check tdd-red-evidence <record.json>: validates the persisted record (command, exit code, failing test, expected, actual) - gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now requires the evidence record + gate verdict, not a nonzero exit or a RED: tag * chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs * fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib - gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B < 49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid) - tests: the row-6 fixture used String.replace (first-occurrence), so the `not ok` line still named the target test and the classifier was right to accept it; replaceAll makes the failure genuinely unrelated - eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs (lint the src/*.cts source, per ADR-457 migration rule) Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line * chore(3770): add changeset * chore(3770): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
580059251a |
fix(#4040): route partially-created .planning to initialization recovery (#4283)
* test(#4040): add failing-first regression tests for partial-init routing Red: init.progress/init.resume/init.new-project payloads carry no partial-init discriminator, and progress.md/resume-project.md/ new-project.md route an interrupted bootstrap (.planning/PROJECT.md + config.json only) to Route F / STATE reconstruction / a hard error. * fix(#4040): route partially-created .planning to initialization recovery A bootstrap interrupted after .planning/PROJECT.md (but before REQUIREMENTS.md/ROADMAP.md/STATE.md) was mis-routed three ways: progress.md read it as between-milestones (Route F) or 'no planning structure', resume-project.md offered STATE.md reconstruction, and new-project.md errored 'already initialized' — a routing loop with no recovery exit. Add a shared buildInitCompletenessFields discriminator (planning_exists / requirements_exists / milestones_exists / init_incomplete) to the init.progress, init.resume and init.new-project payloads, and branch on init_incomplete in progress.md, resume-project.md and new-project.md BEFORE the legacy branches. MILESTONES.md presence excludes the archival between-milestones state, so Route F and the STATE-reconstruction path keep working. Emitted-Drift-Ack-Growth: progress.md — deliberate #4040 growth: new init_incomplete recovery branch (routing text + guard on the no-planning and Route F branches) added ahead of the legacy init_context routes. Emitted-Drift-Ack-Growth: resume-project.md — deliberate #4040 growth: new init_incomplete branch routing an interrupted bootstrap to initialization recovery before the STATE.md-reconstruction branch. Emitted-Drift-Ack-Growth: new-project.md — deliberate #4040 growth: project_exists gate split on init_incomplete so a partial bootstrap resumes initialization instead of erroring. * chore(#4040): add changeset fragment * chore(#4040): backfill PR number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
75ee7b0214 |
enhance(#4273): add phase.tdd-applicable single-owner predicate (#4277)
* enhance(#4273): add phase.tdd-applicable single-owner predicate One query verb computes TDD-applicability for a plan (CLI flag, plan type: tdd frontmatter, a task's tdd="true" attribute, or the workflow.tdd_mode config default), mirroring phase.mvp-mode's precedence-cascade shape. Foundation for epic #4272 Phase 2, which wires both dispatch backends to consume it instead of restating the predicate independently. Also fixes workflow.tdd_mode, workflow.research, and workflow.nyquist_validation, which never reached cmdInitExecutePhase/cmdInitPlanPhase/cmdInitDebug/cmdInitNewMilestone because loadConfig() never populates config.workflow — a dead accessor found while wiring this verb's own config read, fixed inline per the no-defer rule rather than left alongside it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4273): document phase.tdd-applicable's FEATURES.md entry Add a docs/features/ fragment for the new phase.tdd-applicable query verb and regenerate docs/FEATURES.md. docs/COMMANDS.md is left untouched: it documents /gsd-* slash commands only, and the sibling verb phase.tdd-applicable mirrors (phase.mvp-mode) has no formal CLI reference entry anywhere in docs/ either -- only inline prose mentions in docs/reference/workflow-fragments.md -- so there is no COMMANDS.md precedent to extend. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): use PHASE_NOT_FOUND reason code, remove try/finally from tests Two orthogonal code reviews flagged a mistyped error reason and a CONTRIBUTING.md-banned try/finally pattern in the phase.tdd-applicable change; both are corrected here. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): stop whitelisting capability-owned config keys centrally workflow.tdd_mode, workflow.research, and workflow.nyquist_validation are each already owned by their own first-party capability's federated config schema (the tdd/research/nyquist capabilities declare them under their own capability.json `config`), resolved via isCapabilityConfigKey. Adding them to gsd-core/bin/shared/config-schema.manifest.json's central validKeys, as the prior commit in this branch did (mirroring workflow.mvp_mode, which genuinely is central-only), declares the same key in two places at once. That collision breaks capability-loader.cts's loadRegistry composition: gsd-test caught this as 84-85 unrelated failures across capability-cli/capability-command-dispatch/capability-lifecycle test files, every one showing "unknown capability: <id>" for a freshly-installed third-party capability that should have resolved fine. Verified directly (not asserted): reverting only this file, keeping the config-loader.cts tdd_mode/research/nyquist_validation flattening and the init.cts call-site fixes from the prior commit, and re-running the exact capability install + capability set repro from tests/capability-cli.test.cjs's "issue-2322" test locally reproduces the failure with the whitelist entries present and clears it without them. loadConfig() still surfaces all three flattened values correctly with no central whitelist entry (confirmed directly against the compiled module) — the whitelist additions were never required for the #4273 fix to work; they were an incorrect over-application of the mvp_mode precedent to keys that aren't central. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): use getNested for tdd_mode (no legacy top-level fallback), allowlist new test file Both fixes address defects found by a gsd-test bench run: tdd_mode routed through get() invented an undocumented top-level alias that silently outranked the canonical workflow.tdd_mode key, and the new phase-tdd-applicable test file was missing from the file-count allowlist. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4273): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f4bf449296 |
fix(#3850): surface gaps_found VERIFICATION files in audit-uat (#3879)
* fix(#3850): surface gaps_found VERIFICATION files in audit-uat cmdAuditUat admits `human_needed` OR `gaps_found`, but parseVerificationItems had a body only for the first and returned an empty array for the second — standing on a comment deferring to `plan-phase --gaps`, a different command audit-uat never reaches. Since cmdAuditUat pushes a file into `results` only when `items.length > 0`, a `gaps_found` report did not under-report: it vanished, taking its phase's `by_phase` row with it, so a clean-looking total gave the reader no cue anything was skipped. Eligibility now has one owner (the caller) and parseVerificationItems reports what the file says. The closed-entry filter could not be built on extractFrontmatter: its array-item parser keeps only each `- ` entry's FIRST line and has no notion of nested key/value objects, so an entry's `status:`/ `resolution:` siblings never reach its output and a closed entry is indistinguishable from an open one downstream. Rather than grow a competing object-list parser — or change extractFrontmatter, whose blast radius is every frontmatter consumer in the repo — this reads the raw segment BEFORE the flattening, via the existing anchored sliceTopLevelFrontmatterSegments, and hands it to the `## Gaps` machinery that already parses exactly this `- `-opened, indentation- continued shape. The human_needed path is byte-for-byte unchanged: same reader, same display names, same numbering, no resolved-entry filtering — pinned by a test and verified by identical CLI output on base and head. parseGapsItems keeps its narrower `status: resolved` rule so no *-UAT.md behaviour moves. Closes #3850 * chore(#3850): backfill changeset pr number for #3879 * fix(#3850): one parse per entry, one fence parser, one resolved-entry rule Adversarial review on #3879: B1, B2, M3, m5, m8 and n9. B1 — `sliceFrontmatterArrayEntries` hand-rolled a second frontmatter fence regex, which re-asserted the byte-0 rule #2977 removed: a BOM'd file (PowerShell 5.1 `>`/`Out-File` writes one by default) sliced nothing, so a `gaps_found` report vanished from the audit exactly as it did before this fix — this issue's own symptom, on a platform the repo already has a named defect class for. `extractFrontmatter`'s BOM+fence logic is now factored out as `frontmatterRegion` and shared. One fence parser, not two. B2 — the resolved-entry skip paired two DIFFERENT parsers by array index: `parseYamlRegion` is indent-blind, `splitGapsEntries` is indent-anchored. A block sequence written at its key's indent — ordinary, legal YAML — makes them disagree about entry count, and from the first disagreement every index names a different entry, so an OPEN entry inherits a CLOSED one's resolution and is silently dropped. That is the defect this PR exists to fix, reintroduced inside the fix. Display name and sibling fields now come from ONE parse of the raw slice; `frontmatterEntryDisplayName` applies `parseQuotedScalar` exactly as `parseYamlRegion` does, so the string is byte-identical to what `extractFrontmatter` produced. The flattened array remains the #2286 GATE, but is no longer the source of items. `sliceFrontmatterArrayEntries` also takes the LAST duplicate key, matching `parseYamlRegion`'s last-wins assignment. M3 — `frontmatterEntryToUatItem` is the single entry->UatItem mapper both readers use, rather than two copies differing only in `result`. m8 — closed entries are skipped on BOTH statuses. The earlier asymmetry cited an acceptance criterion #3850 does not contain: the issue has no AC section, and its suggested fix (2) states the skip unconditionally, naming a file with 14 of 16 entries resolved. That file is `human_needed`, so the asymmetry left the reporter's own scenario over-reporting by 14. m5 — `sliceTopLevelFrontmatterSegments`' contract doc names both consumers and says the column-0 boundary rule is now a cross-module contract. n9 — the vestigial bare block is gone and its body de-indented. Tests: the B1 BOM case, B2's nested-sequence and bare-bullet repros, a CRLF fixture (M4 — it survived by accident, now pinned) and the unified skip rule. Fail-first verified by running the new tests against the pre-fix build: the BOM, nested-sequence and unified-skip cases are red there. * fix(#3850): read the entries as objects, not as re-parsed display text Rebased onto `next`, which changed the ground this fix stood on. ADR-3473 §8.1 (#3881) replaced the hand-rolled frontmatter scanner with the vendored js-yaml: `parseQuotedScalar` and `parseYamlRegion` no longer exist, and an object entry now flattens to `test: A, resolution: R` rather than to its first line. The original mechanism existed ONLY to work around that lossy first-line flattening — it sliced the raw frontmatter segment and re-parsed each entry by hand so a `resolution:` sibling was visible at all. With a real parser upstream that workaround is obsolete, so it is deleted rather than repaired: `sliceFrontmatterArrayEntries`, `frontmatterEntryDisplayName`, the `splitGapsEntries`/`extractGapEntryFields` reuse and the second fence regex are all gone. `frontmatter.cts` instead exposes `frontmatterObjectListEntries(content, key)` — the same parse `extractFrontmatter` runs (same BOM strip, same byte-0 fence, same anchor/alias and sentinel guards, same ambiguous-colon repair), stopping one step before the display flattening. `flattenObjectListItem` is exposed alongside it so a caller deriving a display name produces the byte-identical string `extractFrontmatter` would have. That collapses the review's blockers into properties of the parse rather than things this fix has to get right: - B1 (BOM) — shares `extractFrontmatter`'s strip; verified through the CLI. - B2 (index pairing) — there is no second reader. Display name and sibling fields come from one object. - M3 (duplicate mapper) — one `frontmatterEntryToUatItem` for both readers. - M4 (CRLF) — js-yaml's, not ours; verified through the CLI. Also confirmed on the rebased base, per review: #3850 still reproduces on `next` after #3707 landed (`total_files: 0`, `total_items: 0` on a `gaps_found` fixture), so this PR is still doing work #3707 did not do. Nothing was dropped as redundant. One behaviour note: `entryField` returns a present value verbatim and treats only whitespace-only as absent. Trimming would rewrite an author's `truth:` on its way to becoming the display name. * fix(#3850): keep every frontmatter list entry at its own row Review round 3's Blocker. `frontmatterObjectListEntries` filtered its result to objects, and filtering COMPACTS: `parseHumanVerificationItems` then numbered the survivors by their position in the compacted array. On a list mixing object and non-object entries the non-object rows disappeared outright and the rest were renumbered — #3850's own vanishing-row defect, reached through entry SHAPE instead of file STATUS. Base never had it: it walked the display array, so every row surfaced at its own position. Renamed to `frontmatterListEntries` and it no longer filters (the name now matches what it returns). Deciding what a non-object entry MEANS is a caller's judgement; dropping it is nobody's. Both readers now walk the DISPLAY array — one element per row, the array #2286 already gates on — and consult the parsed array only for "does this entry carry a closure field?". `parsedEntriesFor` owns that pairing and checks the two lengths agree before trusting an index; all-null is the correct degradation, since over-reporting a closed row is recoverable and closing the wrong one is not. Names stay byte-identical to base for every entry shape, including a nested sequence (`[nested]`, not `["nested"]`). Same class closed in the gaps reader: a non-object `gaps:` entry surfaced nothing at all and now surfaces as `unknown`, which is this module's documented fail-safe direction (`parseGapsItems`) on a false-negative bug. Also restores the shared fence parser round 2 accepted. The ADR-3473 rebase dropped `frontmatterRegion` and left the BOM strip and byte-0 fence rule inlined twice; `extractFrontmatter` now routes through it, so "one fence parser" is enforced rather than asserted in a comment. Minors: `frontmatterEntryToUatItem`'s dead `forcedResult` option deleted and its "shared by both readers" comment corrected — it has one call site, and the two readers differ deliberately, each mirroring its own established sibling (`parseGapsItems` vs #2286). Documented at the divergence. Tests: `B2` asserted a name substring, so it passed while the row was mis-numbered and would have passed through outright loss; it now asserts positions and count. B2b pins the reviewer's 6-entry mixed fixture verbatim, B2c the survivors' file positions across skipped rows, B2d the gaps reader. All four fail-first against the reviewed head; 332/332 green with the fix. * fix(#3850): make status authoritative, and let the two gaps readers agree Round 4 review, all five findings. Major. `isFrontmatterEntryResolved` treated a non-empty `resolution:` as closure regardless of `status:`, so `status: failed` + `resolution: "attempted retry, still failing"` vanished from the report — the silently-vanishing-item defect #3850 exists to close, reached by field combination instead of file status. Closure is now per key, because the two keys have different conventions and one rule cannot serve both: `gaps:` `status: resolved` only, byte-identical to the rule `parseGapsItems` applies to a `## Gaps` markdown section, so one authored entry cannot read closed in one reader and open in the other. `human_verification:` a bare `resolution:` still closes, since that is how verifier-written entries record it — but a readable `status:` that contradicts it wins. A single unified rule was the first draft and is wrong: it closes a frontmatter `gaps:` entry carrying `resolution:` and no `status:`, which `parseGapsItems` surfaces, and `parseVerificationGapsItems`' own docstring claims it mirrors that reader's fail-safe status handling. The contradiction guard is not a judgment call about YAML. It is the rule this codebase already applies to the same field pair: `validateResolution` (probe-core.cts) rejects a populated `resolution:` on a non-resolved status outright — "a populated payload is an authoring mistake ... Reject it so the mistake surfaces." A reporter cannot throw, so it surfaces the item. Minor 1. Direct unit tests for `frontmatterListEntries` and `flattenObjectListItem` in `tests/frontmatter.unit.test.cjs`, the file that historically co-changes with `frontmatter.cts`. They were reachable only through `uat.cts`' readers before. Minor 2. `parsedEntriesFor`'s degrade-to-all-null branch is asserted directly. Verified unreachable through content rather than assumed: both readers enter through `frontmatterRegion`, `extractFrontmatter`'s only extra argument gates a warning, and `normalizeParsedValue`'s `value.map` is 1:1. It is a drift alarm for a future edit to either parser, so the helper is exported for tests rather than left as the one unpinned branch. Minor 3. The vestigial `const skipResolved = true` and its dead conditional are gone. Minor 4. `frontmatterEntryToUatItem` no longer reads `test:`. A `gaps:` entry has no `test:` in its vocabulary — the template's entries carry truth/status/reason/artifacts/missing — so it was speculative support for a field the shape does not have, and it collided with the 1..N row numbers `parseHumanVerificationItems` assigns by array position. Not reading it makes the collision impossible; an offset would have rewritten an authored value, against `entryField`'s verbatim contract. Docs, changeset and the dispatcher docstring all stated the unconditional rule and are corrected — three prior rounds here were comment/code drift. Fail-first proven: restoring the universal rule reddens all three new unit tests and both rewritten properties. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
590edec7a7 |
fix(#3956): require positive evidence for verify artifacts/key-links pass (#4004)
* fix(#3956): require positive evidence for verify artifacts/key-links pass An all-string or path-less must_haves.artifacts / key_links block is item-by-item skipped, leaving zero checked results, yet the pass verdict was computed as `passed === results.length` (0 === 0), so all_passed / all_verified read true with status valid and exit 0: a silent false GREEN over zero acceptance evidence. Add a positive-evidence floor (results.length > 0) to both verdicts, mirroring the no-vacuous-pass rule at src/uat-predicate.cts. A well-formed block, the fully-empty-block error, the parser's string tolerance, and key-links pending (#1202) semantics are all unchanged. Governing: ADR-3473 section 8 / 37C (absence, emptiness and failure must not encode as success) and Decision 3 (failure is a value). * chore(#3956): add changeset for verify vacuous-pass fix * test(#3956): add mixed-block coverage and correct the key-links vacuous-pass comment Addresses review on #4004: - Correct the cmdVerifyKeyLinks positive-evidence-floor comment: only bare-string items are continue-skipped; a from:-less object is NOT skipped (it falls through to a verified:false hard failure), so it was never part of the vacuous-pass surface. The prior comment overclaimed symmetry with the artifacts side. - Add a mixed-block regression test per verb (one bare-string prose bullet + one well-formed entry): the string is skipped, results.length === 1 > 0, and the verdict follows the single real entry — pinning that the floor does not over-reject a partial block. - Tighten the changeset wording to match (all-bare-string, not "no path:/from: key"). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
a788afb120 |
fix(#4010): confine stateReplaceField to same-line whitespace so an empty field's following line survives (#4021)
* fix(#4010): confine stateReplaceField to same-line whitespace so an empty field's following line survives stateReplaceField's bold and plain patterns used `\s*` for the label-to-value gap, which matches the newline after an empty field; `(.*)` then captured the following line and the rebuild discarded it -- silent STATE.md data loss on any `state update` against an empty body field (Status:, Stopped at:, Paused at:), with exit 0 and no warning. Confine the gap to same-line whitespace (`[ \t]*`), mirroring the already-correct read side (stateExtractField, src/state-document.cts:404/:409), and pin the label-value separator to a single space when the label line had none, so an empty field yields `**Status:** value` rather than a glued `**Status:**value`. Non-empty and pipe-table replacements are byte-identical to prior behaviour. ADR-3180 §7.7 makes stateExtractField the same-line-confined owner; this aligns the writer to it. Regression test fails before / passes after and covers bold and plain shapes, LF and CRLF, the non-empty byte-identity guard, and an end-to-end transitionCore characterization at the consumer (ADR-3180 Decision 4(c)). * chore(#4010): add changeset for the stateReplaceField empty-field fix * test(#4010): add boundary and property coverage; scope the changeset's unchanged claim Addresses review on #4021: - Add boundary tests for the shapes the example tests missed: an empty field at end-of-document (no following line, bold + plain), two consecutive empty fields (only the target is filled, the other empty field's line survives), and an empty new value on an empty field (joinFieldReplacement synthesizes no dangling separator and the following line is preserved). - Add a fast-check property over the bold/plain branches and joinFieldReplacement: for any field name, any values (empty fields included), and any new value, replacing one field changes only its own line and never the total line count — the invariant #4010 violated, now guarded directly. - Scope the changeset's "unchanged" claim to ordinary space/tab separators (an exotic vertical-tab/form-feed separator, which no GSD template emits, now normalises to a single space). * test(#4010): pin glued-separator non-empty field, scope joinFieldReplacement JSDoc Round-3 review carried forward a Minor finding: joinFieldReplacement's JSDoc still claimed non-empty replacements are unconditionally "byte-identical to prior behaviour", but a non-empty field written with no label-to-value separator (**Status:**value) gains a single inserted space under the narrowed [ \t]* gap. Round 2 scoped only the changeset prose; the source JSDoc was left making the false unconditional claim. - Scope the JSDoc's byte-identity claim to ordinary space/tab separators and name the no-separator normalization as the one intentional exception. - Add a test pinning the glued-separator case (**Status:**Planning): exactly one space inserted, following line survives, not byte-identical. Emitted .cjs is gitignored (class-1), so no emitted-drift-ack applies. build:lib clean; 74/74 state-document tests pass. Claude-Session: https://claude.ai/code/session_01Mzmut6aeqZ1APfUBAkBZTR --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
515191f07d |
feat(#3677): quick-batch hardening and acceptance (#4240)
* chore(#3677): checkpoint design artifacts (gitignored, dev-only) * test(#3677): add failing regression test for the crash-window duplicate-dispatch gap (RED) Independently re-traces resume-mode.md/planner-wave.md/worktree-dispatch.md/ merge-wave.md and src/quick-batch.cts's resumeBatch (lines 894-899) and confirms the prior research pass's Open Question 1: a coordinator crash between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge) leaves BATCH.json at "pending" with no STATE.md row yet (only written in Step 9), so --resume's eligibility re-derivation would dispatch a second executor into a new worktree for the same item, orphaning the first. This test asserts worktree-dispatch.md's Step 6 excludes an item whose SUMMARY.md already exists from the spawn set, mirroring planner-wave.md's existing PLAN.md-existence check one layer earlier. Fails against the current worktree-dispatch.md, which has no such guard. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1 for the full trace and fix-location rationale. * fix(#3677): guard worktree-dispatch.md against re-dispatching an already-executed item (GREEN) worktree-dispatch.md's Step 6 re-derives eligibility every dispatch round via the same quick-batch resume call resume-mode.md uses, but had no check for "did this item already finish executing" the way planner-wave.md already checks "did this item already get planned" (PLAN.md existence) before re-planning. A coordinator crash between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge) left the item eligible for a second dispatch on --resume, orphaning the first worktree's real, already- committed work and silently losing it once the second executor's SUMMARY.md write clobbered the first at the same item_dir path. Adds a SUMMARY.md-existence exclusion before spawn-plan is computed, symmetric to planner-wave.md's PLAN.md check. The excluded item is not lost: merge-wave.md's own mergeable-wave criterion (status=pending, SUMMARY.md on disk, not yet merged) already picks it up independently of this eligible/spawn list. Workflow-prose-only fix — touches no already-merged/reviewed .cts module. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1 for the fix-location rationale (why not resumeBatch itself). * test(#3677): add real-git coverage for worktree-ownership tampering, scope drift, and submodules Closes the three coverage gaps identified in 40-design.md §2/§3 (#3677, epic #3344 Phase 5's own AC bullets: "arbitrary-worktree ownership attempts", "scope drift", "submodules"): - Arbitrary-worktree ownership tampering: a manifest entry naming a non-agent branch is silently dropped at normalization before any git subprocess runs; a manifest entry naming a plausible agent-branch that was never actually created by this repo's own worktree.create (a genuinely foreign repo/branch) is blocked via base_mismatch. Both leave the foreign location and repoRoot's HEAD provably untouched. - Advisory scope drift: a committed path outside declared files_modified still merges successfully (advisory, never blocking) while surfacing a scope_out_of_declared warning naming the drifted path; an exact declared-scope match produces zero warnings (boundary case). - Real .gitmodules submodule integration: a repo containing a real local git submodule merges cleanly through executeWorktreeWaveCleanupPlan for an unrelated plan; a real gitlink pointer bump (declared) merges cleanly with the superproject tree reflecting the new pinned commit; an undeclared bump is advisory-only and surfaces a scope warning naming vendor/sub, same as any other undeclared modification. No src/*.cts changes — all three gaps were coverage-only; the underlying primitives already behaved correctly (independently verified against real git subprocess output before writing each assertion). * docs(#3677): document how to diagnose a preserved quick-batch worktree Extends the one-sentence "worktree is preserved (never deleted)" mention into a concrete diagnosis procedure: where the preserved directory is, how to read the executor's real commits/diff against the plan's declared files_modified, how to read the item's own SUMMARY.md independent of merge outcome, how to manually merge-and-clean-up or discard, and how to re-run --resume afterward. Also documents that a SUMMARY.md-written-but-still- pending item (the crash-window case fixed in this same PR) needs no manual intervention — --resume routes it straight to the merge step. * chore(#3677): checkpoint final acceptance-evidence mapping (gitignored, dev-only) * fix(#3677): make crash-window duplicate-dispatch guard behaviorally provable and durably recoverable Orthogonal review (Spec finding): the crash-window regression test added earlier this phase only asserted readStep('worktree-dispatch.md') + regex matches against the markdown prose — proving the DOCUMENTATION says the right thing, never that the runtime condition (pending status + on-disk SUMMARY.md + absent STATE row) is actually handled correctly. #3677's own "Alternatives considered" explicitly rejects "document recovery without fault injection" for exactly this reason. Extracts the filtering decision into a pure, independently testable function, filterAlreadyExecuted(eligibleIds, executedIds) in src/quick-batch-dispatch.cts, wired to a new `quick-batch filter-executed` CLI verb (src/quick-batch-command-router.cts) — the same pure-decision- then-CLI-wired pattern computeSpawnPlan/computeMergeOrder already establish. worktree-dispatch.md now calls this verb explicitly instead of only describing the decision in prose. A genuine fixture-based test in tests/quick-batch.test.cjs constructs a REAL BATCH.json (createBatch), writes a REAL SUMMARY.md on disk at the item's real item_dir, calls the REAL resumeBatch, and proves both that resumeBatch alone still reports the item eligible AND that filterAlreadyExecuted (fed a real filesystem check) correctly excludes it. The prior prose-assertion tests are kept — they now prove the workflow markdown is correctly WIRED to the verb — but are no longer the only proof. Self-discovered defect while building that fixture (fixed inline, not deferred): tracing merge-wave.md against /gsd:quick's own prior art (QUICK_WORKTREE_MANIFEST=$(mktemp ...), quick.md:415) showed $QUICK_BATCH_WORKTREE_MANIFEST is a fresh PER-PROCESS temp file. A resumed coordinator correctly does not re-dispatch an already-executed item (this fix), but nothing durably recorded that item's worktree_path/branch/base either — Step 7 in the resumed process would have had no data to build its cleanup-wave entry from. Adds dispatched_worktree/dispatched_branch/ dispatched_base to QuickBatchItem (src/quick-batch.cts) — deliberately NOT a reuse of the pre-existing `worktree` field, whose loadBatch validation requires the path to exist on disk (verified empirically: reusing it made the batch permanently unloadable the moment a legitimately-merged worktree was removed). worktree-dispatch.md persists the triple once a worktree is created; merge-wave.md falls back to it when the ephemeral manifest lacks an entry, clears it after a successful merge, and fails closed rather than guessing if no record exists anywhere. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.1 and §9.3 for the full trace, empirical verification notes, and rejected alternatives (reusing `worktree` directly). * test(#3677): prove the arbitrary-worktree-ownership boundary against two real sibling worktrees Orthogonal review (Security finding): the two existing ownership-tampering tests didn't test ownership — one was trivially rejected by WORKTREE_AGENT_BRANCH_RE's shape check before any git call (proves branch- NAME filtering, not ownership), the other pointed at a wholly separate, never-linked foreign repo, so merge-base failed immediately because the branch didn't exist as a ref at all. Neither exercised the real scenario: a manifest entry whose worktree_path/branch are swapped to point at a DIFFERENT, GENUINELY-REGISTERED sibling worktree of the SAME repoRoot, with a branch name passing the shape check and a base in allowed_bases. Investigated executeWorktreeWaveCleanupPlan (src/worktree-safety.cts) directly: this is NOT a reachable gap. Git enforces branch-per-worktree uniqueness, so a swapped-in entry.branch can only match worktree_path's ACTUAL checked-out branch if it names that sibling's own real, uniquely- generated branch name — which manifest tampering confined to one batch's own record has no way to know (branch names are agent-<quick_id>[-<timestamp>]-shaped, and quick_id allocation is collision-checked GLOBALLY across every existing quick task and batch, not merely within one batch). Adds a stronger test that empirically proves this: two REAL, concurrently- alive sibling worktrees of the same repo (both via real `git worktree add`, both WORKTREE_AGENT_BRANCH_RE-passing, both sharing one merge-base), with worktree_path/branch swapped between them in both directions. Both attempts are blocked via branch_mismatch; both real worktrees, their branches, and one sibling's real uncommitted-to-main commit survive completely untouched. Supplements (does not replace) the original two tests, which still prove distinct, real boundaries. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.2 for the full trace, including the one explicitly-documented (not fixed) trust boundary this investigation surfaced: the primitive defends against fabricated data, not a caller bug that misattributes a real-but-wrong item's own triple to a different item. * chore(#3677): checkpoint design-doc addendum for review pass 2 findings (gitignored, dev-only) * docs(#3677): add changeset for PR 4240 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
d5f8191f66 |
fix(#3730): quick-tasks-migrate — canonical-schema repair, auto-run on first quick (#4216)
* test(#3730): a legacy Quick Tasks table must be migratable to canonical * fix(#3730): quick-tasks-migrate — canonical-schema repair, auto-run on first quick * fix(#3730): review fixes — usage parity, contiguous table span, collision-safe bucket, template-width delimiter Emitted-Drift-Ack-Growth: fast.md — #3730 runs quick-tasks-migrate before the first append (auto-migration on first quick run) Emitted-Drift-Ack-Growth: quick.md — #3730 replaces the match-any-format note with the migration instruction * chore(#3730): backfill changeset pr number * fix(#3730): scope the quick-batch row-48 guard to branches touching quick-batch --------- Co-authored-by: sim <sim@local> |