d0bdd700c28473fb339bf76782306db978cd4280
198 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
da37986cd0 |
fix(#1847): resolve standard tier to claude sonnet 5
Point the sonnet/standard tier at Claude Sonnet 5 (`claude-sonnet-5`,
GA 2026-06-30) across the Anthropic-backed runtimes and provider presets,
replacing the superseded `claude-sonnet-4-6`. Mirrors the change into the
CONFIGURATION.md and settings-advanced.md runtime-defaults tables (the
#3229 catalog↔docs parity gate) plus the pt-BR/zh-CN translations, and
updates the tests that pin the old ID. Regenerates the workflow size
baseline for the (smaller) settings-advanced.md.
Scope is Sonnet only — opus/haiku IDs are untouched. Prepared as a 1.6.1
hotfix off the v1.6.0 tag.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit
|
||
|
|
93e5d2dd84 |
fix(#1525): skip deferred phases on autonomous reruns (#1846)
* fix(#1525): skip deferred phases on autonomous reruns * chore(#1525): add changeset fragment * chore(#1525): fix changeset body * test(#1525): refresh install parity fixtures * test(#1525): shrink autonomous workflow * test(#1525): refresh autonomous baselines * test(#1525): tolerate Windows temp cleanup flake |
||
|
|
50ff7a8707 |
fix(#1776): scope prune phase resolution to ## Current Position (#1832)
* fix(#1776): scope prune phase resolution to ## Current Position cmdStatePrune resolved the current phase by extracting the `Phase` field over the WHOLE STATE.md body. stateExtractField's fallback chain ends in a pipe-table match (`| Phase | N |`), so a STATE.md lacking a `Current Phase` field and a prose `Phase:` line — but carrying an unrelated `Phase`-labelled table row (e.g. a historical verification table) — resolved that stale table cell as the current phase and computed a wrong cutoff (bailing "Only N phases" or pruning at a stale boundary). Resolve the phase via the same canonical chain buildStateFrontmatter uses — frontmatter `current_phase` → `Current Phase` field → prose `Phase: X of Y` — but scope ONLY the prose term to the `## Current Position` section via the fence-aware locateCurrentPosition seam (new exported sliceCurrentPositionSection). Frontmatter and the explicit `Current Phase` field stay document-wide (they are unambiguous); the shared stateExtractField is not narrowed for any other caller. Tests (folded into tests/state-prune.test.cjs): a stray `| Phase | 2 |` table with the real phase in frontmatter no longer drives the cutoff (fail-first on base); template-conformant STATE.md is unchanged; and a fast-check boundary-containment property that a `| Phase | N |` row outside Current Position never leaks into the scoped resolution. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1776): add changeset for prune Current Position scoping Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
dae7f81482 |
fix(#1528): drop next-phase guidance from security-blocked verify-work presentation (#1687)
* fix(#1528): drop next-phase guidance from security-blocked verify-work presentation When security enforcement blocks phase advancement (no SECURITY.md produced), the verify-work presentation told the user advancement was blocked but still offered `/gsd:plan-phase {next}` and `/gsd:execute-phase {next}`, competing with the current-phase fix. Remove those two next-phase lines so the blocked state routes only to the current-phase resolution (secure-phase, ui-review). The post-transition presentation — reached only after the completion contract passes — still offers next-phase planning, which is the correct place for it. Regression coverage added to tests/ui-review-next-guidance.test.cjs: the security-blocked block must not offer next-phase actions, and the post-completion block must still offer them. Regenerated workflow size baseline. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1528): add changeset for security-blocked next-phase fix Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1528): recapture golden-install-parity fixtures for verify-work.md change Rebased onto next; verify-work.md's installed hash changed across all 16 runtime fixtures. Diff confined to the single gsd-core/workflows/verify-work.md key per runtime. Assert mode 16/16 green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
1f649838b8 | Merge branch 'next' into fix/1698-codex-output-last-message | ||
|
|
feb7036399 | chore: sync next package version to 1.7.0-rc.1 | ||
|
|
b5c8a3e590 |
feat(#1827): rebuildCore + rebuild intent + drift-class unit tests (#1829)
Phase 1 of approved feature #1817. Implements the body-structure derivability contract landed in ADR-1817 (Phase 0, PR #1828). Source changes (src/state-transition.cts): - Add `rebuild` as the 11th intent in StateTransitionIntent (ADR-1769 transition set extended per ADR-1817 §1). - Add `phaseInventoryProvider` optional dep + PhaseInventoryRecord type (Leaky-Abstractions guard: pure core stays testable without disk I/O; rebuild skips table reconciliation when the provider is absent). - Add `case 'rebuild':` dispatch arm to transitionCore (the missing-case compile-time guarantee extends to the 11th case). - Implement rebuildCore orchestrator + four drift-class helpers per ADR-1817 §2: * reconcileCurrentPosition — body prose re-derived from frontmatter * reconcileByPhaseTable — **By Phase:** table re-derived from disk inventory via the new dep * stripTemplatePlaceholders — `**Field:** [placeholder]` → `**Field:** (pending)` (no canonical source available) * deduplicateSessionArchive — keep most-recent 3 archived H3 blocks - Implement appendRebuildLogSection per ADR-1817 §3 — every mutation appends a structured entry (timestamp/kind/section/before/after/reason) to `## Rebuild Log`. Idempotency guarantee (ADR-1817 §4): a no-mutation rebuild appends NO log entry, so two successive runs on a clean file are byte-identical. Compiled output (gsd-core/bin/lib/state-transition.cjs): regenerated via `npm run build:lib` (tsc -p tsconfig.build.json). Tests (tests/state-rebuild.test.cjs, 18 cases): - Dispatch + idempotency contract (3 tests, including the load-bearing 'rebuild on a clean file is a no-op' that pins §4). - Current Position prose reconciliation, criterion #1 (3 tests). - Template-placeholder removal, criterion #3 (3 tests). - Session Continuity Archive de-duplication, criterion #4 (4 tests). - **By Phase:** table reconciliation via phaseInventoryProvider, criterion #2 (3 tests, including the Leaky-Abstractions no-op guard). - Regression guard for sync + prune, criterion #7 (2 tests). Verified: node --test tests/state-rebuild.test.cjs → 18/18 pass. Verified: node --test tests/state-transition.test.cjs → 85/85 pass (no regression on the existing 10 transitions). Phase 2 (#1826) wires the CLI surface (cmdStateRebuild + --dry-run + integration tests + docs + changeset). This PR adds the engine only — no user-visible command yet, so no-changelog label applied. |
||
|
|
18995380ce |
feat(#1154): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1738)
* feat(verify-phase): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1154) Carry the edge-probe's existing `backstop` (non-inferable) tier through the plan-phase projection as a structured flat-scalar marker instead of a prose parenthetical, and make verify-phase abstain -> human_needed (never silent-pass) on a backstop truth it cannot confirm with explicit evidence. Truth-axis mirror of #644's prohibition judgment-tier (ADR-550 D4). Engine (deterministic, CI-tested per ADR-550 D5 — never the LLM verdict): - src/probe-core.cts: truthStatement/truthVerification normalizers, projectTruths (conservative serializer), dispositionForUnverifiableTruth (backstop+no-evidence -> unverified/flagged/insufficient_spec; backstop+evidence -> green; inferable -> green, the over-abstention guard). - src/roadmap.cts: coerceTruthToString now reads `statement` first so an object-form backstop truth is surfaced, not dropped (Hyrum backward-compat for truth-readers). Workflow/agent/docs: plan-phase emits the structured marker (flat scalar, ADR-550 #1278); verify-phase + gsd-verifier add the abstain arm; new references/honest-verifier.md; FEATURES/COMMANDS document insufficient_spec; ADR-550 amended (truth-axis D4 mirror). Decisions adopted (trek-e review): insufficient_spec feeds existing human_needed with a distinguishable reason (no new VERIFIER_STATUS); changeset Changed; round-trip parity test; abstain-on-unconfirmed-backstop regression test red-first. Implementation notes (deviations from the issue's proposed file list, verified live): - frontmatter.cts needs no change — its flat parser already round-trips object-form truths. - verify.cts needs no change — it grades artifacts/key_links structurally; truths are LLM-graded at the workflow layer, so consumption lives there + the deterministic helper. - No CJS<->SDK hand-sync — the SDK seam was retired (ADR-0174); src/*.cts is sole source. Regenerated artifacts: golden-install-parity fixtures, INVENTORY-MANIFEST, size baselines. * chore(#1154): add changeset (Changed) for honest verifier User-facing changelog fragment for #1738. Typed `Changed` (not `Added`) per trek-e review condition 3 — the verify behavior shifts for backstop-bearing specs (a confident silent `passed` becomes `human_needed`), which is user-visible even though the schema marker is additive. * docs(#1154): score-formula also excludes abstained insufficient_spec truths (review nit-1) trek-e review nit: the verify-phase score sentence said PRESENT_BEHAVIOR_UNVERIFIED truths were "the only ones excluded" from verified_truths. Post-#1154 an abstained `insufficient_spec` backstop truth is also excluded (it is not ✓ VERIFIED and routes to human_needed). Behavior was already correct; this tightens the wording. Regenerated golden-install-parity fixtures + workflow-size baseline for the touched verify-phase.md. (Nit-2 — a dedicated insufficient_spec_items frontmatter list — is intentionally not taken: the current design is ADR-550-D4-conformant, the abstain cause rides as a distinguishable report reason, and adding it would exceed the approved scope.) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ac001be49e |
fix(#1778): use 1.6 named-flag frontmatter.set form in thread workflow (#1816)
* fix(#1778): use 1.6 named-flag frontmatter.set form in thread workflow The thread workflow's CLOSE and RESUME branches called frontmatter.set with the pre-1.6 fully-positional shape (frontmatter.set <file> <field> <value>). Since 1.6 the dispatcher (gsd-tools.cjs) parses the file positionally and reads field/value from the named flags --field/--value via parseNamedArgs; the positional form leaves field/value undefined, cmdFrontmatterSet errors 'file, field, and value required', and the status/updated writes are silently skipped. Closing a thread never marked it status: resolved and resuming never marked it status: in_progress. Switch all four sites (CLOSE status+updated, RESUME status+updated) to the 1.6 hybrid form that verify-work.md already uses: frontmatter.set <file> --field <field> --value <value> Add a regression test with three guards: (1) behavioral — the named-flag form writes the field while the positional form errors with the documented message and does not mutate the file; (2) workflow parity — no workflow under gsd-core/workflows/ emits the positional form, so a future edit that reintroduces it anywhere fails CI; (3) thread-specific — CLOSE writes status: resolved and RESUME writes status: in_progress via the named flags. * docs(#1778): add changeset fragment for thread workflow frontmatter fix * docs(#1778): fix unclosed inline-code backtick in changeset fragment * fix(#1778): move regression into owning test + regen baselines lint-regression-test-names rejects new bug-NNNN-*.test.cjs files; move the #1778 regression (behavioral named-vs-positional + workflow-parity scan + thread CLOSE/RESUME assertions) into tests/frontmatter-cli.test.cjs, the canonical home for frontmatter CLI regressions, and delete the standalone file. frontmatter-cli.test.cjs already carries the allow-test-rule exemption for workflow .md content tests. gsd-core/workflows/thread.md ships to every runtime and is size-tracked, so recapture the 16 golden-install-parity fixtures (thread.md hash) and the per-file workflow size baseline (thread.md 12400 -> 12464) via UPDATE_GOLDEN=1 and npm run size:baseline. |
||
|
|
fd576528a7 |
fix(#1747): register four search-provider keys in the config schema (#1814)
* fix(#1747): register four search-provider keys in the config schema buildNewProjectConfig emits seven search-provider availability flags and research-provider.cts providerAvailability() consumes all seven, but only three were registered in VALID_CONFIG_KEYS (config-schema.manifest.json). config-loader.cts then printed an 'unknown config key(s)' warning for the four unregistered keys (tavily_search, ref_search, perplexity, jina) on every freshly generated .planning/config.json. Register the four missing keys in the schema manifest and document them alongside brave/exa/firecrawl in CONFIGURATION.md. Add a regression test plus a structural drift guard that requires every config-driven research-provider flag to be in VALID_CONFIG_KEYS, so a future provider addition cannot silently reintroduce the drift. * fix(#1747): move regression into owning test file + add changeset lint-regression-test-names rejects new bug-NNNN-*.test.cjs files; move the #1747 regression (four provider keys in VALID_CONFIG_KEYS + provider-flag drift guard) into tests/bug-2530-valid-config-keys.test.cjs, the canonical home for VALID_CONFIG_KEYS regressions, and delete the standalone file. Add the missing .changeset fragment — config-schema.manifest.json lives under gsd-core/ (user-facing), so changeset-lint requires a fragment. * test(#1747): regenerate golden-install-parity fixtures for schema change Adding four provider keys to config-schema.manifest.json shifts its shipped content hash (65dea848 -> 7d398e94); recapture all 16 runtime fixtures via UPDATE_GOLDEN=1. Each fixture changes exactly one line — the manifest hash. |
||
|
|
2bada4de1a |
fix(#1698): capture codex review via --output-last-message, not stdout
The Codex reviewer in review.md captured the review by redirecting codex exec's stdout to the review file. On Windows, codex writes process-teardown output to stdout after the final agent message, so that noise was appended to a non-empty file and slipped past the `[ ! -s ]` empty-output guard as a silently polluted review (consumed by severity extraction and the plan-review-convergence gate). Capture the final message via codex's own `-o/--output-last-message <FILE>` and discard stdout. The #1115 contract is preserved (stderr to .err, capability-gated $CODEX_BYPASS_FLAG, --ephemeral, --skip-git-repo-check) and the empty-output fallback still fires when codex leaves no/empty output. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
21f0b4316f |
refactor(#1796): ADR-1769 Path A — finish STATE.md preservation consolidation (#1799)
Extract readModifyWriteStateMd's post-sync preservation block into a pure,
field-classification-table-driven applyStatePreservation in the STATE.md
Transition Module. progress / status / stopped_at now join current_phase_name
as table-governed (getFieldClassification), so a preservation-policy change is
a one-row table edit instead of a per-call-site patch.
This realizes the consolidation ADR-1769 / CONTEXT.md already claimed shipped
('Absorbs readModifyWriteStateMd post-sync preservation block') and routes the
#1264 preservation policy through the single field-classification table — the
bug class is now structurally guarded by the table, not just the call-site
shouldResync flag.
Behavior is byte-identical to the pre-amendment inline block (Hyrum-safe — the
15 readModifyWriteStateMd callers' observable preservation is unchanged):
- state/frontmatter/transition + bug regression suite: 847 pass
- phase/milestone/verify (other RMW consumers): 582 pass
- codex (gpt-5.5/high) adversarial review: CLEAN (58,564-case equiv sweep)
ADR-1769 amendment appended documenting #1796.
Closes #1796
|
||
|
|
79c69d443b |
refactor(#1793): ADR-1769 Phase 7 — sync, prune, update migrations (#1760, #1761) (#1794)
Migrate cmdStateSync, cmdStatePrune, cmdStateUpdate (state.cts) onto the STATE.md Transition Module substrate and close the maintenance bug pair #1760/#1761 (ADR-1769, epic #1769). Completes the substrate: all 10 lifecycle/maintenance transitions now route through transitionCore. - Add {kind: 'sync'|'prune'|'update'} to StateTransitionIntent with syncCore, pruneCore, updateCore in src/state-transition.cts. - updateCore: body-only single-field update (strip/reassemble), mirroring the pre-migration cmdStateUpdate contract. - pruneCore: pure content→content section pruning (Decisions / Recently Completed / resolved Blockers / Performance Metrics rows at or below cutoff), byte-identical tokenizeHeadings splicing. Adapter owns currentPhase, dry-run, and STATE-ARCHIVE.md. - syncCore: body writes (Total Plans in Phase, Progress bar, Last Activity) given disk-derived numbers. Adapter owns the disk scan + roadmap scope. - #1760: cmdStatePrune now derives currentPhase from 'Current Phase' OR 'Phase' (the canonical template emits 'Phase: X of Y'), so prune engages on template-conformant STATE.md instead of bailing 'Only 0 phases'. - #1761: cmdStateSync skips the Progress write (percent=null) when a milestone version is set in frontmatter but the ROADMAP has no versioned heading for it (milestone cannot be bounded). Projects without a milestone version are unaffected. Leaves progress untouched rather than silently writing fallback- derived wrong values. - Regression: bug-1760 (prune engages on template field) + bug-1761 (sync leaves progress untouched when unbounded). ADR-1769 marked Accepted. All 82 transition + 515 regression tests pass. Closes #1793 |
||
|
|
064f63b299 |
refactor(#1791): ADR-1769 Phase 6 — patch migration + curated current_phase_name preserve (#1695) (#1792)
Migrate cmdStatePatch (state.cts) onto the STATE.md Transition Module substrate and close the curated-field clobber bug class #1743/#1695 (ADR-1769, epic #1769). - Add {kind: 'patch'} to StateTransitionIntent, with patchCore in src/state-transition.cts. Applies each caller-supplied {field:value} pair via stateReplaceField over the full content (body + frontmatter), tracking updated vs. failed. data.updated/data.failed mirror the CLI output shape. - Collapse cmdStatePatch to a transitionCore dispatch. Field-name validation (security) and the resync-progress decision stay in the adapter. - #1695/#1743 fix: extend the #1230 delta heuristic in readModifyWriteStateMd to the curated current_phase_name, table-driven via getFieldClassification('current_phase_name').preservation === 'preserve-always'. When a write does NOT change the body Phase: source line, the curated frontmatter value wins over syncStateFrontmatter's body re-derivation (which harvests a wrong parenthetical aside — #1695). begin/planned/complete-phase rewrite their body Phase line, so the delta does not fire for them and current_phase_name still advances. - Regression: bug-1695-state-patch-clobbers-phase-name.test.cjs — unrelated patch preserves curated current_phase_name; patching the Phase source advances. All 70 transition + 493 regression tests pass (state/phase/milestone + #905/#397/#3242 lineages). Closes #1791 |
||
|
|
3ceb83329d |
refactor(#1789): ADR-1769 Phase 5 — milestoneComplete migration (#1790)
Migrate the STATE.md write path inside cmdMilestoneComplete (milestone.cts) onto the STATE.md Transition Module substrate (ADR-1769, epic #1769). - Add {kind: 'milestoneComplete'} to StateTransitionIntent, with milestoneCompleteCore in src/state-transition.cts (consulting the field-classification table). Owns the closure write: Status ('<version> milestone complete'), Last Activity, Last Activity Description, a Current Position reset to 'Awaiting next milestone', and an Operator Next Steps reset pointing at the next-milestone command. - Collapse the inline STATE.md transform in cmdMilestoneComplete to a transitionCore dispatch. The adapter retains writeStateMd (lock + steady-state syncStateFrontmatter post-sync) and resolves the runtime-specific next-milestone slash command, injecting it via intent.nextMilestoneCommand. - The two section resets carry their pre-seam allow-adhoc-markdown waivers (regex semantics pinned by existing tests; pending collectSection #1372). - realClock.today() is byte-identical to milestone.cts's local today (both new Date().toISOString().split('T')[0]). - Characterization tests pin field updates, both section resets (replace + insert paths), and frontmatter #1255 parity. All 66 transition + milestone + state tests pass. Closes #1789 |
||
|
|
91704c9fcf |
refactor(#1786): ADR-1769 Phase 4 — plannedPhase + milestoneSwitch migration (#1788)
Migrate cmdStatePlannedPhase and cmdStateMilestoneSwitch (state.cts) onto the STATE.md Transition Module substrate (ADR-1769, epic #1769). - Add {kind: 'plannedPhase'} and {kind: 'milestoneSwitch'} to StateTransitionIntent, with plannedPhaseCore and milestoneSwitchCore in src/state-transition.cts (consulting the field-classification table). - plannedPhaseCore owns the template-aware Status/Last Activity updates, Total Plans in Phase, Last Activity Description, and the Current Position section (via the inlined mutateCurrentPositionForAdvance twin). The adapter keeps the resync:false readModifyWriteStateMd wrapper (#500 RC1). - milestoneSwitchCore owns the new-milestone reset: rebuilt frontmatter (milestone/name/status='planning'/zeroed progress) + Current Position body reset. The adapter keeps acquireStateLock + platformWriteSync (NOT readModifyWriteStateMd — milestoneSwitch rebuilds frontmatter directly). - Collapse both callbacks to transitionCore dispatches. The now-dead updateCurrentPositionFields helper and KNOWN_STATUS_PATTERNS import are removed (their behavior lives in the transition module's mutateCurrentPositionForAdvance). - Characterization tests pin the template-aware preserve-authored invariant, Total Plans, Last Activity narrative, Current Position update, and the full milestone reset (frontmatter + position body + gsd_state_version preserve + Accumulated Context preserve). All 58 transition + 177 state + phase + bug-2630/bug-905 regression tests pass. Closes #1786 |
||
|
|
6f80524be1 |
refactor(#1784): ADR-1769 Phase 3 — completePhase migration onto Transition Module (#1785)
Migrate the STATE.md write path inside cmdPhaseComplete (phase.cts) onto the STATE.md Transition Module substrate (ADR-1769, epic #1769). - Add {kind: 'completePhase'} to StateTransitionIntent + completePhaseCore in src/state-transition.cts. Pure body field mutations (Current Phase shape/name, Status, Current Plan, Last Activity + Description, Completed/Total Phases + Progress percent), consulting the field-classification table for touched keys. Roadmap progress injected via a new optional deps.roadmapProvider. - Collapse the ~90-line inline STATE.md transform in cmdPhaseComplete to a transitionCore dispatch. The adapter retains updatePerformanceMetricsSection (section table) + syncStateFrontmatter (disk-scan post-sync) and the multi-file atomic transaction (STATE is committed with ROADMAP/REQUIREMENTS, so readModifyWriteStateMd is not used here). - deriveProgressFromRoadmap/clampPercent/stateReplaceFieldWithFallback move from the phase.cts call site into the pure core (no circular dep: phase-lifecycle and state-document don't import state.cts). - Characterization tests in tests/state-transition.test.cjs pin the field updates, roadmap progress derivation, #1255 frontmatter parity, and the 'Phase:' fallback. All 44 transition + 193 phase tests pass. No user-visible behavior change. cmdPhaseComplete CLI output and 'phase complete' behavior unchanged. Closes #1784 |
||
|
|
1512e6f415 |
refactor(#1782): ADR-1769 Phase 2 — advancePlan migration onto Transition Module (#1783)
Migrates cmdStateAdvancePlan (~80-line RMW callback) onto the transitionCore dispatch established in Phase 1 (#1775): - src/state-transition.cts: - Add {kind: 'advancePlan'} to StateTransitionIntent union - Extend StateTransitionResult with optional data field for intent-specific output (advanced, currentPlan, totalPlans) - advancePlanCore: parses legacy + compound plan formats, handles advance vs phase-complete branching, strips frontmatter before body mutation (#1255 pattern — codex Phase 2 HIGH finding), uses stateReplaceFieldIfTemplate for template-default-aware field replacement (Knuth invariant), mutates ## Current Position section via mutateCurrentPositionForAdvance - mutateCurrentPositionForAdvance: inlined section mutation (avoids circular dep with state.cjs's updateCurrentPositionFields) - src/state.cts:cmdStateAdvancePlan: collapsed to 30-line dispatch - tests/state-transition.test.cjs: 5 characterization tests (advance, phase-complete, error, compound format, frontmatter #1255) Codex gpt-5.5/high review: 1 HIGH blocking (frontmatter strip — fixed), 1 medium follow-up (StateTransitionResult.data shape discrimination — noted for Phases 3-7). gsd-test: 21893/21893 PASS (Linux docker). Closes #1782 |
||
|
|
744bb7aaee |
refactor(#1771): ADR-1769 Phase 1 — STATE.md Transition Module substrate + beginPhase (#1775)
* refactor(#1771): ADR-1769 Phase 1 — STATE.md Transition Module substrate + beginPhase Lands the Phase 1 substrate per ADR-1769: - src/state-transition.cts (new Module): - Field-classification table (FieldClass enum + FIELD_CLASSIFICATION rows) - STATE_MD_SECTIONS constants block - Pure transitionCore(content, intent, deps) dispatch - beginPhase intent implementation (first-time + #3127 resume paths) - src/state.cts:cmdStateBeginPhase — collapses ~190 lines to a thin dispatch onto transitionCore via readModifyWriteStateMd. The lock, no-op write guard, and #1230 post-sync delta heuristic stay in the RMW seam; the body-mutation policy moves to transitionCore. - tests/state-transition.test.cjs (24 tests): - Substrate invariants (table enum, section constants) - Characterization: 6 first-time body field updates - Characterization: 5 #3127 idempotency-guard resume behaviors - Characterization: 3 Current Position section mutations - Characterization: Current focus body text line (#1104) - Property (RULESET.TESTS.property-based-testing): beginPhase status propagation + FIELD_CLASSIFICATION own-property contract - Resume Current Position mutation (preserves Plan/Phase/Status) No external behavior change. Full state.test.cjs regression (177 tests) plus bug-3127/#3242/#905/#948 pass. Property tests surfaced two pre-existing quirks (state-document.cjs greedy \s* on whitespace-only field values; Object.prototype method leakage on FIELD_CLASSIFICATION lookups for strings like 'toString') — documented in test comments; fix-out-of-scope for Phase 1. Closes #1771 * refactor(#1771): ADR-1769 Phase 1 codex review corrections Addresses 3 blocking findings from codex gpt-5.5/high review: 1. FIELD_CLASSIFICATION shape (state-transition.cts): - Was flat FieldClass enum (collapsed source + preservation) - Now two-column {source, preservation} rows per ADR-1769 §4 - Added missing fields verified via Memtrace against buildStateFrontmatter (state.cts:1633-1653): gsd_state_version, last_updated, last_activity_desc, progress.{total_phases, completed_phases, total_plans, completed_plans, percent} - Field 2-7 preservation dispatch can now consult the table 2. Prototype-pollution hardening (state-transition.cts): - Table is now Object.freeze(Object.assign(Object.create(null), {...})) - getFieldClassification() helper uses Object.hasOwn; returns null for inherited prototype methods (toString/valueOf/__proto__) - Old code: FIELD_CLASSIFICATION['toString'] returned the function 3. STATE_MD_SECTIONS aligned to canonical template: - Verified against gsd-core/templates/state.md via Memtrace - Was: 8 entries including non-template sections (## Session, ## Decisions, ## Operator Next Steps, ## Session Log, ## Roadmap Evolution) - Now: 6 canonical top-level sections (## Project Reference, ## Current Position, ## Performance Metrics, ## Accumulated Context, ## Deferred Items, ## Session Continuity) Also: beginPhase now consults getFieldClassification() per touched field (codex finding: 'table not consulted by transitionCore'). Unknown fields raise immediately — adding a field without a table row is caught at runtime. Repo-hygiene catches from gsd-test (not node --test, which missed these): - gsd-core/bin/lib/state-transition.cjs added to eslint.config.mjs ignore list (ADR-457 tsc-generated) - docs/INVENTORY-MANIFEST.json regenerated via node scripts/gen-inventory-manifest.cjs --write Property test for Object.prototype leakage tightened to verify getFieldClassification() returns null for toString/valueOf/__proto__. Ref #1771 * fix(#1771): cast Object.create(null) to satisfy @typescript-eslint/no-unsafe-assignment ESLint CI failed on src/state-transition.cts:72:14 — Object.create(null) returns `any`, which leaked through Object.assign to the typed `FIELD_CLASSIFICATION` declaration. Adding an explicit cast to `Record<string, FieldClassification>` eliminates the unsafe-assignment while preserving the null-prototype protection codex review recommended. gsd-test: 21888/21888 PASS. * fix(#1771): add 'see #1771' to allow-test-rule exemption per ADR-456 CI lint-allow-test-rule-refs failed: 'New allow-test-rule exemption without an issue ref — add `see #NNN` per ADR-456'. Updated comment on tests/state-transition.test.cjs to reference the Phase 1 issue. * fix(#1771): remove unnecessary allow-test-rule exemption The exemption was added speculatively. The test file does not use readFileSync + .includes()/.match()/.startsWith() on source content — it calls transitionCore() with string literals and verifies results via stateExtractField() and array .includes() on the updated[] array. No exemption needed. |
||
|
|
4ced0a64cc |
feat(#1561): assumption-delta advisory checkpoint (#1767)
* feat(#1561): assumption-delta advisory checkpoint * chore(#1561): backfill changeset PR number (#1767) --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
b0d5ca3379 |
feat(#1517): support custom reviewer instances for /gsd:review (#1766)
* feat(#1517): support custom reviewer instances for /gsd:review Add a bounded review.reviewer_instances config surface so one model-capable adapter (e.g. opencode) can run as several independent reviewer identities in a single /gsd:review pass. Instances participate only via review.default_reviewers, expand before built-in slugs, are available iff their cli is detected, and a non-matching entry is a hard error (typo must be loud). >=2 same-cli instances emit a shared-adapter caveat in REVIEWS.md. Default path with no instances is byte-for-byte unchanged. Single-source instance->cli resolution lives in resolveReviewerSelection / normalizeReviewerInstances (parity-locked in tests/review-reviewer-instances.test.cjs). cli validated against KNOWN_REVIEWER_SLUGS only (never arbitrary shell); model/agent opaque, never shell-interpolated. Closes #1517 * chore(#1517): backfill changeset pr:1766 --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
ce62f2b68d |
refactor(#1763): ADR-1235 agent migration — cut over the trivial-converter group to the descriptor path (#1764)
ADR-1235 step 1: route the trivial-converter runtime group (cursor, windsurf, augment, trae, codebuddy) off the inline install() agent loop onto the descriptor-driven installRuntimeArtifacts path. Establishes the converter-context foundation (pre-converter cross-cutting + no agent-stamp). Agent install output is byte-identical for all 16 runtimes (golden-parity, global + verified local). cline deliberately excluded (local rules-only). Closes #1763. |
||
|
|
a3d3c2a445 |
refactor(#1756): derive getDirName from a documented runtime.localConfigDir descriptor axis (#1757)
ADR-1239 Phase B (parent #1679). getDirName was a hand-maintained 15-branch if-chain mapping each runtime to its local content-rewrite dot-dir. Relocate those values into a documented runtime.localConfigDir descriptor field; derive getDirName from registry.runtimes[id].runtime.localConfigDir (fallback .claude). - 16 capability.json gain runtime.localConfigDir (byte-identical values) - capability-validator.cjs requires it (non-empty dot-dir); registry regenerated - docs/reference/capability-manifest.md documents the field + the three divergent values (copilot=.github, antigravity=.agents, kimi=.kimi-code) - drift-guard test: golden value map + key-set equality both ways Byte-identical install output for all 16 runtimes (golden-parity harness #1730). Closes #1756 Co-authored-by: review-bot <review-bot@gsd> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e075a41c86 |
feat(#1754): CLI version-skew detection — warn when a global install shadows project-local GSD (#1755)
* feat(#1754): CLI version-skew detection — warn when a global install shadows project-local GSD Addresses #1754 (approved-enhancement). Detects when the running gsd-tools.cjs is outside the project root while a project-local install exists — the shadowing scenario from #1748 where a stale global canary CLI (retired @gsd-build/sdk) silently overrides project-local GSD. Implementation (Node CLI entry-point, not shell snippet — avoids bloating 93 workflow files past their size caps): - src/cli-skew-check.cts: pure function checkCliSkew({resolvedPath, projectRoot, projectLocalExists}) → string|null. Compares paths via path.relative; returns a warning when the resolved CLI is outside the project root AND a project-local install exists. Includes @gsd-build/sdk removal hint when the path matches. No I/O (pure), no gsd-sdk literal (avoids bug-2801 lint). - gsd-core/bin/gsd-tools.cjs: wired at startup via the existing findProjectRoot resolver. Non-blocking (try/catch; advisory stderr warning, never gates). - eslint.config.mjs: registers the new ADR-457 generated artifact in the ignores. - tests: 6-case suite (skew/no-skew/legacy/normalization); all green. - Golden fixtures regenerated (UPDATE_GOLDEN=1) for the new compiled artifact. - docs/how-to/update-gsd.md: Diátaxis reference note for the skew warning. Full suite: 3354 pass, 0 regressions (1 pre-existing local AGENTS.md failure). lint:ci green. Closes #1754 * chore(#1754): backfill changeset pr placeholder * chore(#1754): regenerate INVENTORY-MANIFEST for the new cli-skew-check source module --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
cf2e66b39e |
feat(#1708): typed documentation-sourced #853 dispatch-flatten (ADR-1239 Phase B) (#1719)
* feat(#1708): typed documentation-sourced #853 dispatch-flatten Graduate the #853 orchestrator-backgrounding decision from a scattered RUNTIME==='codex' prose check to a typed, documentation-sourced engine decision. Adds a backgroundDispatch dispatch sub-axis (sourced per host: codex+cursor documented true, 9 documented false, 5 undocumented), shouldFlattenDispatch(dispatch) (inline UNLESS background && backgroundDispatch, fail-closed), and a gsd_run query dispatch-should-flatten the plan/execute workflows call. Cursor is newly background-eligible per its docs (inline->background) — a documentation-justified behavior change. No RUNTIME-name residue for this decision. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1708): backgroundDispatch citations in matrix + CONTEXT note Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1708): address review findings on typed dispatch-flatten Code/adversarial review: convert the manager.md/autonomous.md Compound Action preamble from hardcoded 'On Codex' to FLATTEN-based branching (the handlers already use the query; the preamble contradicted them and was wrong for cursor); make shouldFlattenDispatch null-safe + type-honest (accepts raw 'undocumented' registry values); make backgroundDispatch a required descriptor field (matching its siblings, all 16 carry it); strengthen the config.runtime behavioral test; update the bug-853 prose-pin test + comment. Security review clean; Codex confirmed no fail-open. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1708): backfill backgroundDispatch in role:runtime test fixtures Making backgroundDispatch a required descriptor field broke role:runtime fixtures in capability-manifest-version/capability-registry/host-integration-descriptors tests that build a dispatch object without it (caught by full gsd-test, not scoped npm test). Backfill backgroundDispatch:false into the well-formed fixtures; the deliberately-malformed 'required-field' test fixture is left malformed by design. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1708): update fix-1521 dispatch-gating assertion to the FLATTEN gate fix-1521 pinned the codex-specific run_in_background prose that #1708 graduated to the typed dispatch-should-flatten/FLATTEN gate. Update its assertions to verify FLATTEN=false gating (not a runtime name) + that the old RUNTIME===codex gate is gone. Caught by full gsd-test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1708): add changeset for typed dispatch-flatten Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1708): remove stray temp PR-body file Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1708): add issue ref to bug-853 allow-test-rule annotations ADR-456 requires every allow-test-rule exemption to carry a see #NNN reference; the source-text-is-the-product annotations added when migrating the prose assertions lacked it (lint-tests CI gate). Add (see #1708). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
fb5f89db10 |
feat(#1704): destSubpath write-confinement (ADR-1239 Phase B) (#1706)
* feat(#1679): confine install writes within configHome ADR-1239 Phase B write-confinement: a pure assertDestWithinConfigHome(configDir, destSubpath) rejects a destSubpath that escapes configHome (path traversal / NUL byte) at plan-build time on BOTH the install and uninstall plan paths; surface.applySurface and installOpencodeFamilySkills route through it, and _copyStaged carries a defense-in-depth containment check. Security-load-bearing for the Phase C third-party-descriptor loader. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1704): add changeset for destSubpath write-confinement Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1704): fix windows path-portability in confinement test The N1 'accepts a true child subpath' assertion compared against path.join (no drive resolution) while the helper uses path.resolve — on Windows that mismatches the C: drive prefix. Compute the expected via path.resolve to mirror the helper. Windows-CI-only failure (local gsd-test is Mac+Linux). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
30d4b85de5 |
feat(#1684): negotiated host-integration interface (ADR-1239 Phase A) (#1690)
* feat(#1684): add negotiated host-integration interface module ADR-1239 Phase A: a pure, additive, no-I/O module exposing PROTOCOL_VERSION, the 8-axis HOST_INTEGRATION_AXES closed vocabulary, the UNDOCUMENTED fail-closed sentinel, negotiateHostCapabilities (effective subset of host-declared and engine-known), a typed degradation ladder, and host-capability profiles. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1684): validate and document host-integration axes (16 runtimes) Extend validateRuntimeBody to validate the 8 hostIntegration axes (closed enums + undocumented sentinel + dispatch struct + reserved-key guards) and the widened runtime vocabulary; author a documentation-sourced hostIntegration block in all 16 runtime descriptors; regenerate the registry. Every per-CLI value is documented (cited) or the explicit undocumented sentinel. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add host-integration capability matrix and adr amendment New per-CLI, per-axis citation reference (value/source/evidence for all 16 CLIs); ADR-1239 Phase-A-implemented amendment; CONTEXT.md glossary seam entry. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1684): harden dispatch negotiation edge cases Code-review hardening: treat NaN/Infinity maxDepth as missing (fail-closed, +warning); reset nested/background when namedDispatch collapses to false (struct consistency); SAFE_DEFAULTS dispatch floor to read-only; warn on non-finite protocolVersion; symmetric undocumented warnings for dispatch fields. Pure module — no consumers; behaviour fail-closed throughout. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1684): register host-integration.cjs in lint-ignore and inventory New tsc-generated bin/lib artifact: add to the eslint ignore list (ADR-457 — lint the .cts source), regenerate docs/INVENTORY-MANIFEST.json, and add the docs/INVENTORY.md CLI-modules row. Fixes the 3 gsd-test failures (551-eslint-bin-lib-coverage x2 + inventory-manifest-sync). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add changeset fragment for host-integration interface Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add how-to for sourcing a host's integration axes Diataxis how-to guide for adding/updating a host's runtime.hostIntegration axes from authoritative docs, the undocumented-sentinel rule, validation, and extending the closed vocabulary. Completes the Step-5 doc quadrants (reference + explanation + how-to). Indexed in docs/README.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2b215b4163 |
feat(#1688): warn on stale model bake for static-frontmatter runtimes (#1692)
* docs(#1650): fix stale opencode install-path claim in core settings * feat(#1688): warn on stale model bake for static-frontmatter runtimes * chore(#1688): backfill changeset pr field with real PR number * test(#1688): make resolveAgentDir assertions use path.join for windows * docs(#1688): codify windows path-literal-in-assert anti-pattern + align test |
||
|
|
548a986ffc | chore: sync next package version to 1.6.0 | ||
|
|
47906b052d |
fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final) (#1550)
* fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final) BSD/macOS mktemp only substitutes the XXXXXX template when it is the final path component. Templates like `...-XXXXXX.json` / `gsd-pr-body.XXXXXX.md` return a LITERAL `XXXXXX` path (no randomization) on macOS, so concurrent workflow runs collide on the same temp manifest/body file — one run can overwrite or consume another's. Reproduced on macOS: the second call to the suffixed template fails `mkstemp: File exists`. Fix: use a suffixless `XXXXXX` template (so it IS the final component), then rename to add the intended extension — portable across BSD + GNU userlands, no GNU-only `--suffix` flag. Empty-file-then-write semantics are preserved at every site. Affected workflow temp files: - execute-phase.md: gsd-worktree-wave-*.json (wave worktree manifest) - quick.md: gsd-quick-worktree-*.json - spec-phase.md: edge-probe-reqs-*.json - ship.md: gsd-pr-body-*.md - profile-user.md: gsd-profile-answers-*.json, gsd-profile-analysis-*.json The execute-phase.md edit uses a compact intermediate var + trailing comment to stay under the ADR-857 phase-6 size ceiling (93166); regenerated the workflow size baseline accordingly. Validated on macOS: 20 concurrent calls yield 20 unique randomized paths. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1520): add changeset fragment (Fixed) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1520): add fail-first workflow-prose guard for mktemp XXXXXX suffix Repo-wide scan of gsd-core/workflows/**/*.md that fails on any mktemp template whose XXXXXX run is followed by a filename suffix (the BSD/macOS non-randomizing form). Fails on the six pre-fix instances and passes on the fix, and locks the copy-paste-prone idiom out of future workflows. Mirrors the bug-637 hardcoded-$HOME workflow guard. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1520): rename regression test to fix- prefix (regression-test-names lint) New tests/bug-NNNN-*.test.cjs files are banned by the lint-regression-test-names ratchet; use the fix- prefix (matches the fix-1445 precedent). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1520): add issue ref to allow-test-rule exemption (ADR-456 lint) lint-allow-test-rule-refs requires every new `allow-test-rule:` comment to carry a #NNN reference (don't allowlist). Add (#1520) to the source-text exemption. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#1520): abort touched mktemp chains on failure (|| exit 1) Per review: the VAR=$(mktemp …) && mv … && VAR=… chains dropped the issue's suggested failure guard. If mktemp fails, $VAR is empty and the subsequent mv/write lands on an unintended relative path. Add `|| exit 1` to all six touched chains so a mktemp failure aborts the snippet. Regenerated the workflow size baseline for the slightly longer lines. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1520): rebase onto next — regen size baseline + describe rename Resolve the workflow-size-baseline.json conflict from next advancing by regenerating from the current workflow sizes. Also rename the test describe from `bug #1520` to `#1520` (the file uses the fix- prefix) per review nit. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#1520): harmonize profile-user mktemp to ${TMPDIR:-/tmp} (review nit) The two profile-user.md temp sites this PR already rewrites kept a hardcoded /tmp while the four sibling workflows use ${TMPDIR:-/tmp}. Harmonize for consistency and macOS-correctness (some sandboxes have no writable /tmp). Regenerated the workflow size baseline. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1520): regen size baseline after rebase onto next --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
1a46109b97 |
enhance(#1579): deterministic gsd-tools query eval.score verb (#1583)
* feat(#1579): deterministic gsd-tools query eval.score verb Split C of #1573 (pure code, lowest risk). Adds an eval.score query verb (coverage*0.6 + infra*0.4; bands 80/60/40) mirroring the verify.* chain; gsd-eval-auditor consumes it instead of doing weighted arithmetic in-prompt. Non-breaking — additive only. arXiv: 2601.15130 (Plausibility Trap/DPDM), 2507.10281 (Table Agent), 2508.15754 (TIR). * fix(#1579): address review — domain guard, property test, glossary, SKIP_ROOT, inventory/baseline - C3 input-domain: reject out-of-domain eval.score (require 0<=covered<=total; was emitting overall_score>100 / negatives) - C1 property test: add tests/eval.property.test.cjs (fast-check) — determinism, band monotonicity, [0,100] bounds, never-throws - C2 glossary: CONTEXT.md "Eval Scoring Module" entry (source-of-truth path + interface) - C4: add `eval` to SKIP_ROOT_RESOLUTION (pure arithmetic; no .planning/ access) - inventory: register generated eval.cjs/eval-command-router.cjs (INVENTORY-MANIFEST.json + INVENTORY.md rows) - size: regen agent-size baseline for gsd-eval-auditor (reused gsd_run shim + eval.score step) - eslint: ignore generated eval*.cjs (ADR-457 bin/lib migration coverage) * fix(#1579): register eval family in alias-drift gates Add EVAL_COMMAND_ALIASES/EVAL_SUBCOMMANDS to scripts/check-alias-drift.cjs families and to familyArrayKeys in the manifest-coverage test, so the eval family lands under the same drift guard as every sibling family (state/verify/init/phase/phases/validate/roadmap). Addresses trek-e review. check:alias-drift ok; feat-3251 coverage 9/9; eval suites 10/10. * docs(#1579): use half-open verdict band ranges in CLI-TOOLS overall_score is fractional and thresholds are >=80/>=60/>=40, so a score in [79,80) is correctly NEEDS WORK despite the old '60-79' label. Relabel bands as 60-<80 / 40-<60 / 0-<40 to match the code. Addresses trek-e nit. * fix(#1579): validate eval.score CLI inputs Reject unknown infra tokens and fractional counts, and pin the 80-point verdict boundary including rounding-before-banding behavior. |
||
|
|
a63684c222 |
enhance(#1577): WebFetch/WebSearch injection isolation + opt-in blocking (#1585)
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the largest untrusted channel) in gsd-read-injection-scanner; shared untrusted-input-boundary reference @-included by the 8 ingest agents (randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring); opt-in security.injection_blocking (default advisory — non-breaking). arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth). * fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized - A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already in the transcript. The prompt-level data/instruction boundary is the primary control. - A2: registered security.injection_blocking in the config schema + defaults manifests (default false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads. - A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention). - A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale). - A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input. - Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) + drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as the noted pre-existing follow-up. * fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate The new reference quotes injection phrases ('ignore previous instructions', 'you are now…') as examples agents must NOT comply with, tripping the repo's own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red on HEAD). Allowlist it alongside the other security docs (security-model.md, TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS scanner test doesn't scan references/, so only the shell gate needed it. Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15. * fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so no named web-ingress agent is uncovered, keeping the two justified additions (gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10. - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset. - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source- document ingress per the boundary), though it has no web tools. INGEST_AGENTS in the isolation test now asserts all 10; size baselines regenerated (+60 bytes each, both well under the DEFAULT cap); changeset reworded 8 -> 10. Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39. * docs(#1577): document security.injection_blocking + boundary seam trek-e Major 2 + Minor: - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to the Full Schema and a Security Settings subsection, distinguishing it from the workflow.security_* namespace; honest circuit-breaker-not-redactor framing matching ADR-1577 / security-model. - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry. Verified: lint:docs ok; config-field-docs + contributor-standards green. * test(#1577): make read-injection property test git-text, not binary trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as degenerate-edge inputs. The NUL is what actually made git classify it binary (git binary = NUL in first 8K). Replace both with text-safe escapes that keep the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File now diffs/blames line-by-line. Verified: property test 2/2; no NUL/raw-noncharacter bytes remain. * docs(#1577): align untrusted boundary docs Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner. * docs(#1577): align ADR ingest agent count Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
d101daff30 |
fix(#1516): expose adaptive model_profile in /gsd-new-project AI Models prompt (#1654)
* fix(#1516): expose adaptive model_profile in /gsd-new-project AI Models prompt Both onboarding paths (Step 2a auto-mode + Step 5 interactive) enumerated only 4 profiles (Balanced/Quality/Budget/Inherit), omitting 'adaptive' even though the model catalog (model-catalog.json profiles) and docs/CONFIGURATION.md register 5. Mirrors the proven /gsd:settings two-question split (#3784): Q1 routes between Adaptive/Standard-tier/Inherit; Q2 (conditional on Q1=Standard) picks Quality/Balanced/Budget — keeping every AskUserQuestion within the 4-option cap. Both config-new-project example payloads now list adaptive. Regression cases folded into the owning tests/new-project-mvp-prompt.test.cjs (per the lint-regression-test-names ban on new top-level bug-NNNN files): each AI Models prompt makes adaptive reachable, all 5 profiles reachable, 4-option cap honored, both example enums include adaptive, brace balance. Workflow size baseline bumped (new-project.md 62324 -> 66138 bytes; still well under the XL hard cap). * chore(#1516): backfill changeset pr ref to 1654 |
||
|
|
77c7b4fc9d |
fix(#1522): enforce canonical verification before phase transition (#1548)
* fix: require fresh phase verification before transition * no-mistakes(review): Fix canonical verification closeout gates * no-mistakes(review): Fix verify-work frontmatter promotion command * no-mistakes(review): Fix stale verification gates * no-mistakes(review): Fix canonical verification routing gates * no-mistakes(review): Fix verification dependency and runtime routing gates * no-mistakes(review): Block stale verification bypasses * fix: handle large init manager outputs in verification workflows * chore: update changeset pr number * fix(verify-work): use fresh verification.status for stale gate The stale check after UAT used phase_completion.verification_status from session-start INIT while human_needed promotion already queried fresh verification.status. Align the stale gate with the canonical query so mid-session verification refresh is not ignored. * fix(init): skip roadmap-checked phases when selecting next_phase Roadmap-only phases without a disk directory were still promoted to next_phase when their checkbox was already checked. Exclude checkboxComplete phases so progress routing does not point at work the roadmap already marks done. * fix: gaps_found not overridden by stale, transition uses canonical verification - verification.cts: check gaps_found before stale so gap-closure routing is not masked by a newer summary mtime - phase.cts: remove redundant findStaleVerificationSummary — readVerificationStatus already handles stale detection - transition.md: replace raw grep on file content with verification.status query to avoid false-positive blocks from body text matching * ci: retrigger tests after rebase * fix(transition): replace gsd_run advisory check with awk frontmatter extraction The runtime launcher is not defined until the update_roadmap_and_state step bash block (~line 165). The early verify_completion block used gsd_run to query verification.status, which violated the runtime-launcher-parity test: 'preamble appears AFTER the first gsd_run reference'. Replace the gsd_run call with an awk-based frontmatter extractor that reads only the status: field between the two --- fences. This avoids both the preamble-ordering constraint and the original false-positive grep bug where body text like 'previous_status: gaps_found' would match a full-text regex. The phase.complete gate at update_roadmap_and_state is the canonical enforcement point; this early check is advisory only. Also update workflow-size-baseline.json for the updated transition.md size. Fixes: runtime-launcher-parity test (B) Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> * fix: re-check verification under planning lock in phase complete Move readVerificationStatus into withPlanningLock so stale verification cannot slip through when a SUMMARY.md is written between the gate and the roadmap/state mutation. Return the blocked status from the lock callback and emit the error after release to avoid leaving .lock behind. * fix(transition): gate on canonical verification.status including stale Replace awk frontmatter read with verification.status query so transition blocks when summaries are newer than VERIFICATION.md, matching phase.complete and other workflows (autonomous, progress, verify-work). * Fix workflow verification gates for yolo transition and stale routing Require VERIFY_STATUS passed before yolo/interactive transition advance. Route stale verification recovery to verify-work, matching canonical projection. * fix(transition): use verification.status query for stale-aware advisory check The awk-based check read raw frontmatter status: passed, which misses the stale case where summaries are newer than the VERIFICATION.md file even though the frontmatter still says passed. The stale status is computed from file modification times, not stored in frontmatter. Move the preamble to the verify_completion bash block (the first block with a gsd_run call) so gsd_run query verification.status can be used for the advisory check. This gives the full readVerificationStatus logic including mtime-based staleness detection, matching the enforcement gate at phase.complete. Capture full JSON (VERIFY_JSON) so next_action can be included in the advisory output alongside the status. Also update workflow-size-baseline.json for the updated transition.md size. Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> * ci: trigger test matrix for 525b946 Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> * fix(transition): restore awk frontmatter extraction for pre-shim verification check The gsd_run launcher shim is not defined until line ~163 of transition.md, so the verification debt check at line ~80 cannot use gsd_run. Restore the awk-based frontmatter extraction that correctly reads status without needing the runtime, and restore the shim at its proper location before phase.complete. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(#1522): clarify transition verification gate wording * fix(#1522): update transition workflow size baseline * fix(#1522): update workflow-size-baseline after rebase onto next Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> * fix(#1522): guard findStaleVerificationSummary FS calls + thread opts.fs seam (review) Address review blocker B1 on #1548: findStaleVerificationSummary ran fs.readdirSync and two fs.statSync calls unguarded between readVerificationStatus's try/catch sections, so a TOCTOU race (a SUMMARY listed by scanPhasePlans then removed before statSync) or any FS error threw uncaught into callers NOT under the planning lock (init.manager / init.progress / uat-predicate). Wrap the body in try/catch degrading to 'not stale', and thread the injectable opts.fs seam (add statSync to FsLike, pass fsImpl from the caller) for parity with readVerificationStatus's no-throw contract and testability. Also adds the Verification Module glossary entry to CONTEXT.md (review B3). --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
2e406e8e05 | chore: sync next package version to 1.6.0-rc.3 | ||
|
|
bcc5a6d1ba |
fix(#1634): honor capability hook matcher and node-prefix command (#1638)
* fix(#1634): honor capability hook matcher and node-prefix command Capability hook install (applyCapabilitySharedEdits) wrote each settings.json hook entry with no `matcher`, so a tool-scoped hook fired on every tool (a fail-closed guard could then block the whole session), and emitted a bare single-quoted script path so a .js-family hook from a git/tarball source without +x failed with Permission denied on every matching call. - Pass through an optional declared `matcher` (entry-level sibling of `hooks`); absent => omitted (match-all), so existing shipped capabilities are unchanged. - Validate `matcher` in the declaration (non-empty string, no control chars). - Emit `node <quoted-path>` for .js/.cjs/.mjs hooks (mirrors first-party); .sh and others keep the bare quoted path (unchanged). Root cause: the manifest hook schema (validator rule C4) was {event, script} only with no matcher, and applyCapabilitySharedEdits never read or wrote one; the command used shellSingleQuote(absScript) with no node prefix. Regression tests fail-first on both defects (matcher dropped; bare path) and pass after the fix; #1460 command assertions updated for the node prefix. * chore(#1634): backfill changeset pr:1638 * fix(#1634): resolve lint and windows CI failures - validator: replace the control-character range regex with a char-code loop. The literal /[\x00-\x1f\x7f]/ tripped ESLint's no-control-regex rule; char codes are equally precise and lint-clean. Behavior unchanged (still rejects matchers containing ASCII control characters incl. DEL). - test: gate the executable-bit precondition on POSIX. Windows fs does not honor POSIX write modes (a 0o644 write reads back as 0o666), so the precondition is meaningless there and failed the windows-latest lane. The node-prefix assertion — the actual fix — is platform-independent and still runs everywhere. * docs(#1634): amend ADR-894 for optional lifecycle hook matcher The `role: "feature"` `hooks[]` entry now carries an optional `matcher` (settings.json tool-scoping pattern: exact/pipe/wildcard/regex). Document the field in the §2 schema table and record a Grilling-amendments entry: the install path projects a declared matcher onto the emitted settings.json hook entry (absent = match-all, so shipped capabilities are unchanged), and per-runtime matcher projection (ADR-857 D8) stays a separate concern. This amendment ships with the fix that introduced the field rather than as a follow-up. * docs(#1634): record WINDOWS-POSIX-MODE-BIT-ASSERT defect in CONTEXT.md Capture the CI failure pattern from #1634/PR #1638 so it is not repeated: a test that writes a file with a POSIX mode and then asserts statSync().mode & 0o777 === <octal> passes on macOS/Linux but fails on windows-latest (Windows fs does not honor POSIX write modes — reads back 0o666). Added as a machine-greppable DEFECT predicate (symptom/examples/detect/fix-forward/ prevention) next to DEFECT.WINDOWS-TEST-PORTABILITY, with the fix-forward: gate the mode-bit precondition on process.platform !== 'win32' and keep the platform-independent behavioral assertion running everywhere. |
||
|
|
207d8f1697 |
fix(#1626): make the security gate severity-aware via per-threat severity (#1635)
workflow.security_block_on was documented as the minimum threat severity that blocks advancement, but threats carried no severity and the auditor's threats_open count (the SECURITY.md gate field) counted every open threat regardless of severity — so the threshold had no effect, and the auditor's block_on vocabulary (open/unregistered/none) did not even match the config enum (critical/high/medium/low/none). - planner: add a Severity column to the STRIDE threat register; assign severity per threat. - auditor: read severity; reconcile the <config> block_on domain to the severity enum; redefine threats_open as the count of OPEN threats whose severity is at or above block_on (none => 0). Below-threshold opens are reported as non-blocking and excluded from threats_open. - SECURITY.md template + planning-config.md reconciled. No gate-check site changed: threats_open == 0 stays the gate everywhere; only its computation is now severity-filtered. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f9d9dfb4bc |
fix(#1627): scale security rigor by ASVS level (planner disposition + auditor depth) (#1636)
workflow.security_asvs_level was display-only — the planner hardcoded 'mitigate if ASVS L1 requires it' and the auditor only echoed the level, so L2/L3 behaved identically to L1. - New reference gsd-core/references/security-asvs-levels.md defines L1 (opportunistic), L2 (standard), L3 (comprehensive) for both planner threat disposition and auditor verification depth (higher = superset). - planner: disposition now scales with the configured ASVS level (no hardcoded L1) + @-pointer to the reference. - auditor: verification depth scales with asvs_level (L1 grep-presence, L2 boundary/vector check, L3 end-to-end trace + bypass check). - planning-config.md + INVENTORY updated; planner kept under its 48K cap by extracting the goal-backward worked example to planner-guidance.md. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
94be6d5b60 | fix(#1625): resolve security config in secure-phase.md before auditor handoff (#1633) | ||
|
|
fc2a7c0555 | fix(#1615): install Windsurf slash workflows | ||
|
|
c28cccbf85 |
Merge pull request #1595 from davesienkowski/feat/1592-plan-drift-precheck
feat(#1592): add plan:pre codebase-drift pre-check before planner runs |
||
|
|
cbd21092a9 | fix(#1614): install Antigravity skills flat | ||
|
|
ba96c70b14 |
feat(#1602): deterministic coverage-metadata UAT routing for verify-work
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a deterministic classifier that `verify-work` consumes to route deliverables to auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic. - New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage block (extractFrontmatter can't — its `-` items are scalars-only; this is a focused parser, sibling of parseMustHavesBlock), validates each entry, and classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE typed-IR surface. Exposed via `uat classify-coverage --summary <f>`. - Auto-pass is the narrow proven case only: strict-boolean human_judgment:false AND non-empty all-`pass` verification AND zero validation errors. Everything else — judgment, empty/failing verification, malformed entry — routes to the human (fail-safe). A malformed block falls back to legacy prose extraction and surfaces an error; an absent block is byte-identical to pre-#1602. - execute-plan create_summary populates the block (fail-safe default human_judgment:true); verify-work extract_tests consumes it; create_uat_file marks auto-passed entries `source: automated`. - Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY, eslint/gitignore registration, and Diataxis docs (COMMANDS reference + USER-GUIDE explanation) updated. - Behavioral tests via the CLI (no source-grep); parser-robustness regressions for the null-entry/comment-header/mis-indent cases found in adversarial review. Closes #1602 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d1f7ba82f2 |
feat(#1592): add plan:pre codebase-drift pre-check before planner runs
Add a non-blocking, warn-only codebase-drift gate at plan:pre so a stale STRUCTURE.md is surfaced before /gsd:plan-phase spawns the planner, instead of being discovered mid-execution by the existing execute:wave:post gate. Gated on a dedicated workflow.plan_drift_precheck toggle (default true), independent of schema_drift_gate. Never blocks planning, never spawns the mapper agent at plan time. Review feedback (#1595): - Use a documented conventional-commit type (feat, not enhance) per CONTRIBUTING.md / gsd-validate-commit.sh. - Normalize the plan_drift_precheck command references to the colon prose form (/gsd:plan-phase, /gsd:map-codebase) to match plan-phase.md §5.65; registry regenerated from capability.json. - Make the test temp dirs hermetic: drain mkdtemp dirs in an after() hook via the helpers.cleanup() budget (local/no-raw-rmsync-in-tests-compliant). Closes #1592 Claude-Session: https://claude.ai/code/session_016JBiXEAofvB3prJim29XMS |
||
|
|
b2c0086c1b |
fix(#1574): resolve review — copilot instruction file is .github/copilot-instructions.md
GitHub Copilot reads repository-wide instructions only from .github/copilot-instructions.md (confirmed via GitHub Docs), not a root copilot-instructions.md. Aligns getProjectInstructionFile with the installer (runtime-config-adapter-registry installSurface 'copilot-instructions') and cites the docs source in the doc-comment. |
||
|
|
bf9bd1f4e0 | fix(#1529): emit runtime-native instruction file from new-project | ||
|
|
f20dc691c1 | chore: sync next package version to 1.6.0-rc.2 | ||
|
|
e12a2abfd8 |
feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed), enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way to browse or audit parked seeds on demand. This adds a read-only listing, following the established --list → workflow pattern (per the approved scope on - gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the seeds dir, returns { count, seeds[], summary } JSON with each seed's id, slug, status, scope, trigger_when, planted, title. Optional case-insensitive status filter. User-controlled content is sanitized (sanitizeForDisplay) and every path validated (requireSafePath); read-only. Independent of audit.scanSeeds, which only returns unimplemented seeds for the milestone surface. - /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that renders the seed table. Closes #441 * chore(#441): point changeset fragment at PR #722 * test(#441): allowlist list-seeds test in prompt-injection scan The test asserts that list-seeds neutralizes injection payloads (<system>, [INST]) embedded in seed content, so the fixtures legitimately contain those patterns — same as the sibling security tests already on the allowlist. * fix(#441): use canonical /gsd:capture colon form in list-seeds workflow Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new list-seeds workflow used the hyphen form. * docs(#441): sync help full.md + INVENTORY for --list-seeds Adds the --list-seeds entry to the help reference (help/modes/full.md, per bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json. * docs(#441): add --list-seeds how-to + drop phantom statuses Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers): - USER-GUIDE.md Seeds section (how-to): extend the task to cover auditing parked seeds on demand via --list-seeds, including the status filter — kept task-oriented per Diataxis how-to mode. - CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected from the list-seeds filter vocabulary; the system only produces dormant|active|triggered (src/audit.cts scanSeeds). Reference must be factually accurate and complete. * fix(#441): guard non-scalar status frontmatter in cmdListSeeds A seed with a bare `status:` line (extractFrontmatter yields {}) or a `status: [a, b]` value (yields an array) crashed the whole audit list: `(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string. Coerce every frontmatter read through a `fmStr` helper (mirrors the existing `typeof fm.id === 'string'` guard), so a non-scalar status falls back to dormant and non-scalar scope/trigger_when/title can no longer leak a raw array/object into the JSON contract. Title is now capped symmetrically. Adds regression coverage for empty and array `status:` and non-scalar fields. Refs #441 * docs(#441): align list-seeds workflow status vocabulary The load_seeds step listed `implemented` as an example status filter, but the real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds); `implemented` has no producer. Matches the earlier CLI-TOOLS.md correction. Refs #441 * refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested in-process (review minor #1). No behavior change. Filter comparison now matches the raw lowercased status (both sides already normalized) instead of sanitizeForDisplay(status); sanitization is for output, not matching (review nit #3). * test(#441): add fast-check property coverage and count=1 boundary for list-seeds Adds tests/list-seeds.property.test.cjs with four fast-check properties over deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing (review minor #1). Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2). * chore(#441): sync runtime launcher snippet into list-seeds workflow Propagate the current _runtime-launcher.snippet.sh (with non-Claude runtime home probes) into the new list-seeds.md workflow via scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation. * test(#441): record list-seeds.md in workflow size baseline (#1074) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ac40f070ef |
feat(#1318): require external reviewers to verify plan claims against source (#1421)
* feat(#1318): require external reviewers to verify plan claims against source /gsd-review built its external-reviewer prompt from plan text only and never asked reviewers to open the repo and verify claims, so a grounded HIGH could be outvoted by ungrounded LOWs. Add a concise, generic source-grounding block to build_prompt's Review Instructions: treat yourself as running in the working tree, open referenced files, cite path:line + mechanism, trace asserted mechanisms, downgrade to an open question if you have no file access, and know that grounded findings are weighted more heavily. Also clarify that CodeRabbit (a diff-only reviewer that never receives the prompt) must not be weighted as a grounded plan-level verdict in consensus synthesis. Workflow stays under its size cap (baseline bumped deliberately). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1318): add changeset for reviewer source-grounding Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1318): mark changeset docs-exempt (internal reviewer-prompt wording) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#1318): document reviewer source-grounding in COMMANDS.md; drop docs-exempt Review: a user-visible behavioral Changed warrants a docs touch, not a docs-exempt. Add a sentence to the /gsd-review entry in docs/COMMANDS.md (reviewers verify against source, cite file:line, grounded findings weighted higher) and remove the changeset docs-exempt marker so lint:docs passes via docs-updated. Also note the literal build_prompt test anchor. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1318): harden build_prompt fence extraction to be fence-run-aware Addresses maintainer review on PR #1421 (required-before-merge). The buildPromptReviewInstructions() test helper located the closing fence with `src.indexOf('\n```')`, which terminates at the FIRST triple-backtick line — so a build_prompt ```markdown block whose body embeds a fenced code example would truncate mid-content (dropping the `## Review Instructions` section) and give a spurious failure or false pass. Since this feature feeds source/plan content (which routinely contains code fences) to reviewers, that is a live fragility. Rewrite the extraction to be fence-run-aware, mirroring the CommonMark close rule in src/markdown-sectionizer.cts stripFencedCode: parse the opener's backtick run length, then close on the first line with >= that many backticks and only trailing whitespace — so a shorter nested fence is treated as content. Add a fail-first regression test (a 4-backtick outer fence wrapping a nested ```bash block) asserting the trailing `## Review Instructions` still extracts. Test-only change; no production .cts touched. Verified: test file 7/7, empirical fail-first proof the old indexOf logic truncated, full suite 4236/4236, eslint clean. Codex review: approve. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
8748e95ed1 |
fix(#666): pr-branch silently ignored planning.sub_repos (#667)
* fix(pr-branch): handle sub_repos from config with git -C (#666) Adds a `handle_sub_repos` step between `detect_state` and `analyze_commits`. When `planning.sub_repos` is set in config, the workflow now: - Reads sub-repo paths via `gsd_run query config-get sub_repos` - Skips the step entirely when the list is empty/null/[] - Scans each repo with `git -C "$REPO" status --porcelain` - Offers the user all/select/skip choices - For selected repos: creates a PR branch, commits all staged/unstaged changes, pushes, and opens a companion PR via `gh pr create` All git commands use `git -C "$REPO"` — never `cd "$REPO"` — because shell state does not persist between agent-executed commands. Closes #666 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore: update changeset pr number to 667 * fix(pr-branch): address maintainer review — correct seam, behavioral tests, robustness Resolves all three blockers and seven robustness issues raised in PR #667 review: Blockers: - Use `planning.sub_repos` (not top-level `sub_repos`) so config-get actually resolves - Replace prose grep test with behavioral fixture tests using runGsdTools + local bare repo - Extract sub-repo git work into new `cmdPrSubrepo` seam in src/commands.cts; never uses git add -A — stages explicit files only (universal-anti-patterns.md:44) Robustness: - Dirty-repo list persisted via mktemp/cat, not bash arrays (cross-block safe) - Branch name embeds repo slug (${CURRENT_BRANCH}-${REPO_SAFE}-pr) to avoid collision - push --set-upstream so gh pr create finds the branch - Sub-repo base branch resolved via ls-remote with fallback to repo's default branch - Remote slug parsed with /github\.com[:/]/ (handles SSH + HTTPS + .git-less URLs) - rollback() cleans up branch on any mid-sequence failure - node -e replaces jq (always available, no undeclared hard dep) Refs: #666 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(pr-branch): security guard, push timeout, rollback fix, porcelain fix Security (Blocker 1): - Use security.cjs validatePath() in cmdPrSubrepo for symlink-safe workspace containment check — rejects ../escape, absolute paths, and symlink traversal - Add negative regression test: '../escape' repo path must be rejected Robustness: - Push uses timeout: 60_000 ms (network op needs more than the 10 s default) - Capture prevBranchName before checkout -b so rollback uses explicit name instead of git checkout - (fails on fresh single-branch repos) - Porcelain path parse: line.trimStart().slice(2).trim() handles all XY combinations and the execGit global-trim edge case uniformly Tests: 17/17 pass, lint: 0 errors Refs: #666 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(pr-branch): move regression tests to commands.test.cjs, add core.quotePath=false - Move cmdPrSubrepo behavioral + workflow source-invariant tests from standalone bug-666-*.test.cjs into tests/commands.test.cjs under describe('pr-subrepo') per TESTING-SUITES.md policy (no new bug-* files). Adds allow-test-rule: source-text-is-the-product see #666 for the workflow-source-invariant suite. - Add -c core.quotePath=false to git status --porcelain call so non-ASCII filenames (e.g. café) are not C-escaped, keeping slice(2) parse correct. * fix(pr-branch): remove obsolete regression tests for sub-repos handling * fix(pr-branch): update workflow-size-baseline, add dirty-scan timeout - Regenerate tests/workflow-size-baseline.json for pr-branch.md growth (+handle_sub_repos step, +timeout addition). - Add { timeout: 10_000 } to the execFileSync git status --porcelain call in the handle_sub_repos dirty-scan (repo convention: every git subprocess is bounded, never hangs). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore: regenerate INVENTORY-MANIFEST after rebase onto next Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#666): handle rename staging and split changedFiles from filesToStage For git mv renames, the old path no longer exists in the worktree after the move — staging it with git add fails. Split parsing into changedFiles (both paths, for result.files) and filesToStage (new path only for renames; old is already staged by git mv). Also adds porcelain tests for staged renames, non-ASCII filenames, and a fast-check property test. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#666): rollback on push failure in cmdPrSubrepo If push fails the branch only exists locally; rollback cleans it up so the sub-repo is not left in a half-committed state. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#666): do not rollback after commit on push failure; add push-fail regression test Post-commit push failures are network/auth/policy issues — the user's work is already committed on the local branch. Calling rollback() at that point force-deletes the only ref holding the commit (data loss). Leave the branch in place and emit a retry instruction instead. Adds a regression test (pre-receive hook that rejects all pushes) asserting the branch and commit survive a push rejection so the failure path stays covered going forward. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore: regenerate INVENTORY-MANIFEST after rebase onto next Rebased onto current next (#1267 retired core.cjs). Stale tsbuildinfo and a leftover bin/lib/core.cjs build artifact were masking the drift — wiped both, rebuilt clean, and regenerated the manifest. gen-inventory-manifest --check now exits 0. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#666): validate sub-repo paths before git invocation in pr-branch.md The handle_sub_repos workflow ran git -C on raw planning.sub_repos config values at two points before the pr-subrepo seam's validatePath guard ever ran: the dirty-scan detection (git status) and the base-branch resolution (git ls-remote / remote show). A traversal entry could point git outside the workspace; an embedded newline could inject a spurious record into the newline-joined dirty-file output and into the shell-interpolated commit message. Adds a containment check + character allowlist to the dirty-scan node script (reject before any execFileSync), and a defense-in-depth shell case guard on the same value before the second, independent git -C invocation in the base-branch resolution block. Adds a behavioral test that extracts and executes the actual shipped node script from pr-branch.md (not a mirror) against a real traversal target and an embedded-newline entry, asserting neither reaches git or the dirty-file output. Also updates the stale cmdPrSubrepo doc comment: push failures no longer delete the branch (see prior commit), only stage/commit failures do. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#666): make sub-repo traversal scan test genuinely fail-first The outside repo's only change was an untracked file, which the ?? filter excludes — so the repo looked clean even with the guard removed, making the traversal assertion vacuous (it passed against a neutered guard). Commit the file first, then modify it, so the outside repo has a tracked dirty change: without the path guard it WOULD be reported dirty, so the test now fails-first. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#666): symlink-safe (realpath) sub-repo containment in pr-branch.md Finding A from re-review: the workflow guard used path.resolve, which only normalizes '..' textually and does not follow symlinks — so an in-tree symlink whose name has no '..' or '/' (e.g. "evil" -> /outside) passed both the charset filter and the resolve+startsWith check, letting git status / ls-remote / remote show run against a directory outside the workspace. The pr-subrepo seam already used fs.realpathSync (validatePath); this brings the workflow layer to parity. - dirty-scan: realpathSync the root once, and realpathSync each candidate before the containment check; skip on throw. - base-branch resolution: replace the weak `case *..*|/*` guard with a realpath containment check that yields a validated absolute SUB_REPO_DIR, and run git -C against that instead of re-concatenating $ROOT/$REPO_REL. - security test: add a symlink-escape entry and a positive control (legit in-root backend must still be reported). Confirmed fails-first — regressing the scan to path.resolve makes the symlink case leak. Also fixes a misleading-fallback minor: the workflow now checks the seam's exit status and skips the companion-PR step on failure, instead of printing "branch pushed, open PR manually" after a real stage/commit/push failure. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#666): harden pr-branch sub-repo flow against round-12 edge cases Pre-emptive hardening of the workflow changes from the symlink fix: - continue-outside-loop: the "skip companion PR on seam failure" block used a bash `continue`, but the per-sub-repo iteration is prose-driven (the agent loops, not a literal `for`), so `continue` would warn and no-op. Reframed as prose-gated control flow keyed on $SUBREPO_EXIT — no bash loop assumption. - Windows portability: the new symlink security case now degrades gracefully (try/catch around fs.symlinkSync; skip just the symlink assertion when symlink creation lacks privileges) so it doesn't hard-fail on Windows CI. Verified: seam exits 1 on error / 0 on success (error() → process.exit(1), propagated through the shim), so the $SUBREPO_EXIT check is meaningful; bash -n clean on the touched blocks; commands 156/156; lint:ci green; manifest in sync. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |