* fix(#1863): use named flags for state.* calls in executor + workflows
The named-only state-command router (parseNamedArgs) silently drops
positional args, so state.cjs threw its required-arg error and
metrics/decisions/blockers/session continuity were never recorded.
Convert record-metric / add-decision / add-blocker / record-session in
agents/gsd-executor.md to the named-flag form (mirroring execute-plan.md),
and fix the two remaining positional record-session calls in
gsd-core/workflows/milestone-summary.md and forensics.md. Recapture the
golden-install-parity fixtures and size baselines for the edited files.
Also fix a pre-existing detached-rebuild handle leak in
tests/graphify-auto-update.slow.test.cjs: three dispatch tests returned
after observing only the synchronous "running" status without awaiting the
detached rebuild's terminal state. That leak was latent until the new
#1863 regression block's added runtime shifted --test-force-exit timing
and surfaced it as a non-zero chunk exit. The three tests now await
terminal status via the file's existing waitForBuildStatus helper.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1863): add changeset
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Each of the 22 consumer agents now self-loads its configured agent_skills
in its mandatory init step, so .planning/config.json agent_skills.<type>
reaches the agent on every runtime — including Cursor and /gsd-autonomous,
where Skill()-delegated workflow bash init did not reliably execute.
- gsd-core/references/agent-skills-bootstrap.md: shared contract
(query + Read + dedup guard that skips when <agent_skills> is already
in the prompt, so Claude's orchestrator-side injection never doubles)
- 22 agents/gsd-*.md: one self-load line naming the agent's own type
- gsd-core/workflows/autonomous.md: note that delegated agents self-load
- tests/agent-skills-bootstrap.test.cjs: regression + parity (CONSUMER_AGENTS
bijection + fast-check property) — Generative-Fix-Divergence guard
- docs: ADR-1866, CONFIGURATION dual-injection How It Works, INVENTORY
row, Changed changeset
Closes#1866
execute-phase.md and gsd-ai-researcher.md are installed artifacts; their edits
shift install hashes and file sizes. Diff is scoped to those two files' hashes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(verify-phase): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1154)
Carry the edge-probe's existing `backstop` (non-inferable) tier through the
plan-phase projection as a structured flat-scalar marker instead of a prose
parenthetical, and make verify-phase abstain -> human_needed (never silent-pass)
on a backstop truth it cannot confirm with explicit evidence. Truth-axis mirror
of #644's prohibition judgment-tier (ADR-550 D4).
Engine (deterministic, CI-tested per ADR-550 D5 — never the LLM verdict):
- src/probe-core.cts: truthStatement/truthVerification normalizers, projectTruths
(conservative serializer), dispositionForUnverifiableTruth (backstop+no-evidence
-> unverified/flagged/insufficient_spec; backstop+evidence -> green; inferable
-> green, the over-abstention guard).
- src/roadmap.cts: coerceTruthToString now reads `statement` first so an object-form
backstop truth is surfaced, not dropped (Hyrum backward-compat for truth-readers).
Workflow/agent/docs: plan-phase emits the structured marker (flat scalar, ADR-550
#1278); verify-phase + gsd-verifier add the abstain arm; new references/honest-verifier.md;
FEATURES/COMMANDS document insufficient_spec; ADR-550 amended (truth-axis D4 mirror).
Decisions adopted (trek-e review): insufficient_spec feeds existing human_needed with a
distinguishable reason (no new VERIFIER_STATUS); changeset Changed; round-trip parity
test; abstain-on-unconfirmed-backstop regression test red-first.
Implementation notes (deviations from the issue's proposed file list, verified live):
- frontmatter.cts needs no change — its flat parser already round-trips object-form truths.
- verify.cts needs no change — it grades artifacts/key_links structurally; truths are
LLM-graded at the workflow layer, so consumption lives there + the deterministic helper.
- No CJS<->SDK hand-sync — the SDK seam was retired (ADR-0174); src/*.cts is sole source.
Regenerated artifacts: golden-install-parity fixtures, INVENTORY-MANIFEST, size baselines.
* chore(#1154): add changeset (Changed) for honest verifier
User-facing changelog fragment for #1738. Typed `Changed` (not `Added`) per
trek-e review condition 3 — the verify behavior shifts for backstop-bearing specs
(a confident silent `passed` becomes `human_needed`), which is user-visible even
though the schema marker is additive.
* docs(#1154): score-formula also excludes abstained insufficient_spec truths (review nit-1)
trek-e review nit: the verify-phase score sentence said PRESENT_BEHAVIOR_UNVERIFIED
truths were "the only ones excluded" from verified_truths. Post-#1154 an abstained
`insufficient_spec` backstop truth is also excluded (it is not ✓ VERIFIED and routes
to human_needed). Behavior was already correct; this tightens the wording.
Regenerated golden-install-parity fixtures + workflow-size baseline for the touched
verify-phase.md. (Nit-2 — a dedicated insufficient_spec_items frontmatter list — is
intentionally not taken: the current design is ADR-550-D4-conformant, the abstain
cause rides as a distinguishable report reason, and adding it would exceed the
approved scope.)
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking
Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the
largest untrusted channel) in gsd-read-injection-scanner; shared
untrusted-input-boundary reference @-included by the 8 ingest agents
(randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring);
opt-in security.injection_blocking (default advisory — non-breaking).
arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth).
* fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized
- A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a
circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already
in the transcript. The prompt-level data/instruction boundary is the primary control.
- A2: registered security.injection_blocking in the config schema + defaults manifests (default
false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads.
- A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention).
- A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale).
- A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input.
- Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) +
drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as
the noted pre-existing follow-up.
* fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate
The new reference quotes injection phrases ('ignore previous instructions',
'you are now…') as examples agents must NOT comply with, tripping the repo's
own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red
on HEAD). Allowlist it alongside the other security docs (security-model.md,
TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS
scanner test doesn't scan references/, so only the shell gate needed it.
Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15.
* fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer
trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so
no named web-ingress agent is uncovered, keeping the two justified additions
(gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10.
- gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset.
- gsd-assumptions-analyzer reads 5-15 codebase source files (external/source-
document ingress per the boundary), though it has no web tools.
INGEST_AGENTS in the isolation test now asserts all 10; size baselines
regenerated (+60 bytes each, both well under the DEFAULT cap); changeset
reworded 8 -> 10.
Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39.
* docs(#1577): document security.injection_blocking + boundary seam
trek-e Major 2 + Minor:
- docs/CONFIGURATION.md: add the top-level security.injection_blocking key to
the Full Schema and a Security Settings subsection, distinguishing it from
the workflow.security_* namespace; honest circuit-breaker-not-redactor
framing matching ADR-1577 / security-model.
- CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry.
Verified: lint:docs ok; config-field-docs + contributor-standards green.
* test(#1577): make read-injection property test git-text, not binary
trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as
degenerate-edge inputs. The NUL is what actually made git classify it binary
(git binary = NUL in first 8K). Replace both with text-safe escapes that keep
the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File
now diffs/blames line-by-line.
Verified: property test 2/2; no NUL/raw-noncharacter bytes remain.
* docs(#1577): align untrusted boundary docs
Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner.
* docs(#1577): align ADR ingest agent count
Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs.
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
workflow.security_block_on was documented as the minimum threat severity
that blocks advancement, but threats carried no severity and the auditor's
threats_open count (the SECURITY.md gate field) counted every open threat
regardless of severity — so the threshold had no effect, and the auditor's
block_on vocabulary (open/unregistered/none) did not even match the config
enum (critical/high/medium/low/none).
- planner: add a Severity column to the STRIDE threat register; assign
severity per threat.
- auditor: read severity; reconcile the <config> block_on domain to the
severity enum; redefine threats_open as the count of OPEN threats whose
severity is at or above block_on (none => 0). Below-threshold opens are
reported as non-blocking and excluded from threats_open.
- SECURITY.md template + planning-config.md reconciled.
No gate-check site changed: threats_open == 0 stays the gate everywhere;
only its computation is now severity-filtered.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
workflow.security_asvs_level was display-only — the planner hardcoded
'mitigate if ASVS L1 requires it' and the auditor only echoed the level,
so L2/L3 behaved identically to L1.
- New reference gsd-core/references/security-asvs-levels.md defines L1
(opportunistic), L2 (standard), L3 (comprehensive) for both planner
threat disposition and auditor verification depth (higher = superset).
- planner: disposition now scales with the configured ASVS level (no
hardcoded L1) + @-pointer to the reference.
- auditor: verification depth scales with asvs_level (L1 grep-presence,
L2 boundary/vector check, L3 end-to-end trace + bypass check).
- planning-config.md + INVENTORY updated; planner kept under its 48K cap
by extracting the goal-backward worked example to planner-guidance.md.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1478,#1479,#1480): prohibit ungrounded baselines, error-suppressing fallbacks, and stale-artifact authority in verify blocks
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: add changeset for #1478/#1479/#1480 planner verify gate fix
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: correct changeset format
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1478,#1479,#1480): fix test contract violations from new planner dimensions
gsd-planner.md exceeded both the planner-decomposition 48K char limit and
the reachability-check 50K char limit after the new HARD RULE blocks were
added inline. The full rule details already exist in planner-antipatterns.md
(added in the same PR); replace the verbose inline blocks with a single
@-reference pointer to the antipatterns file, reducing the file from 50981
to 49130 chars (under both limits).
Also regenerate tests/agent-size-baseline.json to reflect the new sizes of
gsd-planner.md and gsd-plan-checker.md.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds a region/function-scoped negative-grep idiom to the gsd-planner verification guidance plus a warn-only `validate_plan` check (`scanFileWideNegativeGateConflict`) that flags when a task's file-wide negative grep bans a construct a sibling task legitimately requires elsewhere in the same file. ReDoS-safe (linear, no RegExp on author patterns); region-scoped gates are exempt. Warn-only — never errors, never flips `valid`.
Closes#968
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds mcp__perplexity__* to both researcher profiles (generated source-of-truth) and regenerates the agents; adds a generative dispatch-table↔tools parity guard so future provider drift fails CI. Regenerates the agent-size baseline for the +20-byte frontmatter growth.
Fixes#1284
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* enh(#966): gate behavior-dependent truths on behavioral evidence in gsd-verifier
Introduce a per-truth PRESENT_BEHAVIOR_UNVERIFIED state for must-haves that assert a state transition or a cancellation/cleanup/ordering invariant whose only evidence is symbol presence + wiring. Such truths are excluded from the verified_truths score, reported as a behavior_unverified count, recorded in an always-on behavior_unverified_items frontmatter list, and routed to the existing human_needed sink — so a clean N/N can no longer be reached on symbol presence alone.
The overall-status vocabulary and the src/verification.cts seam are unchanged (the new state is per-truth only); gaps_found keeps decision-tree precedence; override-passed truths still count toward verified_truths. Mirrors the calibration into the shipped verify-phase.md workflow (with an infra/foundation carve-out), the VERIFICATION.md templates, and docs (planning-artifacts.md, AGENTS.md). gsd-verifier.md kept under its 48KB LARGE cap; size baselines regenerated.
Closes#966
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#966): add changeset fragment for gsd-verifier behavior-unverified calibration
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1243): consume Claude plugin-provided skills via native Skill-tool directive + grant Skill to agent_skills-consumer agents
- Relax global skill name validation to accept namespaced form `^[A-Za-z0-9_-]+(:[A-Za-z0-9_-]+)*$`
- Namespaced names (containing colon) on claude runtime emit a Skill-tool load directive instead of a @-include line
- Namespaced names on non-claude runtimes are skipped with a warning
- Bare unresolved names retain existing warn-and-skip behavior (no promotion to directive)
- Grant `Skill` tool to all 22 agent_skills consumer agents; 5 generated agents updated via research-profiles.cjs + regen, 17 hand-authored agents edited directly
- Add 16 TDD tests in describe('bug #1243') covering happy/mixed/precedence/negative/cross-runtime/regression/grant cases
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(#1243): document plugin-provided skills in agent_skills
Update the Agent Skills Injection reference in CONFIGURATION.md with
the three entry forms (project-relative, global:<name>,
global:<plugin>:<skill>), the Claude-only runtime behaviour of the
namespaced form and the warn-skip on other runtimes, the plugin
pre-install prerequisite, and the consumer-agent Skill tool grant.
Add docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md with a
step-by-step guide for installing the plugin, locating the namespaced
skill name, wiring it into agent_skills, and verifying injection.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#1243): align agent_skills docs with emitted block format + mixed-block regression test (code-review)
- Replace two-section mixed-block example (bogus "Load these plugin-provided skills using the Skill tool:" header) with the actual single-section inline format in CONFIGURATION.md and docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md
- Fix quoted warning text in how-to doc to exactly match the emitted string: [agent-skills] WARNING: Plugin-namespaced skill "global:<name>" requires a Skill-tool-capable runtime (claude) — skipping on runtime "<runtime>"
- Replace phantom agent slugs (gsd-checker, gsd-researcher, gsd-advisor, gsd-synthesizer) in CONFIGURATION.md Supported Agent Types with real agents/gsd-*.md examples (gsd-plan-checker, gsd-phase-researcher, gsd-code-reviewer, gsd-ui-auditor, gsd-research-synthesizer)
- Add byte-identical mixed-block regression test: one path-resolvable global skill + one plugin-namespaced skill on claude runtime → asserts r.ir.block === single-section interleaved block, no secondary header
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(#1243): regenerate agent-size baseline for the Skill-tool grant
The 22 agent_skills-consumer agents each grew +7 bytes from adding `Skill`
to their tools list; refresh the committed per-agent size baseline (#1074 guard).
* chore(#1243): add Added changeset fragment
* fix(#1243): traceable allow-test-rule ref + separator-agnostic byte-identical tests (CI)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds the spec-time prohibition probe (spec-phase Step 5.6) — the second adapter of the probe-core resolution model. Surfaces unwritten must-NOT constraints as negative SPEC acceptance criteria with test/judgment verification tiers; fail-closed at verify time. Per ADR-550. Closes#644.
* fix(#1205): roadmapper applies phase_id_convention to generated phase IDs
- Add Phase ID Convention section to <phase_identification> block:
documents sequential (default) vs milestone-prefixed forms, and
instructs the agent to read phase_id_convention from config.json
- Update <output_formats> to show both header and checklist forms for
sequential and milestone-prefixed conventions with examples
(e.g. ### Phase 1-01: Name, - [ ] **Phase 1-01: Name**)
- Add TDD regression test tests/bug-1205-roadmapper-convention.test.cjs
(5 assertions, confirmed fail-first then pass after fix)
- Update tests/agent-size-baseline.json to reflect legitimate growth
- Add .changeset/brave-otters-leap.md (Fixed, pr:0 placeholder)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: backfill changeset pr: 1215 for fix/1205
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1205): move phase_id_convention regression into roadmapper-granularity.test.cjs
lint-regression-test-names rejects new standalone bug-NNNN-*.test.cjs files;
regression cases must live in the owning module's test file.
Move the 5 phase_id_convention assertions (#1205 regression) from the
removed tests/bug-1205-roadmapper-convention.test.cjs into
tests/roadmapper-granularity.test.cjs as a new describe block, alongside
the existing granularity calibration tests. Also update the allow-test-rule
comment to cover both #163 and #1205 surface contracts.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1205): fix lint-allow-test-rule-refs for roadmapper-granularity
- Add issue ref (see #1205) to allow-test-rule comment in
tests/roadmapper-granularity.test.cjs so lint-allow-test-rule-refs
passes (new exemptions require #NNN per ADR-456)
- Prune stale 'source-text-is-the-product' entry from
scripts/lint-allow-test-rule-refs.allowlist.json (ratchet-down;
comment now compliant and no longer needs grandfathering)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(#956): add MemPalace memory capability (ADR-857 feature plug-in)
Adds an opt-in, default-resilient ADR-857 feature capability that wires
MemPalace (local-first memory: MCP server + CLI) into the GSD loop:
deliberate recall before discuss/plan and verbatim + temporal-KG capture
at phase boundaries. Three memory modes (augment default; kg_backend and
replace forward-declared). Master gate mempalace.enabled (default off);
every hook onError:skip, zero gates; absent/disabled MemPalace => loop
unchanged. Transport is rendered-markdown only — MemPalace runs
out-of-process, no third-party code in gsd-core (ADR-857 §7).
Capability: capabilities/mempalace/ (manifest + 2 fragments), skills
commands/gsd/mempalace-{recall,capture}.md, agent
agents/gsd-mempalace-curator.md. Registration: ns-context router,
utility cluster, KNOWN_SKILLS, help full.md, model-catalog, copilot
install list, size baselines; regenerated capability-registry +
inventory manifest. ship:post wired into ship.md (wire-on-demand).
HELD on #1196: this capability also declares hooks at discuss:pre and
discuss:post, which are structurally un-wireable until the host-loop
conformance model covers the discuss phase (discuss-phase.md is not in
HOST_LOOP_FILES). The phase6-capstone-conformance gate therefore fails
on exactly those two orphaned points by design — see #1196. Once #1196
lands, rebase onto next and the gate goes green with no further change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(#956): backfill changeset PR number (#1201)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Extends the #1074 scheme to tests/agent-size-budget.test.cjs, which used the
same assertTightCeiling tier ratchet but was still line-based (never rebased in
#717). Completes the migration — the last part of the #1074 epic.
- Rebase agent sizing from lines to LF-normalized bytes (#717/#683).
- Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests);
add a per-agent baseline (tests/agent-size-baseline.json) as the primary
anti-creep, and byte hard caps (XL 56 KiB / LARGE 48 KiB / DEFAULT 24 KiB),
each above its tier high-water with real headroom. No separate new-file cap:
a net-new agent is DEFAULT-tier, already bounded by the DEFAULT cap.
- Keep the agent-classification tests verbatim.
- scripts/workflow-size.cjs: add generic measureMdFiles(dir, predicate)
(workflows + agents share one byte-measurement path); measureWorkflows now
delegates to it.
- scripts/update-size-baseline.cjs: one 'npm run size:baseline' now regenerates
BOTH the workflow and agent baselines (gsd-* filter for agents).
Rebased onto next after PR 2/3 (#1096) merged: replicate the
scripts/lib/workflow-size.cjs -> scripts/workflow-size.cjs move (PR 1/3) across
the generator and the agent test's require; regenerate the agent baseline
against current agents (a uniform +170 B preamble drift on all 33 since
authoring).
Addresses the #1097 review (trek-e):
- BLOCKER (acceptance criterion 5): document the agent contract in CONTEXT.md.
Adds RULESET.AGENT_SIZE_BUDGET (caps 57344/49152/24576, per-agent baseline,
dual size:baseline, shared measureMdFiles seam) and disambiguates it from the
separate DEFECT.AGENT-FILE-SIZE-CAP-BREACH 45K-CHAR guard (two units, two
purposes).
- Docs: now that #1096's docs/TESTING-SUITES.md "Workflow size budget" section
is in next, fold in the agent coverage here (renamed to "Workflow & agent
size budget"): agent caps + per-agent baseline + the how-to + reference rows,
and the disambiguation from the 45K-char guard.
- Minor (negative proof): add a boundary-fixture test exercising the hard-cap
comparison at cap-1/cap/cap+1 through the real lfByteCount path, so a future
threshold/operator edit can't silently neuter a cap.
- Nit: align the tier test name wording ("stays within") with the <= operator.
Negative proof on a real tracked agent (gsd-planner): baseline catches +10 B;
XL hard cap catches 57,516 > 57,344 with the baseline current.
Closes#1095 (PR 3/3 child); landing this completes the #1074 epic.