1fe85cd43eb9edbe3eec51fcd5a301a2fec5b002
7 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e20744eacb |
enhance(#3884): failure is a value — strict argv, and --pick that signals absence (#3922)
* test(#3884): failing-first coverage for strict argv and absence-signalling --pick ADR-3473 §8.4 says failure is a value. Three families currently encode failure as success, and this commit pins each one RED before the fix lands. Measured on this tree, 2026-08-26: gsd-tools generate-slug "test" --pick nonexistent -> empty stdout, exit 0 (#3365) gsd-tools audit-open --pick nonexistent_field -> dumps the entire human-readable audit report, exit 0 gsd-tools generate-slug "Hello World" --raw --pick bogus -> prints "hello-world", another field's value, exit 0 gsd-tools query state.planned-phase 3 (positional, no --phase) -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted current_phase_name (#3358) tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract ("returns empty string for missing field", success === true). That assertion is replaced by the required behavior rather than deleted. The new parseNamedArgs block calls the spec-object signature that does not exist yet, so it fails today by construction. The 11 existing behavior-lock tests are left untouched here; they are corrected in the implementation commit. C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180 Decision 4(b). A unit assertion on the parser would have passed throughout this defect's life. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3884): failure is a value — strict argv, and --pick that signals absence Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable ways to say "I could not answer". parseNamedArgs (src/command-arg-projection.cts) Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the hub's Result shape instead of a bare Record. Declaring the positional arity is what makes #3358's call site unrepresentable rather than merely detectable: an unrecognized flag or a token past the declared boundary is now InvalidArgs, naming the offending token and listing the accepted flags. The legacy positional-array call shape throws a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale hand-written .cjs call site fails loudly instead of destructuring undefined off a Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a projection over the one parser, not a second parser. Measured before, against a STATE.md with a populated phase-2 block: query state.planned-phase 3 (positional, no --phase) -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted current_phase_name After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical. The flag form is unchanged and still updates STATE.md. --pick <field> (gsd-core/bin/gsd-tools.cjs) extractField returns {found,value}, and the pick block no longer shares one catch between "output was not JSON" and "field was absent". An absent field exits 1 with pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1 with pick_output_not_json instead of dumping the command's entire output. A field that is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer, not a failure, and it is what keeps `--pick count` printing 0 on a fresh project. Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" — a different field's value, confidently, at exit 0. ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0 would demote "could not answer" to "the answer is zero" — the hazard docs/how-to/resolve-unreachable-guard-findings.md already warns against. Guard ledger (ADR-3473 Decision 6) scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls, a nullglob mechanism this change does not touch) is retained in full, as are the shared scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file is not deleted. Call-site audit 45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an if test, && chain, or a pipeline whose status is consumed, and no shell block in workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind a prior found/existence check. No ADR-3409-class "field the command never produces" remains. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows Two review findings, both fixed here rather than recorded as limits. 1. A newline in an untrusted token forged a second stderr line. Before, plain-text mode: $ gsd-tools query state.planned-phase $'foo\nError: forged second line' Error: unexpected positional argument "foo Error: forged second line" After: Error: unexpected positional argument "foo\nError: forged second line" --json-errors mode was never affected — io.error runs that payload through JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the three new InvalidArgs reasons plus the two new --pick diagnostics all interpolate a token that comes straight from argv. Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at every interpolation site — not a copy per site. It is deliberately NOT applied inside error() itself: several callers in this tree emit intentional multi-line diagnostics, and escaping newlines there would mangle them. The available-top-level-keys list needed the same treatment for a reason the review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user document and echoes that document's own keys into the diagnostic. Verified reachable — a frontmatter key containing a newline reaches the key list — so formatKeyForDiagnosticList is guarding a live path, not a hypothetical one. Ordinary keys still render plain and unquoted; a fix that merely dropped the key would also have passed a "one line" assertion, so the test pins the escaped key's presence too. 2. Five behavior-table rows were implemented but nothing pinned them: B7 a dotted path that dies partway B9 bracket syntax on a non-array B10 a negative array index, in and out of range B14 a JSON root that is not an object B17 an @file: payload over 50KB B17 is the load-bearing one. output() writes @file:<path> instead of inline JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no test, a future reordering of those two steps turns every large result into a false pick_output_not_json. The fixture seeds 1200 phase directories and measures the payload at 62474 characters, asserting the spill actually happened rather than assuming it. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): correct the strict-argv surface against a full verification run The first full run came back with 90 failures across 12 files, none in the new tests. They were the argv surface telling me what it actually is. Ten root causes; each classified before anything was changed. I over-implemented, and that is reverted. ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional tokens". It says nothing about a value flag whose value is missing. Making that an error was my design decision, not the rule, and it broke a deliberately recorded contract: `--prd` with no value resolving to null (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5; tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The "requires a value" branch is deleted outright rather than kept behind an option — an unused strictness mode is speculative generality. Unknown-flag and unexpected-positional rejection, which is what §8.4 actually mandates, is unchanged. --wave needed a third flag kind the original design did not anticipate. `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the shipped workflow reconstructs and passes it (execute-phase.md:84), while #2932 records token-PRESENCE semantics: the CLI cares only that the flag appeared, and the value belongs to the workflow layer. That is neither a boolean flag nor a value flag, so `optionalValueFlags` now exists — presence-only in `data`, and the validation cursor consumes a following non-flag token so it is not reported as a stray positional. Every other declared boolean flag was checked against every argument-hint and prose usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one of this shape. Five tests were pinning forms that never worked. tests/adr857-core-without-capabilities.test.cjs passed `init plan-phase --phase 01-stub`, but the documented form is positional (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form is the literal string "--phase". Measured on the pre-fix build against a real .planning/phases/01-stub/ directory: init plan-phase 01-stub -> phase_found=true init plan-phase --phase 01-stub -> phase_found=false The test asserted only exit 0 and key presence, so it had been green while proving nothing about phase resolution. Corrected to the documented form and strengthened to assert phase_found === true. Same class in state.test.cjs (`--plan-count`, a flag that does not exist; the real one is `--plans`), milestone-archive.test.cjs (`init new-milestone --json`, silently ignored), and concurrency-safety.test.cjs (a bare positional field name whose OR-assertion passed because a whole-document dump happens to contain the substring it looked for). Six handlers had no argv validation at all — the same #3358 shape this phase exists to close, found while fixing the rest: init verify-work / phase-op / review / todos / remove-workspace read args[2] with nothing checking the rest, and validate health read --repair/--backfill through a bare args.includes() scan that bypassed the parser entirely. All now go through the seam, so the flag has one owner. tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's local expectation does not override §8, so they are inverted and renamed — a test still called "ignores an unrecognized flag" while asserting rejection would be its own defect. Row C6's point is its PWNED canary; that assertion is kept verbatim and only its exit-status expectation changed, because the hostile token is now rejected rather than absorbed. The blast-radius estimate in 40-design.md is corrected rather than quietly left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was accurate for what the graph can see — parseNamedArgs's callers. It cannot see that those callers' handlers accept argv shapes wider than the code reading args[2] suggests, which is where the real surface was. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert Second full run: 46 failures, down from 90. Four causes, two of them mine. Reverted `validate health` entirely — it was scope creep, and it broke a real flag. ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`. The previous commit routed `validate health` through the parser on the reasoning that a flag should have one owner. That was wrong twice over: §8.4 names parseNamedArgs and count queries, and `validate health` was never a parseNamedArgs call site — it read its flags, just not through the parser, so it had no silent-drop defect to fix. Tightening it omitted `--json`, which the health-diagnostic suites use heavily. The handler is now byte-for-behaviour back to its pre-branch form. `validate context` stays converted: it genuinely was a call site, and its `--json` is now declared rather than read by a second `args.includes` scan. The five handlers that had NO validation at all — init verify-work / phase-op / review / todos / remove-workspace — stay fixed. Those read args[2] with nothing checking the rest, which is the #3358 shape this phase owns. Finished the A2/A3 revert. Three tests still encoded the deleted "a value flag with a missing value is an error" rule, including one added by the previous commit for that rule. All three now assert the reverted null contract, and the ones whose titles said "rejected" are renamed — a test named for a contract it no longer asserts is its own defect. `--wave=` and `--wave --weird` are correctly rejected. Neither is documented in commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and neither is emitted by the shipped prompt layer, so both are unrecognized tokens that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the property it exists for — asserted directly now, at the parser, that `--wave` does not swallow a following flag as its value — and only its exit-status expectation changed. A contradiction inside this branch, surfaced by the audit and resolved the safe way. Two pre-existing #3573 tests call `state begin-phase '2'` and `state planned-phase '2'` with a bare positional, relying on the old permissive parser to ignore it. This branch's own #3358 regression test requires that exact argv to be REJECTED. The two are mutually exclusive. Widening the router to accept a bare positional — mirroring complete-phase — would have silently re-opened #3358, and was verified to do exactly that: with the widened router, `query state.planned-phase 3` returned exit 0 and wrote current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the two #3573 tests move to it. Their assertions were never about the call shape — only that total_phases survives the resync — and both still pass. complete-phase is untouched: its bare positional IS documented, and it keeps the dynamic boundary and the negative-space note that record why. The audit that produced this is in the PR body: for every handler whose declaration changed, the flags it reads anywhere in its body, the flags the shipped surface documents, and the shapes the suite passes, compared. The `--json` miss was a pattern, not an accident — declaring a handler's flags from its parseNamedArgs call alone misses whatever it reads elsewhere. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3884): backfill the changeset PR number Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
95f7c14413 |
fix(#3642): stop the single-section total_phases leak into an absent milestone (#3727)
* test(#3642): failing-first single-section leak rows * fix(#3642): gate the unbounded total on any-milestone-section, not >=2 * test(#3642): rewrite the 3185 wrapper row to the withhold contract * docs(#3642): glossary amendment for the >=1 sibling; changeset * chore(#3642): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
b08af152e4 |
fix(#3573): keep the stored total_phases when the roadmap is absent at state-write time (#3595)
* test(#3573): pin stored-total retention when the roadmap is absent at state-write time Failing-first regression for #3573: with ROADMAP.md absent and a milestone asserted, every state.* write persisted the phase-directory count as progress.total_phases (5 -> 1 in the issue) — only STARTED phases count, quietly defeating #549's single source of truth. Rows pin the stored-value outcome + stderr warning across record-session and begin-phase, the fresh-project doctrine (no milestone asserted -> dir count stays), and the roadmap-present control. * fix(#3573): keep the stored total_phases when the roadmap is absent at state-write time The #3354 withhold covered milestoned-but-unbounded roadmaps but not the roadmap-absent shape: with ROADMAP.md unreadable the #549 heading counter never runs, milestoneBounded is vacuously true, and every state.* write persisted the phase-directory count as progress.total_phases — counting only STARTED phases (5 -> 1 in the issue). When the STATE asserts a milestone (storedMilestone), the stored frontmatter total now wins and a (#3353)-style stderr warning names the condition; with no asserted milestone the disk count stays authoritative (fresh-project doctrine). * fix(#3573): thread stored milestone into the state json read for write/read parity; discriminate the doctrine row; pin planned-phase Review findings: cmdStateJson passed storedMilestone=undefined so the new withhold never fired on the read surface — state json reported the dir count while the persisted file preserved the stored total (exactly the divergence #3354 closed for its shape). The fresh-project doctrine row now uses stored 5 vs dirs 2 so a milestone-gate-less withhold mutant cannot survive it; the third issue-named verb (planned-phase) is pinned. * chore(#3573): add changeset fragment * chore(#3573): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
86a808f8ad |
fix(#3355): pick phase-dir dedup survivor from content, not mtime (#3486)
* fix(#3355): pick phase-dir dedup survivor from content, not mtime * chore(#3355): add changeset fragment * chore(#3355): backfill PR number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
70b5c1a1bf |
fix(#3354): preserve stored total_phases when milestone is unbounded (#3480)
* fix(#3354): preserve stored total_phases when milestone is unbounded * chore(#3354): backfill PR number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
1bc7f7e6b0 |
test(#3339): fold the state/phase/dispatch & model-profile issue-* cluster — Wave 7
Folds 9 legacy issue-*.test.cjs regression files (140 test() blocks) into their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053). LAST of 4 issue-* waves — closes out the 74-file fix-*/issue-* backlog (pending BUG_FILE_RE extension, held for a follow-up commit until Wave 6 is confirmed merged, per the epic's own zero-backlog precondition). - issue-2828-flat-roadmap-total-phases.test.cjs (1) + issue-3204-state- writer-phase-count.test.cjs (21): both target state-document.cjs buildStateFrontmatter via different CLI entrypoints — merged jointly into state-document.test.cjs, 0 dropped. - issue-2945-phase-complete-checkbox-rollback.test.cjs (4) + issue-2949- phase-complete-stage3-sentinel.test.cjs (4): both target phase.cts cmdPhaseComplete; issue explicitly warned of overlap — verified disjoint fixtures/assertions, 0 dropped, merged into phase.test.cjs. - issue-2927-reviewer-lane-overlay-invocation.test.cjs (10) merged into review-lane-descriptor.test.cjs. - issue-2939-dispatch-flatten-maxdepth.test.cjs (9) merged into host-integration.test.cjs, 2 dropped as verified exact duplicates. - issue-2977-frontmatter-bom.test.cjs (5) merged into frontmatter.test.cjs. - issue-2045-third-party-skills-surface.test.cjs (6) merged into capability-loader.test.cjs. - issue-2517-runtime-aware-profiles.test.cjs (80, the largest single fold in the epic) merged into model-resolver.test.cjs, 1 dropped as a verified true duplicate (checked against src/model-resolver.cts logic, not just title similarity). Fixed a genuine eslint irregular-whitespace finding: a literal BOM character embedded in a doc comment (pre-existing content from the original #2977 source, illustrating what a BOM looks like) — replaced with a readable U+FEFF notation. 3 stale doc references found and fixed (docs/adr/2313, 3180, 443). Zero net test-coverage loss. No production code changed. |
||
|
|
4bd6fb066b |
chore(#2880): close ADR-2143 deployment misses — table-regex fingerprint + state-document seam migration (#2889)
* chore(#2880): close ADR-2143 seam misses — widen table-regex fingerprint, migrate state-document onto the seam The no-adhoc-markdown-parsing rule matched only a negated class whose sole member was a pipe ([^|]), so the stricter and more common [^|\n] spelling evaded it entirely -- src/state-document.cts hand-rolled exactly that shape and linted clean. Widen the fingerprint to any negated class excluding a pipe, which is the ADR-2143 section 7 prohibition as written. With the rule fixed, state-document.cts goes red. Replace tableRowPattern with locateFieldRow: a line scan using the markdown-table seam's splitTableRow for cell semantics, returning the value cell's byte range, and splice that range instead of running a whole-document content.replace. An edit now physically cannot cross a row boundary (section 4). Behavior is frozen -- stateReplaceField has 79 dependents across 5 command processes. Characterization tests lock all 14 table-branch rows plus CRLF, extract round-trip and the withFallback caller shape; a fast-check property asserts every non-target line stays byte-identical. Refs #2880, epic #2143 * fix(#2880): address adversarial review — lone-CR rows, field-name padding, quadratic scan, over-broad fingerprint Isolated adversarial review found four defects in the first commit. 1. locateFieldRow split lines on \n only. JS treats a lone \r as a line terminator, so the regex it replaced matched rows separated by bare CR. "| Phase | 3 |\r| Other | 9 |" returned 3 before and null after. Now CR, LF and CRLF are all terminators, byte offsets unchanged. 2. The field name was normalised with trim().toLowerCase(). The old regex embedded it verbatim, so its whitespace had to be absorbed by the row's own padding -- and because the group is a literal-character match rather than a whitespace class, a tab-padded cell does not accept a space-padded name. Replaced with an offset-aligned search reproducing the original backtracking exactly. 3. The widened fingerprint regex had two unbounded [^\]]* around an optional and ran quadratically over every regex source in every linted file: 256000 chars took 23 seconds. Replaced with a single-pass scanner that never rescans; the same input is now ~1ms. 4. The fingerprint also matched non-table idioms such as [^\s|] and [^"|]. Narrowed to a class excluding the pipe plus only \n, \r or \t. Differential fuzz against origin/next: 20000 cases, 0 mismatches. Refs #2880, epic #2143 * test(#2880): drop wall-clock assertion from the ReDoS regression guard local/no-elapsed-assertion flagged the elapsed-time check, and CLAUDE.md bans timing assertions outright as flaky. The 256000-char input stays as the regression guard for the quadratic scan; correctness of the verdict is what is asserted. If the quadratic path returns, the test stops completing and surfaces as a suite timeout rather than a silent pass. Also adds the changeset fragment for #2880. Refs #2880 * fix(#2880): spec-correct case folding, property tests, naming Code-review findings. The field-name comparison used toLowerCase(). The regex it replaced used /i WITHOUT /u, and ECMAScript Canonicalize deliberately does not fold a non-ASCII character onto an ASCII one -- KELVIN SIGN U+212A matched ASCII K where the old code returned null. Replaced with spec-correct Canonicalize, including the multi-character uppercase case (eszett -> SS), which a naive uppercase comparison also gets wrong. Added the fast-check property tests CLAUDE.md requires for parsers: one for the negated-class scanner, one for the field-name fold semantics, each against an independent reference implementation. Both reference impls failed on first run against real bugs, so neither property is vacuous. Renamed p2/p3 to name the exactly-three-pipes invariant, and reduced a duplicated comment to a cross-reference. Differential fuzz vs origin/next: 20000 runs, 0 mismatches, with the harness proven to discriminate the KELVIN case. Refs #2880 * chore(#2880): backfill changeset PR number (#2889) * docs(#2890): correct the local ESLint plugin path in CONTEXT.md CONTEXT.md named the local AST-rule plugin directory as scripts/eslint-rules/, which does not exist. The real location is eslint-rules/ at the repo root -- what eslint.config.mjs actually imports -- and CONTEXT.md's own later entry already says so explicitly, so the file disagreed with itself. Found by a line-by-line audit of all 1036 lines against the live graph; this was the only confirmed inaccuracy. Closes #2890 --------- Co-authored-by: Test <test@example.com> |