* feat(#3414): promote git-cmd.js token-walk into a shared scanner, fix #3169 Phase 3 of epic #3212 (ADR-3212 §4). New src/token-scanner.cts generalizes hooks/lib/git-cmd.js's proven token-walk (#3129 — "has not re-opened"): tokenizeShellLike (quote-aware shell tokenizer, byte-identical port) and indentWidth (bullet-nesting depth). git-cmd.js migrates onto tokenizeShellLike with zero behavior change (parity-asserted against every existing #3129 fixture in tests/worktree-safety.test.cjs's folded block); isGitSubcommand's phases 1-3 (env-prefix skip, executable check, global-option consume) extracted into skipToSubcommand, shared with the new extractBranchArgument (git checkout -b / git branch <name>) — a new capability exercising the seam on the domain the ADR names, not a migration of existing duplicated logic (none existed). Fixes #3169: src/decisions.cts's parseDecisionLines couldn't distinguish a cross-reference bullet nested under an open decision from a fresh malformed declaration attempt. An earlier bold-run-content-classification design was tried and disproven against the repo's own existing FIX-B fixtures (D-02, "no colon no dash") before being adopted — both have identical shape under any content-only rule. Nesting depth (via indentWidth) is the actual distinguishing signal: a bullet indented deeper than the currently-open decision's own bullet is elaboration, folded into its text like a continuation line, never tested against the parse-miss guard. A bullet at the same-or-shallower indent is unchanged. Scope-narrowing disclosed, not silent: of the ADR's four named bugs (#3197, #3169, #2570, #2528), three no longer need this phase's work. were independently fixed and closed since the ADR was authored — #2570's fix is already a correctly-bounded regex per the ADR's own decidability test (no scanner needed); #2528's fix is a deliberate, twice-reviewed non-scanner design (its own code comment records a scanner-based attempt that regressed a symmetric case and was reverted) that this phase does not disturb. Only #3169 required new work. get_impact: isGitSubcommand CRITICAL/196 affected symbols, parseDecisionLines CRITICAL/164 affected symbols (ADR §6 due diligence). Six-gate ripple: .gitignore, eslint.config.mjs, docs/INVENTORY.md, docs/INVENTORY-MANIFEST.json (regenerated), CONTEXT.md glossary. Design: .gsd/phase/chore-3414-tokenizer-first-seam/40-design.md Test matrix: .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3414): add required fast-check property tests per code review TESTING-STANDARDS.md:169 requires at least one fast-check property test for any module that implements parsing — src/token-scanner.cts had none, an orthogonal Standards-axis review finding. Adds two seeded property tests (mirroring Phase 1/2's fast-check-setup.cjs convention): indentWidth counts exactly a generated leading-space run; tokenizeShellLike round-trips a generated array of whitespace/quote-free words joined with single spaces. The design doc's own "no property test needed" rationale was wrong — it argued no algebraic law applied, but the standard is unconditional for parsing modules regardless of whether one "feels" applicable. Corrected in .gsd/phase/chore-3414-tokenizer-first-seam/50-test-matrix.md. Also fixes two Spec-axis wording drifts the same review found between the design doc and the shipped code (doc-only, no behavior change): extractBranchArgument's documented signature dropped an unused subVariants parameter that was never implemented, and the #3169 fail-first fixture description corrected from "15-decision plan via cmdDecisionCoverageVerify" to the actual compact 3-decision analog via the real blocking gate, check.decision-coverage-plan. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3414): add changeset for #3169 fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3414): backfill changeset pr number to 3424 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
69 lines
2.5 KiB
TypeScript
69 lines
2.5 KiB
TypeScript
/**
|
|
* token-scanner.cts — the tokenizer-first seam for stateful grammars (ADR-3212
|
|
* §4, epic #3212 Phase 3, #3414).
|
|
*
|
|
* Source in src/token-scanner.cts, compiled to gsd-core/bin/lib/token-scanner.cjs
|
|
* (gitignored), per the repo's ADR-457 build-at-publish convention.
|
|
*
|
|
* Generalizes the proven `hooks/lib/git-cmd.js` token-walk (#3129 — "has not
|
|
* re-opened") into a primitive other stateful-grammar consumers can share.
|
|
* `git-cmd.js` migrates onto `tokenizeShellLike` with zero behavior change
|
|
* (ADR §6); its own env-prefix/flag/subcommand walk logic stays put — only
|
|
* the character-level tokenizer moves here.
|
|
*
|
|
* `indentWidth` is the primitive `src/decisions.cts`'s #3169 fix needs: a
|
|
* decision-bullet list can NEST (a bullet elaborating on an already-open
|
|
* decision, indented deeper than it), and a per-line regex cannot see that
|
|
* nesting — only tracking indentation relative to the currently-open
|
|
* decision can (ADR §4 criterion 1). See
|
|
* .gsd/phase/chore-3414-tokenizer-first-seam/40-design.md §1.3 for the
|
|
* disproven bold-run-content-classification alternative and why nesting
|
|
* depth, not bullet content, is the actual distinguishing signal.
|
|
*/
|
|
|
|
/**
|
|
* Tokenize a shell-like command string: whitespace-split, with single- and
|
|
* double-quoted spans taken as one token each. No escape handling, no
|
|
* brace/variable expansion — matches `hooks/lib/git-cmd.js`'s original
|
|
* `tokenize()` exactly (parity-asserted in tests/token-scanner.test.cjs).
|
|
*/
|
|
export function tokenizeShellLike(cmd: string): string[] {
|
|
const tokens: string[] = [];
|
|
let i = 0;
|
|
const len = cmd.length;
|
|
|
|
while (i < len) {
|
|
while (i < len && /\s/.test(cmd[i])) i++;
|
|
if (i >= len) break;
|
|
|
|
let token = '';
|
|
while (i < len && !/\s/.test(cmd[i])) {
|
|
if (cmd[i] === "'") {
|
|
i++;
|
|
while (i < len && cmd[i] !== "'") token += cmd[i++];
|
|
if (i < len) i++;
|
|
} else if (cmd[i] === '"') {
|
|
i++;
|
|
while (i < len && cmd[i] !== '"') token += cmd[i++];
|
|
if (i < len) i++;
|
|
} else {
|
|
token += cmd[i++];
|
|
}
|
|
}
|
|
if (token) tokens.push(token);
|
|
}
|
|
|
|
return tokens;
|
|
}
|
|
|
|
/**
|
|
* Count a line's leading whitespace width. Tabs count as one column each
|
|
* (not expanded) — matches `src/decisions.cts`'s existing continuation-line
|
|
* check (`/^[ \t]/`), which never expanded tabs either; this function does
|
|
* not introduce a new tab-width policy.
|
|
*/
|
|
export function indentWidth(line: string): number {
|
|
const match = line.match(/^[ \t]*/);
|
|
return match ? match[0].length : 0;
|
|
}
|