enhance(#3118): close the dead injectables and the shell projection follow-on — Wave 4 (#3124)

* test(#3118): failing-first coverage for the dead injectables and the shell projection

Adds the counter-tests Wave 4 closes against, before any fix:

- antigravityWatermark had zero test references. The four existing tests
  that look like watermark coverage hand the fallback a literal mark and
  never call the producer, so nothing pinned whether a real run's mark is
  correct. Covers all six branches plus the non-object cache classes.
- Pins the fail-open: a transcript read that throws reports lines:0,
  indistinguishable from a genuinely empty transcript, and the consumer
  then replays a previous run's review as this run's.
- Pins the export-line escaping across the repair, persist and win32 bash
  lanes, including the parity assertion that they must not diverge.
- sliceCurrentPositionSection: empty-vs-absent, fenced heading, second
  occurrence, H3, CRLF.
- Proves deps.progressProvider is inert by supplying a throwing stub to
  all ten transition intents.

Verification through the remote runner only.

Refs #3118

* fix(#3118): distinguish an unreadable transcript from an empty one

antigravityWatermark's final read can throw on a transcript that
indisputably exists. It returned lines:0, which is the same value a
genuinely empty transcript produces, so the caller could not tell the
two apart.

antigravityTranscriptFallback derives its skip from that count. A mark
of {convId:'c1', lines:0} for a conversation that pre-dates the run
makes it skip nothing and return the last PLANNER_RESPONSE in a
transcript written before this run started — a previous review
presented as this one's, which is exactly what the function's own
'never stale' docstring promises cannot happen.

The unreadable case now sets unreadable:true and the fallback declines
for a same-conv-id unreadable mark. An absent or empty transcript is
untouched: those genuinely have zero prior lines.

* fix(#3118): escape the export line for the file it lands in, not the echo

Three lanes emit export PATH="<dir>:$PATH". repair escaped it with
escapePosixDoubleQuoted; persist and the win32 Git Bash lane escaped it
with escapeSingleQuotedShellLiteral instead.

The single-quoting is correct for the echo, so nothing runs when the
user pastes the command. But the bytes appended to ~/.bashrc are the
export line itself, and inside double quotes in an rc file a $(...) or
a backtick in the directory name is command substitution that runs on
every new shell. Those characters are legal in a path on both POSIX and
Windows, so the path was reachable.

projectPathExportLine is now the single source of that line and escapes
for its final rc-file context; each lane still applies its own transport
escaping on top. fish keeps the single-quote escaper — its value really
does stay single-quoted.

The cmd.exe lane interpolated into a cmd double-quoted string with no
cmd-level escaping, so a quote closed the region and &cmd& ran. A quote
is reserved on Windows and cannot appear in a real path, so there is no
correct command to suggest: the win32 lanes now fail closed for one.

Metacharacter-free paths render byte-identically on every lane.

* fix(#3118): drop a stray carriage return and a deps field nobody reads

locateCurrentPosition subtracted a fixed one byte to exclude the newline
before the next heading, which assumes LF. On a CRLF document the slice
kept an unpaired trailing carriage return. It now walks back over the
newline and over a preceding carriage return if there is one.

StateTransitionDeps also required a progressProvider that 33 sites
supplied and no site ever called. A required field nothing reads widens
the module's interface without changing its implementation, which is the
shape epic #3051 cites as its reason for refusing blanket injection.
Removed along with the ProgressRecord alias that existed only as its
return type; state-document.cts's unrelated interface of the same name
is untouched.

* fix(#3118): stop an empty span duplicating bytes, and name the empty results

Three findings from the isolated review pass.

locateCurrentPosition could return end < start when the section was
empty and the next heading followed with no blank line between. Every
mutator splices with slice(0,start) + body + slice(end), so an inverted
span duplicated the region between them — a blank line silently
inserted into STATE.md on every transition, two bytes on CRLF. The span
is now clamped, and an empty section is a zero-length span, which is
what it always meant.

The win32 fail-closed path left the installer printing 'Add it with one
of:' with nothing under it. An empty shellActions folded two different
facts together, so projectPathActionProjection now carries a frozen
PATH_ACTION_REASON and the installer branches on it. Two empty results
with different causes staying distinguishable is the subject of the
epic this belongs to.

fish_add_path parses a leading dash as an option, so a directory named
-v printed 'No paths to add' instead of being added. Verified against
fish 4.8.1: the end-of-options separator fixes it.

Replaces the console-prose test the second fix first arrived with — a
regex over captured stdout is what CONTRIBUTING prohibits, and the
typed reason is the surface it asks for instead.

* fix(#3118): escape TOML control characters, and stop a test name overstating

Five findings from the two review axes.

escapeTomlDoubleQuotedString escaped only backslash and quote. TOML
basic strings also require U+0000-U+0008, U+000A-U+001F and U+007F to be
escaped, so a value carrying a raw newline or NUL wrote a config.toml no
parser accepts — rejecting the whole file, not just that value. Four of
its call sites write real config. Tab stays raw; the grammar exempts it.

The byte-identity test claimed every lane was unchanged for an ordinary
path, which is false: fish now takes the end-of-options separator on
every path, not only hostile ones. Renamed, and the one intended delta
now has its own named test instead of hiding inside a claim that read
as broader than it was.

Also: exact-equality assertions in place of substring checks that could
pass on a subtly wrong escape, newline and null-byte cases for all five
quoting primitives, and a temp dir registered with t.after so it is
removed when an assertion fails.

* docs(#3118): add the changeset fragments

* fix(#3118): degrade instead of throwing on a null conversation cache

A cache file whose whole content is the literal null — what a truncated
or zeroed write leaves behind — made both antigravityWatermark and
antigravityTranscriptFallback throw. JSON.parse('null') succeeds, so the
try/catch wrapping the parse never fired, and resolveConvId then called
hasOwnProperty on null.

Both functions advertise the opposite; the existing test next to them is
named 'a missing cache or transcript degrades to empty, never throws'.
Parsing successfully is not the same fact as the payload being usable,
and a guard that only wraps the parse cannot tell them apart.

resolveConvId is now total for any non-object input, so one guard covers
both callers. Caught by the null case in this wave's own cache matrix.

* test(#3118): correct a stale fish expectation and a parity comparison

The pre-existing 'POSIX persist mode escapes single quotes' test pinned
fish_add_path without the end-of-options separator this wave adds, so it
asserted behavior that is no longer correct. A repo-wide scan found one
such hardcoded expectation; every other site derives its expectation
from the projection.

The new parity test compared the token from a POSIX path against the
win32 lane, which posix-normalizes its input first — two different
inputs, so the tokens differed for a reason that had nothing to do with
the parity it claims to check. It now derives the win32 expectation from
the same input the lane receives.

* docs(#3118): reword a comment the injection scanner reads as an instruction

The scanner pattern act\s+as\s+(?:a|an|the)\s+ carries no word
boundary, so 'the same fact as the payload' matched on the tail of
'fact'. Reworded per the documented remedy for this collision.

The missing boundary is a scanner defect rather than a prose problem —
any contributor writing 'fact as the' trips it — but the pattern is
gate plumbing, which the sibling epic owns, so it is surfaced rather
than changed here.

* chore(#3118): backfill changeset pr number to 3124

* chore(#3118): backfill changeset pr number to 3124

* fix(#2784): make the negation scan single-pass and index it correctly

Three defects in the negation suppression added by #3127, all in one
block, none of which had a test.

The pair scan was verbs.some(nouns.some(...)) with a slice and a split
per pair, so it grew cubically with clause length: 1.1ms before that PR
and 8462ms after, on 800 verb+noun pairs in one clause. api-coverage's
property test generates documents large enough to reach the runner's
600s file cap, which is why it hangs as 'fail 0, cancelled 1' rather
than failing an assertion. Every (verb, noun) window is a subset of the
single widest one, so one scan of that window answers the same question
in a linear pass. Verified equivalent against the old predicate over
20,000 generated clauses.

Both checks also subtracted clause.start from offsets that collectTerm-
Matches already returns clause-local. The first clause on a line has
start 0 so it worked there and nowhere else: later clauses went
negative, and slice reads a negative index from the end, so suppression
silently examined unrelated text.

The comment claimed 'without any API integration' was suppressed. It is
not — the qualifier sits outside the two-word lookback and the noun
precedes the verb. Widening the window would trade a false positive
that costs one declaration line for a false negative that slips a real
integration past a blocking gate, so the behavior stands and the
comment now says so. Pinned by a test.

The qualifier sets were also rebuilt for every line of every document.
This commit is contained in:
Tom Boucher
2026-08-06 23:57:05 -04:00
committed by GitHub
parent 0ccc18dd3c
commit 0e6fa2e2cf
20 changed files with 1296 additions and 86 deletions

View File

@@ -257,6 +257,31 @@ const SURFACE_DESCRIPTOR_WORDS = new Set([
* is deliberately absent (an external API IS external). */
const INTERNAL_DESCRIPTORS = new Set(['internal', 'in-house', 'local', 'first-party', 'private']);
/** #2784: negation suppression. A clause that pairs an integration verb with an
* API noun but the verb itself is directly negated (e.g. "does not integrate",
* "integrates no external API") is suppressed. Two windows are checked: a
* negation qualifier within 2 words directly before the verb, or "no"/"zero"/
* "none" between the verb and a following noun.
* KNOWN LIMIT (deliberate, not a bug to fix later): a negation further than 2
* words before the verb, or a clause where the noun precedes the verb, is NOT
* suppressed — e.g. "Ships without any API integration." still reports
* detected:true, because "without" sits outside the verb's 2-word lookback
* and the noun precedes the verb. This is intentional: detectApiIntegration
* is fail-closed — an unsuppressed false positive costs a one-line
* COVERAGE.md declaration, while widening the suppression window risks a
* false negative that silently lets a real external-API phase past a
* blocking gate.
* Hoisted to module scope (#3127 regression fix): this was previously
* allocated fresh on every source line, which is wasted work on documents
* with many lines. */
const NEGATION_QUALIFIERS = new Set([
'no', 'not', 'without', 'zero', 'neither', 'nor', 'none', "don't", "doesn't", "didn't", "won't", "can't", "cannot",
]);
/** Negation tokens checked BETWEEN a verb and a later noun (narrower than
* NEGATION_QUALIFIERS — "not" and "without" are checked only immediately
* before the verb, via NEGATION_QUALIFIERS above). */
const NEGATION_NOUN_TOKENS = new Set(['no', 'zero', 'none']);
/** A capitalized compound modifier ("Resolver-only", "Read-only", "E-commerce"
* — lowercase letter right after the hyphen) is an adjective phrase, not a
* service name. Real hyphenated services capitalize the second segment
@@ -484,23 +509,37 @@ export function detectApiIntegration(
//
// #2784: negation suppression. A clause that pairs an integration verb with
// an API noun but the verb itself is directly negated (e.g. "does not
// integrate", "integrates no external API", "without any API integration")
// is suppressed. The check is scoped to the verb's immediate context (the
// word directly before the verb, or the word directly between verb and noun)
// — NOT a blanket clause-wide scan, because "without changing runtime
// dependencies" in a long clause does NOT negate the integration.
const NEGATION_QUALIFIERS = new Set([
'no', 'not', 'without', 'zero', 'neither', 'nor', 'none', "don't", "doesn't", "didn't", "won't", "can't", "cannot",
]);
// integrate", "integrates no external API") is suppressed. The check is
// scoped to the verb's immediate context (the word directly before the
// verb, or the word directly between verb and noun) — NOT a blanket
// clause-wide scan, because "without changing runtime dependencies" in a
// long clause does NOT negate the integration.
// KNOWN LIMIT (deliberate, not a bug to fix later): a negation further than
// 2 words before the verb, or a clause where the noun precedes the verb, is
// NOT suppressed — e.g. "Ships without any API integration." is NOT
// suppressed today (pinned by a test in tests/api-coverage.test.cjs).
// detectApiIntegration is fail-closed by design: an unsuppressed false
// positive costs a one-line COVERAGE.md declaration, while widening the
// window trades that for a silent false negative on a blocking gate.
if (verbRe && nounRe) {
for (const clause of clauses) {
const verbs = collectTermMatches(verbRe, clause.text);
if (verbs.length === 0) continue;
// #2784: check if any verb is immediately preceded by a negation
// qualifier (within 2 words before the verb match).
//
// OFFSET NOTE: `v.start`/`n.start` (from collectTermMatches below and
// above) are already CLAUSE-LOCAL — collectTermMatches was called with
// `clause.text`, not the full line — and so is `clauseText`
// (`clause.text.toLowerCase()`). They must be used AS-IS to index into
// `clauseText`; do not re-base them against `clause.start` (that field
// is the clause's offset within the LINE, a different coordinate space,
// used only to map line-level spans like `extraNouns`/`masked` into a
// clause). Subtracting `clause.start` here double-offsets the slice
// bounds for every clause after the first on a line (#3127 follow-up).
const clauseText = clause.text.toLowerCase();
const hasNegatedVerb = verbs.some((v) => {
const before = clauseText.slice(Math.max(0, v.start - clause.start - 20), v.start - clause.start);
const before = clauseText.slice(Math.max(0, v.start - 20), v.start);
const beforeWords = before.split(/\s+/).filter(Boolean).slice(-2);
return beforeWords.some((w: string) => NEGATION_QUALIFIERS.has(w.replace(/[^a-z']/g, '')));
});
@@ -513,15 +552,63 @@ export function detectApiIntegration(
}
}
if (nounTerms.size === 0) continue;
// Check for negation between verb and noun
const hasNegatedNoun = verbs.some((v) => {
return nouns.some((n) => {
if (n.start <= v.start) return false;
const between = clauseText.slice(v.start - clause.start + v.term.length, n.start - clause.start);
// Check for negation between verb and noun.
//
// #3127 regression: the original form of this check was
// O(verbs × nouns), re-slicing and re-splitting the clause text for
// every (verb, noun) pair — effectively cubic in clause length (a
// clause of N repeated "integrate api" pairs did O(N^2) pair checks,
// each doing an O(N) slice/split). On a clause with 800 repeated
// pairs this took ~8.5s; fast-check's property test then generated
// documents large enough to hang the whole test file past node:test's
// 600s timeout. It ALSO subtracted `clause.start` from `v.start`/
// `n.start` before slicing `clauseText` — but `v.start`/`n.start` are
// already local to `clause.text` (collectTermMatches was called with
// clause.text, not the full line), and `clauseText` is exactly
// `clause.text.toLowerCase()`. So that subtraction double-offset the
// slice bounds for every clause after the first on a line, sliding
// (and for negative results, JS's negative-index slice() wraparound
// non-monotonically re-mapping) the window to characters unrelated to
// the verb/noun pair — an independent latent bug, fixed here as part
// of establishing a well-defined O(1) predicate (a piecewise/clamped
// window has no single "widest span" to reason about at all).
//
// EXACT-EQUIVALENCE, single pass: the predicate is "does any pair
// (v, n) with n.start > v.start have a negation token in the span
// (v.end, n.start)". Every such span is a SUBSET of the widest
// possible span for a given noun: [min(v.end) over verbs valid for
// that noun, n.start). And since that window only widens as a
// noun's start increases (more verbs become valid, and the noun
// bound itself grows), the single widest span across the WHOLE
// clause is anchored at the noun with the maximum start, using the
// minimum verb-end among verbs valid for THAT noun (not the global
// minimum verb-end, which could belong to a verb that starts after
// this noun and so is never a valid pairing with it — a mismatch
// that would either miss or falsely include a negation). If that one
// substring contains no negation token, no narrower pair-specific
// substring can either; if it does, the (minVerb, maxNoun) pair
// itself contains it. This drops the check to O(verbs + nouns).
let hasNegatedNoun = false;
if (nouns.length > 0) {
let minVerbStart = Infinity;
for (const v of verbs) if (v.start < minVerbStart) minVerbStart = v.start;
let maxNounStart = -Infinity;
for (const n of nouns) if (n.start > maxNounStart) maxNounStart = n.start;
if (maxNounStart > minVerbStart) {
let minQualifyingVerbEnd = Infinity;
for (const v of verbs) {
if (v.start < maxNounStart) {
const vEnd = v.start + v.term.length;
if (vEnd < minQualifyingVerbEnd) minQualifyingVerbEnd = vEnd;
}
}
const between = clauseText.slice(minQualifyingVerbEnd, maxNounStart);
const betweenWords = between.split(/\s+/).filter(Boolean);
return betweenWords.some((w: string) => ['no', 'zero', 'none'].includes(w.replace(/[^a-z']/g, '')));
});
});
hasNegatedNoun = betweenWords.some((w: string) =>
NEGATION_NOUN_TOKENS.has(w.replace(/[^a-z']/g, '')),
);
}
}
if (hasNegatedVerb || hasNegatedNoun) continue;
for (const vTerm of new Set(verbs.map((t) => t.term))) {
for (const nTerm of nounTerms) emitPair(vTerm, nTerm, rawLine);