* test(#3118): failing-first coverage for the dead injectables and the shell projection Adds the counter-tests Wave 4 closes against, before any fix: - antigravityWatermark had zero test references. The four existing tests that look like watermark coverage hand the fallback a literal mark and never call the producer, so nothing pinned whether a real run's mark is correct. Covers all six branches plus the non-object cache classes. - Pins the fail-open: a transcript read that throws reports lines:0, indistinguishable from a genuinely empty transcript, and the consumer then replays a previous run's review as this run's. - Pins the export-line escaping across the repair, persist and win32 bash lanes, including the parity assertion that they must not diverge. - sliceCurrentPositionSection: empty-vs-absent, fenced heading, second occurrence, H3, CRLF. - Proves deps.progressProvider is inert by supplying a throwing stub to all ten transition intents. Verification through the remote runner only. Refs #3118 * fix(#3118): distinguish an unreadable transcript from an empty one antigravityWatermark's final read can throw on a transcript that indisputably exists. It returned lines:0, which is the same value a genuinely empty transcript produces, so the caller could not tell the two apart. antigravityTranscriptFallback derives its skip from that count. A mark of {convId:'c1', lines:0} for a conversation that pre-dates the run makes it skip nothing and return the last PLANNER_RESPONSE in a transcript written before this run started — a previous review presented as this one's, which is exactly what the function's own 'never stale' docstring promises cannot happen. The unreadable case now sets unreadable:true and the fallback declines for a same-conv-id unreadable mark. An absent or empty transcript is untouched: those genuinely have zero prior lines. * fix(#3118): escape the export line for the file it lands in, not the echo Three lanes emit export PATH="<dir>:$PATH". repair escaped it with escapePosixDoubleQuoted; persist and the win32 Git Bash lane escaped it with escapeSingleQuotedShellLiteral instead. The single-quoting is correct for the echo, so nothing runs when the user pastes the command. But the bytes appended to ~/.bashrc are the export line itself, and inside double quotes in an rc file a $(...) or a backtick in the directory name is command substitution that runs on every new shell. Those characters are legal in a path on both POSIX and Windows, so the path was reachable. projectPathExportLine is now the single source of that line and escapes for its final rc-file context; each lane still applies its own transport escaping on top. fish keeps the single-quote escaper — its value really does stay single-quoted. The cmd.exe lane interpolated into a cmd double-quoted string with no cmd-level escaping, so a quote closed the region and &cmd& ran. A quote is reserved on Windows and cannot appear in a real path, so there is no correct command to suggest: the win32 lanes now fail closed for one. Metacharacter-free paths render byte-identically on every lane. * fix(#3118): drop a stray carriage return and a deps field nobody reads locateCurrentPosition subtracted a fixed one byte to exclude the newline before the next heading, which assumes LF. On a CRLF document the slice kept an unpaired trailing carriage return. It now walks back over the newline and over a preceding carriage return if there is one. StateTransitionDeps also required a progressProvider that 33 sites supplied and no site ever called. A required field nothing reads widens the module's interface without changing its implementation, which is the shape epic #3051 cites as its reason for refusing blanket injection. Removed along with the ProgressRecord alias that existed only as its return type; state-document.cts's unrelated interface of the same name is untouched. * fix(#3118): stop an empty span duplicating bytes, and name the empty results Three findings from the isolated review pass. locateCurrentPosition could return end < start when the section was empty and the next heading followed with no blank line between. Every mutator splices with slice(0,start) + body + slice(end), so an inverted span duplicated the region between them — a blank line silently inserted into STATE.md on every transition, two bytes on CRLF. The span is now clamped, and an empty section is a zero-length span, which is what it always meant. The win32 fail-closed path left the installer printing 'Add it with one of:' with nothing under it. An empty shellActions folded two different facts together, so projectPathActionProjection now carries a frozen PATH_ACTION_REASON and the installer branches on it. Two empty results with different causes staying distinguishable is the subject of the epic this belongs to. fish_add_path parses a leading dash as an option, so a directory named -v printed 'No paths to add' instead of being added. Verified against fish 4.8.1: the end-of-options separator fixes it. Replaces the console-prose test the second fix first arrived with — a regex over captured stdout is what CONTRIBUTING prohibits, and the typed reason is the surface it asks for instead. * fix(#3118): escape TOML control characters, and stop a test name overstating Five findings from the two review axes. escapeTomlDoubleQuotedString escaped only backslash and quote. TOML basic strings also require U+0000-U+0008, U+000A-U+001F and U+007F to be escaped, so a value carrying a raw newline or NUL wrote a config.toml no parser accepts — rejecting the whole file, not just that value. Four of its call sites write real config. Tab stays raw; the grammar exempts it. The byte-identity test claimed every lane was unchanged for an ordinary path, which is false: fish now takes the end-of-options separator on every path, not only hostile ones. Renamed, and the one intended delta now has its own named test instead of hiding inside a claim that read as broader than it was. Also: exact-equality assertions in place of substring checks that could pass on a subtly wrong escape, newline and null-byte cases for all five quoting primitives, and a temp dir registered with t.after so it is removed when an assertion fails. * docs(#3118): add the changeset fragments * fix(#3118): degrade instead of throwing on a null conversation cache A cache file whose whole content is the literal null — what a truncated or zeroed write leaves behind — made both antigravityWatermark and antigravityTranscriptFallback throw. JSON.parse('null') succeeds, so the try/catch wrapping the parse never fired, and resolveConvId then called hasOwnProperty on null. Both functions advertise the opposite; the existing test next to them is named 'a missing cache or transcript degrades to empty, never throws'. Parsing successfully is not the same fact as the payload being usable, and a guard that only wraps the parse cannot tell them apart. resolveConvId is now total for any non-object input, so one guard covers both callers. Caught by the null case in this wave's own cache matrix. * test(#3118): correct a stale fish expectation and a parity comparison The pre-existing 'POSIX persist mode escapes single quotes' test pinned fish_add_path without the end-of-options separator this wave adds, so it asserted behavior that is no longer correct. A repo-wide scan found one such hardcoded expectation; every other site derives its expectation from the projection. The new parity test compared the token from a POSIX path against the win32 lane, which posix-normalizes its input first — two different inputs, so the tokens differed for a reason that had nothing to do with the parity it claims to check. It now derives the win32 expectation from the same input the lane receives. * docs(#3118): reword a comment the injection scanner reads as an instruction The scanner pattern act\s+as\s+(?:a|an|the)\s+ carries no word boundary, so 'the same fact as the payload' matched on the tail of 'fact'. Reworded per the documented remedy for this collision. The missing boundary is a scanner defect rather than a prose problem — any contributor writing 'fact as the' trips it — but the pattern is gate plumbing, which the sibling epic owns, so it is surfaced rather than changed here. * chore(#3118): backfill changeset pr number to 3124 * chore(#3118): backfill changeset pr number to 3124 * fix(#2784): make the negation scan single-pass and index it correctly Three defects in the negation suppression added by #3127, all in one block, none of which had a test. The pair scan was verbs.some(nouns.some(...)) with a slice and a split per pair, so it grew cubically with clause length: 1.1ms before that PR and 8462ms after, on 800 verb+noun pairs in one clause. api-coverage's property test generates documents large enough to reach the runner's 600s file cap, which is why it hangs as 'fail 0, cancelled 1' rather than failing an assertion. Every (verb, noun) window is a subset of the single widest one, so one scan of that window answers the same question in a linear pass. Verified equivalent against the old predicate over 20,000 generated clauses. Both checks also subtracted clause.start from offsets that collectTerm- Matches already returns clause-local. The first clause on a line has start 0 so it worked there and nowhere else: later clauses went negative, and slice reads a negative index from the end, so suppression silently examined unrelated text. The comment claimed 'without any API integration' was suppressed. It is not — the qualifier sits outside the two-word lookback and the noun precedes the verb. Widening the window would trade a false positive that costs one declaration line for a false negative that slips a real integration past a blocking gate, so the behavior stands and the comment now says so. Pinned by a test. The qualifier sets were also rebuilt for every line of every document.
This commit is contained in:
@@ -131,6 +131,31 @@ describe('detectApiIntegration — pure detector (#1562)', () => {
|
||||
});
|
||||
}
|
||||
|
||||
// ── #2784/#3127: negation suppression + its #3127 perf regression ────────
|
||||
// PR #3127 added clause-local negation suppression (hasNegatedVerb /
|
||||
// hasNegatedNoun) but the noun-side check was O(verbs × nouns) with a
|
||||
// slice+split per pair — effectively cubic in clause length. A clause with
|
||||
// a few hundred repeated "integrate api" pairs took seconds; fast-check's
|
||||
// property test then generated documents large enough to hang the whole
|
||||
// file past node:test's 600s timeout (fail 0 / cancelled 1). These two
|
||||
// tests pin the O(verbs + nouns) fix at a size (~500 pairs) chosen so that
|
||||
// a re-regression to the old quadratic-per-pair behavior makes this file
|
||||
// exceed the runner's own timeout — that hang, not a wall-clock assertion
|
||||
// (CLAUDE.md forbids those), is the failure signal for a re-regression.
|
||||
test('suppresses a negated pair regardless of how many terms precede it', () => {
|
||||
const pairs = Array(500).fill('integrate api').join(' ');
|
||||
const scope = pairs + ' but we integrate no api here';
|
||||
const r = detectApiIntegration(scope);
|
||||
assert.strictEqual(r.detected, false);
|
||||
});
|
||||
|
||||
test('still detects a genuine pair in a very long clause', () => {
|
||||
const pairs = Array(500).fill('integrate api').join(' ');
|
||||
const r = detectApiIntegration(pairs);
|
||||
assert.strictEqual(r.detected, true);
|
||||
assert.ok(r.signals.length > 0);
|
||||
});
|
||||
|
||||
test('fenced code blocks are stripped — trigger inside a code fence does not fire', () => {
|
||||
const scope = [
|
||||
'Refactor the helpers.',
|
||||
@@ -165,6 +190,74 @@ describe('detectApiIntegration — pure detector (#1562)', () => {
|
||||
assert.strictEqual(detectApiIntegration(splitLine).detected, false);
|
||||
assert.strictEqual(detectApiIntegration('integrates the api').detected, true);
|
||||
});
|
||||
|
||||
// ── #3127 follow-up: hasNegatedVerb offset regression coverage ───────────
|
||||
// The `hasNegatedVerb` window previously subtracted `clause.start` a SECOND
|
||||
// time from an already clause-local `v.start`, which only accidentally
|
||||
// worked for the first clause on a line (clause.start === 0) and silently
|
||||
// mis-windowed (via negative-index slice wraparound) every later clause.
|
||||
test('suppresses a pair when the verb is directly negated', () => {
|
||||
const r = detectApiIntegration('This phase does not integrate any external API.');
|
||||
assert.strictEqual(r.detected, false);
|
||||
});
|
||||
|
||||
test('suppresses a pair when a negation sits between verb and noun', () => {
|
||||
const r = detectApiIntegration('This phase integrates no external API.');
|
||||
assert.strictEqual(r.detected, false);
|
||||
});
|
||||
|
||||
test('still detects a genuine integration', () => {
|
||||
const r = detectApiIntegration('We integrate the Stripe API.');
|
||||
assert.strictEqual(r.detected, true);
|
||||
const compound = r.signals.find((s) => s.verb !== '(surface)');
|
||||
assert.ok(compound, 'expected a compound verb+noun signal');
|
||||
assert.strictEqual(compound.verb, 'integrate');
|
||||
assert.strictEqual(compound.noun, 'api');
|
||||
});
|
||||
|
||||
test('a distant negation does not suppress a genuine pair', () => {
|
||||
// #2784's own documented example: "without changing runtime dependencies"
|
||||
// inside a long clause must NOT negate an unrelated genuine pairing later
|
||||
// in the SAME clause — the negation-window checks are scoped to the
|
||||
// verb's immediate context, not a blanket clause-wide scan.
|
||||
const r = detectApiIntegration(
|
||||
'The migration proceeds without changing runtime dependencies while we integrate the Stripe API for payment processing across the whole checkout flow.'
|
||||
);
|
||||
assert.strictEqual(r.detected, true);
|
||||
});
|
||||
|
||||
test('suppression works for a clause that is not the first on the line', () => {
|
||||
// THE REGRESSION TEST for the double-offset bug: before the fix, this
|
||||
// failed (wrongly detected === true) because the negation window for the
|
||||
// second clause was computed against the wrong index base (v.start was
|
||||
// re-based against clause.start even though it was already clause-local),
|
||||
// producing an empty/garbage "before" window via negative-index slice()
|
||||
// wraparound so the verb-adjacent negation was never seen.
|
||||
const r = detectApiIntegration('We ship a thing; this phase does not integrate an external API.');
|
||||
assert.strictEqual(r.detected, false);
|
||||
});
|
||||
|
||||
test('suppression in an early clause does not mask a genuine pair in a later clause', () => {
|
||||
// Intended semantics: negation suppression is scoped PER CLAUSE (matching
|
||||
// the "no cross-clause binding" design elsewhere in this module — see the
|
||||
// CLAUSE_BOUNDARY_CHARS note). A negated pair in clause 1 must not swallow
|
||||
// an independent, unnegated, genuine pair in clause 2: each clause is
|
||||
// evaluated on its own, and `detected` is true if ANY clause fires.
|
||||
const r = detectApiIntegration('This phase integrates no external API; we also integrate the Stripe API.');
|
||||
assert.strictEqual(r.detected, true);
|
||||
});
|
||||
|
||||
test('a negation outside the lookback window does not suppress — fail-closed by design', () => {
|
||||
// This pins DELIBERATE fail-closed behavior, not an aspiration: "without"
|
||||
// sits outside hasNegatedVerb's 2-word lookback before the verb, and the
|
||||
// noun ("integration") precedes the verb ("Ships") so hasNegatedNoun's
|
||||
// verb-to-noun window never applies either. Widening the lookback to catch
|
||||
// this phrase would trade a cheap false positive (a one-line COVERAGE.md
|
||||
// declaration) for a silent false negative — a real external-API phase
|
||||
// slipping past a blocking gate — which is the dangerous direction.
|
||||
const r = detectApiIntegration('Ships without any API integration.');
|
||||
assert.strictEqual(r.detected, true);
|
||||
});
|
||||
});
|
||||
|
||||
// ──────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
Reference in New Issue
Block a user