* enhance(#4155): invalidate verification results when covered inputs change readVerificationStatus() now recomputes a deterministic sha256 fingerprint over a VERIFICATION.md's declared covered_files (phase PLAN/SUMMARY, requirements, implementation files in the verified change set) and returns stale on any mismatch, fail-closed when a covered file is missing, unreadable, or escapes the project root. Legacy reports with no fingerprint metadata keep the prior SUMMARY-mtime staleness check unchanged. The verifier computes covered_digest via the new verification.fingerprint CLI command rather than by hand, since a digest is deterministic math, not an LLM-estimated value. * chore(#4155): backfill fork PR number in changeset * fix(#4155): trim gsd-verifier.md fingerprint instructions to fit LARGE tier byte cap * fix(#4155): address CodeRabbit findings on fingerprint fail-closed behavior Partial fingerprint metadata (one of covered_files/covered_digest present, the other missing or malformed) now fails closed to stale instead of silently downgrading to the legacy mtime-only check. computeCoveredDigest also canonicalizes with realpathSync before re-confining, so an in-root symlink whose target escapes the project root can no longer produce a matching digest. gsd-verifier.md restores the completeness requirement and checklist item trimmed by the earlier size-budget fix, within the LARGE tier byte cap. * chore(#4155): acknowledge gsd-verifier.md growth for the #4155 fingerprint instructions Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the covered-input fingerprint instructions and frontmatter fields the #4155 verification staleness mechanism requires; trimmed to stay within the LARGE tier byte cap * fix(#4155): address gemini adversarial review findings computeCoveredDigest now threads the caller-supplied opts.fs seam through its confinement and read paths instead of always using raw node:fs — a caller like planning-inspect.cts's containmentEnforcingVerificationFs (GAP 2, #2790 follow-up) was silently bypassed for covered-input reads. The project-root anchor itself still canonicalizes through real fs (it is a trusted value the caller derived, not attacker-influenced covered-input data); only per-file candidate reads go through the injected seam. Covered-file paths are now canonicalized (./ prefixes, redundant slashes, internal .. segments) before becoming dedup/sort/hash keys or confinement subjects — closes both a spurious-stale false positive (two spellings of the same file hashing differently) and a confinement gap (an internal .. segment that doesn't start the string). gsd-verifier.md now states covered-file paths are project-root-relative, not phaseDir-relative, closing an ambiguity that would have made a real verifier agent's first fingerprint invocation fail closed. defaultFsImpl's methods now late-bind through fs.<method> rather than capturing function references at module load — the earlier direct-capture form was invisible to existing tests' t.mock.method(fs, 'statSync', ...) seams, a real regression caught by the full suite (not the reviewer). * fix(#4155): catch a plan/summary added to the phase dir after verification but never declared The content digest only recomputes hashes for paths the verifier actually declared in covered_files — it had no way to notice a plan or summary added to the phase directory after verification if that new file was never declared, silently regressing behind the legacy mtime check it replaces (which scans the live directory, not a declared list). findUncoveredCurrentArtifact re-scans the live phase directory for every current *-PLAN.md/*-SUMMARY.md and requires each to be represented in covered_files, closing that gap; a directory scan failure fails closed to stale rather than silently skipping the check. CONTEXT.md's Verification Module entry corrected to describe the fingerprint path's stricter fail-closed FS-error contract (routes to stale) instead of the module's original degrade-to-safe one (missing / not-stale), which only the legacy path still keeps. * refactor(#4155): extract canonicalizeCoveredFiles, add real nested-project e2e test computeCoveredDigest and cmdVerificationFingerprint each normalized/deduped/ sorted covered_files independently — one shared helper now backs both (gemini review's ponytail-lens finding). Adds one CLI-to-readVerificationStatus test against a genuine .planning/phases/NN-x/ project with an implementation file outside .planning/ entirely, closing the review finding that prior #4155 unit fixtures put phaseDir directly under an ownerless tmpdir (findProjectRoot falls back to phaseDir itself there) and never exercised real multi-level path resolution. * fix(#4155): route computeCoveredDigest through real fs, fail closed on unreadable plans/ Two independent review rounds (opus critical-reviewer + opus ponytail + agy, run twice) found two instances of the same fail-open class: - computeCoveredDigest's per-file reads routed through the caller's injected fsImpl. planning-inspect.cts passes a `.planning/`-confined containment fs into readVerificationStatus's opts.fs, so any covered implementation file outside `.planning/` (mandatory per the issue) made the confinement wrapper throw, which was caught and turned into a stale digest -- reporting every fingerprinted phase permanently stale via `planning.inspect`, regardless of actual drift. Per-file reads now always use real node:fs, matching the pre-existing treatment of root canonicalization; the realRel-vs-realRoot check is the real confinement boundary for this data and needs no seam. - allCurrentArtifactsCovered's try/catch never fired (scanPhasePlans reports readdir failures via a `scope` field, it never throws), so an unreadable nested plans/ dir was silently treated as "zero artifacts, all covered" instead of failing closed. Now branches on scope !== SCOPE.COMPLETE. Also, per ponytail's second-round findings: reverted an unwarranted FINGERPRINT_VERSION bump and digest length-prefix from the first fix (no v1 digest has ever existed -- the feature is unreleased -- and the prefix closed a collision that grants no capability beyond what a writer of covered_files already has more cheaply); removed a verifier-facing escape-hatch instruction whose own example was a case that should trigger staleness, not bypass it; corrected CONTEXT.md references to the renamed allCurrentArtifactsCovered and a stale "unconditional" rescan claim; simplified the isStale derivation, removed dead FsLike members, and tightened test coverage. Regression tests for both fail-open bugs are included and were each confirmed to fail against the pre-fix code before the fix landed. full test suite: 2558/2560 pass, 2 skipped, 0 fail * fix(#4155): trim gsd-verifier.md under the LARGE size cap Fork CI caught what my local runs missed: the superseded/nested-plans instruction added earlier pushed gsd-verifier.md to 49299 bytes, 147 over the LARGE tier's 49152-byte hard cap (tests/agent-size-budget.test.cjs). Tightened the #4155 instruction's wording and dropped a redundant inline comment tag; no content lost. * chore(#4155): point changeset at the upstream PR number pr: 19 was the fork PR opened for internal review-lane CI; now that open-gsd/gsd-core#4290 exists, the changeset field must match it per CONTRIBUTING.md's release-notes convention. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
This commit is contained in:
committed by
GitHub
parent
5ad9a36f35
commit
5869febb16
@@ -24,7 +24,7 @@ Module owning phase create, rename, complete, remove, list, and plan-index opera
|
||||
Module owning phase-effort estimation and its calibration against measured reality (ADR-2629, epic #1952). Pure — no I/O, no config reads; the CLI seam (`src/estimate-cli.cts`, verbs `estimate-check` / `estimate-calibration`) owns reading `.planning/config.json` and `.planning/estimation-calibration.json`. Interface: `parseEstimate`/`renderEstimate` (the PLAN.md `estimate: {tokens, tasks, confidence}` block), `parseActuals`/`renderActuals` (the SUMMARY.md `actuals: {tokens, tasks, commits}` block), `deriveConfidence(sampleCount) → low|med|high`, `classifyAgainstBudget(estimate, budget) → {overBudget, ratio, recommendation, budgetValid}`, `computeCalibration(samples) → {factor, sampleCount, applied, confidence, clamped}`, `applyCalibration`, `parseCalibrationDocument`/`renderCalibrationDocument`, `extractFrontmatterBlock` (leading-`---`-anchored scalar-block reader; stays hand-rolled for its TYPE CONTRACT — it returns numeric-looking values as numbers so `parseEstimate`/`parseActuals` see the types they validate, whereas the migrated `extractFrontmatter`, ADR-3473 §8.1/#3881, resolves every scalar as a string under js-yaml's FAILSAFE_SCHEMA; migrating this function onto the shared parser is follow-on work under ADR-3473 §8.1, not done), `calibrationBasis` (returns `estimate.raw_tokens` when present, else `tokens` — calibration must measure actual/raw or the loop un-corrects itself), and `measureTokens` (a re-export of `prompt-budget`'s `estimateTokens`). **Domain terms: _raw_ vs _calibrated_ tokens** — the same token count in two mutually incompatible states, carried by the compile-time brands `RawTokens` (the planner's uncorrected projection, and the only legal calibration denominator) and `CalibratedTokens` (the projection with the project's factor applied, and the only figure meaningful against the budget), constructed at trust boundaries via `asRawTokens` / `asCalibratedTokens` (#2671). The brands erase at compile time — the emitted `.cjs`, the CLI JSON, and both frontmatter schemas are unchanged — and exist because mixing the two states was NOT catchable at runtime: both are positive integers of the same magnitude, and the mix-up shipped twice past a green ~26,800-test suite (#2631 factor², #2632 self-defeating loop). `asRawTokens` refuses a `CalibratedTokens` by design; the single legitimate crossover (a pre-#2632 plan whose `tokens` IS the raw projection) lives behind one commented assertion in `calibrationBasis`. Compile fixtures: `tests/fixtures/brand-typing/`. **_smart zone_** — the usable prefix of a model's context window before output quality degrades, expressed as the configurable `workflow.smart_zone_tokens` budget (default 100000, a *policy default* rather than a benchmark constant since the effective ceiling is model/task-dependent); **_estimate/actuals_** — a projected phase cost recorded at plan time and the measured cost recorded at completion, both on the **same `estimateTokens` scale** so their ratio measures the miss rather than a difference between two measurement methods. Two invariants: (1) every signal is **exogenous** — the correction routes on a measured actual/estimate ratio and `confidence` routes on a calibration sample count, never on a model's self-assessment (this project measured self-rated confidence and found it weak — `gsd-core/references/honest-verifier.md:25-29`; see `.out-of-scope/general-purpose-agent-prompt-skills.md`); (2) the over-budget flag is **advisory** — a warning plus a split recommendation, never a block. Calibration is median-of-ratios, clamped to `[0.5, 3.0]`, and inert below 3 samples. CLI seam verbs: `estimate-check` (classify one figure; `--calibrated` when the input already has the factor applied — omitting it squares the correction), `estimate-calibration` (report the current factor), `estimate-calibrate` (#2632 — pair every completed phase's PLAN `estimate` with its SUMMARY `actuals`, rebuild `.planning/estimation-calibration.json` idempotently, and report the result; this is what closes the loop). Source of truth: `gsd-core/bin/lib/phase-estimation.cjs` and `src/estimate-cli.cts`. Test anchors: `tests/phase-estimation.test.cjs`, `tests/estimate-calibrate.test.cjs`.
|
||||
|
||||
### Verification Module
|
||||
Module owning the canonical phase-verification status projection shared by phase transition, progress, manager, autonomous, and closeout readiness paths. `readVerificationStatus(phaseDir, opts?)` reads the first `*-VERIFICATION.md` frontmatter `status`, maps it through `VERIFICATION_ROUTING_TABLE`, and fail-closes — only `{passed}` satisfies the canonical gate; `missing`/`unknown`/`gaps_found`/`human_needed`/`stale` all route away from "complete" (#1522). `findStaleVerificationSummary` flags a SUMMARY newer than the VERIFICATION file (status `stale`). Both honor a no-throw, degrade-to-safe contract (any FS error → `missing` / not-stale) and an injectable `opts.fs` seam. `isPhaseComplete(phaseDir, deps?)` is the single canonical owner of "is phase P complete?" (ADR-3180 §7.4, issue #3186, disk-strict per #2957): it wraps `readVerificationStatus`, calling it UNCONDITIONALLY — plan count is never a precondition, so a zero-plan phase with a passing `*-VERIFICATION.md` is complete (#3168) — and returns `{ value: { complete, verification }, scope }`; `complete` is exactly `verification.status === 'passed'`. A ROADMAP checkbox carries no machine authority and is never consulted. `cmdPhaseComplete`, `buildPhaseCompletionProjection`, and `buildStateFrontmatter` all route through it. Source of truth: `gsd-core/bin/lib/verification.cjs` (generated from `src/verification.cts`).
|
||||
Module owning the canonical phase-verification status projection shared by phase transition, progress, manager, autonomous, and closeout readiness paths. `readVerificationStatus(phaseDir, opts?)` reads the first `*-VERIFICATION.md` frontmatter `status`, maps it through `VERIFICATION_ROUTING_TABLE`, and fail-closes — only `{passed}` satisfies the canonical gate; `missing`/`unknown`/`gaps_found`/`human_needed`/`stale` all route away from "complete" (#1522). Staleness is one of TWO mutually exclusive checks, selected by whether the report DECLARES a fingerprint (#4155, i.e. either `covered_files` or `covered_digest` is present — a report that declares only one, or an empty/malformed pair, fails closed to `stale` rather than silently downgrading to the weaker legacy check): a report with a well-formed pair is checked by `computeCoveredDigest(projectRoot, coveredFiles)` — a deterministic sha256-of-sha256s over the SORTED, de-duplicated, `./`/`..`-canonicalized covered set, hashed by file BYTES always through the real `node:fs`, never a caller-injected `opts.fs` seam (`covered_files` spans the whole `projectRoot`, not just `.planning/`, so a caller-scoped containment wrapper narrower than `projectRoot` would reject legitimate covered files; the `realRel`-vs-`realRoot` re-check inside `computeCoveredDigest` is the actual confinement boundary, against `projectRoot`) — recomputed and compared to the stored digest. `allCurrentArtifactsCovered` additionally re-scans the LIVE phase directory (via `scanPhasePlans`, short-circuited so it only runs once the digest itself has matched, and fails closed to `false` on any scope other than `SCOPE.COMPLETE`) for every current `*-PLAN.md`/`*-SUMMARY.md` and requires each to be represented in `covered_files`: a plan or summary added to the phase AFTER verification, never declared, would otherwise pass the digest check untouched. ANY mismatch — digest, an uncovered current artifact, or a covered path that is missing, unreadable, or escapes `projectRoot` (`findProjectRoot(phaseDir)`, canonicalized via real `fs`, never the injected seam — it is a trusted caller-derived anchor, not attacker-influenced covered-input data) — routes to `stale`, fail-closed (`computeCoveredDigest` returns `null` for the whole set rather than a partial digest). A report with no fingerprint metadata at all (every report predating #4155) falls back to the legacy `findStaleVerificationSummary`, which flags a SUMMARY newer than the VERIFICATION file — unchanged. The verifier computes `covered_digest` via the CLI (`verification.fingerprint <phaseDir> <file>...`, routed in `verification-command-router.cjs`), never by hand — this is deterministic math, not an LLM-estimated value. The legacy path alone keeps the module's original no-throw, degrade-to-safe contract (any FS error → `missing` / not-stale); the fingerprint path's contract is stricter (any FS error → `stale`), matching #4155's fail-closed design. `isPhaseComplete(phaseDir, deps?)` is the single canonical owner of "is phase P complete?" (ADR-3180 §7.4, issue #3186, disk-strict per #2957): it wraps `readVerificationStatus`, calling it UNCONDITIONALLY — plan count is never a precondition, so a zero-plan phase with a passing `*-VERIFICATION.md` is complete (#3168) — and returns `{ value: { complete, verification }, scope }`; `complete` is exactly `verification.status === 'passed'`. A ROADMAP checkbox carries no machine authority and is never consulted. `cmdPhaseComplete`, `buildPhaseCompletionProjection`, and `buildStateFrontmatter` all route through it. Source of truth: `gsd-core/bin/lib/verification.cjs` (generated from `src/verification.cts`).
|
||||
|
||||
`SEAM.verification-isphasecomplete.owns=isPhaseComplete(phaseDir, deps?) is the single canonical owner of "is phase P complete?" (ADR-3180 §7.4, issue #3186)`
|
||||
`SEAM.verification-isphasecomplete.enforced-by=test:tests/verification-status.test.cjs`
|
||||
|
||||
Reference in New Issue
Block a user