* feat(#3243): sync installed codex .toml model/effort to the passive posture Implements ADR-2313 D7, and owns the Codex .toml typed IR that Phase 1's review assigned to this phase. The IR exists for a structural reason, not tidiness: this phase has to PARSE these files, and a parser kept bug-compatible with a separate renderer is the generative-fix-divergence shape this epic already dealt with once for the model predicate. So Phase 2's parsing MOVES here rather than being copied — agent-install-check now imports it, and its test file passing unchanged is the proof the extraction altered nothing. The load-bearing property is byte-identical round-trip: render(parse(x)) === x. Without it a sync silently reformats a user's file — line endings, key order, BOM, trailing newline — turning a two-line repair into a whole-file diff in their dotfile repo. The IR keeps original lines and removes targeted ones rather than reconstructing from parsed fields, which is what makes that property hold. It also reconciles a real contradiction between Phase 2 and ADR-2313. An unterminated developer_instructions block: the reader excludes the rest of the file, deliberately failing toward a false positive, because misreading prose as a pin only wastes a user's time. The writer must refuse, because proceeding on a malformed document rewrites it. A false positive is the safe direction for a reader and the dangerous one for a writer. So the parse reports the fact and the two consumers branch on it — one parse, one truth, two policies, instead of two parsers that agree today. The sync leaves a legal real-Codex pin and its coupled effort untouched, reported skipped rather than synced; strips a stale Anthropic or tier model and an orphaned effort; keeps dry-run as the default; refuses any file whose parse fails; and skips symlinks exactly as the Claude path already did. The Claude path itself is byte-identical. PARSE_REASON.NO_HEADER from the ADR's illustrative snippet is deliberately not implemented — a missing header is legal, not an error, so it would be a dead enum member that the enum-lock test then pins. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3243): preserve per-line endings and make the codex write atomic Two findings from an isolated review, both in the write path. BLOCKER: mixed line endings broke the byte-identical round-trip. `eol` was a single whole-file flag and split(/\r?\n/) discarded each line's own terminator, so render re-joined with ONE style and normalized every line — even with zero strips performed. A file with one CRLF line and the rest LF came back fully converted. That falsified the A14 guarantee, violated the design's "must not silently rewrite every line", and made the CONTEXT.md glossary claim wrong. It was untested because A12 and B15 only cover PURE CRLF; no mixed-ending fixture existed anywhere. Fixed by keeping each line's terminator alongside its content, so render is a plain concatenation and a strip removes only the target line and its own terminator. `eol` survives as informational metadata that render never reads. Seven fixtures added for the paths nothing exercised: mixed endings unmodified and with a strip, a lone \r, a file ending on the block's closing ''' with no newline, multiple trailing newlines, a BOM-only file, and an empty file. MINOR, but it contradicted this phase's own contract: the write was in-place open-truncate, so a failure between truncate and completion leaves a truncated .toml — exactly what ADR-2313 says must never happen. The Codex path now writes a sibling temp file and renames over the target, which is atomic on one filesystem, with cleanup on failure. It uses the repo's existing retryRenameSync rather than a hand-rolled rename, and deliberately NOT platformWriteSync, whose normalizeContent would mangle the very CRLF and trailing-newline bytes the round-trip property exists to preserve. The Claude path keeps its in-place write untouched. It has the same shape, but changing it is not this phase's business and its tests must stay byte-identical. B20 previously mocked writeFileSync to throw BEFORE touching anything, so it proved nothing about a mid-write failure — its passing comment was true only because of how the mock was built. It now performs a real truncated write wherever writeFileSync is called, catching both the naive direct-to-target path and the new temp path, and asserts the target is byte-identical afterwards with no stray temp file left. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3243): preserve the trailing-newline state when stripping a last line Caught by B17, one of this phase's own tests — the suite working, not a test problem. Content is reconstructed as the concatenation of lines[k] + terminators[k], so a file with no trailing newline has '' as its last terminator. removeLine spliced out both arrays at the same index, which is right for a middle line but wrong for the last one: it dropped the empty terminator and left the PREVIOUS line's newline in place. A file ending `...\nmodel = "sonnet"` with no trailing newline came back as `...\n`, gaining a newline the user never wrote. The new last line now inherits the removed line's terminator, so a removal leaves the file exactly as if that line had never been written. Removing the only line yields an empty file rather than a stray terminator. Both stripModel and stripReasoningEffort funnel through the one removeLine, confirmed rather than assumed, so a single fix covers both — including the row-B7 shape where a stale model and its orphaned effort are removed in sequence and the second removal targets the last line. Two of the four new cases are honestly not red-first and say so in their comments: removing a last line that HAS a trailing newline only exposes the bug under mixed EOL, since uniform files coincidentally have equal terminators on both sides; and removing the only line already degenerated correctly through Array.slice. They are kept as guards for the new branch rather than dressed up as catches. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3243): document the codex repair path and close the loop How-to: the Codex-400 entry added in Phase 2 told users to re-run the installer, because that was the only repair available then. It now leads with `effort sync` and keeps the reinstall as the alternative, with the reason to prefer one — a reinstall regenerates the agent files wholesale, so anyone who hand-edited theirs loses those edits. Detect, preview, apply is now one continuous path in one place. Reference: docs/COMMANDS.md had no `effort sync` entry at all — the same gap `validate agents` had in Phase 2, found the same way. The entry documents BOTH runtimes, because the command genuinely forks on runtime and describing only the new half would misdescribe it. The write flag is `--apply`. The design doc and test matrix both said `--no-dry-run` throughout, which does not exist — verified against the actual arg parser in gsd-tools.cjs before writing. Documenting a flag that does not exist is worse than documenting nothing, because it fails at the moment someone needs it. Both surfaces state that only the targeted lines are removed and every other byte is preserved. That is a user-visible guarantee rather than an implementation note: it is the difference between a two-line diff and a reformatted file in someone's dotfile repo, it is what the IR's round-trip property exists to deliver, and writing it down makes it a contract a future change has to break knowingly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3243): inherit the trailing-newline state, not the line ending style My previous rule was subtly wrong and this phase's own test caught it. "The new last line inherits the removed line's terminator" copies the removed line's STYLE as well as its presence. A26 uses mixed endings on purpose — line one terminated \r\n, the model line terminated \n — so inheriting silently rewrote line one's ending to \n. That is precisely the defect class the mixed-EOL blocker fix existed to eliminate, reintroduced one layer down by the fix for it. The correct rule inherits the EMPTINESS only. If the removed line had no terminator, the new last line loses its own, preserving "this file has no trailing newline". Otherwise the new last line keeps its own terminator: it is already a newline, and already the right style for that line. A26's assertion moved too, and that deserves saying plainly rather than burying: it previously encoded my wrong rule. Changing a test to match the implementation is usually the mistake, so it was checked from first principles instead — a file whose first line ends \r\n and whose last line ends \n, with that last line removed entirely, must be the first line with its own \r\n intact. The new expectation is what the user's file should actually look like; the old one was wrong. A29 adds the interaction nothing covered: the compounding case (strip a stale model, then its orphaned effort, the second removal landing on the last line) with non-uniform endings either side. The two fixes meet there and nothing exercised the meeting point. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3243): drop the phantom trailing line from the IR representation Root cause, not another patch on the removal rule. Three consecutive fixes there each surfaced the next issue, which was the signal that the data model was wrong. splitPreservingTerminators left a phantom empty final entry for any file ending in a newline: "a\nb\n" became lines ['a','b','']. So for the common case the real last content line was NOT the last array element, removeLine's isLastLine check never matched it, and every rule I gave was reasoning about the wrong element. What hid it: render was already a plain concatenation, so a phantom empty line with an empty terminator contributes nothing to the output. A14's byte-identical round-trip could never have caught it — the defect is byte-neutral until a removal shifts the index arithmetic under it. That is worth recording, because "the round-trip test is green" was exactly the reassurance that kept the search pointed elsewhere. The representation is now 1:1 — terminators[i] follows lines[i] and may be '' — with no phantom, verified across empty, no-trailing-newline, trailing-newline, blank-line and mixed-CRLF inputs. render stays a plain concat and needs no special cases. With the phantom gone the removal rule is correct as stated and finally applies to the genuinely last element. Consumers checked rather than assumed: the block-range detector and header scanner are agnostic to array shape, and Phase 2's reader uses its own independent split, so tests/agent-install-check.test.cjs is untouched and still passes unchanged. One test expectation was wrong and is corrected rather than quietly adjusted: A18 asserted a 7-element terminators array whose trailing '' was the phantom itself. It now asserts the six real terminators, which is what the invariant lines.length === terminators.length requires. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3243): backfill changeset pr number (#3296) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
392 lines
20 KiB
TypeScript
392 lines
20 KiB
TypeScript
/**
|
|
* Codex Agent TOML — typed IR for `~/.codex/agents/<agent>.toml` (#3243, ADR-2313).
|
|
*
|
|
* A genuine leaf: node builtins only. This is a **document model**, not a policy —
|
|
* it knows how to parse/render/strip two known keys (`model`,
|
|
* `model_reasoning_effort`) from a Codex agent `.toml`. It does NOT know which
|
|
* `model` values are illegal for Codex (that predicate — Anthropic-flavored
|
|
* detection — stays in `model-catalog.cts`; callers decide what to strip).
|
|
*
|
|
* Moved here (not copied) from `agent-install-check.cts` (#3242, Phase 2), which
|
|
* wrote the hard half: block-range detection, BOM stripping, TOML value
|
|
* unquoting, and the lenient header scan. That module's behavior is UNCHANGED —
|
|
* it imports `stripBOM`/`scanTomlLines` from here and its regression suite
|
|
* (`tests/agent-install-check.test.cjs`) is the proof.
|
|
*
|
|
* ── The reconciliation (40-design.md) ──────────────────────────────────────
|
|
*
|
|
* Phase 2's reader and this phase's writer disagree on how to handle an
|
|
* unterminated `developer_instructions` block, deliberately:
|
|
*
|
|
* - The READER (`scanTomlLines`, used directly by `checkCodexModelPosture`)
|
|
* stays LENIENT: an unterminated block still excludes "the rest of the
|
|
* file" from the header scan (findDeveloperInstructionsBlockRange's
|
|
* existing fallback), because misreading prompt prose as a pin is only a
|
|
* false positive — it wastes a user's time, nothing more.
|
|
* - The WRITER (`parseCodexAgentToml`, used by the Codex sync) is STRICT: an
|
|
* unterminated block makes the whole document `{ok:false}`, because a
|
|
* writer that proceeds on a malformed document risks rewriting it.
|
|
*
|
|
* One block-range detector, two call sites, two policies — never two detectors
|
|
* that could silently drift from each other.
|
|
*/
|
|
|
|
/** Frozen reason enum for a failed {@link parseCodexAgentToml}. */
|
|
export const PARSE_REASON = Object.freeze({
|
|
UNTERMINATED_BLOCK: 'unterminated_block',
|
|
});
|
|
|
|
export type ParseReason = (typeof PARSE_REASON)[keyof typeof PARSE_REASON];
|
|
|
|
/**
|
|
* The typed IR. Carries the original `lines` (never re-tokenized once parsed)
|
|
* plus the detected BOM/trailing-newline metadata so {@link renderCodexAgentToml}
|
|
* can reproduce the source **byte-identically** when nothing was stripped (the
|
|
* load-bearing round-trip property — see 50-test-matrix.md row A14), even when
|
|
* the source mixes line-ending styles (`\r\n`, `\n`, a lone `\r`) within one
|
|
* file. `stripModel`/`stripReasoningEffort` remove a targeted line AND its own
|
|
* terminator from `lines`/`terminators`; every other line's content is
|
|
* untouched, and so is its own terminator — EXCEPT when the removed line was
|
|
* the file's last line AND the removed line had no terminator (the source had
|
|
* no trailing newline), in which case the new last line's terminator is
|
|
* cleared to `''` too (see {@link removeLine}) so the file's trailing-newline
|
|
* status is preserved rather than silently changed. When the source DID end
|
|
* with a newline, the new last line keeps its OWN terminator unchanged —
|
|
* copying the removed line's terminator there would corrupt a mixed-EOL
|
|
* source by silently changing the new last line's own ending style.
|
|
*/
|
|
export interface CodexAgentDoc {
|
|
/** Content lines (BOM-stripped, terminator-free). Paired 1:1 by index with `terminators`. */
|
|
lines: string[];
|
|
/**
|
|
* Each line's OWN terminator (`'\r\n'`, `'\n'`, `'\r'`, or `''` for a line
|
|
* with none — the last line of a source with no trailing newline). Never
|
|
* collapsed to one whole-file style: {@link renderCodexAgentToml} rejoins
|
|
* `lines[i] + terminators[i]` so a mixed-EOL source round-trips exactly.
|
|
*/
|
|
terminators: string[];
|
|
/**
|
|
* The line-ending style that appears in the source, informational only —
|
|
* `renderCodexAgentToml` does NOT use this field (it rejoins each line with
|
|
* its own `terminators[i]` instead). Recorded as `'\r\n'` if any `\r\n`
|
|
* appears anywhere in the source, else `'\n'`, purely for callers that want
|
|
* a human-readable summary (e.g. logging); never a rendering input.
|
|
*/
|
|
eol: '\n' | '\r\n';
|
|
/** Whether the source began with a UTF-8 BOM (U+FEFF). */
|
|
hadBOM: boolean;
|
|
/** Whether the source ended with a line terminator (any of `\r\n`/`\n`/`\r`). */
|
|
trailingNewline: boolean;
|
|
/** The `developer_instructions = '''...'''` block's line range, or `{start:-1,end:-1}` if absent. */
|
|
blockRange: { start: number; end: number };
|
|
/** The resolved `model` value (last occurrence outside the block), or null. */
|
|
model: string | null;
|
|
/** Line index of the `model` key, or null if absent. */
|
|
modelLineIndex: number | null;
|
|
/** The resolved `model_reasoning_effort` value (last occurrence outside the block), or null. */
|
|
reasoningEffort: string | null;
|
|
/** Line index of the `model_reasoning_effort` key, or null if absent. */
|
|
reasoningEffortLineIndex: number | null;
|
|
}
|
|
|
|
export type ParseCodexAgentTomlResult =
|
|
| { ok: true; doc: CodexAgentDoc }
|
|
| { ok: false; reason: ParseReason };
|
|
|
|
// The UTF-8 BOM codepoint, spelled as an escape rather than the literal
|
|
// character so the source file never carries an invisible codepoint.
|
|
const BOM_CHAR = String.fromCharCode(0xfeff);
|
|
|
|
// Strips a leading UTF-8 BOM (U+FEFF), which fs.readFileSync(..., 'utf8') does not
|
|
// strip on its own, and unwraps a TOML basic/literal string value's surrounding
|
|
// quotes so `model = "sonnet"` yields `sonnet`, not `"sonnet"`.
|
|
export function stripBOM(content: string): string {
|
|
return content.charCodeAt(0) === 0xfeff ? content.slice(1) : content;
|
|
}
|
|
|
|
export function unquoteTomlValue(rawValue: string): string {
|
|
const trimmed = rawValue.trim();
|
|
const quoted = trimmed.match(/^"([^"]*)"/) ?? trimmed.match(/^'([^']*)'/);
|
|
return quoted ? quoted[1] : trimmed;
|
|
}
|
|
|
|
// The `developer_instructions` block is a TOML multi-line literal string
|
|
// (`developer_instructions = '''...'''`) that `generateCodexAgentToml` always
|
|
// emits after the header fields. Prompt prose inside that block discusses models
|
|
// constantly, so a `model = ...`-shaped line inside it must never be read as a
|
|
// live pin — but the block can legally appear anywhere in the file (a
|
|
// hand-reordered agent can move `model` after it), and another key's *value* can
|
|
// legally contain the literal text `developer_instructions = '''` (e.g. a
|
|
// `description` field quoting it) without that being the real block opener. So
|
|
// instead of truncating the file at the first textual occurrence of the marker
|
|
// anywhere in the content, this locates the block by its anchored line-start
|
|
// opener (`^\s*developer_instructions\s*=\s*'''`, never a mid-line/mid-value
|
|
// match) and its closing `'''` line, and excludes only the lines between them —
|
|
// every other line in the file, before AND after the block, is scanned.
|
|
//
|
|
// If no opener is found, nothing is excluded (the whole file is scanned) and
|
|
// `terminated` is trivially true. If the block IS opened but never closed before
|
|
// EOF (malformed file), `terminated` is false: the lenient reader (scanTomlLines)
|
|
// still treats the rest of the file as inside the block (the safe direction for
|
|
// a reader — see module header comment); the strict writer (parseCodexAgentToml)
|
|
// reads `terminated` and refuses instead. The emitter always uses `'''` (a TOML
|
|
// literal string), never a `"""` basic multi-line string, so only `'''` is
|
|
// treated as the block delimiter here.
|
|
export function findDeveloperInstructionsBlockRange(
|
|
lines: string[],
|
|
): { start: number; end: number; terminated: boolean } {
|
|
const openIndex = lines.findIndex((line) => /^\s*developer_instructions\s*=\s*'''/.test(line));
|
|
if (openIndex === -1) {
|
|
return { start: -1, end: -1, terminated: true };
|
|
}
|
|
const afterOpenMarker = lines[openIndex].replace(/^\s*developer_instructions\s*=\s*'''/, '');
|
|
if (afterOpenMarker.includes("'''")) {
|
|
// Same-line block: developer_instructions = '''one line'''
|
|
return { start: openIndex, end: openIndex, terminated: true };
|
|
}
|
|
for (let i = openIndex + 1; i < lines.length; i++) {
|
|
if (lines[i].includes("'''")) {
|
|
return { start: openIndex, end: i, terminated: true };
|
|
}
|
|
}
|
|
return { start: openIndex, end: lines.length - 1, terminated: false };
|
|
}
|
|
|
|
/** Header-scan result shared by the lenient reader ({@link scanTomlLines}). */
|
|
export interface HeaderScanResult {
|
|
model: string | null;
|
|
hasReasoningEffort: boolean;
|
|
}
|
|
|
|
interface HeaderLineInfo {
|
|
model: string | null;
|
|
modelLineIndex: number | null;
|
|
reasoningEffort: string | null;
|
|
reasoningEffortLineIndex: number | null;
|
|
}
|
|
|
|
// Line-oriented scan of every line OUTSIDE the `developer_instructions` block
|
|
// (see findDeveloperInstructionsBlockRange). Full-key-name anchoring —
|
|
// `^([A-Za-z_][\w]*)\s*=` for a bare key, or `^"([^"]*)"\s*=` / `^'([^']*)'\s*=`
|
|
// for TOML's legal quoted-key forms, normalized to the same key name — means
|
|
// `model_verbosity` / `model_reasoning_effort` never satisfy a `model` probe,
|
|
// and vice versa; `#`-prefixed lines (after trimming leading whitespace) are
|
|
// treated as comments, never live pins. Shared by both scanTomlLines (the
|
|
// lenient reader, boolean-only for reasoning effort) and parseCodexAgentToml
|
|
// (the strict writer, which also needs the effort's value and both keys' line
|
|
// indices so stripModel/stripReasoningEffort can remove exactly one line).
|
|
function scanHeaderLines(lines: string[], blockStart: number, blockEnd: number): HeaderLineInfo {
|
|
let model: string | null = null;
|
|
let modelLineIndex: number | null = null;
|
|
let reasoningEffort: string | null = null;
|
|
let reasoningEffortLineIndex: number | null = null;
|
|
for (let i = 0; i < lines.length; i++) {
|
|
if (blockStart !== -1 && i >= blockStart && i <= blockEnd) continue;
|
|
const trimmed = lines[i].trim();
|
|
if (trimmed === '' || trimmed.startsWith('#')) continue;
|
|
const match = trimmed.match(/^(?:"([^"]*)"|'([^']*)'|([A-Za-z_][\w]*))\s*=\s*(.*)$/);
|
|
if (!match) continue;
|
|
const key = match[1] ?? match[2] ?? match[3];
|
|
const rawValue = match[4];
|
|
if (key === 'model') {
|
|
model = unquoteTomlValue(rawValue);
|
|
modelLineIndex = i;
|
|
} else if (key === 'model_reasoning_effort') {
|
|
reasoningEffort = unquoteTomlValue(rawValue);
|
|
reasoningEffortLineIndex = i;
|
|
}
|
|
}
|
|
return { model, modelLineIndex, reasoningEffort, reasoningEffortLineIndex };
|
|
}
|
|
|
|
/**
|
|
* The LENIENT reader entry point (Phase 2, moved verbatim in behavior). Never
|
|
* fails: an unterminated block falls back to "rest of file is inside the
|
|
* block" via {@link findDeveloperInstructionsBlockRange}'s own fallback.
|
|
* `content` is expected already BOM-stripped (callers pass `stripBOM(raw)`).
|
|
*/
|
|
export function scanTomlLines(content: string): HeaderScanResult {
|
|
const lines = content.split(/\r?\n/);
|
|
const { start, end } = findDeveloperInstructionsBlockRange(lines);
|
|
const { model, reasoningEffort } = scanHeaderLines(lines, start, end);
|
|
return { model, hasReasoningEffort: reasoningEffort !== null };
|
|
}
|
|
|
|
// Splits `content` into `{lines, terminators}` where `terminators[i]` is the
|
|
// terminator that FOLLOWS `lines[i]` (`'\r\n'`, `'\r'`, `'\n'`, or `''` for a
|
|
// line with none — only possible as the file's last line). The two arrays are
|
|
// always the same length and there is NEVER a phantom trailing entry: a
|
|
// source ending in a terminator (the common case) yields exactly as many
|
|
// lines as it has content lines, not one more. `render` is then a plain
|
|
// `lines[i] + terminators[i]` concatenation with no special-casing of "the
|
|
// last line" — see `renderCodexAgentToml`.
|
|
//
|
|
// `String#split` with a capturing group interleaves the delimiters into the
|
|
// result array — `"a\r\nb\nc".split(/(\r\n|\r|\n)/)` yields
|
|
// `["a","\r\n","b","\n","c"]` — so even indices are line content and odd
|
|
// indices are that line's terminator. When `content` ends WITH a terminator,
|
|
// `split` appends one extra empty-string element after the last real
|
|
// terminator (e.g. `"a\n".split(...)` → `["a","\n",""]`); that trailing `""`
|
|
// is not a real line, it is `split`'s "nothing after the last delimiter"
|
|
// marker, so the loop below stops before consuming it instead of recording it
|
|
// as a phantom empty final line (the defect this replaced — see A29: a doc
|
|
// with a phantom last element made every removal rule reason about the wrong
|
|
// element for any trailing-newline-terminated file, the common case). `\r\n`
|
|
// is tried before the bare `\r` alternative so a CRLF is never misread as a
|
|
// lone-CR line followed by an empty LF-terminated line.
|
|
function splitPreservingTerminators(content: string): { lines: string[]; terminators: string[] } {
|
|
if (content === '') return { lines: [], terminators: [] };
|
|
const parts = content.split(/(\r\n|\r|\n)/);
|
|
const lastIndex = parts.length - 1;
|
|
const lines: string[] = [];
|
|
const terminators: string[] = [];
|
|
for (let i = 0; i < parts.length; i += 2) {
|
|
if (i === lastIndex && parts[i] === '') break; // split's post-terminator marker, not a real line
|
|
lines.push(parts[i]);
|
|
terminators.push(parts[i + 1] ?? '');
|
|
}
|
|
return { lines, terminators };
|
|
}
|
|
|
|
/**
|
|
* The STRICT parse entry point (Phase 3, the writer's half of the
|
|
* reconciliation). Returns `{ok:false, reason:UNTERMINATED_BLOCK}` rather than
|
|
* guessing when the `developer_instructions` block is opened but never closed.
|
|
* On success, `doc` carries enough (the original `lines`/`terminators`, BOM/
|
|
* trailing-newline flags, and the two resolved values with their line indices)
|
|
* for {@link renderCodexAgentToml} to reproduce the source byte-identically —
|
|
* including a source with mixed line-ending styles — and for
|
|
* {@link stripModel}/{@link stripReasoningEffort} to remove exactly one line
|
|
* and its own terminator.
|
|
*/
|
|
export function parseCodexAgentToml(content: string): ParseCodexAgentTomlResult {
|
|
const hadBOM = content.charCodeAt(0) === 0xfeff;
|
|
const stripped = stripBOM(content);
|
|
// Informational only — see CodexAgentDoc.eol's docstring. Never used by
|
|
// renderCodexAgentToml.
|
|
const eol: '\n' | '\r\n' = stripped.includes('\r\n') ? '\r\n' : '\n';
|
|
const trailingNewline = /(\r\n|\r|\n)$/.test(stripped);
|
|
const { lines, terminators } = splitPreservingTerminators(stripped);
|
|
const { start, end, terminated } = findDeveloperInstructionsBlockRange(lines);
|
|
if (start !== -1 && !terminated) {
|
|
return { ok: false, reason: PARSE_REASON.UNTERMINATED_BLOCK };
|
|
}
|
|
const { model, modelLineIndex, reasoningEffort, reasoningEffortLineIndex } = scanHeaderLines(lines, start, end);
|
|
const doc: CodexAgentDoc = {
|
|
lines,
|
|
terminators,
|
|
eol,
|
|
hadBOM,
|
|
trailingNewline,
|
|
blockRange: { start, end },
|
|
model,
|
|
modelLineIndex,
|
|
reasoningEffort,
|
|
reasoningEffortLineIndex,
|
|
};
|
|
return { ok: true, doc };
|
|
}
|
|
|
|
/**
|
|
* Renders `doc` back to a string. For an unmodified doc this is
|
|
* byte-identical to the original `parseCodexAgentToml` input (matrix row
|
|
* A14) — it never re-derives line content, only rejoins each line with its
|
|
* OWN recorded terminator (`terminators[i]`, never the whole-file `eol`) and
|
|
* re-prepends a BOM if one was present. This is a plain concatenation of the
|
|
* surviving `[line, terminator]` pieces, so a source with mixed `\r\n`/`\n`/
|
|
* lone-`\r` line endings round-trips exactly, and a strip
|
|
* ({@link stripModel}/{@link stripReasoningEffort}) removes only the target
|
|
* line and its own terminator — every other line's ending is untouched.
|
|
*/
|
|
export function renderCodexAgentToml(doc: CodexAgentDoc): string {
|
|
let body = '';
|
|
for (let i = 0; i < doc.lines.length; i++) {
|
|
body += doc.lines[i] + (doc.terminators[i] ?? '');
|
|
}
|
|
return doc.hadBOM ? BOM_CHAR + body : body;
|
|
}
|
|
|
|
// Removes exactly one line (by index) — AND its own terminator — from
|
|
// `doc.lines`/`doc.terminators`, re-indexing the block range and the OTHER
|
|
// key's line index so a subsequent strip/render still sees a consistent doc.
|
|
// Never touches any other line's content or terminator.
|
|
//
|
|
// The one exception is when `index` names the file's LAST line: a plain
|
|
// slice-out would drop the removed line's terminator but leave the
|
|
// *previous* line's terminator standing in its place, which silently
|
|
// invents (or drops) a trailing newline the source never had — a middle-line
|
|
// removal never has this problem because the terminator that survives (the
|
|
// one that WAS between the previous line and the removed one) is exactly the
|
|
// terminator the new neighbors should have between them. For a last-line
|
|
// removal, the file's trailing-newline-or-not status lives in whether the
|
|
// REMOVED line's own terminator was empty (that is what `trailingNewline`
|
|
// was computed from) — so the new last line inherits the removed line's
|
|
// EMPTINESS only: if the removed terminator was `''`, the new last line's
|
|
// terminator is cleared to `''` too. If the removed terminator was
|
|
// non-empty, the source already ended with a newline and the new last line
|
|
// already has the right one (its OWN, unchanged) — overwriting it with the
|
|
// removed line's terminator would silently change the new last line's own
|
|
// ending style on a mixed-EOL source (see A26). Removing the only remaining
|
|
// line is the degenerate case: there is no new last line, so the result is
|
|
// the empty document.
|
|
function removeLine(doc: CodexAgentDoc, index: number, which: 'model' | 'reasoningEffort'): CodexAgentDoc {
|
|
const isLastLine = index === doc.lines.length - 1;
|
|
let lines: string[];
|
|
let terminators: string[];
|
|
if (doc.lines.length === 1) {
|
|
lines = [];
|
|
terminators = [];
|
|
} else if (isLastLine) {
|
|
lines = doc.lines.slice(0, index);
|
|
terminators = doc.terminators.slice(0, index);
|
|
// Inherit the removed line's EMPTINESS, never its STYLE: if the removed
|
|
// line had no terminator (the source had no trailing newline), the new
|
|
// last line's terminator becomes '' too. Otherwise the source DID end
|
|
// with a newline, and the new last line already has the right one — its
|
|
// OWN terminator (already carried over by the slice above), which may
|
|
// differ in style from the removed line's (a mixed-EOL source) — so it is
|
|
// left unchanged rather than overwritten.
|
|
if (doc.terminators[index] === '') {
|
|
terminators[terminators.length - 1] = '';
|
|
}
|
|
} else {
|
|
lines = doc.lines.slice(0, index).concat(doc.lines.slice(index + 1));
|
|
terminators = doc.terminators.slice(0, index).concat(doc.terminators.slice(index + 1));
|
|
}
|
|
const reindex = (i: number | null): number | null => (i === null ? null : i > index ? i - 1 : i);
|
|
const blockRange = { ...doc.blockRange };
|
|
if (blockRange.start !== -1) {
|
|
if (blockRange.start > index) blockRange.start -= 1;
|
|
if (blockRange.end > index) blockRange.end -= 1;
|
|
}
|
|
return {
|
|
...doc,
|
|
lines,
|
|
terminators,
|
|
blockRange,
|
|
model: which === 'model' ? null : doc.model,
|
|
modelLineIndex: which === 'model' ? null : reindex(doc.modelLineIndex),
|
|
reasoningEffort: which === 'reasoningEffort' ? null : doc.reasoningEffort,
|
|
reasoningEffortLineIndex: which === 'reasoningEffort' ? null : reindex(doc.reasoningEffortLineIndex),
|
|
};
|
|
}
|
|
|
|
/**
|
|
* Returns a new doc with the `model` line removed (a no-op copy if there was
|
|
* no `model` line). Every other byte — comments, other keys, the
|
|
* `developer_instructions` block, line endings, BOM — is untouched.
|
|
*/
|
|
export function stripModel(doc: CodexAgentDoc): CodexAgentDoc {
|
|
if (doc.modelLineIndex === null) return doc;
|
|
return removeLine(doc, doc.modelLineIndex, 'model');
|
|
}
|
|
|
|
/**
|
|
* Returns a new doc with the `model_reasoning_effort` line removed (a no-op
|
|
* copy if there was none). Every other byte is untouched.
|
|
*/
|
|
export function stripReasoningEffort(doc: CodexAgentDoc): CodexAgentDoc {
|
|
if (doc.reasoningEffortLineIndex === null) return doc;
|
|
return removeLine(doc, doc.reasoningEffortLineIndex, 'reasoningEffort');
|
|
}
|