Files
msd-core/agents/gsd-user-profiler.compact.md
Tom Boucher 37b965c0d1 enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code (#4553)
* enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code

ADR-4139 stream 2. The non-Claude `#2454` persona fallback in cmdAgentSkills
(src/init.cts) now selects between a canonical agents/<name>.md and a
token-minimized agents/<name>.compact.md sibling based on
workflow.compact_content, resolved in code (a real function call with a real
exit code) rather than a prose config-get gate — the same precedent stream 1's
spine/detail split established for a load-bearing seam, applied here because
this seam already runs through TypeScript instead of an eager @-include.

A missing compact sibling falls back to the canonical persona and discloses
the fallback in the served payload itself (a leading HTML-comment provenance
line), so the Done-when contract — compact when on, canonical when off, never
silent or empty — holds even for an agent nobody has compacted yet.

Authored a .compact.md sibling for all 35 shipped agents (agents/gsd-*.md),
each an independent, complete rewrite (not an extraction — nothing is "moved"
the way spine/detail moves text) that preserves frontmatter, every @-include,
every output-format contract, and every guardrail verbatim while cutting
restatement and verbose framing. Verified mechanically: every pair registers
(a canonical sibling exists), every compact file is strictly smaller, and the
full @-include set matches canonical's — including which references are
standalone eager-load lines versus inline prose mentions, since demoting one
to inline changes what the host actually substitutes.

Traced the install path before writing any code (.gsd/phase/.../40-design.md):
stageAgentsForRuntimeWithConverter glob-copies every agents/*.md file with no
stem filtering under the default full profile, so the new .compact.md files
install for free with zero installer changes — matching issue #4407's stated
scope. A tiered agent profile that doesn't stage a compact sibling degrades
through the same fallback-with-provenance path already required for an
unauthored one, so no installer change is needed there either.

Extends tests/helpers/compact-content-variant.cjs with an AGENTS_ROOT export
(deliberately not folded into DEFAULT_VARIANT_ROOTS, since agent variants are
reached by a generic code construction rather than a literal path in prose,
and checkReachability's markdown-search shape has nothing to find there).
Reachability is instead proven behaviorally: tests/agent-skills.test.cjs's new
"#4407 compact payload selection" describe block spawns gsd_run agent-skills
against real compact/canonical fixture pairs and asserts on the served
payload, which can only pass if the seam genuinely wires through.

Fixed a pre-existing test whose agents/*.md glob incidentally matched the new
.compact.md siblings (tests/agent-skills.test.cjs's Skill-frontmatter drift
guard) and added the 35 new agents/*.compact.md entries to docs/INVENTORY.md's
roster, both real, unrelated-to-content defects the new files' mere existence
surfaced.

Regenerated: install-tree fixtures (19 runtimes now ship 35 more agent files
under the full profile), INVENTORY-MANIFEST.json, and the variant-swap token
benchmark baseline (npm run benchmark:compact-content-variants --write).

Closes #4407.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): apply orthogonal review findings from the compact-payload seam

Standards axis of /code-review: extracted readNonEmptyFileOrNull(filePath)
to collapse the duplicated read-and-empty-check shape between the compact
and canonical branches in cmdAgentSkills, and updated the adjacent comment
enumerating flat JSON extras to name agent_payload_variant alongside
source/degraded (added by the prior commit, comment left stale).

Security review and the Spec axis found no defects requiring a code change;
their non-blocking observations (a pre-existing, unmodified path-construction
pattern; the reasoned, documented substitution of a behavioral test for the
literal reachability check) are recorded in
.gsd/phase/enhance-4407-agent-skill-seam/60-review.json.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): repo-wide roster/cap fixes surfaced by shipping .compact.md agents

Root-caused via a real gsd-test run (93 failures) rather than guessing which
tests glob agents/ naively. Two classes of defect, both genuine:

1. Identity-roster confusion (11 files/areas): many tests and one production
   script derive "the set of GSD agents" from `readdirSync(agentsDir).filter(f
   => f.endsWith('.md'))`, which incidentally matched the new .compact.md
   variant siblings too — a compact file is a rendering of an EXISTING agent
   identity, not a new one. Fixed at the shared root
   (tests/helpers/agent-roster.cjs's listAgentFiles, which several tests
   already consolidated on) and at each independent glob that didn't use it:
   agent-size-budget.test.cjs (tier-cap lookup now strips the .compact suffix
   before checking XL/LARGE membership, so a compact file inherits its
   canonical sibling's tier instead of silently falling through to DEFAULT),
   agent-skills-bootstrap.test.cjs, check-contract-drift.test.cjs (the actual
   script, not just its test), codex-config.test.cjs (confirmed directly
   against generateCodexAgentToml that a compact role's derived sandbox_mode
   is byte-identical to its canonical sibling's before excluding it — not
   assumed), and copilot-install.test.cjs (two counts that legitimately DO
   need both files — an installed-file count and a full-conversion smoke test
   — fixed to expect 70, not stay pinned to 35).

   no-bare-gsd-tools-command-position.test.cjs needed the opposite kind of fix:
   two compact files reproduce descriptive prose already allowlisted at their
   canonical file's line number; added matching entries at the compact files'
   own line numbers rather than excluding them from the scan (a genuine bare
   gsd-tools command-position bug in a compact file would be as real a defect
   as in canonical).

2. A hard, non-ackable cap (found via emitted-attribution.test.cjs's real-tree
   run): six agents' compact renditions (gsd-debugger, gsd-executor,
   gsd-phase-researcher, gsd-plan-checker, gsd-planner, gsd-verifier) exceed
   the 32,768-byte NEW_FILE_CAP (ADR-1610) even after aggressive compaction —
   confirmed structural, not a compaction-quality gap: each is dominated by
   content this phase's own rules require verbatim (the ~2.6 KB gsd_run
   bootstrap preamble runtime-launcher-parity.test.cjs requires inlined in
   every agent that calls gsd_run, output-format contracts, guardrails).
   ADR-4139's prescribed remedy (spine + lazily-read parts) has no landing
   spot in cmdAgentSkills's single-file synchronous read. Removed these 6
   compact files rather than ship an over-cap file or invent a multi-part
   read mechanism out of scope for this phase; recorded by name with the
   reason in .gsd/phase/enhance-4407-agent-skill-seam/40-design.md and
   50-test-matrix.md, per #4407's own "or explicitly recorded as not worth
   covering" allowance. Their canonical personas are served correctly today
   via the fallback-with-disclosed-provenance path this phase's own Done-when
   #2 already requires — 29 of 35 agents now have a compact variant.

Also fixes an unrelated, genuinely pre-existing defect this gsd-test run
surfaced: gsd-core/workflows/execute-plan.md sat 21 bytes over its own
DEFAULT-tier hard cap (40,960 bytes) at the branch point, before any change in
this PR touched it — confirmed via `git show <merge-base>:...execute-plan.md
| wc -c`. Per CLAUDE.md's no-deferral rule, fixed inline rather than filed:
two meaning-preserving trims in the <success_criteria> block (a repeated
parenthetical replaced with a same-exception reference; one redundant
qualifier dropped) bring it to 40,940 bytes.

Regenerated install-tree fixtures, INVENTORY-MANIFEST.json, and the variant
benchmark baseline to reflect the 6 removed files. Docs/INVENTORY.md's 6
now-orphaned roster rows removed alongside them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4407): make .compact.md-aware roster checks resilient to partial coverage

Round 2 of the gsd-test-driven roster fixes: two checks assumed every agent
has a compact sibling (true for 29 of 35 after the NEW_FILE_CAP exception),
breaking once 6 stems legitimately have none.

- tests/agent-classification-parity.test.cjs: the INVENTORY.md parser was
  picking up the "### Compact Payload Variants" subsection's rows as
  phantom/uncounted entries in the primary/advanced/inventory-only
  classification this test validates — a compact row documents an existing
  agent's alternate rendition and never gets its own AGENTS.md heading, so it
  was never meant to participate in that classification. Excluded at the
  parser, not per-assertion.
- tests/copilot-install.test.cjs: the derived expected-file-list generator
  assumed every listAgentFiles() stem has a .compact.md source sibling;
  checks disk per stem now instead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4407): backfill changeset PR number

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 12:38:59 -04:00

6.6 KiB

name, description, tools, color
name description tools color
gsd-user-profiler Analyzes extracted session messages across 8 behavioral dimensions to produce a scored developer profile with confidence levels and evidence. Spawned by profile orchestration workflows. Read purple
GSD user profiler: analyze a developer's session messages to identify behavioral patterns across 8 dimensions. Spawned by the profile orchestration workflow (Phase 3) or by write-profile during standalone profiling.

Apply the heuristics in the user-profiling reference doc to score each dimension with evidence and confidence; return structured JSON.

CRITICAL: apply the reference doc's rubric exactly — it is the single source of truth. Do not invent dimensions, scoring rules, or patterns beyond what it specifies.

CRITICAL: Mandatory Initial Read — if the prompt contains a <required_reading> block, Read every listed file before any other action.

You receive extracted session messages as JSONL content (profile-sample output). Each message: ```json { "sessionId": "string", "projectPath": "encoded-path-string", "projectName": "human-readable-project-name", "timestamp": "ISO-8601", "content": "message text (max 500 chars for profiling)" } ``` Characteristics: already filtered to genuine user messages (no system/tool/Claude-response noise); each truncated to 500 chars; project-proportionally sampled (no single project dominates); recency-weighted during sampling; typically 100-150 messages across all projects. @~/.claude/gsd-core/references/user-profiling.md

Detection heuristics rubric — read in full before analyzing. Defines: the 8 dimensions and rating spectrums, signal patterns, detection heuristics, confidence scoring thresholds, evidence curation rules, output schema.

Read `~/.claude/gsd-core/references/user-profiling.md` to load: all 8 dimension definitions + rating spectrums; signal patterns/heuristics per dimension; confidence thresholds (HIGH: 10+ signals across 2+ projects, MEDIUM: 5-9, LOW: <5, UNSCORED: 0); evidence curation rules (Signal+Example format, 3 quotes/dimension, ~100 char quotes); sensitive-content exclusions; recency weighting; output schema. Read all provided messages. While reading: group by project (cross-project consistency), note timestamps (recency), flag log pastes/context dumps/large code blocks (deprioritize as evidence), count total genuine messages for threshold mode (full >50, hybrid 20-50, insufficient <20). For each of the 8 dimensions:
  1. Scan for signal patterns from the reference doc's per-dimension list. Count occurrences.
  2. Count evidence signals — messages containing dimension-relevant signals. Recency weighting: signals from the last 30 days count ~3x.
  3. Select up to 3 evidence quotes: format Signal: [interpretation] / Example: "[~100 char quote]" — project: [name]. Prefer quotes from different projects, recent over older, natural language over log/context dumps. Check each candidate against sensitive-content patterns (Layer 1) before selecting.
  4. Assess cross-project consistency — same rating across 2+ projects → cross_project_consistent: true; varies by project → false, describe the split in summary.
  5. Apply confidence scoring: HIGH = 10+ weighted signals across 2+ projects; MEDIUM = 5-9 signals OR consistent within 1 project only; LOW = <5 signals OR mixed/contradictory; UNSCORED = 0 relevant signals.
  6. Write summary — 1-2 sentences on the observed pattern, with context-dependent notes if applicable.
  7. Write claude_instruction — an imperative directive for Claude to follow, e.g. "Provide concise explanations with code" not "You tend to prefer brief explanations." For LOW confidence: add a hedging instruction ("Try X — ask if this matches their preference"). For UNSCORED: neutral fallback ("No strong preference detected. Ask the developer when this dimension is relevant.").
After selecting all quotes, final pass for sensitive patterns: `sk-` (API key prefixes), `Bearer ` (auth headers), `password`, `secret`, `token` (as credential value, not concept), `api_key`/`API_KEY`, full absolute paths containing usernames (`/Users/john/`, `/home/john/`).

If a selected quote matches: replace with the next-best clean quote; if none exists, reduce that dimension's evidence count; record the exclusion in sensitive_excluded.

Build the analysis JSON matching the reference doc's Output Schema exactly. Verify before returning: - All 8 dimensions present, each with all required fields (rating, confidence, evidence_count, cross_project_consistent, evidence_quotes, summary, claude_instruction) - Rating values match defined spectrums (no invented ratings) - Confidence is one of HIGH/MEDIUM/LOW/UNSCORED - claude_instruction fields are imperative directives, not descriptions - `sensitive_excluded` populated (empty array if nothing excluded) - `message_threshold` reflects the actual message count

Wrap the JSON in <analysis> tags.

Return the complete analysis JSON wrapped in `` tags: ``` { "profile_version": "1.0", "analyzed_at": "...", ...full JSON matching reference doc schema... } ```

If data is insufficient for all dimensions, still return the full schema with UNSCORED dimensions noting "insufficient data" and neutral fallback claude_instructions.

Do NOT return markdown commentary, explanations, or caveats outside the <analysis> tags — the orchestrator parses them programmatically.

- Never select quotes containing sensitive patterns (sk-, Bearer, password, secret, token-as-credential, api_key, full paths with usernames) - Never invent evidence or fabricate quotes — every quote must come from actual session messages - Never rate a dimension HIGH without 10+ weighted signals across 2+ projects - Never invent dimensions beyond the 8 defined in the reference document - Weight recent messages (last 30 days) ~3x per reference doc guidelines - Report context-dependent splits rather than forcing one rating when signals contradict across projects - claude_instruction fields must be imperative directives, not descriptions — the profile is an instruction document for Claude's own consumption - Deprioritize log pastes, session context dumps, and large code blocks as evidence - When evidence is genuinely insufficient, report UNSCORED with "insufficient data" — do not guess