Files
msd-core/gsd-core/references/ai-frameworks.md
Tom Boucher 463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00

11 KiB
Raw Blame History

AI Framework Decision Matrix

Reference used by gsd-framework-selector and gsd-ai-researcher. Distilled from official docs, benchmarks, and developer reports (2026).


Quick Picks

Situation Pick
Simplest path to a working agent (OpenAI) OpenAI Agents SDK
Simplest path to a working agent (model-agnostic) CrewAI
Production RAG / document Q&A LlamaIndex
Complex stateful workflows with branching LangGraph
Multi-agent teams with defined roles CrewAI
Code-aware autonomous agents (Anthropic) Claude Agent SDK
"I don't know my requirements yet" LangChain
Regulated / audit-trail required LangGraph
Enterprise Microsoft/.NET shops AutoGen/AG2
Google Cloud / Gemini-committed teams Google ADK
Pure NLP pipelines with explicit control Haystack

Framework Profiles

CrewAI

  • Type: Multi-agent orchestration
  • Language: Python only
  • Model support: Model-agnostic
  • Learning curve: Beginner (role/task/crew maps to real teams)
  • Best for: Content pipelines, research automation, business process workflows, rapid prototyping
  • Avoid if: Fine-grained state management, TypeScript, fault-tolerant checkpointing, complex conditional branching
  • Strengths: Fastest multi-agent prototyping, 5.76x faster than LangGraph on QA tasks, built-in memory (short/long/entity/contextual), Flows architecture, standalone (no LangChain dep)
  • Weaknesses: Limited checkpointing, coarse error handling, Python only
  • Eval concerns: Task decomposition accuracy, inter-agent handoff, goal completion rate, loop detection

LlamaIndex

  • Type: RAG and data ingestion
  • Language: Python + TypeScript
  • Model support: Model-agnostic
  • Learning curve: Intermediate
  • Best for: Legal research, internal knowledge assistants, enterprise document search, any system where retrieval quality is the #1 priority
  • Avoid if: Primary need is agent orchestration, multi-agent collaboration, or chatbot conversation flow
  • Strengths: Best-in-class document parsing (LlamaParse), 35% retrieval accuracy improvement, 20-30% faster queries, mixed retrieval strategies (vector + graph + reranker)
  • Weaknesses: Data framework first — agent orchestration is secondary
  • Eval concerns: Context faithfulness, hallucination, answer relevance, retrieval precision/recall

LangChain

  • Type: General-purpose LLM framework
  • Language: Python + TypeScript
  • Model support: Model-agnostic (widest ecosystem)
  • Learning curve: Intermediate–Advanced
  • Best for: Evolving requirements, many third-party integrations, teams wanting one framework for everything, RAG + agents + chains
  • Avoid if: Simple well-defined use case, RAG-primary (use LlamaIndex), complex stateful workflows (use LangGraph), performance at scale is critical
  • Strengths: Largest community and integration ecosystem, 25% faster development vs scratch, covers RAG/agents/chains/memory
  • Weaknesses: Abstraction overhead, p99 latency degrades under load, complexity creep risk
  • Eval concerns: End-to-end task completion, chain correctness, retrieval quality

LangGraph

  • Type: Stateful agent workflows (graph-based)
  • Language: Python + TypeScript (full parity)
  • Model support: Model-agnostic (inherits LangChain integrations)
  • Learning curve: Intermediate–Advanced (graph mental model)
  • Best for: Production-grade stateful workflows, regulated industries, audit trails, human-in-the-loop flows, fault-tolerant multi-step agents
  • Avoid if: Simple chatbot, purely linear workflow, rapid prototyping
  • Strengths: Best checkpointing (every node), time-travel debugging, native Postgres/Redis persistence, streaming support, chosen by 62% of developers for stateful agent work (2026)
  • Weaknesses: More upfront scaffolding, steeper curve, overkill for simple cases
  • Eval concerns: State transition correctness, goal completion rate, tool use accuracy, safety guardrails

OpenAI Agents SDK

  • Type: Native OpenAI agent framework
  • Language: Python + TypeScript
  • Model support: Optimized for OpenAI (supports 100+ via Chat Completions compatibility)
  • Learning curve: Beginner (4 primitives: Agents, Handoffs, Guardrails, Tracing)
  • Best for: OpenAI-committed teams, rapid agent prototyping, voice agents (gpt-realtime), teams wanting visual builder (AgentKit)
  • Avoid if: Model flexibility needed, complex multi-agent collaboration, persistent state management required, vendor lock-in concern
  • Strengths: Simplest mental model, built-in tracing and guardrails, Handoffs for agent delegation, Realtime Agents for voice
  • Weaknesses: OpenAI vendor lock-in, no built-in persistent state, younger ecosystem
  • Eval concerns: Instruction following, safety guardrails, escalation accuracy, tone consistency

Claude Agent SDK (Anthropic)

  • Type: Code-aware autonomous agent framework
  • Language: Python + TypeScript
  • Model support: Claude models only
  • Learning curve: Intermediate (18 hook events, MCP, tool decorators)
  • Best for: Developer tooling, code generation/review agents, autonomous coding assistants, MCP-heavy architectures, safety-critical applications
  • Avoid if: Model flexibility needed, stable/mature API required, use case unrelated to code/tool-use
  • Strengths: Deepest MCP integration, built-in filesystem/shell access, 18 lifecycle hooks, automatic context compaction, extended thinking, safety-first design
  • Weaknesses: Claude-only vendor lock-in, newer/evolving API, smaller community
  • Eval concerns: Tool use correctness, safety, code quality, instruction following

AutoGen / AG2 / Microsoft Agent Framework

  • Type: Multi-agent conversational framework
  • Language: Python (AG2), Python + .NET (Microsoft Agent Framework)
  • Model support: Model-agnostic
  • Learning curve: Intermediate–Advanced
  • Best for: Research applications, conversational problem-solving, code generation + execution loops, Microsoft/.NET shops
  • Avoid if: You want ecosystem stability, deterministic workflows, or "safest long-term bet" (fragmentation risk)
  • Strengths: Most sophisticated conversational agent patterns, code generation + execution loop, async event-driven (v0.4+), cross-language interop (Microsoft Agent Framework)
  • Weaknesses: Ecosystem fragmented (AutoGen maintenance mode, AG2 fork, Microsoft Agent Framework preview) — genuine long-term risk
  • Eval concerns: Conversation goal completion, consensus quality, code execution correctness

Google ADK (Agent Development Kit)

  • Type: Multi-agent orchestration framework
  • Language: Python + Java
  • Model support: Optimized for Gemini; supports other models via LiteLLM
  • Learning curve: Intermediate (agent/tool/session model, familiar if you know LangGraph)
  • Best for: Google Cloud / Vertex AI shops, multi-agent workflows needing built-in session management and memory, teams already committed to Gemini, agent pipelines that need Google Search / BigQuery tool integration
  • Avoid if: Model flexibility is required beyond Gemini, no Google Cloud dependency acceptable, TypeScript-only stack
  • Strengths: First-party Google support, built-in session/memory/artifact management, tight Vertex AI and Google Search integration, own eval framework (RAGAS-compatible), multi-agent by design (sequential, parallel, loop patterns), Java SDK for enterprise teams
  • Weaknesses: Gemini vendor lock-in in practice, younger community than LangChain/LlamaIndex, less third-party integration depth
  • Eval concerns: Multi-agent task decomposition, tool use correctness, session state consistency, goal completion rate

Haystack

  • Type: NLP pipeline framework
  • Language: Python
  • Model support: Model-agnostic
  • Learning curve: Intermediate
  • Best for: Explicit, auditable NLP pipelines, document processing with fine-grained control, enterprise search, regulated industries needing transparency
  • Avoid if: Rapid prototyping, multi-agent workflows, or you want a large community
  • Strengths: Explicit pipeline control, strong for structured data pipelines, good documentation
  • Weaknesses: Smaller community, less agent-oriented than alternatives
  • Eval concerns: Extraction accuracy, pipeline output validity, retrieval quality

Decision Dimensions

By System Type

System Type Primary Framework(s) Key Eval Concerns
RAG / Knowledge Q&A LlamaIndex, LangChain Context faithfulness, hallucination, retrieval precision/recall
Multi-agent orchestration CrewAI, LangGraph, Google ADK Task decomposition, handoff quality, goal completion
Conversational assistants OpenAI Agents SDK, Claude Agent SDK Tone, safety, instruction following, escalation
Structured data extraction LangChain, LlamaIndex Schema compliance, extraction accuracy
Autonomous task agents LangGraph, OpenAI Agents SDK Safety guardrails, tool correctness, cost adherence
Content generation Claude Agent SDK, OpenAI Agents SDK Brand voice, factual accuracy, tone
Code automation Claude Agent SDK Code correctness, safety, test pass rate

By Team Size and Stage

Context Recommendation
Solo dev, prototyping OpenAI Agents SDK or CrewAI (fastest to running)
Solo dev, RAG LlamaIndex (batteries included)
Team, production, stateful LangGraph (best fault tolerance)
Team, evolving requirements LangChain (broadest escape hatches)
Team, multi-agent CrewAI (simplest role abstraction)
Enterprise, .NET AutoGen AG2 / Microsoft Agent Framework

By Model Commitment

Preference Framework
OpenAI-only OpenAI Agents SDK
Anthropic/Claude-only Claude Agent SDK
Google/Gemini-committed Google ADK
Model-agnostic (full flexibility) LangChain, LlamaIndex, CrewAI, LangGraph, Haystack

Anti-Patterns

  1. Using LangChain for simple chatbots — Direct SDK call is less code, faster, and easier to debug
  2. Using CrewAI for complex stateful workflows — Checkpointing gaps will bite you in production
  3. Using OpenAI Agents SDK with non-OpenAI models — Loses the integration benefits you chose it for
  4. Using LlamaIndex as a multi-agent framework — It can do agents, but that's not its strength
  5. Defaulting to LangChain without evaluating alternatives — "Everyone uses it" ≠ right for your use case
  6. Starting a new project on AutoGen (not AG2) — AutoGen is in maintenance mode; use AG2 or wait for Microsoft Agent Framework GA
  7. Choosing LangGraph for simple linear flows — The graph overhead is not worth it; use LangChain chains instead
  8. Ignoring vendor lock-in — Provider-native SDKs (OpenAI, Claude) trade flexibility for integration depth; decide consciously

Combination Plays (Multi-Framework Stacks)

Production Pattern Stack
RAG with observability LlamaIndex + LangSmith or Langfuse
Stateful agent with RAG LangGraph + LlamaIndex
Multi-agent with tracing CrewAI + Langfuse
OpenAI agents with evals OpenAI Agents SDK + Promptfoo or Braintrust
Claude agents with MCP Claude Agent SDK + LangSmith or Arize Phoenix