Use when an execution-layer dev (frontend-dev / backend-dev / ai-agent-dev / ml-engineer / miniapp-dev) has finished feature code + Unit tests and is about to enter code-review. Provides the SIT scope, environment, AC-driven integration walk, and evidence sink (progress/<role>.md). SIT is now a dev-owned step, not a separate QA stage.
npx skills add https://github.com/pcliangx/AppGenesisForge --skill agf-running-sit-tests
Use this skill when:
SIT verifies that independently-developed components compose correctly — frontend ↔ backend ↔ DB ↔ external services. It is NOT:
pytest / vitest on the branch before SIT)You just wrote the Unit tests, so you have the clearest picture of the unit-vs-integration boundary. If a failure reproduces by running just the backend unit tests with mocks, it's a unit-level miss, not a SIT finding; fold it back into the unit suite rather than writing it up as a SIT defect.
maindocs/changes/<change>/tasks.md(AC↔scenario 映射,[ADR-012);旧 feature fallback docs/prd/[feature]-[date].md.env.local with SIT-mode flags configured (or .env.sit if a dedicated SIT config exists)*.msw.ts),非手写——mock 与 OpenAPI 契约同源(见 ADR-006 / coding.md 契约纪律)If any precondition fails: SendMessage product-lead, do not proceed.
Default SIT environment is local docker-compose, brought up via the root Makefile(本地开发一键 SSOT,依赖管理走 uv——见 ADR-000;不要手写 pip/alembic/uvicorn 命令绕开 uv.lock):
# from repo root
make dev # postgres + backend (uv) + frontend (pnpm) 一键起栈
make migrate # apply latest schema (uv run alembic)
For LLM-dependent features, set provider env vars per agf-wiring-multi-llm-sdk skill. Use a dedicated SIT API key with a hard daily spend cap so a runaway test doesn't drain the budget.
Walk every AC from docs/changes/<change>/tasks.md(旧 feature fallback docs/prd/[feature].md). For each AC at the integration layer:
> Verify, don't assume. Don't write "Passed" because the code looks right. Run the action, capture the actual response, compare. Per .claude/standards/coding.md "Verify before assert" — paste the actual command output into the progress entry.
For each AC verdict:
curl -i output or browser DevTools Network exportpsql -c "SELECT ..." before/after dumpsKeep evidence inline in the SIT 证据 section of progress/<role>.md (small payloads). Large/binary artifacts → store under progress/evidence/[feature]/ and reference by path.
SIT no longer produces a standalone report under docs/qa/. All evidence lives in progress/<role>.md under the SIT 证据 section of the task entry — pass = single AC-tagged line (✅ AC-N (integration): <一句话>), fail/blocked expands命令 + 输出 + 偏差.
Format authority: .claude/standards/ac-lifecycle.md → 完整条目格式 (5-section format 状态 / Skills / SIT 证据 / 质量门 / 下一步 + the SIT 证据 block rules). The progress file is archived into docs/qa/[feature]-process-log.md after UAT sign-off (product-lead), so SIT evidence survives without a separate report artifact.
完成 SIT 自跑后,先自检再报告——跑 bash .claude/scripts/agf-advisory.sh progress/<role>.md(advisory 机筛统一入口,ADR-026 D2)机筛 placeholder / 漏证据 / pass 含失败 token / 质量门矛盾(advisory,不阻断),把 flag 的修掉再 SendMessage,省一轮 code-review 打回。这是全链路唯一一次机筛:reviewer 的 SIT Audit 不重跑 advisory(ADR-011 决策 2 + ADR-026 D2)。然后:
progress/<role>.md 条目路径 + 时间戳progress/<role>.md 的 SIT 段(如实记录 fail / blocked),但 SendMessage 写明阻塞原因与影响范围,由 product-lead 决定本轮是否就修.claude/agents/code-reviewer.md "SIT Audit" section) and you'll be sent backToolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup
Use when implementing any feature or bugfix, before writing implementation code
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes
Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always
Expert guidance for systematic backtesting of trading strategies. Use when developing, testing, stress-testing, or validating quantitative trading strategies. Covers "beating ideas to death" methodology, parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results. Applicable when user asks about backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development.
Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding assays, expression testing, thermostability measurements, enzyme activity assays, or protein sequence optimization. Also use for submitting experiments via API, tracking experiment status, downloading results, optimizing protein sequences for better expression using computational tools (NetSolP, SoluProt, SolubleMPNN, ESM), or managing protein design workflows with wet-lab validation.
This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
Take pcliangx/agf-running-sit-tests from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.