mcpbeat Sign in

Dd Unblock Pr Agent Skill

Load when investigating a failing PR CI pipeline or checking PR health. Attributes each CI failure as flaky, infra, or regression, proposes a targeted action, and reports code coverage.

2k tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
967
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/DataDog/pup --skill dd-unblock-pr

The instruction itself

10 sections, as written by the author

Unblock PR

One-line summary: Investigate a failing PR CI pipeline — attribute each failure as flaky, infra, or regression and propose a targeted action.

Requires: dd-pup skill (pup CLI installed and authenticated), dd-triage-flaky-test skill (for flaky failure deep investigation).


Input

| Parameter | Description |

|---|---|

| PR branch | The branch under investigation (e.g. my-feature-branch) |

| Repository | Lowercase, no-schema URL (e.g. github.com/org/repo). Derive from git remote get-url origin if not provided. |


Workflow

STEP 0 — Parse Input

Derive repository ID and default branch from git if not provided:

# Repository ID: fully lowercase, no-schema URL (the API rejects mixed-case)
git remote get-url origin
# Strip protocol and trailing .git, then lowercase the result
# e.g. https://github.com/DataDog/my-repo.git → github.com/datadog/my-repo

# Default branch
git symbolic-ref refs/remotes/origin/HEAD
# Strip refs/remotes/origin/ prefix — fall back to main if unset

STEP 1 — Get PR CI Summary (run both in parallel)

Pipeline failures (job level):

pup cicd events search \
  --query "@ci.status:error @git.branch:<branch> @git.repository.id_v2:\"<repo>\"" \
  --level job \
  --from 24h \
  --limit 50

Test failures:

pup cicd tests search \
  --query "@test.status:fail @git.branch:<branch> @git.repository.id_v2:\"<repo>\"" \
  --from 24h \
  --limit 50

Run both queries in parallel. Collect all distinct @test.service values from test event results. If more than one distinct service is found, note each separately in the triage brief — do not collapse them into a single service filter. If pipeline results contain only infrastructure job types (build, lint, deploy) with no test-runner output, discard test search results and skip to STEP 3.

STEP 1.5 — Fetch Code Coverage (run in parallel with STEP 1)

This step runs unconditionally — coverage context is valuable whether CI is red or green.

The --repo value must be fully lowercase (the API rejects mixed-case). Normalize before calling:

repo_lower=$(echo "<repo>" | tr '[:upper:]' '[:lower:]')
pup code-coverage branch-summary \
  --repo "$repo_lower" \
  --branch "<branch>"

If the command returns no data or exits with an error, report "No data available" for Coverage in the PR Health section.

Note: Code quality and security violation counts are not available in pup — those lines always show "No data available".

STEP 2 — Blame Guard per Failing Job

First check whether @error_classification.domain / @error_classification.type are present on job events from STEP 1 — if populated, use them as primary classification signals.

For each failing job where classification is still needed, run both checks in parallel:

Default branch check — was this job already failing before this PR?

pup cicd events aggregate \
  --query "@ci.status:error @ci.job.name:\"<job>\" @git.branch:<default-branch> @git.repository.id_v2:\"<repo>\"" \
  --compute count \
  --from 24h

Blast radius check — is this job failing on other branches too?

pup cicd events aggregate \
  --query "@ci.status:error @ci.job.name:\"<job>\" @git.repository.id_v2:\"<repo>\"" \
  --compute count \
  --group-by "@git.branch" \
  --from 24h

Performance fallback: if the blast radius query is slow or times out, skip it and rely on the default branch check alone.

STEP 3 — Classify Each Failure

Priority order:

  • If @error_classification.domain / @error_classification.type present → use as primary signal
  • If test failure AND test appears in flaky tests with flaky_test_state:active:
   pup cicd flaky-tests search \
     --query "flaky_test_state:active @test.name:\"<test-name>\" @git.repository.id_v2:\"<repo>\""

flaky

  • Use blame guard results:

| Failing on default branch? | Failing on ≥3 other branches? | Classification |

|---|---|---|

| Yes | Yes | infra (pre-existing, widespread) |

| Yes | No | infra (pre-existing on default branch) |

| No | No | regression (introduced by this PR) |

| No | Yes | flaky (intermittent, cross-branch) |

| Insufficient data | — | unknown |

STEP 4 — Produce Triage Brief

One entry per failing job:

PR CI Triage Brief
==================
Branch:   <branch>
Repo:     <repo>

Job: <job-name>
  Classification:  <flaky | infra | regression | unknown>
  Evidence:        <1 key data point — error message, pipeline count, or test result>
  Confidence:      <high | medium | low>
  Recommended:     <action>

[repeat for each failing job]

Overall: <N> failures — <e.g. "1 regression, 1 flaky, 1 infra">

PR Health
=========
Coverage:   <X>% on <branch> | No data available
Quality:    No data available
Security:   No data available

All three lines always appear.

STEP 5 — Propose Actions

regression → Prompt user to investigate their code changes. No write action available.

flaky → Load dd-triage-flaky-test skill for deep investigation. That skill will:

  • Attempt an agent-native fix using flaky_category + stack trace
  • Propose quarantine via pup test-optimization flaky-tests update if a quick fix isn't possible

infra → Before proposing a retry, assess whether the failure is transient:

  • Check @error_classification.type and error message for signals like timeout, runner unavailable, network error, quota exceeded — these indicate transient failures where a retry is likely to help
  • If the error is deterministic (build misconfiguration, missing secret, explicit test assertion failure), a retry is unlikely to help — note this and suggest investigating the root cause
  • If the failure is pre-existing on the default branch, inform the user — a retry will likely fail again; await the upstream fix instead

If transient and GitHub Actions: extract the run ID from @ci.pipeline.url (e.g. https://github.com/org/repo/actions/runs/<run_id>):

gh run rerun <run_id> --failed

For other providers, share @ci.pipeline.url and direct to the provider UI for retry.

unknown → Suggest checking raw job logs via the CI provider UI or @ci.pipeline.url from the pipeline event.

Other skills for the same job

different authors, same section of the catalogue
Webapp Testing
by anthropics
vendor ×12

Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.

6k tokens scripts
Finishing A Development Branch
by ZhanlinCui
×7

Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup

1k tokens
Test Driven Development
by w95
×7

Use when implementing any feature or bugfix, before writing implementation code

2k tokens
Systematic Debugging
by ratacat
×7

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes

10k tokens scripts
Verification Before Completion
by ZhanlinCui
×6

Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims; evidence before assertions always

1k tokens
Backtest Expert
by BaggaT236
×3

Expert guidance for systematic backtesting of trading strategies. Use when developing, testing, stress-testing, or validating quantitative trading strategies. Covers "beating ideas to death" methodology, parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results. Applicable when user asks about backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development.

15k tokens scripts
Adaptyv
by christophacham
×3

Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding assays, expression testing, thermostability measurements, enzyme activity assays, or protein sequence optimization. Also use for submitting experiments via API, tracking experiment status, downloading results, optimizing protein sequences for better expression using computational tools (NetSolP, SoluProt, SolubleMPNN, ESM), or managing protein design workflows with wet-lab validation.

16k tokens
Aeon
by christophacham
×3

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.

19k tokens

How to use it

Copy the folder

Take datadog/dd-unblock-pr from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.