mcpbeat Sign in

Debug Issue With Datadog Agent Skill

| Establish root cause by combining Datadog telemetry with the Langfuse repo. Use when investigating or triaging a user report, Linear or GitHub issue, incident, or pasted production error.

7k tokens
context cost
the whole folder, loaded on every use
5
files
instructions only
0
copies elsewhere
how many repositories repackaged it
32482
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/langfuse/langfuse --skill debug-issue-with-datadog

What comes with it

21 303 bytes besides the instruction
references/datadog-playbook.md
references/intake.md
references/output-template.md
references/repo-debug-map.md

The instruction itself

6 sections, as written by the author

Debug Issue with Datadog

Use this skill whenever the task is investigative rather than

implementational: a user, customer, or oncall has surfaced a problem and you

need to figure out *what is actually happening in production* and *where in the

code it lives*. The deliverable is an analysis, not a patch — though the

analysis should make the right patch obvious.

When to Apply

  • A Linear issue (typically with an LFE-XXXX ID) describes a production

failure, error spike, or customer report.

  • A GitHub issue or pasted incident/error report needs triage.
  • A monitor alerted and you need to understand *why* before deciding what to

fix.

  • Existing tickets under the "Make monitoring useful again" project (parent

LFE-8837) and similar — these expect the structured analysis output below.

If the task is "implement this fix" rather than "figure out what's broken",

this is the wrong skill — go to backend-dev-guidelines or the relevant

package guide.

Workflow

Read the inputs first, then plan the Datadog sweep, then read the code, then

write the analysis. Do not skip ahead to suggested patches before the data

supports them.

  • Intake. Pull every signal already available in the report. See

references/intake.md. For a Linear URL/ID, fetch

the issue *and* its comments via the Linear MCP — the description is often

updated inline as triage proceeds. For a GitHub issue, use gh issue view.

For pasted text, treat it as the description.

If intake contains an alert identity (a Datadog monitor ID or title, an

incident.io alert/INC reference, or an on-call page), first apply

incident-alert-tickets — a

documented cause section may resolve the investigation before any sweep.

  • Scope the sweep. From the intake, pick the affected subsystem and time

window. Use references/repo-debug-map.md

to translate "PostHog integration", "ingestion failures", "evals stuck",

etc. into the Datadog filters and source files you should be looking at.

  • Run the broad Datadog sweep. Default to the full sweep in

references/datadog-playbook.md: APM

spans, error logs, metrics, and monitors — split across prod-eu

and prod-us (and prod-hipaa / prod-jp when relevant). Always check

regional disparity first; it usually rules whole hypotheses in or out.

Use datadog-query-recipes for

reusable tenant, public API, queue consumer, and cross-environment query

shapes.

  • Cluster the errors. Group by (projectId, error.message) or

(error.type, error.message). Treat each distinct cluster as its own

hypothesis — Langfuse incidents commonly have *multiple* coexisting root

causes, not one.

  • Map clusters to code. For each cluster, open the relevant handler file

from the repo-debug map and read enough of it to confirm or refute the

hypothesis. Cite specific files and line ranges in the output.

  • Write the analysis using

references/output-template.md.

  • Deliver. Default: print the analysis in chat. If the user asked for it,

also save under the workflow they specified (file, Linear comment via, etc.).

If the investigation was anchored to an alert identity, also offer the

human-gated write-back from

incident-alert-tickets: append the

established root cause as a dated cause section, or create the monitor's

ticket.

Datadog MCP Usage Notes

Two Datadog MCP servers are typically available — one bound to the EU site

(datadoghq.eu) and one to the US site (datadoghq.com). Always run

region-relevant queries against both unless intake clearly localizes the

incident. The prod-eu / prod-us env tags live on each side respectively.

  • Span search filter pattern:

service:worker resource_name:"process posthog-integration-project" status:error

  • Log search filter pattern:

service:worker env:prod-eu @langfuse.project.id:cm1r6u… status:error

  • For high-volume queries, prefer aggregate_spans / aggregate_events

grouped by (error.message, projectId) over fetching individual traces.

  • Always link to the Datadog UI for the queries you ran (final section of the

output template).

See references/datadog-playbook.md for the

full set of starter queries and parameter shapes.

Output Expectations

From the output template:

  • Header: data source, time window, region split (EU vs US table).
  • Hotspots: per-projectId (or per-cluster) error counts.
  • Root cause by error class: each cluster gets a short hypothesis with

reasoning, distinguishing primary causes from symptoms.

  • Suggested patches: P0/P1/P2 grouped, with concrete file paths and short code

sketches. Reference the actual handler in worker/src/features/** or

web/src/**.

  • Dashboards: paste the Datadog query URLs at the end.

Findings come first, recommendations last. If the data is thin, say so

explicitly and propose what would need to be true to confirm each hypothesis —

do not invent root causes.

Cross-References

  • Per-monitor knowledge base — look up documented causes before the sweep,

record new ones after (human-gated):

incident-alert-tickets

  • Production telemetry query recipes, tenant/public API usage, and queue

consumer measurements:

datadog-query-recipes

  • Backend layout, queue contracts, instrumentation patterns:

backend-dev-guidelines

  • ClickHouse-related findings (memory ceilings, JOIN spills, slow queries):

clickhouse-best-practices

  • Once a fix is identified and you switch to implementation, hand off to the

package AGENTS.md for the affected directory.

Other skills for the same job

different authors, same section of the catalogue
MCP Builder
by anthropics
vendor ×13

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

30k tokens scripts
Changelog Generator
by frostant
×9

Automatically creates user-facing changelogs from git commits by analyzing commit history, categorizing changes, and transforming technical commits into clear, customer-friendly release notes. Turns hours of manual changelog writing into minutes of automated generation.

774 tokens
Finishing A Development Branch
by ZhanlinCui
×7

Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup

1k tokens
MCP Builder
by JayZeeDesign
×7

Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).

37k tokens scripts
Vercel React Native Skills
by vercel-labs
vendor ×6

React Native and Expo best practices for building performant mobile apps. Use when building React Native components, optimizing list performance, implementing animations, or working with native modules. Triggers on tasks involving React Native, Expo, mobile performance, or native platform APIs.

39k tokens
Vercel React Best Practices
by ratacat
×5

React and Next.js performance optimization guidelines from Vercel Engineering. This skill should be used when writing, reviewing, or refactoring React/Next.js code to ensure optimal performance patterns. Triggers on tasks involving React components, Next.js pages, data fetching, bundle optimization, or performance improvements.

34k tokens
Next Best Practices
by vercel-labs
vendor ×4

Next.js best practices - file conventions, RSC boundaries, data patterns, async APIs, metadata, error handling, route handlers, image/font optimization, bundling

20k tokens
Using Git Worktrees
by ZhanlinCui
×4

Use when starting feature work that needs isolation from current workspace or before executing implementation plans - creates isolated git worktrees with smart directory selection and safety verification

1k tokens

How to use it

Copy the folder

Take langfuse/debug-issue-with-datadog from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.