datadog-labs/agent-observability-replay-trace
>- Use when a developer wants to iterate on ONE specific Agent Observability / LLM Obs trace whose output they didn't like — re-running that trace against their LOCAL code, seeing a concise diff of the old vs new output, and looping (change code → replay → diff) until satisfied. Invoked as a trace"; "this trace's output is wrong, fix it and re-run"; "re-run trace <id> with <change>"; pasting a trace id from the Agent Observability UI with a description of what to fix. It fetches the trace via the datadog-llmo MCP or the pup CLI, edits code, re-runs the app to emit a NEW trace, and diffs the two — no local server, no browser. For agents traced with ddtrace / LLM Obs (Python first-class), with JSON-serializable entry agent-observability-replay-experiment), building an experiment from a dataset/CSV, writing evaluators, root-causing failed traces, or RUM/HTTP session replay.
npx skills add https://github.com/datadog-labs/agent-skills --skill agent-observability-replay-trace
A fast iteration loop on a single production trace: take a trace whose output a developer didn't like,
optionally change the code, re-run it against their LOCAL code, and show a concise diff of old vs new
output — repeating until they're happy. Assumes nothing about the project's layout.
Invoked from the developer's coding agent: /agent-observability-replay-trace <trace-id> [<changes to test>].
With no modification, do the replay + diff only (a reproduce/regression check), then offer to enter the loop.
**This file is the workflow spine — terse on purpose. The depth lives in references/details.md (trace
backend + pup flags, the runner contract, export mode, polling, the trace-link scoping fix) and
references/local-setup.md (making a deployed-only app locally runnable). Read details.md before you touch
pup or generate the runner.**
Writing code — keep comments minimal to none. Everything you generate or edit (the annotation, the
runner's ENTRYPOINTS entries, a local harness, iteration edits) should match the surrounding code and
carry no unnecessary comments — don't narrate what the code plainly does; add a comment only for a
genuinely non-obvious *why*.
This is a live loop. At every decision point present the choices as an AskUserQuestion selector (the
plan-mode-style menu), not a plain question that ends your turn. Two gates: (a) after you propose code
changes, before replaying; (b) after each diff. The selector's free-text option lets the user type detail
(what to refine) inline — act on it directly, don't ask a follow-up. Keep re-presenting after every replay
until they pick "stop here".
ddtrace / LLM Obs (an ml_app + a discoverable entrypoint). Python is first-class;other languages work but you write the runner to the contract in their SDK/build tooling.
binary "is it runnable?"; deployed-only apps often still expose a plain callable).
datadog-llmo MCP (used when present) or the pup CLI (fallback, andthe easier install if you have neither) (step 0).
DD_API_KEY + DD_SITE + provider key(s). Not DD_APP_KEY — plain trace, not anExperiment (that's agent-observability-replay-experiment).
traces cannot be deleted** — a mis-scoped replay (wrong ml_app) *permanently* pollutes the production app's
dashboards/eval sets. That's why the <ml_app>-local isolation (steps 4/6/7) is load-bearing, not tidy.
Warn before the first replay.
Pick, in order: (1) the MCP if mcp__datadog-llmo-mcp__* tools are present — the default (slightly
richer for reads: structured tree + content_info); (2) else pup if installed and pup auth targets
the app's org; (3) else the user has neither → guide the pup install (it's easier to set up than the
MCP, so recommend pup here):
brew tap datadog-labs/pack && brew install datadog-labs/pack/pup
pup auth login
(MCP alternative: claude mcp add --scope user --transport http "datadog-llmo-mcp" "https://mcp.datadoghq.com/api/unstable/mcp-server/mcp?toolsets=llmobs"; see https://docs.datadoghq.com/bits_ai/mcp_server/setup/.) Don't proceed without a backend.
The backend↔operation mapping and **pup's exact flags/gotchas are in details.md — read that section before
using pup. Two pup musts: (1) results come back at data.spans[] *or* top-level spans[]**
(varies by version/--no-agent) — parse whichever is present, or you get zero hits on an ingested
trace (a silent false negative, step 7); (2) check token expiry (pup auth status), not just that auth
exists — expiry mid-loop looks like "trace not found."
<trace-id> + optional free-text modification (everything after the id); none → diff-only mode. Determine
the ml_app from the project (LLMObs.enable(ml_app=…) / DD_LLMOBS_ML_APP) or the trace; confirm if
ambiguous.
Fetch via the backend; note total_duration_ms (drives step 7), the trace_url, and
metadata.replay_input/replay_entrypoint if present. **Locate the baseline field — it's not always the
root output:** the value the developer dislikes may be a tool-call input or an intermediate output several
levels deep, and the app may post-process it before the span records it. Pick the field the code change can
actually move, or the delta drowns in noise.
If the root span fans out into repeated sibling subtrees (a batch/map over N parallel sub-runs), the
change under test is usually visible in a single branch — replaying the whole root costs ~N× spend and
time for no extra signal. Offer to replay one representative branch; log what you skipped. **Pick deliberately: the cheapest
branch that reached the terminal / side-effecting tool** (most branches are no-ops that prove nothing), and
reconstruct its input from the child span's input, not the root's. Full root only if the change is
inherently cross-branch.
metadata.replay_entrypoint if present; else infer from the root span (name/kind) + codeand confirm with the user.
metadata.replay_input if present; else derive a suggested input (prefer the codesignature — the rendered prompt is lossy) and have the user confirm/edit.
Ask "what is the innermost callable seam for this root span, and can I call it directly with JSON?" — not
"is the app runnable?". Already directly callable → skip, continue. Buried under a handler/service
(deployed-only, no local __main__, live-infra coupling) → follow references/local-setup.md (detect →
propose → approve → build). The common middle case — a deployed service whose core logic is *already* a plain
callable (ports-and-adapters) — just extract/call that seam; full local-setup is overkill.
self-describe. Stamp it at span start, not the success/deferred-finish path — a failed run must still
carry replay_input (those are the ones you most want to replay):
LLMObs.annotate(span=span, metadata={"replay_entrypoint": "<stable id>", "replay_input": <extractor>})
No replay_output — the original trace is the baseline. (Non-Python annotate APIs differ — e.g. Go
span.Annotate(llmobs.WithAnnotatedMetadata(...)); see details.md.)
per-call ml_app overrides** (Go llmobs.WithMLApp; Python ml_app= on a decorator or in
LLMObs.annotate). Those beat the init-level -local, so the app's spans can still land in
production — tracer-level config is not proof of isolation. If any exist, the app's ml_app must
resolve from env so -local wins.
details.md (load env →derive <ml_app>-local → dispatch one entrypoint on JSON → flush on every exit path incl. errors →
refuse to start unless ml_app ends in -local → print the -local ml_app). Python: copy
scripts/replay_runner_template.py and fill ENTRYPOINTS. Other languages: write to the contract —
don't assume the Python API carries over (Go APIs + export-mode gotchas in details.md), and where the
language has no in-process dotenv add a run wrapper (artifact c) that sources the project env, unsets
ambient provider vars, and exports the -local override. Infer + confirm the run command; follow the
host repo's build-file conventions (Bazel/Gazelle → cmd/<name>/, run Gazelle, build before replay).
Make the code changes, show the developer the diff of your changes, then an AskUserQuestion selector:
Replay now / Adjust the changes first / Cancel. Only replay on "Replay now".
Before the first replay: warn (re-running is real — model spend + real writes), and **sanitize the
environment. The coding agent's own env** (ANTHROPIC_API_KEY / ANTHROPIC_BASE_URL set by Claude
Code, and other provider keys) can make the app's SDK bypass its configured model gateway — a fidelity
gap invisible in the diff. Unset ambient provider vars by default and report that you did (don't
just ask); grep the app for its own ambient-key guards. Also **verify the credential's org matches the
trace's org** — a mismatch ships the replay somewhere you can't query (looks like ingest lag).
On confirmation, record t0 and run — source the project's env file, never inline secrets (the marker
tag is fine on the command; DD_API_KEY=<value> inline is blocked by the permission classifier and leaks to
history/transcript — use the wrapper/env-file):
DD_TAGS=replay_run_id:<unique-id> <run cmd or wrapper> --entrypoint <id> --input-file <path>
The runner emits under <ml_app>-local (idempotent, so replays never pollute production) and prints that
name — poll for the new trace under it.
max(120s, ~3 × total_duration_ms).replay_run_id tagunder <ml_app>-local (pup: --query "replay_run_id:<id>", plain key:value). **Before ever reporting
"not found," re-query with no tag filter** (just <ml_app>-local + window): if that returns spans, your
filter/parse/scope is wrong — not ingestion. A false "no trace" reads as normal and invites a wasteful
re-run.
--query/tag matching can return a spanwhose real ml_app is a *different* app (the --ml-app filter gets ignored). Read ml_app off every
returned span and assert it ends in -local before reporting a clean replay — otherwise you report
"clean replay under -local" while the trace is actually in production (which you can't undo). This false
*confidence* is worse than the false negative. Don't hard-fail on timeout; offer to keep waiting.
Concise summary of how the new output differs from the old — meaningful differences only. Note live-world
drift; and because any nondeterministic agent varies run-to-run, default to two replays (diff-only mode
too, not just model-facing edits) and use replay-to-replay comparison — if the two local runs differ
from each other about as much as from production, the delta is sampling variance, not your change. If the
replay disables a side-effecting integration (dry-run), that integration's subtree is absent — **exclude
it from both sides** before comparing span counts, or the structural diff is junk. Lead the diff with both
trace links:
trace_url verbatim — but under fan-out (you replayed one branch) link the **branchspan**, not the whole-root url.
ml_app=<ml_app>-local or it opens empty — and the trace_url is anorg-switch wrapper (…/switch_to_user/<id>?next=<encoded /llm/traces …>&flow=org_switch), so **inject
ml_app=<ml_app>-local into the decoded next query and re-encode; do NOT append to the outer URL**
(mechanics in details.md). Browser-unverifiable from here — confirm once it opens non-empty.
Harness-failure gate (before the diff): if a replay reveals the harness is wrong — trace landed under
the wrong ml_app, no trace after the step-7 sanity checks, missing flush, or auth/org misrouted — **do NOT
proceed to a diff on bad data.** Stop and present a selector to fix the harness (re-scope ml_app / add flush
/ fix env) and re-replay.
Otherwise, after the diff, an AskUserQuestion selector: Looks good — stop here (finish; leave the edits
in the working tree) / Make more changes (free-text inline → back to step 5). Re-present after every
replay; end only on "stop here".
references/details.md — trace backend + pup exact flags, the runner contract (+ Go, export mode),polling + the false-negative sanity check, the trace-link scoping fix, limitations. Read before pup / the runner.
references/local-setup.md — making a deployed-only app locally runnable (step 3.5). Read when that gap shows.scripts/replay_runner_template.py — the Python runner to copy + fill.Take datadog-labs/agent-observability-replay-trace from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference brew.
Without those the skill loads but fails at the first command.