| Seed reproducible local Langfuse data in ClickHouse and Postgres. Use for complex traces, long sessions, v3/v4 events, bulk list data, or frontend rendering and performance tests; never use ad hoc scripts or raw inserts.
npx skills add https://github.com/langfuse/langfuse --skill seed-test-data
One-shot deterministic test data for local Langfuse. The CLI handles env
loading, preflight checks, ClickHouse/Postgres writes, readback verification,
and prints UI deep links plus a machine-readable JSON summary (last stdout
line).
pnpm run seed -- doctor
Prints PASS/WARN/FAIL per dependency (Postgres, migrations, project,
ClickHouse, v4 dev tables, Redis, MinIO, web app) with the exact fix command
for every failure. Do not debug Docker/ClickHouse manually before running
this.
| I need... | Command |
| --------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| A very complex observation tree (v3) | pnpm run seed -- trace-tree --observations 5000 --depth 12 --breadth 500 |
| The same tree readable in the v4 events UI | add --v4 (writes events_full; events_core fills via MV) |
| Async parents whose subtree outlives their own span (subtree wall-clock duration badge) | add --async-parents to trace-tree (root + hub end immediately while children keep running) |
| A realistic agent flow over a timeline (graph view + scrubbable timeline) | pnpm run seed -- agent-timeline --turns 6 --v4 (LangGraph refine loop planner→retriever→generator→critic→loop, staggered in time; add --timing-only for the pure timing fallback) |
| A demo-grade, real-looking agent trace (videos, screenshots, docs) | pnpm run seed -- support-agent --v4 --id-prefix <hex> (one fixed, fully handcrafted support-copilot refund run: guardrails, parallel context fan-out, 3-turn ReAct loop with real payloads/costs; deterministic — reseed with a FRESH prefix for a clean take; the prefix is the trace id, so a hex prefix reads like production) |
| A plain trace with no agentic types (collapsed-by-default graph panel) | add --plain to trace-tree (SPAN/GENERATION/EVENT only) |
| An extremely DEEP single-chain trace (tree depth = observation count; layout stress) | pnpm run seed -- deep-chain --v4 (1401 sequential generations, each the sole child of the previous — the mis-parented-instrumentation shape from LFE-10959 that collapses tree/timeline layouts; --observations N to change depth) |
| A super tough session (v3 legacy session view) | pnpm run seed -- long-session --traces 300 --observations-per-trace 8 |
| Diverse v4 session shapes (chat / coding-agent / mixed) for the session-detail view | pnpm run seed -- session-shapes --shape all (the agent shape has I/O on AGENT/TOOL with no GENERATION — pre-LFE-10520 the "first generation" default rendered empty cards for it; the current "All observations with I/O" default renders it correctly; v4 on by default) |
| Many traces for list/filter performance | pnpm run seed -- many-traces --count 100000 --days 14 |
| Long-window v4 traffic with cost/latency/token OUTLIERS (outlier chart strip, LFE-14451) | pnpm run seed -- outlier-traffic --days 90 (diurnal base load + deterministic spikes + hour-long ×8-latency incidents; root AGENT + GENERATION carrying cost + TOOL per trace; v4 on by default) |
| Scores with spaces in the name (filter/grammar testing) | pnpm run seed -- scored-traces --traces 24 --v4 |
| Lots of scores on every node (dense score badges, tree-row overflow testing) | add --scores-per-node 12 to trace-tree (N distinct scores per observation; try --depth 2 --breadth 44 for many tall sibling rows) |
| Varied human-annotation queues (annotate UI / keyboard testing) | pnpm run seed -- annotation-queue --core-items 12 (creates a "core types" queue covering every score-field render path + an "edge cases" queue with archived/stale/partial scores and observation/session/deleted/completed items) |
| Huge/malformed/unicode payloads | pnpm run seed -- trace-tree --payload-bytes 1000000 --payload-style malformed (styles: json, text, malformed, unicode, bignum, base64) |
| Big integers beyond 2^53-1 (number-precision testing) | pnpm run seed -- trace-tree --observations 1 --payload-style bignum |
| Huge base64 data-URI in ChatML IO (multimodal crash shape, LFE-10152) | pnpm run seed -- trace-tree --observations 30 --payload-bytes 20000000 --payload-style base64 --v4 (one unbroken multi-MB base64 token in trace + root-observation IO; max 50 MB) |
| See all scenarios and flags | pnpm run seed -- list --json |
| Predict without writing | add --dry-run |
traceIds, sessionIds, counts,verified (ClickHouse readback), links (UI deep links). Use --json to
suppress progress logs. Non-zero exit = data did not land; the error
includes a fix: line.
--seed (default 42) and flags → same ids (ids nevercontain dates), with timestamps anchored to the current UTC day. Re-running
within the same day overwrites in place; a later-day re-run updates the
same ids with re-anchored timestamps (the previous day's rows persist
under their old dates until then). Independent copies come only from
--id-prefix.
7a88fb47-b4e2-43b8-a06c-a5ce950dc53a(login [email protected] / password); override with --project.
links in the browser to verify visually. The v4events-backed UI is the per-user "Fast (Preview)" sidebar toggle, or
LANGFUSE_MIGRATION_V4_WRITE_MODE=events_only server-side.
Add a scenario in packages/shared/scripts/seeder/scenarios/: a plain
function using the deterministic Rng (never Math.random), register it in
scenarios/index.ts, and update the table in
packages/shared/scripts/seeder/AGENTS.md and this skill. Scenario names,
flags, and JSON keys are additive-only contracts. Design rationale:
packages/shared/scripts/seeder/README.md.
Efficient database search tool for bioRxiv preprint server. Use this skill when searching for life sciences preprints by keywords, authors, date ranges, or categories, retrieving paper metadata, downloading PDFs, or conducting literature reviews.
Access BRENDA enzyme database via SOAP API. Retrieve kinetic parameters (Km, kcat), reaction equations, organism data, and substrate-specific enzyme information for biochemical research and metabolic pathway analysis.
Access ClinPGx pharmacogenomics data (successor to PharmGKB). Query gene-drug interactions, CPIC guidelines, allele functions, for precision medicine and genotype-guided dosing decisions.
Query NCBI ClinVar for variant clinical significance. Search by gene/position, interpret pathogenicity classifications, access via E-utilities API or FTP, annotate VCFs, for genomic medicine.
Access COSMIC cancer mutation database. Query somatic mutations, Cancer Gene Census, mutational signatures, gene fusions, for cancer research and precision oncology. Requires authentication.
Query Ensembl genome database REST API for 250+ species. Gene lookups, sequence retrieval, variant analysis, comparative genomics, orthologs, VEP predictions, for genomic research.
Query openFDA API for drugs, devices, adverse events, recalls, regulatory submissions (510k, PMA), substance identification (UNII), for FDA regulatory data analysis and safety research.
Query NCBI Gene via E-utilities/Datasets API. Search by symbol/ID, retrieve gene info (RefSeqs, GO, locations, phenotypes), batch lookups, for gene annotation and functional analysis.
Take langfuse/seed-test-data from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.