Map a MetricFlow semantic_model or metric into ktx semantic layer sources. Covers the MetricFlow to ktx primitive table, `extends:` inheritance flattening, metric-type handling (simple / derived / ratio / cumulative / conversion), `model: ref('x')` resolution, and four worked examples. Load when the turn contains `.yml`/`.yaml` files with top-level `semantic_models:` or `metrics:`.
npx skills add https://github.com/Kaelio/ktx --skill metricflow_ingest
A MetricFlow semantic_model maps to an SL source; MetricFlow measures map to ktx measures; MetricFlow entities map to ktx joins; MetricFlow metrics (top-level) map to ktx measures OR to cross-model derived measures. Files in one WorkUnit are ALWAYS part of the same logical entity (a connected component, possibly spanning extends: + cross-model metric refs). Flatten inheritance and cross-file references at write time.
| MetricFlow | ktx form | Notes |
|---|---|---|
| semantic_model: X { model: ref('t') } with measures + dimensions | Overlay named X with measures, computed-only columns, column_overrides, joins | The model: ref resolves to a manifest table. |
| semantic_model: X { model: source('s','t') } | Overlay named X over table t. | Same shape; source() still resolves to a physical table. |
| semantic_model: X { model: <literal> } with no manifest entry | Standalone with explicit sql:, grain:, columns: | Happens when the dbt manifest isn't available. |
| semantic_model: Y { extends: X } | Merge Y's measures/dimensions/entities into X's overlay, or write a single overlay named for the most-derived child (Y) containing both X's and Y's primitives | Do not emit a second overlay for X - flatten. |
| measures: [{ name, agg, expr }] | measures: [{ name, expr: "<agg>(<expr>)" }] | Aggregation inlined. agg: count_distinct → count(distinct ...). |
| entities: [{ name, type: primary }] | grain: [<entity_name-or-expr>] on the overlay/standalone | Primary/unique entities drive grain. |
| entities: [{ name, type: foreign }] | joins: entry joining to the primary-entity's semantic_model | Only when a matching primary is discoverable. |
| metrics: [{ type: simple, type_params: { measure: X } }] | If the base measure is labeled/described by the metric: in-place edit to the existing measure. Otherwise leave as-is. | Same-name metrics can absorb metadata. |
| metrics: [{ type: simple, filter: <jinja> }] | New measure on the same source, with the filter translated to SQL and attached via filter: | Translate Jinja {{ Dimension('x__y') }} to the column name y. |
| metrics: [{ type: derived, type_params: { expr, metrics } }] | Derived measure on whichever source owns the referenced measures, with expr: referencing measure names | If the metric spans models, still write it once on the source owning the "primary" measure (the one the agent judges most central). Mention the cross-model chain in the description. |
| metrics: [{ type: ratio, type_params: { numerator, denominator } }] | Same as derived; expr: "numerator / NULLIF(denominator, 0)" if no explicit expr | Safe-division by default. |
| metrics: [{ type: cumulative, type_params: { window, grain_to_date } }] | Standalone source with a window-function SQL; reference the resulting column as a normal measure | ktx SL has no first-class cumulative primitive (spec Non-goals). |
| metrics: [{ type: conversion }] | Flag for human - do NOT write. Emit a wiki note describing the intended semantics. | No ktx equivalent in v1. |
| Metric not mappable | Wiki page <metric_name>-definition.md with the full YAML body quoted | Capture the intent even if we can't emit SL. |
Type map: MetricFlow time to ktx time; categorical to string; number to number; boolean to boolean. Follow expr over name when both differ - expr is the physical column.
Verify each MetricFlow model source table with entity_details before producing
the corresponding sl_write_source.
Before writing a wiki page or SL source on any topic:
discover_data({query: "<topic>"}) - see what wikis, SL sources, and rawtables already exist. Prefer updating existing pages over creating new ones.
Before emitting any schema.table or schema.table.column into a wiki body,
SL source, tables: frontmatter, sl_refs, or emit_unmapped_fallback:
entity_details({connectionId, targets: [{display: "<identifier>"}]}) -confirm the identifier resolves; inspect native types, FK/PK, and
sampleValues.
check whether they appear in entity_details sampleValues for the relevant
column. If sampleValues is short or the sample may have missed real values,
run a sql_execution probe with the same warehouse connection id:
sql_execution({connectionId, sql: "SELECT DISTINCT <col> FROM <ref> LIMIT 50"}).
sql_execution({connectionId, sql: "SELECT 1 FROM <ref> LIMIT 0"}).If it errors, the identifier is fictional.
[unverified - from <rawPath>] in the wiki body,citing the exact raw path that mentioned it.
emit_unmapped_fallback with no_physical_table, includethe failing probe error in clarification.
<schema>.<table> placeholder strings from these instructionsinto output.
extends:Within one WorkUnit, multiple semantic_models linked by extends: are guaranteed to be present (the chunker groups them). Resolve inheritance before writing:
extends: chain upward, accumulating measures, dimensions, entities.The spec's worked example has orders, orders_ext (extends orders), and metrics/orders_final.yml (defines revenue referencing both). The right output is ONE overlay named orders_ext (or orders if the team's naming favors the base) containing order_count, gross_amount, refund_amount, and a derived revenue measure. Provenance tags point to all three source files.
model: ref resolutionThe model: field on a semantic_model is a string like ref('table_name'), source('src','table_name'), or a literal. Resolve:
ref('x') → table name x. Verify via sl_discover(x).source('s','t') → table name t. Verify via sl_discover(t).ref(...) / source(...)) → treat as the table name directly.If sl_discover errors because no such table exists, use discover_data and
entity_details to find the warehouse target. If a SQL probe is still needed,
call sql_execution with the same warehouse connection id, for example:
sql_execution({connectionId: "warehouse", sql: "SELECT 1 FROM analytics.orders LIMIT 0"}).
Never invent column names - every column in computed columns:, column_overrides:, grain:, and
sql: must be sourced from raw files, entity_details, or a successful SQL
probe.
After every sl_write_source, call sl_validate. The warehouse will reject invented columns with Unrecognized name: <name> - treat as a hard failure and re-read the schema.
ktx SL has no first-class window: or grain_to_date: primitive in v1 (spec Non-goals). Translate a MetricFlow cumulative metric to a standalone SL source with a window-function SQL:
# MetricFlow input:
metrics:
- name: cum_revenue_7d
type: cumulative
type_params:
measure: gross_amount
window: 7 days
# ktx standalone output:
name: cum_revenue_7d
source_type: sql
sql: |
SELECT
ordered_at,
SUM(amount) OVER (ORDER BY ordered_at RANGE BETWEEN INTERVAL '7' DAY PRECEDING AND CURRENT ROW) AS cum_revenue_7d,
order_id
FROM analytics.orders
grain: [order_id]
columns:
- {name: ordered_at, type: time, role: time}
- {name: cum_revenue_7d, type: number}
- {name: order_id, type: string}
measures:
- {name: cum_revenue_7d, expr: "max(cum_revenue_7d)"}
Pick the time column based on the semantic_model's defaults.agg_time_dimension (e.g. ordered_at). If the MetricFlow config omits it, probe the base table for time-typed columns and choose the most obvious. After writing the standalone SQL source, call emit_unmapped_fallback with rawPath set to the MetricFlow file path, reason: "cumulative_metric_unsupported", and fallback: "sql_standalone".
metrics:
- name: signup_to_first_order
type: conversion
type_params:
conversion_type_params:
entity: customer
base_measure: signup_count
conversion_measure: first_order_count
window: 30 days
Do NOT emit SL for this. Instead:
wiki/global/<metric_name>-intent.md quoting the full YAML body and a one-line explanation of the intended semantics (base event → conversion event within window).emit_unmapped_fallback with rawPath set to the MetricFlow file path, reason: "conversion_metric_unsupported", and fallback: "flagged".When ktx SL gains conversion primitives, re-ingesting will find the prior wiki note (via priorProvenance) and replace it with an SL source.
Every overlay/standalone/wiki page emitted from a MetricFlow source carries HTML-comment provenance tags. When one overlay derives from multiple files (e.g. an extends chain), emit one tag per contributing file:
# <!-- from: raw-sources/conn-1/metricflow/<syncId>/models/orders.yml#L1-20 -->
# <!-- from: raw-sources/conn-1/metricflow/<syncId>/models/orders_ext.yml#L1-12 -->
# <!-- from: raw-sources/conn-1/metricflow/<syncId>/metrics/orders_final.yml#L1-10 -->
name: orders_ext
...
Line ranges (#L<start>-<end>) point to the exact YAML span within the file (the semantic_models: entry for its own name). Use read_raw_span to identify those ranges before writing.
# MetricFlow:
semantic_models:
- name: orders
model: ref('orders')
entities:
- {name: order_id, type: primary}
measures:
- {name: order_count, agg: count, expr: order_id}
- {name: gross_amount, agg: sum, expr: amount}
# ktx overlay at <connId>/orders.yaml:
# <!-- from: raw-sources/.../models/orders.yml#L1-10 -->
name: orders
descriptions:
user: Order fact table.
measures:
- {name: order_count, expr: "count(order_id)"}
- {name: gross_amount, expr: "sum(amount)"}
grain: [order_id]
# MetricFlow:
# models/orders.yml
semantic_models:
- name: orders
model: ref('orders')
measures:
- {name: order_count, agg: count, expr: order_id}
- {name: gross_amount, agg: sum, expr: amount}
# models/orders_ext.yml
semantic_models:
- name: orders_ext
model: ref('orders_ext')
extends: orders
measures:
- {name: refund_amount, agg: sum, expr: refund_amt}
# metrics/orders_final.yml
metrics:
- name: revenue
type: derived
type_params:
expr: gross_amount - refund_amount
metrics:
- {name: gross_amount}
- {name: refund_amount}
# ktx overlay at <connId>/orders_ext.yaml (one file; inheritance flattened):
# <!-- from: raw-sources/.../models/orders.yml#L1-10 -->
# <!-- from: raw-sources/.../models/orders_ext.yml#L1-8 -->
# <!-- from: raw-sources/.../metrics/orders_final.yml#L1-10 -->
name: orders_ext
descriptions:
user: Extended order fact including refund handling; `revenue` = gross - refund.
measures:
- {name: order_count, expr: "count(order_id)"}
- {name: gross_amount, expr: "sum(amount)"}
- {name: refund_amount, expr: "sum(refund_amt)"}
- {name: revenue, expr: "gross_amount - refund_amount"}
grain: [order_id]
# models/sales.yml
semantic_models:
- name: sales
model: ref('sales')
measures:
- {name: revenue, agg: sum, expr: revenue_cents}
# models/costs.yml
semantic_models:
- name: costs
model: ref('costs')
measures:
- {name: cost, agg: sum, expr: cost_cents}
# metrics/margin.yml
metrics:
- name: margin
type: derived
type_params:
expr: revenue - cost
metrics: [{name: revenue}, {name: cost}]
Because the WorkUnit bundles all three files (cross-component union via the metric), write the derived measure on ONE of the two sources - pick the source whose domain "owns" the metric (here, sales - margin is inherently a sales metric). Cross-source references aren't native in ktx SL; treat the metric's operands as already-resolvable in the target source's query context OR emit a standalone SQL that joins the two tables:
# <connId>/sales.yaml
# <!-- from: .../models/sales.yml#L1-8 -->
# <!-- from: .../models/costs.yml#L1-8 -->
# <!-- from: .../metrics/margin.yml#L1-8 -->
name: sales
measures:
- {name: revenue, expr: "sum(revenue_cents)"}
# <connId>/margin.yaml - standalone because it spans two tables
# <!-- from: .../models/sales.yml#L1-8 -->
# <!-- from: .../models/costs.yml#L1-8 -->
# <!-- from: .../metrics/margin.yml#L1-8 -->
name: margin
source_type: sql
sql: |
SELECT s.period_id, s.revenue_cents, COALESCE(c.cost_cents, 0) AS cost_cents
FROM analytics.sales s
LEFT JOIN analytics.costs c ON c.period_id = s.period_id
grain: [period_id]
columns:
- {name: period_id, type: string}
- {name: revenue_cents, type: number}
- {name: cost_cents, type: number}
measures:
- {name: revenue, expr: "sum(revenue_cents)"}
- {name: cost, expr: "sum(cost_cents)"}
- {name: margin, expr: "sum(revenue_cents) - sum(cost_cents)"}
Also write a wiki page at wiki/global/margin-metric.md explaining the cross-source origin.
metrics:
- name: paid_order_count
type: simple
type_params:
measure: order_count
filter: "{{ Dimension('orders__status') }} = 'paid'"
# <connId>/orders.yaml
measures:
- {name: order_count, expr: "count(order_id)"}
- {name: paid_order_count, expr: "count(order_id)", filter: "status = 'paid'"}
Translate {{ Dimension('orders__status') }} to the bare column name status (the table alias prefix is implicit within the SL source's scope).
NGS analysis toolkit. BAM to bigWig conversion, QC (correlation, PCA, fingerprints), heatmaps/profiles (TSS, peaks), for ChIP-seq, RNA-seq, ATAC-seq visualization.
Automate customer engagement workflows including broadcast triggers, message analytics, segment management, and newsletter tracking through Customer.io via Composio
NGS analysis toolkit. BAM to bigWig conversion, QC (correlation, PCA, fingerprints), heatmaps/profiles (TSS, peaks), for ChIP-seq, RNA-seq, ATAC-seq visualization.
Query live pathogen genomic surveillance data through the GenSpectrum LAPIS API to find which viral lineages are circulating now, how fast they are growing, and what mutations they carry. Use whenever a question depends on the current state of a pathogen population rather than on remembered facts - which SARS-CoV-2 variant is dominant, whether a Pango lineage is still designated or has been withdrawn, what clade or genotype of H5N1 is in a host or region, whether a PCR primer or assay target still matches circulating sequence, or how a lineage's prevalence has moved week to week. Triggers include "variant surveillance", "genomic surveillance", "what variant is circulating", "dominant variant", "Pango lineage", "lineage prevalence", "growth advantage", "SARS-CoV-2 variant", "XFG", "clade 2.3.4.4b", "H5N1 genotype", "influenza clade", "RSV/mpox/measles/dengue lineage", "CoV-Spectrum", "LAPIS", "Nextclade", "pango-designation", and any request to report what a pathogen population looks like today.
NGS analysis toolkit. BAM to bigWig conversion, QC (correlation, PCA, fingerprints), heatmaps/profiles (TSS, peaks), for ChIP-seq, RNA-seq, ATAC-seq visualization.
Create a comprehensive product strategy using the 9-section Product Strategy Canvas — vision, segments, costs, value propositions, trade-offs, metrics, growth, capabilities, and defensibility. Use when building a product strategy, creating a strategic plan, or defining product direction.
Build a marketing performance report with key metrics, trend analysis, wins and misses, and prioritized optimization recommendations. Use when wrapping a campaign, when preparing weekly, monthly, or quarterly channel summaries for stakeholders, or when you need data translated into an executive summary with next-period priorities.
Calculate SaaS revenue, retention, and growth metrics. Use when diagnosing momentum, churn, expansion, or product-market-fit signals.
Take kaelio/metricflow_ingest from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.