> Fetch weather/climate data via Earth2Studio data sources for specific variables and times. Do NOT use for inference pipelines, model discovery, or installation.
npx skills add https://github.com/NVIDIA/skills --skill earth2studio-data-fetch
Guide a user through downloading weather/climate data via Earth2Studio data source
APIs. Identifies compatible sources by checking the lexicon, verifies variable
support, and produces a working fetch script outputting an xarray DataArray.
uv pip install earth2studio or equivalent)~/.cdsapirc)You are helping a user download specific weather/climate data using
Earth2Studio's data source APIs. Your job is to identify which data source(s)
can provide the requested variables, verify compatibility via the lexicon
system, and produce a working fetch script.
Data source APIs, available variables, and the lexicon evolve between releases.
Before recommending a data source or writing a fetch script:
and constructor arguments.
that data source.
Live doc references (fetch only what the user's request requires):
<https://nvidia.github.io/earth2studio/modules/datasources_analysis.html>
<https://nvidia.github.io/earth2studio/modules/datasources_forecast.html>
<https://nvidia.github.io/earth2studio/modules/datasources_dataframe.html>
<https://github.com/NVIDIA/earth2studio/blob/main/earth2studio/lexicon/base.py>
<https://github.com/NVIDIA/earth2studio/tree/main/earth2studio/lexicon>
Extract from what the user has said (ask follow-ups if needed, cap at 3
questions):
(e.g. t2m, u500, z850, tp, msl). If the user uses plain language
("500 hPa geopotential height"), map it to the E2Studio name by checking
the live base.py E2STUDIO_VOCAB.
discrete times?
Based on the request type, narrow candidates:
Analysis/reanalysis (historical state at a specific time):
IFS/IFS_ENS (ECMWF), ARCO/CDS/WB2ERA5/NCAR_ERA5 (ERA5 reanalysis),
GOES/MRMS/JPSS (observational)
Forecast (predictions from an initialization time with lead times):
AIFS_FX, CFS_FX
Key differentiators to surface:
history; reanalysis (ERA5 via ARCO/CDS/WB2) goes back decades
WB2ERA5_32x64 is 5.625° global
This is critical. Each data source has a lexicon file that defines which
E2Studio variables it can provide.
To verify:
https://github.com/NVIDIA/earth2studio/blob/main/earth2studio/lexicon/<source>.py
(e.g. gfs.py, hrrr.py, cds.py, arco.py, wb2.py)
source's VOCAB dict
it — try another
The lexicon VOCAB maps Earth2Studio variable names → source-specific
identifiers. If a variable key exists in the VOCAB, the source supports it.
Present the results clearly: *"GFS supports t2m, u500, z850. HRRR also
supports these but is limited to North America. ARCO (ERA5) supports all
three and has data back to 1959."*
Present the viable options with tradeoffs:
| Source | Variables | Coverage | Resolution | Time Range |
|--------|-----------|----------|------------|------------|
| ... | ... | ... | ... | ... |
Let the user pick. If there's one obvious choice, recommend it and ask for
confirmation.
Write a Python script that uses the selected data source to fetch the
requested data. The script structure depends on whether it's an analysis or
forecast source.
Analysis source pattern:
import datetime
from earth2studio.data import <SourceClass>
# Initialize data source
ds = <SourceClass>()
# Fetch data
# Analysis sources use: ds(time, variable) -> xr.DataArray
time = [datetime.datetime(YYYY, M, D, H)] # or array of times
variable = ["var1", "var2"] # E2Studio variable names
data = ds(time, variable)
Forecast source pattern:
import datetime
from earth2studio.data import <SourceClass>
# Initialize data source
ds = <SourceClass>()
# Forecast sources use: ds(time, lead_time, variable) -> xr.DataArray
time = [datetime.datetime(YYYY, M, D, H)] # initialization time
lead_time = [datetime.timedelta(hours=H)] # or array of lead times
variable = ["var1", "var2"]
data = ds(time, lead_time, variable)
Always fetch the specific data source's API doc page to confirm the exact
constructor arguments and call signature before writing the script — they can
vary (some need auth tokens, cache paths, specific parameters).
Include in the script:
print(data), data.shape, data.coords)After delivering the script, mention:
discover skill
EARTH2STUDIO_CACHE)
Owns: identifying data sources for a user's variable/time request,
verifying variable support via lexicon, generating data fetch scripts,
explaining analysis vs. forecast source differences.
Does not own: installation (earth2studio-install), model selection
(earth2studio-discover), inference pipelines, custom data source creation
(point to extend examples), data source authentication setup beyond what
the docs describe.
Typical invocation:
> "I need 500 hPa geopotential height and 2m temperature from ERA5
> for January 1, 2020 at 00Z."
The skill would:
z500, t2m(GCS, S3, CDS API)
DataArrayFile/DataSetFile directly
sources in a single call
variables; always verify via lexicon
are generally faster
| Error | Cause | Solution |
|-------|-------|----------|
| KeyError: '<var>' | Not in lexicon | Check lexicon; try another source |
| FileNotFoundError / 404 | Time not available | Verify temporal coverage |
| CDS API timeout | Queue congestion | Retry or use ARCO for ERA5 |
| ModuleNotFoundError | Not installed | uv pip install earth2studio |
| Empty DataArray | Time/var mismatch | Check datetime and variable name |
Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.
Access NCBI GEO for gene expression/genomics data. Search/download microarray and RNA-seq datasets (GSE, GSM, GPL), retrieve SOFT/Matrix files, for transcriptomics and expression analysis.
Bayesian modeling with PyMC. Build hierarchical models, MCMC (NUTS), variational inference, LOO/WAIC comparison, posterior checks, for probabilistic programming and inference.
Multi-objective optimization framework. NSGA-II, NSGA-III, MOEA/D, Pareto fronts, constraint handling, benchmarks (ZDT, DTLZ), for engineering design and optimization problems.
Statistical modeling toolkit. OLS, GLM, logistic, ARIMA, time series, hypothesis tests, diagnostics, AIC/BIC, for rigorous statistical inference and econometric analysis.
Add unsigned integer (uint) type support to PyTorch operators by updating AT_DISPATCH macros. Use when adding support for uint16, uint32, uint64 types to operators, kernels, or when user mentions enabling unsigned types, barebones unsigned types, or uint support.
Convert PyTorch AT_DISPATCH macros to AT_DISPATCH_V2 format in ATen C++ code. Use when porting AT_DISPATCH_ALL_TYPES_AND*, AT_DISPATCH_FLOATING_TYPES*, or other dispatch macros to the new v2 API. For ATen kernel files, CUDA kernels, and native operator implementations.
Write docstrings for PyTorch functions and methods following PyTorch conventions. Use when writing or updating docstrings in PyTorch code.
Take nvidia/earth2studio-data-fetch from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip, uv.
Without those the skill loads but fails at the first command.