nvidia/earth2studio-data-fetch
> Fetch weather/climate data via Earth2Studio data sources for specific variables and times. Do NOT use for inference pipelines, model discovery, or installation.
npx skills add https://github.com/NVIDIA/skills --skill earth2studio-data-fetch
Guide a user through downloading weather/climate data via Earth2Studio data source
APIs. Identifies compatible sources by checking the lexicon, verifies variable
support, and produces a working fetch script outputting an xarray DataArray.
uv pip install earth2studio or equivalent)~/.cdsapirc)You are helping a user download specific weather/climate data using
Earth2Studio's data source APIs. Your job is to identify which data source(s)
can provide the requested variables, verify compatibility via the lexicon
system, and produce a working fetch script.
Data source APIs, available variables, and the lexicon evolve between releases.
Before recommending a data source or writing a fetch script:
and constructor arguments.
that data source.
Live doc references (fetch only what the user's request requires):
<https://nvidia.github.io/earth2studio/modules/datasources_analysis.html>
<https://nvidia.github.io/earth2studio/modules/datasources_forecast.html>
<https://nvidia.github.io/earth2studio/modules/datasources_dataframe.html>
<https://github.com/NVIDIA/earth2studio/blob/main/earth2studio/lexicon/base.py>
<https://github.com/NVIDIA/earth2studio/tree/main/earth2studio/lexicon>
Extract from what the user has said (ask follow-ups if needed, cap at 3
questions):
(e.g. t2m, u500, z850, tp, msl). If the user uses plain language
("500 hPa geopotential height"), map it to the E2Studio name by checking
the live base.py E2STUDIO_VOCAB.
discrete times?
Based on the request type, narrow candidates:
Analysis/reanalysis (historical state at a specific time):
IFS/IFS_ENS (ECMWF), ARCO/CDS/WB2ERA5/NCAR_ERA5 (ERA5 reanalysis),
GOES/MRMS/JPSS (observational)
Forecast (predictions from an initialization time with lead times):
AIFS_FX, CFS_FX
Key differentiators to surface:
history; reanalysis (ERA5 via ARCO/CDS/WB2) goes back decades
WB2ERA5_32x64 is 5.625° global
This is critical. Each data source has a lexicon file that defines which
E2Studio variables it can provide.
To verify:
https://github.com/NVIDIA/earth2studio/blob/main/earth2studio/lexicon/<source>.py
(e.g. gfs.py, hrrr.py, cds.py, arco.py, wb2.py)
source's VOCAB dict
it — try another
The lexicon VOCAB maps Earth2Studio variable names → source-specific
identifiers. If a variable key exists in the VOCAB, the source supports it.
Present the results clearly: *"GFS supports t2m, u500, z850. HRRR also
supports these but is limited to North America. ARCO (ERA5) supports all
three and has data back to 1959."*
Present the viable options with tradeoffs:
| Source | Variables | Coverage | Resolution | Time Range |
|--------|-----------|----------|------------|------------|
| ... | ... | ... | ... | ... |
Let the user pick. If there's one obvious choice, recommend it and ask for
confirmation.
Write a Python script that uses the selected data source to fetch the
requested data. The script structure depends on whether it's an analysis or
forecast source.
Analysis source pattern:
import datetime
from earth2studio.data import <SourceClass>
# Initialize data source
ds = <SourceClass>()
# Fetch data
# Analysis sources use: ds(time, variable) -> xr.DataArray
time = [datetime.datetime(YYYY, M, D, H)] # or array of times
variable = ["var1", "var2"] # E2Studio variable names
data = ds(time, variable)
Forecast source pattern:
import datetime
from earth2studio.data import <SourceClass>
# Initialize data source
ds = <SourceClass>()
# Forecast sources use: ds(time, lead_time, variable) -> xr.DataArray
time = [datetime.datetime(YYYY, M, D, H)] # initialization time
lead_time = [datetime.timedelta(hours=H)] # or array of lead times
variable = ["var1", "var2"]
data = ds(time, lead_time, variable)
Always fetch the specific data source's API doc page to confirm the exact
constructor arguments and call signature before writing the script — they can
vary (some need auth tokens, cache paths, specific parameters).
Include in the script:
print(data), data.shape, data.coords)After delivering the script, mention:
discover skill
EARTH2STUDIO_CACHE)
Owns: identifying data sources for a user's variable/time request,
verifying variable support via lexicon, generating data fetch scripts,
explaining analysis vs. forecast source differences.
Does not own: installation (earth2studio-install), model selection
(earth2studio-discover), inference pipelines, custom data source creation
(point to extend examples), data source authentication setup beyond what
the docs describe.
Typical invocation:
> "I need 500 hPa geopotential height and 2m temperature from ERA5
> for January 1, 2020 at 00Z."
The skill would:
z500, t2m(GCS, S3, CDS API)
DataArrayFile/DataSetFile directly
sources in a single call
variables; always verify via lexicon
are generally faster
| Error | Cause | Solution |
|-------|-------|----------|
| KeyError: '<var>' | Not in lexicon | Check lexicon; try another source |
| FileNotFoundError / 404 | Time not available | Verify temporal coverage |
| CDS API timeout | Queue congestion | Retry or use ARCO for ERA5 |
| ModuleNotFoundError | Not installed | uv pip install earth2studio |
| Empty DataArray | Time/var mismatch | Check datetime and variable name |
Take nvidia/earth2studio-data-fetch from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip, uv.
Without those the skill loads but fails at the first command.