mcpbeat Sign in

Modal Agent Skill

Use when running Python or GPU workloads serverlessly on Modal — modal.App, inline container Images, gpu= on @app.function, Volumes for weight caching, Cron schedules, ASGI endpoints, modal run vs serve vs deploy. NOT managed prediction APIs with no container of your own (that is replicate); NOT SSH-able GPU boxes rented by the hour (that is runpod).

9k tokens
context cost
the whole folder, loaded on every use
6
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
105
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/ericrisco/rsc-harness --skill modal

What comes with it

20 692 bytes besides the instruction
evals/README.md
evals/cases.yaml
references/images-gpu-cookbook.md
references/web-and-scaling.md
scripts/verify.sh

The instruction itself

14 sections, as written by the author

Modal runs your Python on remote containers without you ever writing a Dockerfile or a YAML

file. The mental model: infrastructure is declared inline as Python decorators. A

modal.App is the deployable unit; each @app.function runs in its own container built from

a modal.Image you describe in code; you attach a GPU, a Volume, or a Secret as keyword

arguments and the platform provisions, scales to zero, and tears down for you. There is no

control plane to babysit — the source file *is* the infra.

Pinned stack: modal 1.4.3 (released 2026-05-18), Python 3.10–3.14 (>=3.10,<3.15).

Install with pip install modal then modal setup to authenticate. Everything below uses

the Modal 1.0+ API; several pre-1.0 forms were removed and are called out as Bad→Good.

Not this skill

Modal owns the serverless-container-as-decorators surface and its CLI lifecycle; the *contents*

of your function belong elsewhere.

| The job | Goes to |

|---|---|

| Calling a managed prediction API with no container of your own | replicate / together-fireworks / fal |

| Renting a persistent, SSH-able GPU box by the hour/week | runpod |

| FastAPI design (routing, Pydantic, deps) independent of host | fastapi |

| Writing a Dockerfile for a registry / k8s / Compose | docker |

| General Python language/runtime questions | python |

| RAG / LLM pipeline orchestration logic itself | llm-pipeline |

Decision: which entrypoint?

| You want… | Use | Persists after exit? |

|---|---|---|

| Run a function once and exit (script, batch) | modal run app.py + @app.local_entrypoint() | No (ephemeral) |

| Hot-reload dev loop for a web endpoint | modal serve app.py | No (dies on Ctrl-C) |

| A persistent named deployment (prod, schedules, endpoints) | modal deploy app.py | Yes |

| Fan out work across many containers | .map() / .starmap() / .spawn() inside an entrypoint | n/a |

Rule: schedules and live web endpoints require modal deploy. modal run exits when the

entrypoint returns, so a Cron defined under modal run never fires. modal serve is for the

dev loop only — it watches your files and redeploys on save, but the app vanishes when you

stop it.

The minimal app skeleton

import modal

app = modal.App("hello-modal")

# The image is the container spec. Build it once, reuse across functions.
image = modal.Image.debian_slim(python_version="3.12").uv_pip_install("requests")


@app.function(image=image)
def fetch(url: str) -> int:
    import requests  # imported INSIDE the function: it lives in the remote image, not locally

    return len(requests.get(url).content)


@app.local_entrypoint()
def main() -> None:
    # Runs on your laptop; .remote() ships the call to a Modal container.
    print(fetch.remote("https://modal.com"))

Run it: modal run app.py. Bad = wiring infra with argparse + a bash launcher + a

hand-rolled Dockerfile. Good = the decorators above; the app, image, and scaling are all

declared in the one file. Note the in-function import: dependencies you uv_pip_install exist

in the *remote* image, so import them inside the function (or guard top-level imports), not at

module top where your laptop would need them too.

Images — pin, layer, cache

Build images by chaining methods on modal.Image. Rules, each with its why:

  • Prefer .uv_pip_install(...) over .pip_install(...) — it resolves and installs with

uv, materially faster image builds.

  • Pin versions.uv_pip_install("torch==2.5.1", "transformers==4.46.0"). Unpinned

deps make builds non-reproducible and silently drift on rebuild.

  • Order layers stable→volatile — system packages and big wheels first, your fast-changing

code last. Modal caches each layer; a change busts that layer and everything after it.

  • Add your own code with .add_local_dir(...) / .add_local_python_source(...), not by

pip-installing your repo. These are applied last so editing your source doesn't rebuild torch.

  • .from_registry("...") when you need a specific base image; .apt_install("ffmpeg")

for system binaries; .run_commands(...) for arbitrary build steps.

image = (
    modal.Image.debian_slim(python_version="3.12")
    .apt_install("ffmpeg")                                   # stable: rarely changes
    .uv_pip_install("torch==2.5.1", "transformers==4.46.0")  # heavy wheels, pinned
    .add_local_python_source("my_pkg")                       # volatile: your code, applied last
)

references/images-gpu-cookbook.md for vLLM / torch+CUDA

/ diffusers recipes and the download-once weight-cache pattern.

GPU — it's a string now

In Modal 1.0+ the GPU is a string on the decorator. The old modal.gpu.H100() objects

were removed.

  • Single GPU: gpu="H100".
  • Count via colon: gpu="A100:2" (two A100s in one container).
  • Memory variant: gpu="A100-80GB" (also A100-40GB).
  • Fallback list (first available wins): gpu=["H100", "A100", "any"].
  • Supported types: T4, L4, A10, L40S, A100(-40GB/-80GB), RTX-PRO-6000, H100, H200, B200.
# Bad — removed API, raises at import.
# @app.function(gpu=modal.gpu.A100())

# Good — string form.
@app.function(image=image, gpu="A100-80GB", timeout=600)
def embed(texts: list[str]) -> list[list[float]]: ...

Pick the smallest GPU that fits: T4/L4 for cheap inference and small models, A10/L40S

mid-range, A100/H100 for training and large-model serving, H200/B200 for frontier-scale.

GPU time is billed per second a container is alive — never attach a GPU to a CPU-only job, and

keep scaledown_window tight so idle GPU containers don't burn money.

Scaling & lifecycle

Tune these keyword args on @app.function, each with its why:

| Param | Effect | Why |

|---|---|---|

| min_containers=N | Keep N warm instances always running | Kills cold starts for latency-sensitive endpoints (costs idle compute) |

| buffer_containers=N | Pre-warm N extra beyond current load | Smooths bursty traffic |

| scaledown_window=300 | Seconds an idle container lingers before shutdown | Reuse hot containers across nearby calls; lower = cheaper, higher = warmer |

| timeout=600 | Max seconds a single call may run | Caps runaway jobs |

| retries=3 | Auto-retry failed inputs | Survives transient failures in .map() fan-outs |

Migration note: keep_warmmin_containers and container_idle_timeout

scaledown_window in the 1.0 migration. The old names are gone.

Concurrency within a container is now its own decorator: @modal.concurrent(max_inputs=N)

stacked under @app.function (it replaces the old allow_concurrent_inputs= argument). Use it

so one container handles N simultaneous requests instead of one-per-container.

Volumes & Secrets

A Volume is a distributed filesystem you mount into containers to persist data across runs —

the canonical use is caching downloaded model weights so cold starts skip the re-download.

weights = modal.Volume.from_name("hf-cache", create_if_missing=True)


@app.function(image=image, gpu="H100", volumes={"/cache": weights})
def serve_model():
    # Reader: refresh the view so you see writes from other containers.
    weights.reload()
    # ... load model from /cache ...


@app.function(image=image, volumes={"/cache": weights})
def download_weights():
    # ... write files into /cache ...
    weights.commit()  # WITHOUT this, writes are NOT durable across containers

Gotcha: writers must call vol.commit() to persist; readers call vol.reload() to see

another container's committed writes. Forgetting commit() is the #1 "my cache is empty"

bug — the files existed in that container and vanished with it.

Secrets land as environment variables in the container:

@app.function(image=image, secrets=[modal.Secret.from_name("hf-token")])
def pull():
    import os

    token = os.environ["HF_TOKEN"]  # value injected from the named Modal Secret

Never bake a token into the image (.run_commands("export TOKEN=...")) — it's recorded in

layer history. Use a Secret. The cookbook above also carries the HF/OpenAI secret patterns.

Web endpoints

Stack a web decorator under @app.function. Pick by surface:

| Decorator | Use for | Needs |

|---|---|---|

| @modal.fastapi_endpoint() | A single GET/POST function-as-URL | fastapi[standard] in image |

| @modal.asgi_app() | A full FastAPI/Starlette app you return | fastapi[standard] |

| @modal.wsgi_app() | A Flask/Django WSGI app | the framework |

| @modal.web_server(port=8000) | Your own server process (e.g. vLLM) on a port | the server |

Decorator stack order matters: @app.function is outermost (top), then optional

@modal.concurrent, then the web decorator innermost (bottom, closest to def).

@app.function(image=image, gpu="H100", min_containers=1, scaledown_window=300)
@modal.concurrent(max_inputs=10)   # middle
@modal.asgi_app()                  # innermost
def web():
    from fastapi import FastAPI

    api = FastAPI()

    @api.get("/health")
    def health():
        return {"ok": True}

    return api

Develop with modal serve app.py (hot-reload); ship with modal deploy app.py (stable URL).

For custom domains, proxy-auth tokens, batching (@modal.batched), and concurrency tuning →

references/web-and-scaling.md. For the FastAPI app's *own*

design (routes, Pydantic, deps), that's fastapi — this skill only mounts it.

Scheduled jobs

# Fixed wall-clock time, with timezone — survives redeploys at the same clock time.
@app.function(schedule=modal.Cron("0 6 * * *", timezone="America/New_York"))
def nightly_report(): ...


# Interval relative to deploy time.
@app.function(schedule=modal.Period(hours=5))
def every_five_hours(): ...

Gotcha: Period is measured from deploy time and resets on every redeploy — redeploy

at 4:59 and your "every 5 hours" clock restarts. Cron is wall-clock stable; prefer it for

"run at 6am" semantics. Either way you must modal deploy (not modal run) for the

schedule to live on the platform.

Parallelism

Fan a function out across containers without managing a pool:

@app.local_entrypoint()
def main():
    urls = ["https://a.com", "https://b.com", "https://c.com"]
    # .map: one arg per call, results in input order.
    sizes = list(fetch.map(urls))
    # .starmap: each item is an argument tuple. .spawn: fire-and-forget -> handle.get() later.
    handle = fetch.spawn("https://slow.com")
    print(sizes, handle.get())

.map(iterable) returns results in order by default; pass order_outputs=False to yield as

they complete (faster when latencies vary). Combine with retries= on the function so a

single bad input doesn't sink the batch.

Anti-patterns

| Anti-pattern | Do instead |

|---|---|

| "I'll use gpu=modal.gpu.A100() like the old docs" | Removed in 1.0. Use the string gpu="A100-80GB". |

| "Attach a GPU, it might speed up this CPU job" | GPU is billed per second alive. CPU-only job → no gpu=. |

| "My files are written, the Volume will keep them" | Not without vol.commit() (writer) / vol.reload() (reader). |

| "Pin later — uv_pip_install('torch') is fine for now" | Unpinned deps drift; builds aren't reproducible. Pin every version. |

| "modal run it, the endpoint/schedule will stay up" | run is ephemeral; it exits. Use modal deploy for anything persistent. |

| "Order the decorators however — Modal figures it out" | @app.function outermost, web decorator innermost. Wrong order errors. |

| "Bake the HF token into the image with run_commands" | Leaks into layer history. Use modal.Secret.from_name(...). |

| "Just call the model via a managed API through Modal" | If you write no container, that's a managed-API job → replicate. |

| "I need a box to SSH into for a week" | That's a persistent rental → runpod, not Modal's scale-to-zero. |

| "Set min_containers high so it's always fast" | Idle warm containers cost money 24/7. Tune scaledown_window first. |

Verify

scripts/verify.sh [TARGET] statically lints the nearest emitted Modal

*.py: it requires a modal.App(, fails if the removed modal.gpu. object form appears,

checks that any web decorator sits under an @app.function, and that any Volume uses

from_name(..., create_if_missing=...). It runs python -c "import modal" only if modal is

installed (skip-pass otherwise), needs no Modal credentials, and exits 0 on an empty target.

Project grounding (02-DOCS)

In a project with a 02-DOCS/ layer (the harness wiki), read

02-DOCS/wiki/stack/modal.md first, then record this app's real Modal choices there — GPU types,

image base, Volume names, schedule, endpoint shape — and index it in 02-DOCS/wiki/index.md. No

02-DOCS/? Skip silently.

How to use it

Copy the folder

Take ericrisco/modal from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.

Install what it needs

The instructions reference pip. Without those the skill loads but fails at the first command.