Vast.ai Python SDK — high-level API for GPU instances, volumes, serverless endpoints, and billing.
npx skills add https://github.com/vast-ai/vast-cli --skill vastai-sdk
vastai / vastai_sdk)The vastai package provides a Python SDK for managing GPU instances, volumes, serverless endpoints, and billing on Vast.ai. The vastai_sdk package is a backward-compatibility shim that re-exports vastai.
pip install vastai
For serverless and async support:
pip install "vastai[serverless]"
The SDK reads the API key from ~/.vast_api_key by default. You can also pass it explicitly:
from vastai import VastAI
vast = VastAI() # reads ~/.vast_api_key
vast = VastAI(api_key="YOUR_API_KEY") # explicit key
Get your API key from https://console.vast.ai/manage-keys/
The old vastai_sdk import still works:
from vastai_sdk import VastAI # equivalent to: from vastai import VastAI
from vastai import VastAI
vast = VastAI(api_key=None, server_url=None, retry=3, raw=False, quiet=False)
# List all your instances
instances = vast.show_instances()
# Get a single instance
instance = vast.show_instance(id=12345)
# Search GPU offers
offers = vast.search_offers(query='gpu_name=RTX_4090 num_gpus>=4 reliability>0.99')
# Create an instance from an offer
result = vast.create_instance(id=<offer_id>, image="pytorch/pytorch:latest", disk=50)
# Lifecycle
vast.start_instance(id=12345)
vast.stop_instance(id=12345)
vast.reboot_instance(id=12345)
vast.destroy_instance(id=12345)
# Label an instance
vast.label_instance(id=12345, label="my-training-run")
# Get SSH connection string
ssh_url = vast.ssh_url(id=12345) # returns "ssh -p PORT user@host"
scp_url = vast.scp_url(id=12345) # returns scp-compatible URL
Interruptible (spot) instances are priced below on-demand instances, but can be interrupted at any time by another user with a lower bid. Note: vast.search_offers(type='bid', ...) exposes min_bid, but vast.create_instance(...) defaults to on-demand at dph_total unless you pass bid_price=<floor>. Always pass bid_price after a type='bid' search, otherwise the instance will be rented as an on-demand instance/price instead of as an interruptible.
When outbid, the instance moves to stopped (not destroyed) and storage charges continue. Resume by raising the bid via vast.change_bid(id=..., price=...).
# Search GPU offers (use help(vast.search_offers) for full query syntax)
offers = vast.search_offers(query='gpu_name=RTX_3090 num_gpus>=2')
# Search volume offers
volumes = vast.search_volumes(query='...')
# Search network volumes
net_vols = vast.search_network_volumes()
# Search templates
templates = vast.search_templates()
# Search invoices
invoices = vast.search_invoices()
# copy() takes vast URLs: "[C.|V.]id:path", "cloud_service[.id]:path", or "local:path"
vast.copy("local:./data/", "C.12345:/workspace/data/") # Local → instance
vast.copy("C.12345:/workspace/results/", "local:./out/") # Instance → local
vast.copy("12345:/workspace/", "67890:/workspace/") # Instance → instance (legacy format)
vast.copy("s3.101:/data/", "C.12345:/workspace/") # Cloud service → instance
vast.copy("V.1234:/file", "C.5678:/workspace/") # Volume → instance
vast.copy("V.1234:/file", "s3.101:/workspace/") # Volume → cloud service
vast.cancel_copy(dst_id=12345) # Cancel an in-progress copy
# Cloud sync via a saved cloud connection (see the UI settings page for connection IDs)
vast.cloud_copy(src="./data", dst="s3://bucket/path", instance=12345,
connection=<conn_id>, transfer="Instance To Cloud")
vast.cancel_sync(dst_id=12345)
Volume copy is currently only supported for copying to other volumes, instances, or cloud services, not local. Do not use /root or / as a destination directory — it breaks ssh permissions on the instance and future copies fail. See https://vast.ai/docs/gpu-instances/data-movement#constraints.
# List all deployments
deployments = vast.show_deployments()
# Get a deployment
deployment = vast.show_deployment(id=42)
# Delete a deployment
vast.delete_deployment(id=42)
machines = vast.show_machines()
machine = vast.show_machine(id=10)
vast.list_machine(id=10, price_gpu=0.30)
vast.unlist_machine(id=10)
keys = vast.show_ssh_keys()
vast.create_ssh_key(ssh_key="ssh-rsa AAAA...")
vast.delete_ssh_key(id=5)
members = vast.show_members()
vast.invite_member(email="[email protected]", role="developer")
vast.remove_member(id=7)
SyncClient provides typed, synchronous access to GPU offers and instances.
from vastai import SyncClient
client = SyncClient(api_key="YOUR_API_KEY") # or reads ~/.vast_api_key
# Search offers with structured filters
offers = client.search(
num_gpus=2,
gpu_name="RTX_4090",
min_reliability=0.99,
max_dph_total=2.0,
)
# Create an instance
instance = client.create_instance(
offer_id=<id>,
image="pytorch/pytorch:latest",
disk_gb=50,
)
# List your instances
instances = client.show_instances() # returns list[SyncInstance]
# Destroy an instance
client.destroy_instance(instance_or_id=12345)
AsyncClient provides async access to GPU offers and instances. Use as an async context manager.
import asyncio
from vastai import AsyncClient
async def main():
async with AsyncClient(api_key="YOUR_API_KEY") as client:
# Search offers
offers = await client.search(num_gpus=1, gpu_name="A100")
# Create instance
instance = await client.create_instance(offer_id=<id>, image="ubuntu:22.04")
# List instances
instances = await client.show_instances() # returns list[AsyncInstance]
# Destroy instance
await client.destroy_instance(instance_or_id=instance.id)
asyncio.run(main())
For inference endpoints (requires pip install "vastai[serverless]"):
import asyncio
from vastai import Serverless
async def main():
serverless = Serverless() # reads ~/.vast_api_key
# Get an endpoint
endpoint = await serverless.get_endpoint("my-endpoint")
# Make a request
response = await serverless.request("/v1/completions", {
"model": "Qwen/Qwen3-8B",
"prompt": "Who are you?",
"max_tokens": 100,
"temperature": 0.7,
})
text = response["response"]["choices"][0]["text"]
print(text)
asyncio.run(main())
# Find cheapest 4x RTX 4090 and launch a job
from vastai import VastAI
vast = VastAI()
offers = vast.search_offers(query='gpu_name=RTX_4090 num_gpus=4 reliability>0.99')
cheapest = min(offers, key=lambda o: o['dph_total'])
result = vast.create_instance(id=cheapest['id'], image="pytorch/pytorch:latest", disk=100)
print(f"Launched instance: {result['new_contract']}")
# Use help() to explore method signatures
help(vast.search_offers)
help(vast.create_instance)
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
Automatically creates user-facing changelogs from git commits by analyzing commit history, categorizing changes, and transforming technical commits into clear, customer-friendly release notes. Turns hours of manual changelog writing into minutes of automated generation.
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
React Native and Expo best practices for building performant mobile apps. Use when building React Native components, optimizing list performance, implementing animations, or working with native modules. Triggers on tasks involving React Native, Expo, mobile performance, or native platform APIs.
React and Next.js performance optimization guidelines from Vercel Engineering. This skill should be used when writing, reviewing, or refactoring React/Next.js code to ensure optimal performance patterns. Triggers on tasks involving React components, Next.js pages, data fetching, bundle optimization, or performance improvements.
Next.js best practices - file conventions, RSC boundaries, data patterns, async APIs, metadata, error handling, route handlers, image/font optimization, bundling
Use when starting feature work that needs isolation from current workspace or before executing implementation plans - creates isolated git worktrees with smart directory selection and safety verification
Take vast-ai/vastai-sdk from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip.
Without those the skill loads but fails at the first command.