mcpbeat

Using Model Endpoint

xuzhougeng/using-model-endpoint

Invoke an already configured model endpoint from a supported Wisp execution context and capture the bounded inference as a Run. Use only when the endpoint URL and authentication are already available inside that context; this skill does not register or manage services.

624 tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
859
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/xuzhougeng/wisp-science --skill using-model-endpoint

The instruction itself

2 sections, as written by the author

Use an existing model endpoint

Wisp can record a bounded client invocation as a Run, but it does not register

or manage the endpoint. Require all of the following:

  • a selected local, wsl:<distro>, or ssh:<alias> context;
  • a concrete endpoint URL reachable from that context;
  • authentication already configured by the user in that execution environment

or the endpoint client's own external configuration;

  • a documented request and response schema;
  • a finite request timeout and a concrete output path.

Do not ask the user to paste secrets into the command, project files, or chat.

Wisp exposes no credential accessor to the Agent and does not inject keyring

values into run_in_context commands.

Invocation workflow

  • Write a small deterministic client such as runs/call_endpoint.py. Read the

URL and credential variable names at runtime; never embed secret values.

  • Validate its request against the endpoint's documented schema.
  • For SSH, stage the client and small inputs with input_paths. Keep large

inputs at an existing absolute remote path.

  • Submit one invocation with run_in_context and register the response with

output_specs:

{
  "context_id": "ssh:gpu-box",
  "title": "Existing endpoint inference",
  "command": "source ~/miniforge3/etc/profile.d/conda.sh && conda activate endpoint-client && python call_endpoint.py --input request.json --output /home/me/wisp-results/endpoint/response.json",
  "timeout_secs": 300,
  "input_paths": ["runs/call_endpoint.py", "data/request.json"],
  "output_specs": [
    {
      "glob": "ssh://gpu-box/home/me/wisp-results/endpoint/response.json",
      "kind": "json",
      "residency": "remote"
    }
  ]
}
  • Replace all example context and paths. Call monitor_run once when waiting

is useful, get_run once for a snapshot, or cancel_run to stop.

Local and WSL Runs are capped at 300 seconds and do not accept input_paths.

Keep their client and outputs in host-visible project paths. If endpoint setup,

tunnelling, health management, or deployment is required, stop and load

managed-model-endpoints for the explicit current boundary.

How to use it

Copy the folder

Take xuzhougeng/using-model-endpoint from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.