mcpbeat Sign in

Local Models Skill for Claude

Run quick, offline, private LLM tasks on local models via llama.cpp, reusing models already downloaded by Ollama. Use for cheap/bulk text work (summarize, classify, extract JSON, anonymize PII, translate, proofread, keywords), local embeddings, and offline image description — and prefer it over a cloud API whenever a task is privacy-sensitive, must run offline, is high-volume/low-stakes, or just needs a fast throwaway answer. Provides an `lm` CLI wrapper plus an OpenAI-compatible local server.

5k tokens
context cost
the whole folder, loaded on every use
4
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
337
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/glebis/claude-skills --skill local-models

What comes with it

14 280 bytes besides the instruction
references/serving-and-embeddings.md
scripts/lm
scripts/ollama_blob.py

The instruction itself

8 sections, as written by the author

local-models

Quick access to local LLMs through llama.cpp, reusing the GGUF models already

pulled by Ollama (no re-download for text and embeddings). Everything runs on the

machine — no API key, no network, no per-token cost.

When to use this skill

Reach for local models instead of a cloud API when the task is:

  • Privacy-sensitive — redacting PII, processing personal notes, health data, secrets-adjacent text. The data never leaves the machine.
  • Offline — no network available, or the user explicitly wants local-only.
  • High-volume / low-stakes — classifying or tagging hundreds of items, where a small model is good enough and cloud cost/latency would add up.
  • A fast throwaway — a quick summary, translation, or "what is this" where round-tripping to a frontier model is overkill.

Prefer a frontier (Claude) model when the task needs strong reasoning, long

context, careful code, or high accuracy — these local models are small (0.6–4B).

The core trick: reuse Ollama's models

Ollama stores model weights as extension-less GGUF blobs under

~/.ollama/models/blobs/. These are ordinary GGUF files — llama.cpp loads

them directly. scripts/ollama_blob.py reads Ollama's manifests and resolves a

friendly name (e.g. qwen2.5:3b) to its weights blob path. No conversion, no

duplicate downloads.

Usage

The entry point is scripts/lm. Run scripts/lm help for the full list. Invoke

it with an absolute path, e.g. ~/ai_projects/claude-skills/local-models/scripts/lm.

lm models                       # list local models (text / vision / embed)
lm ask [MODEL] "PROMPT"         # one-shot prompt (default qwen2.5:3b)
lm chat [MODEL]                 # interactive REPL

# Text presets — accept a file path, inline text, OR stdin:
lm summarize  report.md
cat notes.txt | lm tldr
lm keywords   article.txt
lm anonymize  transcript.txt         # → [NAME] [EMAIL] [PHONE] [ADDRESS] ...
lm proofread  draft.md
lm translate  German "Good morning"
lm classify   "praise,complaint,question"  feedback.txt   # → one label
lm extract    "invoice_number, total, due_date"  invoice.txt   # → JSON

# Vision (downloads model+projector once via HuggingFace — see note below):
lm describe-image photo.jpg
lm tag-image      screenshot.png
lm vision photo.jpg "What brand is the shoe?"

# Embeddings & serving:
lm embed "text to embed"             # → OpenAI-style JSON vector
lm serve qwen2.5:3b 8080             # OpenAI-compatible server on :8080

Output is clean (just the answer) — the wrapper drives llama-completion in

single-turn mode and strips the chat-template scaffolding and llama.cpp logs.

Choosing a model

Defaults are tuned for clean, fast output and can be overridden per call:

  • General text presets → qwen2.5:3b (LM_TEXT_MODEL)
  • Classify / extract → qwen2.5:3b (LM_REASON_MODEL), run at temperature 0
  • Embeddings → jeffh/intfloat-multilingual-e5-large:f16 (LM_EMBED_MODEL)
  • Other envs: LM_NTOK (max tokens), LM_VISION_HF (vision repo), LM_DEBUG=1 (show llama.cpp logs)

Pass an explicit model as the first argument to ask/chat/embed/serve

(e.g. lm ask qwen3:4b "...").

Critical gotchas

  • Ollama's gemma3 GGUF does NOT load in stock llama.cpp. It fails with

key not found in model: gemma3.attention.layer_norm_rms_epsilon because

Ollama writes custom metadata keys mainline llama.cpp doesn't read. Use a

qwen* model instead, or pull a community gemma3 GGUF via -hf. This is why

the defaults are qwen, not gemma3.

  • qwen3:4b emits <think>…</think> reasoning blocks before its answer.

Fine for ask/chat, but it pollutes preset output (JSON, labels) — the

presets default to qwen2.5:3b to avoid this.

  • Vision has no Ollama blob to reuse. Ollama did not store an mmproj

(vision projector) for qwen2.5vl, and llama.cpp needs one. So the vision

commands use llama-mtmd-cli -hf ggml-org/Qwen2.5-VL-3B-Instruct-GGUF, which

downloads model+projector (~2–3 GB) into ~/.cache/llama.cpp on first use,

then runs offline. Warn the user before the first vision call.

  • Each one-shot call reloads the model (a few seconds for these small

models). For many sequential calls, start a server once with lm serve and

hit http://localhost:8080/v1/chat/completions — see

references/serving-and-embeddings.md.

Reference material

  • references/serving-and-embeddings.md —

running llama-server as an OpenAI-compatible endpoint (and pointing the llm

CLI or any OpenAI client at it), plus local embeddings / RAG patterns with

llama-embedding.

Requirements

  • llama.cpp installed (brew install llama.cpp) — provides llama-completion,

llama-mtmd-cli, llama-embedding, llama-server.

  • Ollama with at least one pulled model (for the blob-reuse path). python3 for

the resolver. No API keys.

Other skills for the same job

different authors, same section of the catalogue
Doc Coauthoring
by anthropics
vendor ×10

Guide users through a structured workflow for co-authoring documentation. Use when user wants to write documentation, proposals, technical specs, decision docs, or similar structured content. This workflow helps users efficiently transfer context, refine content through iteration, and verify the doc works for readers. Trigger when user mentions writing docs, creating proposals, drafting specs, or similar documentation tasks.

4k tokens
Changelog Generator
by frostant
×9

Automatically creates user-facing changelogs from git commits by analyzing commit history, categorizing changes, and transforming technical commits into clear, customer-friendly release notes. Turns hours of manual changelog writing into minutes of automated generation.

774 tokens
Test Driven Development
by w95
×7

Use when implementing any feature or bugfix, before writing implementation code

2k tokens
Writing Plans
by ZhanlinCui
×4

Use when you have a spec or requirements for a multi-step task, before touching code

816 tokens
Writing Skills
by ZhanlinCui
×4

Use when creating new skills, editing existing skills, or verifying skills work before deployment

26k tokens scripts
Crafting Effective Readmes
by softaworks
×3

Use when writing or improving README files. Not all READMEs are the same — provides templates and guidance matched to your audience and project type.

15k tokens
Humanizer
by softaworks
×3

| Remove signs of AI-generated writing from text. Use when editing or reviewing text to make it sound more natural and human-written. Based on Wikipedia's inflated symbolism, promotional language, superficial -ing analyses, vague attributions, em dash overuse, rule of three, AI vocabulary words, negative parallelisms, and excessive conjunctive phrases.

6k tokens
Opentrons Integration
by christophacham
×3

Official Opentrons Protocol API for OT-2 and Flex robots. Use when writing protocols specifically for Opentrons hardware with full access to Protocol API v2 features. Best for production Opentrons protocols, official API compatibility. For multi-vendor automation or broader equipment control use pylabrobot.

9k tokens scripts

How to use it

Copy the folder

Take glebis/local-models from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.

Install what it needs

The instructions reference brew. Without those the skill loads but fails at the first command.