artokun/local-llm-free
Run the ComfyUI agent locally for FREE — no subscription, no API key, fully offline — using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama. Use when the user asks about running locally, running for free, offline use, avoiding API costs, Ollama setup, or which local model to pick.
npx skills add https://github.com/artokun/comfyui-mcp --skill local-llm-free
The answer to "can I run this for free / offline / without an API key" is
yes: the panel's Ollama backend drives the full live-canvas agent on a
local model — and we ship models fine-tuned specifically for comfyui-mcp.
artokun/gemma4-comfyui-mcp is Google's Gemma 4 QLoRA-fine-tuned on **1,055
server-verified tool-use trajectories** generated against a live ComfyUI —
covering the full 178-tool surface (113 MCP tools + 65 panel live-canvas
tools). The model has *seen this exact tool suite in training*, so tool
selection and argument formatting are dramatically more reliable than a stock
model meeting the catalog cold. Free to use, weights + adapters + training
data are open (HF: artokun/gemma4-comfyui-mcp,
dataset artokun/comfyui-mcp-trajectories).
(macOS/Windows installers, or curl -fsSL https://ollama.com/install.sh | sh on Linux).
ollama pull artokun/gemma4-comfyui-mcp:e4b # DEFAULT — ~3.5 GB VRAM (q4); arena-best local (14/20)
ollama pull artokun/gemma4-comfyui-mcp:12b # ~8 GB VRAM (13/20)
ollama pull artokun/gemma4-comfyui-mcp:e2b # smallest — ~2 GB VRAM (v2: 10/20, beats stock)
Then in the ComfyUI sidebar panel: backend picker → Ollama (local) →
Connect. :e4b is the built-in default — zero further config once pulled.
(Override via the panel's model picker or COMFYUI_MCP_OLLAMA_MODEL.)
| GPU VRAM free | Recommend |
| --- | --- |
| ~2-3 GB | :e2b (v2: 10/20 — beats stock e2b's 8; handles the foundation flows, expect misses on long multi-step builds) |
| ~4-7 GB | :e4b (the default sweet spot — best local model on the arena, 14/20) |
| 8 GB+ | :12b (13/20; steadier on long multi-step tasks) |
agent generates and edits workflows fine but can't visually critique its
own outputs. Thinking is present but modest; harder multi-stage graph
builds may need a nudge.
pair these models with compact tool mode (--compact) — full docs:
https://comfyui-mcp.artokun.io/docs/local-llms
Take artokun/local-llm-free from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.