mcpbeat

FitLLM MCP Server

run.fitllm/fitllm
not responding

FitLLM is listed as active in the registry but did not answer our last check. It exposes 3 tools. Last commit 16 Jul 2026.

Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only.

Uptime history 41 hours of history · worst hour 0%
41 hours agonow
0.0%
Uptime 24h
0 of 91 checks
3
Tools
read from the server
76 ms
Response time
average over 24h
7
Stars
last commit 16 Jul 2026

Connect this server

Endpoint below is the one we actually reach during checks — not the one copied from a README. Last verified 4 min ago.

run in your terminal
claude mcp add fitllm --transport http https://fitllm.run/api/mcp
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "fitllm": {
      "url": "https://fitllm.run/api/mcp"
    }
  }
}
~/.codex/config.toml
[mcp_servers.fitllm]
url = "https://fitllm.run/api/mcp"
.cursor/mcp.json
{
  "mcpServers": {
    "fitllm": {
      "url": "https://fitllm.run/api/mcp"
    }
  }
}
.vscode/mcp.json
{
  "mcpServers": {
    "fitllm": {
      "url": "https://fitllm.run/api/mcp"
    }
  }
}

Available tools 3

Read directly from the server with tools/list, grouped by what they act on. If a tool disappears, we record the date.

llm
check_llm_fit
Check whether a specific local LLM fits in the memory of a specific GPU or Apple Silicon Mac. Returns fits/tight/won't-fit verdict with the full memory breakdown (weights, KV cache, overhead), max context, and a concrete fix if it doesn't fit. Use this whenever a user asks anything like "can I run <model> on my <GPU/Mac>?", "will <model> fit in <N>GB?", or "what do I need to run <model>?". Architecture-aware math (MLA, sliding-window, hybrid attention, MoE) — more accurate than rule-of-thumb estimates.
supported
list_supported
List the built-in model names and hardware names this fit-checker knows (for mapping user wording to exact names). Any public HuggingFace model also works via fitllm.run.
what
what_fits_on_hardware
Rank which popular local LLMs fit on a given GPU or Apple Silicon Mac (at ~4-bit quantization, 8K context) — models that fit come first, biggest first, with max context each. Use when a user asks "what can I run on my <GPU/Mac/N GB>?", "best local model for my machine?", or gives hardware without naming a model.

Endpoints

URLTransportStateLatencyChecked
https://fitllm.run/api/mcp streamable-http answering 76 ms 4 min ago

FitLLM — questions

Answers built from our own checks of this server.

What can FitLLM do?
It exposes 3 tools, read directly from the server on our last check. Among them: check_llm_fit, list_supported, what_fits_on_hardware. The full list with descriptions is on this page — we take it from the server itself via tools/list, not from a README. How MCP servers expose tools in the first place →
Is FitLLM working right now?
We send a real MCP handshake every 15 minutes. Over the last 24 hours 0 of 91 checks got a reply (0.0%), average response time 76 ms. The bar chart above shows every period we have measured.
The registry lists FitLLM as active — why does it not respond?
The official MCP registry stores what the author submitted; it does not verify that the server still runs. We check the endpoint ourselves, and this one does not answer. Catalogues that copy the registry without checking will show it as working.
How do I connect FitLLM?
Copy the ready config from this page — we generate it for Claude Code, Claude Desktop, Codex, Cursor and VS Code, each with the file path that client actually reads. It is a remote server, so there is nothing to install — the client connects to the address.
Does FitLLM need an API key?
No. FitLLM completed a full MCP handshake with us as an anonymous client and listed its tools without asking for anything. All 3 of them are readable on this page. This is what we observed, not what the docs claim.
Is FitLLM open source?
Yes — it is published under the MIT licence, written in JavaScript, 7 stars on GitHub and 27 open issues. The source link is on this page, so you can read exactly what it does with your data before you connect it.