Inferbench runs on your own machine — the client starts it, so there is no endpoint to ping. 224 installs a week from pypi. Last commit 25 Aug 2026.
Benchmarks local LLM inference speed (tokens/sec) on your own hardware via MCP tools.
We read the source, 11 h ago · rules 3dff92dd89df
What this server is able to do. For an MCP server this is often the job itself — a terminal server runs commands because that is what it is for. Listed so you know what you are plugging in, not as an accusation.
этот файл ставится пользователю, но в репозитории его нет
subprocess.run( # noqa: S603, S607 -- fixed argv, binary name only, no shell
execFileSync(BINARY, ["--version"], { stdio: "ignore" });
Is this your server and something here is wrong? Tell us — corrections are free and do not require a plan.
We found places where it runs commands, builds paths or queries from values it is given. None of that is a flaw by itself — it becomes one when the code changes, and code changes quietly between releases. We re-read it on every one.
This server runs on your own machine — install it with the package manager and the client starts it for you. Package name taken from the official registry entry.
claude mcp add inferbench -- uvx inferbench-cli
{
"mcpServers": {
"inferbench": {
"args": [
"inferbench-cli"
],
"command": "uvx"
}
}
}
[mcp_servers.inferbench]
command = "uvx"
args = ["inferbench-cli"]
{
"mcpServers": {
"inferbench": {
"args": [
"inferbench-cli"
],
"command": "uvx"
}
}
}
{
"mcpServers": {
"inferbench": {
"args": [
"inferbench-cli"
],
"command": "uvx"
}
}
}
Benchmark local LLM models — speed, quality & hardware fitness verdict from any MCP client
Dated SIU, the benchmark price of AI inference work: one free tool, three paid via x402.
GPU and LLM inference benchmarks, hardware evidence, deployment recommendations, and launch configs.
OpenAI-compatible LLM MCP (7 tools); chat via balance key or x402 USDC on Base
MarBoba IDP catalog as MCP tools: projects, APIs, runbooks, on-call, SLOs. Bring your own LLM.
Will a local LLM run on your hardware? GGUF quant, buy-vs-rent-vs-API cost, used-GPU prices.
PathCourse Health inference SKUs as x402 paid MCP tools, billed per call in USDC on Base.
76% fewer MCP tokens in our standard benchmark. Same upstream call. Exact recovery.
Answers built from our own checks of this server.