mcpbeat

The Aggregate — LLM benchmark aggregate MCP Server

ai.theaggregate/the-aggregate
answering

The Aggregate — LLM benchmark aggregate is answering right now. Last checked 7 min ago. It exposes 8 tools.

Fused LLM rankings: one IRT/Elo scale across ~5,000 public benchmark leaderboards, updated daily.

Uptime history 43 hours of history
43 hours agonow
100.0%
Uptime 24h
91 of 91 checks
8
Tools
read from the server
127 ms
Response time
average over 24h
open, no key
Access
streamable-http

Connect this server

Endpoint below is the one we actually reach during checks — not the one copied from a README. Last verified 7 min ago.

run in your terminal
claude mcp add the-aggregate --transport http https://theaggregate.ai/mcp
~/Library/Application Support/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "the-aggregate": {
      "url": "https://theaggregate.ai/mcp"
    }
  }
}
~/.codex/config.toml
[mcp_servers.the-aggregate]
url = "https://theaggregate.ai/mcp"
.cursor/mcp.json
{
  "mcpServers": {
    "the-aggregate": {
      "url": "https://theaggregate.ai/mcp"
    }
  }
}
.vscode/mcp.json
{
  "mcpServers": {
    "the-aggregate": {
      "url": "https://theaggregate.ai/mcp"
    }
  }
}

Available tools 8

Read directly from the server with tools/list, grouped by what they act on. If a tool disappears, we record the date.

about
about_the_aggregate
What this data is: how the IRT fusion works, current coverage counts, update cadence, and how to cite it.
benchmark
get_benchmark
One benchmark in depth: what it measures, the original source leaderboard URL, IRT stats (difficulty, noise, model coverage), skill weights, and the current top models on it.
benchmarks
search_benchmarks
Find benchmarks in the aggregate by (partial) name. Returns model coverage, difficulty on the Elo scale, and the benchmark page URL.
compare
compare_models
Head-to-head between 2-4 models: aggregate ranks, Elo gap with a significance note based on the standard errors, and notable benchmarks they share.
leaderboard
get_leaderboard
Top of the cross-benchmark aggregate ranking: every model placed on one Elo scale by an IRT model fit over ~5,000 public benchmark leaderboards. Supports paging via limit/offset.
model
get_model
One model in depth: aggregate rank, Elo with standard error, provider, what it is, cost per task where known, and its most notable benchmark results (with percentiles).
models
search_models
Find ranked models by (partial) name or provider. Returns rank, Elo and the model page URL.
prediction
get_prediction_duel
Guesswork — the public prediction duel: every day frontier LLMs and The Aggregate's own IRT model predict newly scraped benchmark scores before seeing them, and the errors are scored. Returns the current monthly standings, wins and losses included.

Endpoints

URLTransportStateLatencyChecked
https://theaggregate.ai/mcp streamable-http answering 204 ms 7 min ago

The Aggregate — LLM benchmark aggregate — questions

Answers built from our own checks of this server.

What can The Aggregate — LLM benchmark aggregate do?
It exposes 8 tools, read directly from the server on our last check. Among them: about_the_aggregate, compare_models, get_benchmark, get_leaderboard, get_model, get_prediction_duel and 2 more. The full list with descriptions is on this page — we take it from the server itself via tools/list, not from a README. How MCP servers expose tools in the first place →
Is The Aggregate — LLM benchmark aggregate working right now?
We send a real MCP handshake every 15 minutes. Over the last 24 hours 91 of 91 checks got a reply (100.0%), average response time 127 ms. The bar chart above shows every period we have measured.
How do I connect The Aggregate — LLM benchmark aggregate?
Copy the ready config from this page — we generate it for Claude Code, Claude Desktop, Codex, Cursor and VS Code, each with the file path that client actually reads. It is a remote server, so there is nothing to install — the client connects to the address.
Does The Aggregate — LLM benchmark aggregate need an API key?
No. The Aggregate — LLM benchmark aggregate completed a full MCP handshake with us as an anonymous client and listed its tools without asking for anything. All 8 of them are readable on this page. This is what we observed, not what the docs claim.
How fast is The Aggregate — LLM benchmark aggregate?
It answers our handshake in 127 ms on average, which is faster than 78% of all working MCP servers we measure. That puts it in the quick quarter of the ecosystem. The comparison comes from our own checks across the whole registry, every 15 minutes.