mcpbeat Sign in

Auto Arena Skill for Claude

> Automatically evaluate and compare multiple AI models or agents without pre-existing test data. Generates test queries from a task description, collects responses from all target endpoints, auto-generates evaluation rubrics, runs pairwise comparisons via a judge model, and produces win-rate rankings with reports and charts. Supports checkpoint resume, incremental endpoint addition, and judge model hot-swap. Use when the user asks to compare, benchmark, or rank multiple models or agents on a custom task, or run an arena-style evaluation.

2k tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
763
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/agentscope-ai/OpenJudge --skill auto-arena

The instruction itself

22 sections, as written by the author

Auto Arena Skill

End-to-end automated model comparison using the OpenJudge AutoArenaPipeline:

  • Generate queries — LLM creates diverse test queries from task description
  • Collect responses — query all target endpoints concurrently
  • Generate rubrics — LLM produces evaluation criteria from task + sample queries
  • Pairwise evaluation — judge model compares every model pair (with position-bias swap)
  • Analyze & rank — compute win rates, win matrix, and rankings
  • Report & charts — Markdown report + win-rate bar chart + optional matrix heatmap

Prerequisites

# Install OpenJudge
pip install py-openjudge

# Extra dependency for auto_arena (chart generation)
pip install matplotlib

Gather from user before running

| Info | Required? | Notes |

|------|-----------|-------|

| Task description | Yes | What the models/agents should do (set in config YAML) |

| Target endpoints | Yes | At least 2 OpenAI-compatible endpoints to compare |

| Judge endpoint | Yes | Strong model for pairwise evaluation (e.g. gpt-4, qwen-max) |

| API keys | Yes | Env vars: OPENAI_API_KEY, DASHSCOPE_API_KEY, etc. |

| Number of queries | No | Default: 20 |

| Seed queries | No | Example queries to guide generation style |

| System prompts | No | Per-endpoint system prompts |

| Output directory | No | Default: ./evaluation_results |

| Report language | No | "zh" (default) or "en" |

Quick start

CLI

# Run evaluation
python -m cookbooks.auto_arena --config config.yaml --save

# Use pre-generated queries
python -m cookbooks.auto_arena --config config.yaml \
  --queries_file queries.json --save

# Start fresh, ignore checkpoint
python -m cookbooks.auto_arena --config config.yaml --fresh --save

# Re-run only pairwise evaluation with new judge model
# (keeps queries, responses, and rubrics)
python -m cookbooks.auto_arena --config config.yaml --rerun-judge --save

Python API

import asyncio
from cookbooks.auto_arena.auto_arena_pipeline import AutoArenaPipeline

async def main():
    pipeline = AutoArenaPipeline.from_config("config.yaml")
    result = await pipeline.evaluate()

    print(f"Best model: {result.best_pipeline}")
    for rank, (model, win_rate) in enumerate(result.rankings, 1):
        print(f"{rank}. {model}: {win_rate:.1%}")

asyncio.run(main())

Minimal Python API (no config file)

import asyncio
from cookbooks.auto_arena.auto_arena_pipeline import AutoArenaPipeline
from cookbooks.auto_arena.schema import OpenAIEndpoint

async def main():
    pipeline = AutoArenaPipeline(
        task_description="Customer service chatbot for e-commerce",
        target_endpoints={
            "gpt4": OpenAIEndpoint(
                base_url="https://api.openai.com/v1",
                api_key="sk-...",
                model="gpt-4",
            ),
            "qwen": OpenAIEndpoint(
                base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
                api_key="sk-...",
                model="qwen-max",
            ),
        },
        judge_endpoint=OpenAIEndpoint(
            base_url="https://api.openai.com/v1",
            api_key="sk-...",
            model="gpt-4",
        ),
        num_queries=20,
    )
    result = await pipeline.evaluate()
    print(f"Best: {result.best_pipeline}")

asyncio.run(main())

CLI options

| Flag | Default | Description |

|------|---------|-------------|

| --config | — | Path to YAML configuration file (required) |

| --output_dir | config value | Override output directory |

| --queries_file | — | Path to pre-generated queries JSON (skip generation) |

| --save | False | Save results to file |

| --fresh | False | Start fresh, ignore checkpoint |

| --rerun-judge | False | Re-run pairwise evaluation only (keep queries/responses/rubrics) |

Minimal config file

task:
  description: "Academic GPT assistant for research and writing tasks"

target_endpoints:
  model_v1:
    base_url: "https://api.openai.com/v1"
    api_key: "${OPENAI_API_KEY}"
    model: "gpt-4"
  model_v2:
    base_url: "https://api.openai.com/v1"
    api_key: "${OPENAI_API_KEY}"
    model: "gpt-3.5-turbo"

judge_endpoint:
  base_url: "https://api.openai.com/v1"
  api_key: "${OPENAI_API_KEY}"
  model: "gpt-4"

Full config reference

task

| Field | Required | Description |

|-------|----------|-------------|

| description | Yes | Clear description of the task models will be tested on |

| scenario | No | Usage scenario for additional context |

target_endpoints.\<name\>

| Field | Default | Description |

|-------|---------|-------------|

| base_url | — | API base URL (required) |

| api_key | — | API key, supports ${ENV_VAR} (required) |

| model | — | Model name (required) |

| system_prompt | — | System prompt for this endpoint |

| extra_params | — | Extra API params (e.g. temperature, max_tokens) |

judge_endpoint

Same fields as target_endpoints.<name>. Use a strong model (e.g. gpt-4, qwen-max) with low temperature (~0.1) for consistent judgments.

query_generation

| Field | Default | Description |

|-------|---------|-------------|

| num_queries | 20 | Total number of queries to generate |

| seed_queries | — | Example queries to guide generation |

| categories | — | Query categories with weights for stratified generation |

| endpoint | judge endpoint | Custom endpoint for query generation |

| queries_per_call | 10 | Queries generated per API call (1–50) |

| num_parallel_batches | 3 | Parallel generation batches |

| temperature | 0.9 | Sampling temperature (0.0–2.0) |

| top_p | 0.95 | Top-p sampling (0.0–1.0) |

| max_similarity | 0.85 | Dedup similarity threshold (0.0–1.0) |

| enable_evolution | false | Enable Evol-Instruct complexity evolution |

| evolution_rounds | 1 | Evolution rounds (0–3) |

| complexity_levels | ["constraints", "reasoning", "edge_cases"] | Evolution strategies |

evaluation

| Field | Default | Description |

|-------|---------|-------------|

| max_concurrency | 10 | Max concurrent API requests |

| timeout | 60 | Request timeout in seconds |

| retry_times | 3 | Retry attempts for failed requests |

output

| Field | Default | Description |

|-------|---------|-------------|

| output_dir | ./evaluation_results | Output directory |

| save_queries | true | Save generated queries |

| save_responses | true | Save model responses |

| save_details | true | Save detailed results |

report

| Field | Default | Description |

|-------|---------|-------------|

| enabled | false | Enable Markdown report generation |

| language | "zh" | Report language: "zh" or "en" |

| include_examples | 3 | Examples per section (1–10) |

| chart.enabled | true | Generate win-rate chart |

| chart.orientation | "horizontal" | "horizontal" or "vertical" |

| chart.show_values | true | Show values on bars |

| chart.highlight_best | true | Highlight best model |

| chart.matrix_enabled | false | Generate win-rate matrix heatmap |

| chart.format | "png" | Chart format: "png", "svg", or "pdf" |

Interpreting results

Win rate: percentage of pairwise comparisons a model wins. Each pair is evaluated in both orders (original + swapped) to eliminate position bias.

Rankings example:

  1. gpt4_baseline       [################----] 80.0%
  2. qwen_candidate      [############--------] 60.0%
  3. llama_finetuned      [##########----------] 50.0%

Win matrix: win_matrix[A][B] = how often model A beats model B across all queries.

Checkpoint & resume

The pipeline saves progress after each step. Interrupted runs resume automatically:

  • --fresh — ignore checkpoint, start from scratch
  • --rerun-judge — re-run only the pairwise evaluation step (useful when switching judge models); keeps queries, responses, and rubrics intact
  • Adding new endpoints to config triggers incremental response collection; existing responses are preserved

Output files

evaluation_results/
├── evaluation_results.json     # Rankings, win rates, win matrix
├── evaluation_report.md        # Detailed Markdown report (if enabled)
├── win_rate_chart.png          # Win-rate bar chart (if enabled)
├── win_rate_matrix.png         # Matrix heatmap (if matrix_enabled)
├── queries.json                # Generated test queries
├── responses.json              # All model responses
├── rubrics.json                # Generated evaluation rubrics
├── comparison_details.json     # Pairwise comparison details
└── checkpoint.json             # Pipeline checkpoint

API key by model

| Model prefix | Environment variable |

|-------------|---------------------|

| gpt-*, o1-*, o3-* | OPENAI_API_KEY |

| claude-* | ANTHROPIC_API_KEY |

| qwen-*, dashscope/* | DASHSCOPE_API_KEY |

| deepseek-* | DEEPSEEK_API_KEY |

| Custom endpoint | set api_key + base_url in config |

Additional resources

  • Full config examples: cookbooks/auto_arena/examples/
  • Documentation: Auto Arena Guide

Other skills for the same job

different authors, same section of the catalogue
XLSX
by anthropics
vendor ×15

Comprehensive spreadsheet creation, editing, and analysis with support for formulas, formatting, data analysis, and visualization. When Claude needs to work with spreadsheets (.xlsx, .xlsm, .csv, .tsv, etc) for: (1) Creating new spreadsheets with formulas and formatting, (2) Reading or analyzing data, (3) Modify existing spreadsheets while preserving formulas, (4) Data analysis and visualization in spreadsheets, or (5) Recalculating formulas

5k tokens scripts
XLSX
by w95
×7

Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .csv, or .tsv file (e.g., adding columns, computing formulas, formatting, charting, cleaning messy data); create a new spreadsheet from scratch or from other data sources; or convert between tabular file formats. Trigger especially when the user references a spreadsheet file by name or path — even casually (like \"the xlsx in my downloads\") — and wants something done to it or produced from it. Also trigger for cleaning or restructuring messy tabular data files (malformed rows, misplaced headers, junk data) into proper spreadsheets. The deliverable must be a spreadsheet file. Do NOT trigger when the primary deliverable is a Word document, HTML report, standalone Python script, database pipeline, or Google Sheets API integration, even if tabular data is involved.

3k tokens
Raffle Winner Picker
by frostant
×5

Picks random winners from lists, spreadsheets, or Google Sheets for giveaways, raffles, and contests. Ensures fair, unbiased selection with transparency.

949 tokens
Fda Database
by christophacham
×4

Query openFDA API for drugs, devices, adverse events, recalls, regulatory submissions (510k, PMA), substance identification (UNII), for FDA regulatory data analysis and safety research.

32k tokens scripts
Matlab
by christophacham
×4

MATLAB and GNU Octave numerical computing for matrix operations, data analysis, visualization, and scientific computing. Use when writing MATLAB/Octave scripts for linear algebra, signal processing, image processing, differential equations, optimization, statistics, or creating scientific visualizations. Also use when the user needs help with MATLAB syntax, functions, or wants to convert between MATLAB and Python code. Scripts can be executed with MATLAB or the open-source GNU Octave interpreter.

25k tokens
Umap Learn
by ComeOnOliver
×4

UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.

14k tokens
D3 Viz
by chrisvoncsefalvay
×3

Creating interactive data visualisations using d3.js. This skill should be used when creating custom charts, graphs, network diagrams, geographic visualisations, or any complex SVG-based data visualisation that requires fine-grained control over visual elements, transitions, or interactions. Use this for bespoke visualisations beyond standard charting libraries, whether in React, Vue, Svelte, vanilla JavaScript, or any other environment.

20k tokens
Alphafold Database
by christophacham
×3

Access AlphaFold 200M+ AI-predicted protein structures. Retrieve structures by UniProt ID, download PDB/mmCIF files, analyze confidence metrics (pLDDT, PAE), for drug discovery and structural biology.

7k tokens

How to use it

Copy the folder

Take agentscope-ai/auto-arena from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.

Install what it needs

The instructions reference pip. Without those the skill loads but fails at the first command.