> Automatically evaluate and compare multiple AI models or agents without pre-existing test data. Generates test queries from a task description, collects responses from all target endpoints, auto-generates evaluation rubrics, runs pairwise comparisons via a judge model, and produces win-rate rankings with reports and charts. Supports checkpoint resume, incremental endpoint addition, and judge model hot-swap. Use when the user asks to compare, benchmark, or rank multiple models or agents on a custom task, or run an arena-style evaluation.
npx skills add https://github.com/agentscope-ai/OpenJudge --skill auto-arena
End-to-end automated model comparison using the OpenJudge AutoArenaPipeline:
# Install OpenJudge
pip install py-openjudge
# Extra dependency for auto_arena (chart generation)
pip install matplotlib
| Info | Required? | Notes |
|------|-----------|-------|
| Task description | Yes | What the models/agents should do (set in config YAML) |
| Target endpoints | Yes | At least 2 OpenAI-compatible endpoints to compare |
| Judge endpoint | Yes | Strong model for pairwise evaluation (e.g. gpt-4, qwen-max) |
| API keys | Yes | Env vars: OPENAI_API_KEY, DASHSCOPE_API_KEY, etc. |
| Number of queries | No | Default: 20 |
| Seed queries | No | Example queries to guide generation style |
| System prompts | No | Per-endpoint system prompts |
| Output directory | No | Default: ./evaluation_results |
| Report language | No | "zh" (default) or "en" |
# Run evaluation
python -m cookbooks.auto_arena --config config.yaml --save
# Use pre-generated queries
python -m cookbooks.auto_arena --config config.yaml \
--queries_file queries.json --save
# Start fresh, ignore checkpoint
python -m cookbooks.auto_arena --config config.yaml --fresh --save
# Re-run only pairwise evaluation with new judge model
# (keeps queries, responses, and rubrics)
python -m cookbooks.auto_arena --config config.yaml --rerun-judge --save
import asyncio
from cookbooks.auto_arena.auto_arena_pipeline import AutoArenaPipeline
async def main():
pipeline = AutoArenaPipeline.from_config("config.yaml")
result = await pipeline.evaluate()
print(f"Best model: {result.best_pipeline}")
for rank, (model, win_rate) in enumerate(result.rankings, 1):
print(f"{rank}. {model}: {win_rate:.1%}")
asyncio.run(main())
import asyncio
from cookbooks.auto_arena.auto_arena_pipeline import AutoArenaPipeline
from cookbooks.auto_arena.schema import OpenAIEndpoint
async def main():
pipeline = AutoArenaPipeline(
task_description="Customer service chatbot for e-commerce",
target_endpoints={
"gpt4": OpenAIEndpoint(
base_url="https://api.openai.com/v1",
api_key="sk-...",
model="gpt-4",
),
"qwen": OpenAIEndpoint(
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
api_key="sk-...",
model="qwen-max",
),
},
judge_endpoint=OpenAIEndpoint(
base_url="https://api.openai.com/v1",
api_key="sk-...",
model="gpt-4",
),
num_queries=20,
)
result = await pipeline.evaluate()
print(f"Best: {result.best_pipeline}")
asyncio.run(main())
| Flag | Default | Description |
|------|---------|-------------|
| --config | — | Path to YAML configuration file (required) |
| --output_dir | config value | Override output directory |
| --queries_file | — | Path to pre-generated queries JSON (skip generation) |
| --save | False | Save results to file |
| --fresh | False | Start fresh, ignore checkpoint |
| --rerun-judge | False | Re-run pairwise evaluation only (keep queries/responses/rubrics) |
task:
description: "Academic GPT assistant for research and writing tasks"
target_endpoints:
model_v1:
base_url: "https://api.openai.com/v1"
api_key: "${OPENAI_API_KEY}"
model: "gpt-4"
model_v2:
base_url: "https://api.openai.com/v1"
api_key: "${OPENAI_API_KEY}"
model: "gpt-3.5-turbo"
judge_endpoint:
base_url: "https://api.openai.com/v1"
api_key: "${OPENAI_API_KEY}"
model: "gpt-4"
| Field | Required | Description |
|-------|----------|-------------|
| description | Yes | Clear description of the task models will be tested on |
| scenario | No | Usage scenario for additional context |
| Field | Default | Description |
|-------|---------|-------------|
| base_url | — | API base URL (required) |
| api_key | — | API key, supports ${ENV_VAR} (required) |
| model | — | Model name (required) |
| system_prompt | — | System prompt for this endpoint |
| extra_params | — | Extra API params (e.g. temperature, max_tokens) |
Same fields as target_endpoints.<name>. Use a strong model (e.g. gpt-4, qwen-max) with low temperature (~0.1) for consistent judgments.
| Field | Default | Description |
|-------|---------|-------------|
| num_queries | 20 | Total number of queries to generate |
| seed_queries | — | Example queries to guide generation |
| categories | — | Query categories with weights for stratified generation |
| endpoint | judge endpoint | Custom endpoint for query generation |
| queries_per_call | 10 | Queries generated per API call (1–50) |
| num_parallel_batches | 3 | Parallel generation batches |
| temperature | 0.9 | Sampling temperature (0.0–2.0) |
| top_p | 0.95 | Top-p sampling (0.0–1.0) |
| max_similarity | 0.85 | Dedup similarity threshold (0.0–1.0) |
| enable_evolution | false | Enable Evol-Instruct complexity evolution |
| evolution_rounds | 1 | Evolution rounds (0–3) |
| complexity_levels | ["constraints", "reasoning", "edge_cases"] | Evolution strategies |
| Field | Default | Description |
|-------|---------|-------------|
| max_concurrency | 10 | Max concurrent API requests |
| timeout | 60 | Request timeout in seconds |
| retry_times | 3 | Retry attempts for failed requests |
| Field | Default | Description |
|-------|---------|-------------|
| output_dir | ./evaluation_results | Output directory |
| save_queries | true | Save generated queries |
| save_responses | true | Save model responses |
| save_details | true | Save detailed results |
| Field | Default | Description |
|-------|---------|-------------|
| enabled | false | Enable Markdown report generation |
| language | "zh" | Report language: "zh" or "en" |
| include_examples | 3 | Examples per section (1–10) |
| chart.enabled | true | Generate win-rate chart |
| chart.orientation | "horizontal" | "horizontal" or "vertical" |
| chart.show_values | true | Show values on bars |
| chart.highlight_best | true | Highlight best model |
| chart.matrix_enabled | false | Generate win-rate matrix heatmap |
| chart.format | "png" | Chart format: "png", "svg", or "pdf" |
Win rate: percentage of pairwise comparisons a model wins. Each pair is evaluated in both orders (original + swapped) to eliminate position bias.
Rankings example:
1. gpt4_baseline [################----] 80.0%
2. qwen_candidate [############--------] 60.0%
3. llama_finetuned [##########----------] 50.0%
Win matrix: win_matrix[A][B] = how often model A beats model B across all queries.
The pipeline saves progress after each step. Interrupted runs resume automatically:
--fresh — ignore checkpoint, start from scratch--rerun-judge — re-run only the pairwise evaluation step (useful when switching judge models); keeps queries, responses, and rubrics intactevaluation_results/
├── evaluation_results.json # Rankings, win rates, win matrix
├── evaluation_report.md # Detailed Markdown report (if enabled)
├── win_rate_chart.png # Win-rate bar chart (if enabled)
├── win_rate_matrix.png # Matrix heatmap (if matrix_enabled)
├── queries.json # Generated test queries
├── responses.json # All model responses
├── rubrics.json # Generated evaluation rubrics
├── comparison_details.json # Pairwise comparison details
└── checkpoint.json # Pipeline checkpoint
| Model prefix | Environment variable |
|-------------|---------------------|
| gpt-*, o1-*, o3-* | OPENAI_API_KEY |
| claude-* | ANTHROPIC_API_KEY |
| qwen-*, dashscope/* | DASHSCOPE_API_KEY |
| deepseek-* | DEEPSEEK_API_KEY |
| Custom endpoint | set api_key + base_url in config |
Comprehensive spreadsheet creation, editing, and analysis with support for formulas, formatting, data analysis, and visualization. When Claude needs to work with spreadsheets (.xlsx, .xlsm, .csv, .tsv, etc) for: (1) Creating new spreadsheets with formulas and formatting, (2) Reading or analyzing data, (3) Modify existing spreadsheets while preserving formulas, (4) Data analysis and visualization in spreadsheets, or (5) Recalculating formulas
Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .csv, or .tsv file (e.g., adding columns, computing formulas, formatting, charting, cleaning messy data); create a new spreadsheet from scratch or from other data sources; or convert between tabular file formats. Trigger especially when the user references a spreadsheet file by name or path — even casually (like \"the xlsx in my downloads\") — and wants something done to it or produced from it. Also trigger for cleaning or restructuring messy tabular data files (malformed rows, misplaced headers, junk data) into proper spreadsheets. The deliverable must be a spreadsheet file. Do NOT trigger when the primary deliverable is a Word document, HTML report, standalone Python script, database pipeline, or Google Sheets API integration, even if tabular data is involved.
Picks random winners from lists, spreadsheets, or Google Sheets for giveaways, raffles, and contests. Ensures fair, unbiased selection with transparency.
Query openFDA API for drugs, devices, adverse events, recalls, regulatory submissions (510k, PMA), substance identification (UNII), for FDA regulatory data analysis and safety research.
MATLAB and GNU Octave numerical computing for matrix operations, data analysis, visualization, and scientific computing. Use when writing MATLAB/Octave scripts for linear algebra, signal processing, image processing, differential equations, optimization, statistics, or creating scientific visualizations. Also use when the user needs help with MATLAB syntax, functions, or wants to convert between MATLAB and Python code. Scripts can be executed with MATLAB or the open-source GNU Octave interpreter.
UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.
Creating interactive data visualisations using d3.js. This skill should be used when creating custom charts, graphs, network diagrams, geographic visualisations, or any complex SVG-based data visualisation that requires fine-grained control over visual elements, transitions, or interactions. Use this for bespoke visualisations beyond standard charting libraries, whether in React, Vue, Svelte, vanilla JavaScript, or any other environment.
Access AlphaFold 200M+ AI-predicted protein structures. Retrieve structures by UniProt ID, download PDB/mmCIF files, analyze confidence metrics (pLDDT, PAE), for drug discovery and structural biology.
Take agentscope-ai/auto-arena from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip.
Without those the skill loads but fails at the first command.