mcpbeat Sign in

HTTP Load Profiler Skill for Claude

Run stepped HTTP load tests with ab/wrk, ramping concurrency levels to collect p50/p90/p99 latency, detect performance inflection points, and recommend optimal concurrency. Triggered by requests like 'load test this URL', 'benchmark my API', 'find the max concurrency', or mentions of p99 latency, throughput saturation, or capacity planning.

5k tokens
context cost
the whole folder, loaded on every use
3
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
4468
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/zebbern/claude-code-guide --skill http-load-profiler

What comes with it

15 729 bytes besides the instruction
LICENSE
scripts/http_benchmark.py

The instruction itself

10 sections, as written by the author

HTTP Load Profiler — Stepped Concurrency Load Test + Inflection Point Analysis

Run stepped concurrency load tests against HTTP services, automatically collect latency percentiles, and detect performance inflection points.

Features

  • Dual engine support: Auto-detects wrk (preferred) or ab (Apache Bench); manual override available
  • Stepped concurrency: Ramps up through user-defined concurrency levels (default: 1 → 10 → 50 → 100 → 200 → 500)
  • Latency percentiles: Collects p50 / p90 / p99 latency at each level
  • Inflection point detection: Automatically identifies four types of performance inflection points
  • p99 latency accelerating (increase exceeds 2x the previous step's increase)
  • Throughput efficiency dropping significantly (RPS per connection drops > 40%)
  • Throughput saturated while latency spikes (RPS growth < 10%, p99 growth > 50%)
  • Error rate surging (exceeds 1% and doubles from previous step)
  • Optimal concurrency recommendation: Automatically suggests the best concurrency level based on inflection points
  • Zero Python dependencies: Pure standard library implementation

Quick Start

# Basic usage — run default stepped load test against target URL
python3 scripts/http_benchmark.py https://example.com/api/health

# Custom concurrency steps and duration per step
python3 scripts/http_benchmark.py https://example.com/api/health -s 5,20,50,100,300 -d 15

# Specify ab as the engine
python3 scripts/http_benchmark.py https://example.com/ -t ab

# JSON-only output (for programmatic parsing)
python3 scripts/http_benchmark.py https://example.com/api/health --json

# Use ab with a specific number of requests per step
python3 scripts/http_benchmark.py https://example.com/ -t ab -n 5000

# Save JSON report to a file
python3 scripts/http_benchmark.py https://example.com/api/health --json > report.json

Parameters

| Parameter | Short | Default | Description |

|-----------|-------|---------|-------------|

| url | — | (required) | Target URL (http:// or https://) |

| --steps | -s | 1,10,50,100,200,500 | Concurrency steps (comma-separated positive integers) |

| --duration | -d | 10 | Duration per step in seconds (used directly by wrk; ab estimates request count from this) |

| --requests | -n | concurrency×100 | Total requests per step when using ab |

| --tool | -t | auto-detect | Specify load testing tool: wrk or ab |

| --threads | — | min(concurrency, CPU cores) | Thread count for wrk |

| --json | — | false | Output JSON only |

Output Format

Human-readable (default)

Tool: wrk
Target URL: https://example.com/api/health
Concurrency steps: [1, 10, 50, 100, 200, 500]
Duration per step: 10s

----------------------------------------------------------------------------------
  Conc. |        RPS |  Avg(ms) |  P50(ms) |  P90(ms) |  P99(ms) |   Errors | Inflection
----------------------------------------------------------------------------------
     1 |      245.3 |      4.1 |      3.8 |      5.2 |      8.1 |   0.00% |
    10 |     2301.5 |      4.3 |      4.0 |      5.8 |      9.3 |   0.00% |
    50 |     9876.2 |      5.1 |      4.6 |      7.2 |     12.5 |   0.00% |
   100 |    14523.1 |      6.9 |      5.8 |     10.3 |     22.7 |   0.00% |
   200 |    15102.3 |     13.2 |     10.1 |     22.5 |     58.3 |   0.12% |  ◀
   500 |    14890.5 |     33.6 |     28.3 |     55.2 |    132.1 |   1.35% |  ◀
----------------------------------------------------------------------------------

Inflection point analysis:
  ▶ Concurrency 200:
    - p99 latency accelerating: 22.7ms → 58.3ms (increase 35.6ms, previous step increase 10.2ms)
    - Throughput saturated with latency spike: RPS grew only 3.9% while p99 latency grew 156.8%
  ▶ Concurrency 500:
    - Error rate surging: 0.12% → 1.35%

Recommended optimal concurrency: 100

JSON format (--json)

{
  "url": "https://example.com/api/health",
  "tool": "wrk",
  "duration_per_step": 10,
  "steps": [
    {
      "concurrency": 1,
      "rps": 245.3,
      "avg_latency_ms": 4.1,
      "p50_ms": 3.8,
      "p90_ms": 5.2,
      "p99_ms": 8.1,
      "total_requests": 2453,
      "errors": 0
    }
  ],
  "inflection_points": [
    {
      "concurrency": 200,
      "step_index": 4,
      "reasons": ["p99 latency accelerating: ..."]
    }
  ],
  "recommended_concurrency": 100
}

Prerequisites

At least one of wrk or ab must be installed:

# Ubuntu / Debian
sudo apt-get install wrk          # recommended
sudo apt-get install apache2-utils # ab

# macOS
brew install wrk
# ab is pre-installed on macOS

Inflection Point Detection Algorithm

For each concurrency level, the following metrics are compared against the two preceding levels:

  • p99 latency acceleration: Triggers when the current p99 increase exceeds 2x the previous step's increase
  • Throughput efficiency: Triggers when RPS per connection drops > 40% from the previous step
  • Saturation detection: Triggers when RPS growth < 10% while p99 growth > 50%
  • Error rate: Triggers when rate exceeds 1% and doubles from the previous step

Recommended optimal concurrency: The concurrency level one step before the first inflection point. If no inflection point is found, the level with the highest RPS is selected.

Important Notes

  • Load testing generates real traffic against the target service — do not run against production services without authorization
  • wrk provides more accurate latency percentiles than ab (wrk uses HdrHistogram)
  • ab does not support a duration parameter; the script approximates timing via total request count
  • A minimum of 10 seconds per step is recommended for stable results

Other skills for the same job

different authors, same section of the catalogue
AgentDB Learning Plugins
by Microck
×1

Create and train AI learning plugins with AgentDB's 9 reinforcement learning algorithms. Includes Decision Transformer, Q-Learning, SARSA, Actor-Critic, and more. Use when building self-learning agents, implementing RL, or optimizing agent behavior through experience.

3k tokens
Agentdb Learning Plugins
by ComeOnOliver
×1

Create and train AI learning plugins with AgentDB's 9 reinforcement learning algorithms. Includes Decision Transformer, Q-Learning, SARSA, Actor-Critic, and more. Use when building self-learning agents, implementing RL, or optimizing agent behavior through experience.

7k tokens
Langchain Architecture
by wshobson

Design LLM applications using LangChain 1.x and LangGraph for agents, memory, and tool integration. Use when building LangChain applications, implementing AI agents, or creating complex LLM workflows.

5k tokens
Nemoclaw Maintainer Policies
by NVIDIA
vendor

Provide read-only NemoClaw maintainer policy. Use for questions about Issue Type, labels, Project fields, release labels, triage, duplicates, blocked items, and maintainer decisions. Trigger keywords - maintainer policy, workflow policy, project workflow, issue type, labels, label taxonomy, needs labels, project status, blocked issue, duplicate issue, daily release label, release train, triage policy.

25k tokens
I4h Workflow Dataset Teleop
by NVIDIA
vendor

Record episodes for an agentic env via teleoperation (keyboard, SO-ARM leader, or VR) into HDF5. Use when the user wants to teleop or record human demos.

5k tokens
Detecting Data Anomalies
by foryourhealth111-pixel

| Investigate outliers, rare events, spikes, and suspicious records in datasets. Use as an explicit anomaly-analysis helper when you want concrete anomaly-detection workflow guidance, not generic data validation or end-to-end ML ownership.

4k tokens scripts
LLM Council
by aiwithremy

Run any question, idea, or decision through a council of 5 AI advisors who independently analyze it, peer-review each other anonymously, and synthesize a final verdict. Based on Karpathy's LLM Council methodology. MANDATORY TRIGGERS: 'council this', 'run the council', 'war room this', 'pressure-test this', 'stress-test this', 'debate this'. STRONG TRIGGERS (use when combined with a real decision or tradeoff): 'should I X or Y', 'which option', 'what would you do', 'is this the right move', 'validate this', 'get multiple perspectives', 'I can't decide', 'I'm torn between'. Do NOT trigger on simple yes/no questions, factual lookups, or casual 'should I' without a meaningful tradeoff (e.g. 'should I use markdown' is not a council question). DO trigger when the user presents a genuine decision with stakes, multiple options, and context that suggests they want it pressure-tested from multiple angles.

5k tokens
Matlab Prepare Signal Data
by matlab

| Use this skill when conditioning, loading, preparing, or labeling signal gaps, remove drift, deoutlier, denoise, resample/align a time base) BEFORE analysis; building a `signalDatastore` pipeline; creating a `labeledSignalSet` for Signal Labeler; deriving labels (filename, folder, in-file, ROI, time-frequency ROI); stratified train/val/test splits; framing long signals; parallel processing; and shaping datastore output for `trainnet`. Triggers include "clean up this signal", "remove drift / detrend", "fill gaps", "remove spikes / outliers", "denoise", "resample to a uniform rate", "align channels", "labels from filenames", "stratified split", "prepare for Signal Labeler", and function names like `fillgaps`, `fillmissing`, `detrend`, `filloutliers`, `smoothdata`, `resample`, `synchronize`, `signalDatastore`, `labeledSignalSet`, `filenames2labels`, `folders2labels`, `splitlabels`, `framesig`, `framelbl`, `createDatastores`.

53k tokens

How to use it

Copy the folder

Take zebbern/http-load-profiler from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.

Install what it needs

The instructions reference brew, apt. Without those the skill loads but fails at the first command.