mcpbeat Sign in

SEO Schema Agent Skill

Detect existing JSON-LD structured data on a page, validate against Google's rich-result requirements, and generate missing schema markup (Article, Product, LocalBusiness, FAQPage, BreadcrumbList). Produces paste-ready JSON-LD script blocks. Use when the user asks for "schema markup", "structured data", "JSON-LD", "rich results", "schema validation", or "fix the schema on this page".

5k tokens
context cost
the whole folder, loaded on every use
7
files
instructions only
0
copies elsewhere
how many repositories repackaged it
108
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/seranking/seo-skills --skill seo-schema

What comes with it

10 474 bytes besides the instruction
references/google-rich-results.md
templates/article.json
templates/breadcrumb-list.json
templates/faq-page.json
templates/local-business.json
templates/product.json

What it tells the agent to use

found in the instruction text
WebFetch fetches pages from the network

The instruction itself

5 sections, as written by the author

> Example output: examples/seo-schema-budgetbytes-slow-cooker-chicken-noodle-soup-20260514/SCHEMA.md

Schema Markup

Detect, validate, and generate Schema.org JSON-LD for a page. Output is paste-ready <script> blocks the user can drop into their CMS or page template, plus a validation report on what's currently present and what's broken.

Prerequisites

  • Required for detect/validate paths: mcp__firecrawl-mcp__firecrawl_scrape (raw HTML access). WebFetch returns markdown only — every <script type="application/ld+json"> block is stripped before the skill ever sees it. Without Firecrawl, the skill can still generate new schema from intent detection (steps 4–6) but cannot detect or validate what's already on the page (steps 2–3, 7).
  • Optional: SE Ranking MCP server (used in step 7 for benchmarking competitor schema).
  • User provides: a target URL. Optionally a hint about page intent ("this is a product page", "this is a how-to") if the URL pattern doesn't make it obvious.

Process

  • Fetch HTML mcp__firecrawl-mcp__firecrawl_scrape (preferred) or degrade
  • Cost note. Firecrawl: 1 credit for the target URL, +10 credits if step 7 (competitor benchmark) runs (1 per top-10 SERP result). User may pass --no-firecrawl to force the degraded path (generate-only mode) for credit conservation.
  • If Firecrawl available: scrape the target URL. For SPAs, pass waitFor: 2000 (or a CSS selector for the main content) so the JS-rendered DOM is captured. Use the response's html for JSON-LD parsing in step 2 and metadata for canonical/robots cross-reference.
  • If Firecrawl unavailable: skip steps 2, 3, and 7 entirely (they all need raw HTML). Steps 4–6 still run — the skill becomes "generate-only", producing recommended JSON-LD blocks from intent detection without comparing to what's on the page. Surface clearly in SCHEMA.md: Existing-schema detection: skipped — Firecrawl required (WebFetch returns markdown only). Install via extensions/firecrawl/install.sh.
  • Even with Firecrawl: if JSON-LD blocks appear only after JS render, flag in the output: "JS-rendered schema may not be detected by all crawlers — server-side render JSON-LD where possible."
  • Detect existing schema (requires Firecrawl HTML from step 1)
  • From the returned html: extract every <script type="application/ld+json"> block.
  • Parse each as JSON. Report syntax errors.
  • List each detected @type.
  • Also detect Microdata (itemscope/itemprop) and RDFa (typeof/property) — flag as legacy and recommend migration to JSON-LD (Google's stated preference).
  • If step 1 degraded: skip this step. Record Existing-schema detection skipped in 01-detected.md.
  • Validate against Google's spec
  • Load references/google-rich-results.md.
  • For each detected @type, check required and recommended properties.
  • Surface common errors: missing @context, dates not in ISO 8601, prices as numbers instead of strings, availability as plain text instead of schema.org URL, telephone not in international format.
  • Detect page intent
  • From URL pattern (/blog/, /products/, /contact/, /how-to/, /faq/).
  • From <title> and <h1> tone.
  • From content signals (numbered list of steps → HowTo; visible Q&A blocks → FAQPage; price + buy button → Product; address + hours → LocalBusiness).
  • If multiple intents detected, generate schema for each.
  • Generate missing JSON-LD
  • For each detected intent without matching valid schema, load the relevant template from templates/:
  • article.json — for editorial/blog content
  • product.json — for product/SKU pages
  • local-business.json — for brick-and-mortar landing pages
  • faq-page.json — for explicit Q&A blocks (gov/health allowlist only — see references/google-rich-results.md)
  • breadcrumb-list.json — for any page with breadcrumb navigation
  • Fill template fields from the live HTML (title → headline, h2s → mainEntity questions, etc.).
  • Mark any field that couldn't be auto-filled as {REPLACE: ...} so the user knows to complete it.
  • Don't generate HowTo — Google retired HowTo rich results in September 2023 (mobile + desktop). The schema can still ship for semantic clarity, but expect zero rich-result uplift; flag this in the recommendation rationale rather than treating HowTo as a live option.
  • Validate generated JSON-LD
  • Re-run the same validation rubric from step 3 on the generated blocks.
  • Surface any required fields still marked {REPLACE: ...}.
  • Optional: benchmark against top SERP results DATA_getSerpResults + mcp__firecrawl-mcp__firecrawl_scrape
  • Identify the page's primary keyword (from <title> or user input).
  • Pull top 10 organic results.
  • If Firecrawl available: scrape each of the top 10 (10 Firecrawl credits). For each, parse JSON-LD blocks from the returned html and list detected @types. This produces real schema data, not inferences from markdown.
  • If Firecrawl unavailable: skip the benchmark — WebFetch's markdown strips all schema blocks, so any "detection" from it would be guesswork. Write Competitor benchmark skipped — Firecrawl required to read JSON-LD from competitor pages. into 04-competitor-benchmark.md.
  • Surface "schema types used by 6+ of the top 10 that this page is missing." High-signal addition list. (Only emitted when benchmark ran.)
  • Synthesise SCHEMA.md
  • Validation report (existing schema, pass/fail per block).
  • Recommended additions (with rationale linking back to step 4 detection or step 7 benchmark).
  • Generated <script> blocks ready to paste.

Output format

Create a folder seo-schema-{target-slug}-{YYYYMMDD}/ with:

seo-schema-{target-slug}-{YYYYMMDD}/
├── 01-detected.md           (existing schema, validation results)
├── 02-recommended.md        (which types this page should add and why)
├── 03-generated/
│   ├── article.jsonld
│   ├── faq-page.jsonld
│   └── ... (per generated type)
├── 04-competitor-benchmark.md  (only if step 7 ran)
└── SCHEMA.md                (deliverable: paste-ready blocks + install instructions)

SCHEMA.md follows this shape:

# Schema Markup: {URL}

> Snapshot dated {YYYY-MM-DD}.

## Currently present
- `Article` — valid ✓
- `BreadcrumbList` — invalid ✗ (missing `position` on item 2)
- ...

## Recommended additions
- `FAQPage` — page has 6 visible Q&A blocks but no FAQ schema. Adding this is eligible for FAQ rich results (subject to Google's 2024+ tightening — see references/google-rich-results.md).
- `HowTo` — ...

## Paste these into the `<head>` of the page

### FAQPage
\`\`\`html
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "FAQPage",
  "mainEntity": [...]
}
</script>
\`\`\`

### ... (per generated block)

## Validation pass
- All generated blocks parse cleanly ✓
- All required fields filled ({n} {REPLACE: ...} placeholders remain — see below)
- {REPLACE: ...} placeholders to fill manually:
  - article.jsonld → `image` (need a hero image URL ≥ 1200×800)
  - ...

## Install
1. Copy each `<script>` block above.
2. Paste into the `<head>` of the relevant page (or the global `<head>` template, gated by page type).
3. Test with [Google's Rich Results Test](https://search.google.com/test/rich-results).
4. Submit the URL to GSC for re-indexing if changes are critical.

Tips

  • JSON-LD goes in <head> or top of <body>. Don't bury it.
  • Test every generated block in Google's Rich Results Test before shipping. The validation in step 3/6 follows the published spec but Google's actual test is authoritative.
  • references/google-rich-results.md is dated. If it's >6 months old when you run this skill, flag staleness in the output and recommend the user verify against current docs.
  • Don't mark up content that isn't visibly on the page. Google penalises hidden-content schema. If a page doesn't actually have FAQs visible, don't generate FAQPage schema.
  • For Article schema, image is required. If the page has no obvious hero image, leave the {REPLACE: hero image URL} placeholder rather than guessing.
  • The skill is read-mostly on the SE Ranking side: zero SE Ranking credits unless step 7 (competitor benchmark) is requested — that adds ~5–10 SE Ranking credits for DATA_getSerpResults. Firecrawl costs are separate: 1 credit for the target URL, +10 credits when step 7 runs.
  • Verify after deploy: once the generated schema is pasted into your CMS and re-deployed, re-run this skill on the same URL — the new run's "Currently present" section reflects the live state and confirms the schema actually rendered (vs sitting in the CMS but not yet pushed). Ad-hoc alternative: invoke seo-firecrawl on the URL and grep META.md for the expected @types.

Other skills for the same job

different authors, same section of the catalogue
Internal Comms
by anthropics
vendor ×13

A set of resources to help me write all kinds of internal communications, using the formats that my company likes to use. Claude should use this skill whenever asked to write some sort of internal communications (status reports, leadership updates, 3P updates, company newsletters, FAQs, incident reports, project updates, etc.).

6k tokens
Competitive Ads Extractor
by frostant
×10

Extracts and analyzes competitors' ads from ad libraries (Facebook, LinkedIn, etc.) to understand what messaging, problems, and creative approaches are working. Helps inspire and improve your own ad campaigns.

2k tokens
Lead Research Assistant
by frostant
×8

Identifies high-quality leads for your product or service by analyzing your business, searching for target companies, and providing actionable contact strategies. Perfect for sales, business development, and marketing professionals.

2k tokens
Developer Growth Analysis
by frostant
×6

Analyzes your recent Claude Code chat history to identify coding patterns, development gaps, and areas for improvement, curates relevant learning resources from HackerNews, and automatically sends a personalized growth report to your Slack DMs.

4k tokens
App Store Optimization
by alirezarezvani
×3

Complete App Store Optimization (ASO) toolkit for researching, optimizing, and tracking mobile app performance on Apple App Store and Google Play Store

55k tokens scripts
Deeptools
by christophacham
×3

NGS analysis toolkit. BAM to bigWig conversion, QC (correlation, PCA, fingerprints), heatmaps/profiles (TSS, peaks), for ChIP-seq, RNA-seq, ATAC-seq visualization.

21k tokens scripts
Pymatgen
by christophacham
×3

Materials science toolkit. Crystal structures (CIF, POSCAR), phase diagrams, band structure, DOS, Materials Project integration, format conversion, for computational materials science.

26k tokens scripts
Enhance Prompt
by google-labs-code
vendor ×2

Transforms vague UI ideas into polished, Stitch-optimized prompts. Enhances specificity, adds UI/UX keywords, injects design system context, and structures output for better generation results.

3k tokens

How to use it

Copy the folder

Take seranking/seo-schema from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.