mcpbeat Sign in

Paper Verification Agent Skill

> Use when the user wants to verify paper claims against code or data, audit numerical accuracy, check formula-code alignment, or validate citation accuracy. Triggers on phrases like "verify claims", "check numbers", "do the numbers match", "formula vs code", "audit the paper", or "cross-check results".

1k tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
356
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/fcakyon/phd-skills --skill paper-verification

The instruction itself

9 sections, as written by the author

Paper Verification Methodology

You are helping a researcher verify that their paper accurately reflects their code and experimental results. This is the most critical quality control step in academic writing.

Verification Dimensions

1. Numerical Accuracy Audit

For every number in the paper (dataset sizes, metric values, percentages, counts):

  • Extract the number and its context from the .tex file
  • Trace it to its source: code output, result file, log, or tracking system
  • Verify the value matches exactly (watch for rounding, percentage vs decimal)
  • Flag any number that cannot be traced to a source

Template:

| Paper claim | Location (.tex) | Source file/code | Source value | Match? |
|-------------|-----------------|-----------------|-------------|--------|
| "13,999 frames" | abstract L3 | len(glob(labels/*.json)) | ? | ? |
| "4.2% improvement" | Table 2 | eval_results.json | ? | ? |

Common numerical errors:

  • Rounding inconsistencies (3.14 in text, 3.1415 in table)
  • Stale numbers from earlier experiments not updated after re-runs
  • Percentage vs absolute confusion
  • Off-by-one in dataset counts (headers counted, or not)

2. Terminology Consistency Audit

  • Extract all defined terms from the methods section
  • Search for each term across ALL sections
  • Flag any inconsistent usage:
  • Same concept, different names (e.g., "tag head" vs "classification head")
  • Same name, different meanings across sections
  • Defined but never used, or used but never defined

3. Code-Paper Alignment

For each method described in the paper:

  • Find the corresponding code (function, class, module)
  • Compare the paper's description with the actual implementation
  • Check specifically:
  • Algorithm steps match code flow
  • Hyperparameters in text match config/code defaults
  • Architecture descriptions match model code
  • Loss functions in equations match loss code
  • Training procedures match training scripts

Common mismatches:

  • Paper describes an idealized version, code has edge cases not mentioned
  • Hyperparameters changed during development but paper not updated
  • Paper describes a method that was later modified or removed from code

4. Formula-Code Verification

For each equation in the paper:

  • Identify the equation and its variables
  • Find the code that implements it
  • Map each mathematical operation to its code equivalent
  • Verify:
  • Summation bounds match loop bounds
  • Division operations handle edge cases
  • Normalization factors match
  • Gradient flow matches (detach, no_grad)
  • Reduction operations (mean vs sum) match

5. Citation Fact-Checking Protocol

For each citation in the paper:

Step 1: Extract the claim and the cited paper

Step 2: Verify BibTeX metadata against DBLP:

  • Author names (exact spelling, correct order)
  • Paper title (exact, from published version not preprint)
  • Venue and year (confirmed against actual publication)

Step 3: For cited claims with specific numbers:

  • Locate the exact table/figure in the cited paper
  • Verify the number matches what the citing paper states
  • If the number cannot be confirmed, suggest qualitative language instead

Step 4: Check for common citation errors:

  • Citing preprint when published version exists
  • Wrong year (submission vs publication)
  • Author name misspellings
  • Citing for a claim the paper doesn't actually make

Verification Process

  • Read the full paper (or specified sections)
  • Build the verification table for each dimension
  • For each entry, read the source and verify
  • Produce a prioritized issue list:
  • HIGH: Incorrect numbers, wrong claims, missing citations
  • MEDIUM: Terminology inconsistencies, stale but close numbers
  • LOW: Minor formatting, optional improvements

Output Format

Produce a structured verification report:

  • Summary: X issues found (Y high, Z medium, W low)
  • Numerical audit table: each number with source and match status
  • Terminology issues: inconsistent terms with locations
  • Code-paper mismatches: description vs implementation gaps
  • Citation issues: metadata errors and unverified claims
  • Suggested fixes: specific text replacements for each issue

Other skills for the same job

different authors, same section of the catalogue
Content Research Writer
by frostant
×10

Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section. Transforms your writing process from solo effort to collaborative partnership.

4k tokens
Lead Research Assistant
by frostant
×8

Identifies high-quality leads for your product or service by analyzing your business, searching for target companies, and providing actionable contact strategies. Perfect for sales, business development, and marketing professionals.

2k tokens
Notebooklm
by ZhanlinCui
×6

Use this skill to query your Google NotebookLM notebooks directly from Claude Code for source-grounded, citation-backed answers from Gemini. Browser automation, library management, persistent auth. Drastically reduced hallucinations through document-only responses.

26k tokens scripts
Biorxiv Database
by christophacham
×4

Efficient database search tool for bioRxiv preprint server. Use this skill when searching for life sciences preprints by keywords, authors, date ranges, or categories, retrieving paper metadata, downloading PDFs, or conducting literature reviews.

9k tokens scripts
Openalex Database
by christophacham
×4

Query and analyze scholarly literature using the OpenAlex database. This skill should be used when searching for academic papers, analyzing research trends, finding works by authors or institutions, tracking citations, discovering open access publications, or conducting bibliometric analysis across 240M+ scholarly works. Use for literature searches, research output analysis, citation analysis, and academic database queries.

13k tokens scripts
Uspto Database
by christophacham
×4

Access USPTO APIs for patent/trademark searches, examination history (PEDS), assignments, citations, office actions, TSDR, for IP analysis and prior art searches.

21k tokens scripts
Denario
by christophacham
×3

Multiagent AI system for scientific research assistance that automates research workflows from data analysis to publication. This skill should be used when generating research ideas from datasets, developing research methodologies, executing computational experiments, performing literature searches, or generating publication-ready papers in LaTeX format. Supports end-to-end research pipelines with customizable agent orchestration.

11k tokens
Hypogenic
by christophacham
×3

Automated LLM-driven hypothesis generation and testing on tabular datasets. Use when you want to systematically explore hypotheses about patterns in empirical data (e.g., deception detection, content analysis). Combines literature insights with data-driven hypothesis testing. For manual hypothesis formulation use hypothesis-generation; for creative ideation use scientific-brainstorming.

7k tokens

How to use it

Copy the folder

Take fcakyon/paper-verification from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.