Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls. Never use for ranking people or consequential decisions.
npx skills add https://github.com/K-Dense-AI/scientific-agent-skills --skill scholar-evaluation
Provide developmental, evidence-traceable feedback on a scholarly work:
paper, draft, protocol, literature synthesis, or research idea. Use
qualitative judgment first. Optional scores only describe how submitted
evidence maps to a predeclared bounded rubric.
This skill also audits whether a low-stakes assessment process documents its
construct, provenance, rater quality, uncertainty, traceability, sensitivity,
fairness, accessibility, privacy, and human governance.
Never use this skill to automate, recommend, materially influence, or score:
Never rank people. Never reduce a person to a composite score. Never infer
ability, character, integrity, protected traits, future performance, or worth.
A nominal human-in-the-loop does not remove this boundary.
If asked for a prohibited use, stop. Offer developmental comments on a
scholarly work or a process-only audit that does not process applications,
compare people, recommend an outcome, or advise a decision.
Do not issue publication-readiness, accept/reject, or “top-tier” judgments.
Read references/responsible_assessment.md before any organizational use.
The referenced ScholarEval project is an **experimental
literature-grounded research-idea evaluation framework**, not validated
psychometrics.
The verified primary record is Moussa et al., *ScholarEval: Research Idea
Evaluation Grounded in Literature*, arXiv:2510.16234v2, revised 2026-02-28.
It reports a retrieval-augmented soundness/contribution framework, a
117-idea four-discipline dataset, coverage experiments, and a user study.
Do not generalize those results to person assessment, consequential decisions,
all disciplines, or this skill's rubric. No peer-reviewed publication status
was verified during the dated review. See references/source_ledger.md.
Do not score or infer quality from:
The rubric validator rejects common proxy-measure criteria.
If a qualified reviewer mentions an indicator descriptively outside the
scoring tools, record its exact purpose, source, coverage, field and time
effects, uncertainty, missingness, biases, gaming risk, and why it does not
directly measure quality. Never hide indicators inside an opaque composite.
Bundled scripts accept only strict local JSON/CSV containing pseudonymous IDs,
bounded ratings, statuses, uncertainty, and local references.
Do not put raw private applications, CVs, letters, reviewer identities,
contact details, protected attributes, or source-document text in inputs,
outputs, logs, examples, or prompts. Keep source content in the authorized
records system and use opaque local references.
Allowed classifications are:
syntheticpublic_scholarly_workdeidentified_low_stakesNo script searches the web, loads environment files, reads credentials, calls a
model, executes supplied text, deserializes executable objects, or launches a
process.
Use Bash only to invoke the documented local python3 commands.
Record:
scholarly_work;Stop on a prohibited decision context or unnecessary private data.
State:
Start with values and disciplinary context, not available metrics.
Begin with assets/rubric_template.json, then obtain qualified disciplinary,
assessment-methods, stakeholder, accessibility, privacy, and fairness review.
The template deliberately records content validity as not_established.
Do not change that status without documented evidence for the exact intended
use.
Validate structure:
PYTHONDONTWRITEBYTECODE=1 python3 scripts/validate_rubric.py \
--rubric assets/rubric_template.json
Read references/evaluation_framework.md for construct, anchor, validity, and
rater guidance.
Reviewers may read an authorized work outside the scripts. Record only stable
local locators and claim references in
assets/evidence_manifest_template.json.
For every criterion, distinguish:
missing from not_applicable; andFailure to find prior work does not prove novelty.
Use assets/evaluation_template.json. Each criterion must be:
rated with an anchor score, bounded uncertainty, evidence IDs, and a localrationale reference;
missing with null score/uncertainty and a rationale reference; ornot_applicable with null score/uncertainty and a rationale reference.Do not encode missing or not-applicable as zero. Raters should train, calibrate,
disclose conflicts, rate independently, and document disagreement.
Bounded scoring, without labels or recommendation:
PYTHONDONTWRITEBYTECODE=1 python3 scripts/calculate_scores.py \
--rubric assets/rubric_template.json \
--evaluation assets/evaluation_template.json
Evidence traceability:
PYTHONDONTWRITEBYTECODE=1 python3 scripts/check_traceability.py \
--rubric assets/rubric_template.json \
--evaluation assets/evaluation_template.json \
--evidence assets/evidence_manifest_template.json
Inter-rater agreement:
PYTHONDONTWRITEBYTECODE=1 python3 scripts/summarize_agreement.py \
--rubric assets/rubric_template.json \
--ratings assets/ratings_template.csv
Weight sensitivity requires two or more distinct scholarly-work evaluation
files:
PYTHONDONTWRITEBYTECODE=1 python3 scripts/weight_sensitivity.py \
--rubric assets/rubric_template.json \
--evaluation /tmp/work-a-evaluation.json \
--evaluation /tmp/work-b-evaluation.json
Process controls:
PYTHONDONTWRITEBYTECODE=1 python3 scripts/check_process.py \
--process assets/process_checklist_template.json
The checklist template is intentionally unconfirmed and fails closed.
Instructions and exact schemas are in references/local_tooling.md.
Lead with criterion-level evidence, not the composite. For each criterion:
rated, missing, or not_applicable;Generate an empty-reference scaffold if useful:
PYTHONDONTWRITEBYTECODE=1 python3 scripts/generate_report_scaffold.py \
--rubric assets/rubric_template.json \
--evaluation assets/evaluation_template.json \
--output /tmp/developmental-report-scaffold.json
The scaffold does not read source documents or draft findings.
Before releasing an organizational report, a qualified accountable human
committee must verify:
Document dissent. Do not imply consensus, validity, or precision beyond the
evidence. Periodically evaluate the evaluation and retire harmful criteria.
references/responsible_assessment.md — safety, metrics, governance,accessibility, privacy, and bias.
references/evaluation_framework.md — ScholarEval boundary, construct,criteria, anchors, validity, and interpretation.
references/local_tooling.md — strict schemas, formulas, commands, andoutput behavior.
references/source_ledger.md — authoritative sources and publication-statusverification dated 2026-07-23.
references/security_validation.md — baseline remediation, validation, andresidual security-scan record.
assets/rubric_template.json — bounded rubric template.assets/evaluation_template.json — rating template.assets/evidence_manifest_template.json — traceability template.assets/process_checklist_template.json — fail-closed process checklist.assets/ratings_template.csv — synthetic agreement data.Creates detailed, sectionized implementation plans through research, stakeholder interviews, and multi-LLM review. Use when planning features that need thorough pre-implementation analysis.
This skill should be used when scientists need help with research problem selection, project ideation, troubleshooting stuck projects, or strategic scientific decisions. Use this skill when users ask to pitch a new research idea, work through a project problem, evaluate project risks, plan research strategy, navigate decision trees, or get help choosing what scientific problem to work on. Typical requests include "I have an idea for a project", "I'm stuck on my research", "help me evaluate this project", "what should I work on", or "I need strategic advice about my research".
Research ideation partner. Generate hypotheses, explore interdisciplinary connections, challenge assumptions, develop methodologies, identify research gaps, for creative scientific problem-solving.
Manage and trigger pre-built Zapier workflows and MCP tool orchestration. Use when user mentions workflows, Zaps, automations, daily digest, research, search, lead tracking, expenses, or asks to "run" any process. Also handles Perplexity-based research and Google Sheets data tracking.
Guides researchers through structured ideation frameworks to discover high-impact research directions. Use when exploring new problem spaces, pivoting between projects, or seeking novel angles on existing work.
Loop 2 of the Three-Loop Integrated Development System. META-SKILL that dynamically compiles Loop 1 plans into agent+skill execution graphs. Queen Coordinator selects optimal agents from 86-agent registry and assigns skills (when available) or custom instructions. 9-step swarm with theater detection and reality validation. Receives plans from research-driven-planning, feeds to cicd-intelligent-recovery. Use for adaptive, theater-free implementation.
Loop 1 of the Three-Loop Integrated Development System. Research-driven requirements analysis with iterative risk mitigation through 5x pre-mortem cycles using multi-agent consensus. Feeds validated, risk-mitigated plans to parallel-swarm-implementation. Use when starting new features or projects requiring comprehensive planning with <3% failure confidence and evidence-based technology selection.
Find, compare, and research trade shows, exhibitions, expos, and industry events by vertical, region, date, or audience. Use this skill whenever the user wants to discover which trade shows exist for their industry, compare multiple events side-by-side, decide which shows are worth attending or exhibiting at, look up event dates and venues, research exhibitor counts or visitor profiles, or plan an annual trade show calendar. Also triggers on questions like 'what are the best shows for [industry]', 'when is [show name]', 'should we go to [event] or [event]', 'find me exhibitions in Germany for packaging', 'trade show calendar 2026', 'exhibition calendar Europe', 'B2B trade shows', 'what industry events should I attend', 'upcoming trade fairs', or even vague requests like 'we need to get in front of more buyers — what events should we be at'. If the user mentions any specific trade show by name (CES, MEDICA, Hannover Messe, Interpack, SXSW, Bauma, etc.) and wants information about it, use this skill.
Take k-dense-ai/scholar-evaluation from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.