dongshuyan/academic-humanizer
>- Draft, audit, or minimally revise English- or Chinese-language academic prose to reduce formulaic, vacuous, mechanically repetitive, or process-leaking language while preserving claims, evidence strength, logical relations, manuscript-wide terminology identity, document-level pattern variation, and scholarly register. Use for papers, abstracts, grants, cover letters, and reviewer responses when the user asks to de-AI, humanize, audit AI-like phrasing, or rewrite text without changing meaning. English is primary; Chinese is supported. Not for detector evasion, policy circumvention, pure translation, non-academic copy, or adding facts, citations, examples, or author experiences that the source does not contain.
npx skills add https://github.com/dongshuyan/compass-skills --skill academic-humanizer
Improve academic prose by removing observable writing defects, not by imitating
imperfection or optimizing an authorship detector. Preserve the author's facts,
argument, uncertainty, and disciplinary voice. This skill does not guarantee how
any reader or detector will classify a text.
This skill is agent-agnostic. Its core behavior is defined by SKILL.md and
references/; Python is optional and supports reproducible diagnostics.
<skill-dir> from the directory containing this SKILL.md.<python> mean an available Python 3 launcher, such as python3, py -3,or python.
<input-file> mean a user-authorized local text file. Quote paths thatcontain spaces and use the host shell's path separator.
agent name, or path separator.
agents/openai.yaml is optional interface metadata. Core behavior does notdepend on a particular agent runtime.
directly.
Read these before drafting or editing:
locked spans, deletion safety, and the internal claim ledger.
declared aliases, coined names, intentional distinctions, and the internal
terminology ledger. Always load it for multi-span or manuscript-level work.
local-to-document audit, distribution map, scope limits, and whole-document
repair. Always load it for multi-sentence work.
forms in both languages.
English and Chinese. Always load it; this is a cross-language semantic rule.
English rules or
Chinese rules.
Read worked examples on first use, after changing a
rule, or whenever fact preservation, contrast, or over-correction is uncertain.
Read metrics specification before running
scripts/metrics.py; its output is descriptive evidence only.
asks to de-AI or humanize text.
Do not create another routing tree for paper section or discipline. Methods,
Results, Discussion, reviewer responses, and grants use the same contracts; the
whitelist handles legitimate register differences. Ask one direct question only
when the requested genre changes what counts as acceptable and context does not
resolve it.
Route on editable prose, excluding fenced code, formulas, block quotations, and
a trailing reference list. Use orthographic tokens: each CJK character is one
token and each contiguous Latin word is one token. This keeps embedded terms such
as Transformer or ImageNet from outweighing the Chinese sentence around them:
r = CJK tokens / (CJK tokens + Latin word tokens)
r >= 0.5: Chinese branch.r < 0.5: English branch.English terms in Chinese prose and Chinese terms in English prose remain
verbatim. If Python is available and the route is genuinely unclear, optionally
run <python> "<skill-dir>/scripts/metrics.py" "<input-file>" --route.
Routing is internal and never appears in the clean artifact.
Earlier rows win. References may elaborate this table but must not define a
second priority order.
| Priority | Constraint | Operational meaning |
|---|---|---|
| C0 | Artifact boundary | Process instructions, editor narration, and tool residue never enter the artifact. C0 applies only to process-layer text; it never authorizes deletion of real content. |
| C1 | Semantic fidelity | Every output claim maps to the source bundle; every material source claim remains represented. No added facts, relations, examples, citations, motivations, or limitations. |
| C2 | Locked-span protection | Quotations, formulas, code, references, citation keys, statistical notation, proper nouns, and requested verbatim text remain unchanged. |
| C3 | Terminology identity | One scientific concept uses one canonical term across the editable manuscript. Preserve declared full-name/abbreviation pairs, necessary grammatical forms, and intentional distinctions; never infer identity from similarity alone. |
| C4 | Academic register | Preserve functional hedging, passive voice, nominalization, discourse markers, and Chinese scholarly morphology. |
| C5 | Argument structure | Preserve causal strength, contrast, concession, addition, chronology, scope, and paragraph-level reasoning. Surface connectives may change when the relation survives. |
| C6 | Document patterning | Audit recurrence, clustering, dispersion, positional regularity, sentence rhythm, and rhetorical-function saturation across the complete editable scope. A count is evidence, never a verdict. |
| C7 | Local style repair | Apply language-specific rules only to locally unsupported, vacuous, mechanical, or stacked defects. |
Examples of conflict resolution:
is absent from the source: C1 blocks the addition.
C1 and C5 preserve the result and its relation to adjacent sentences.
the canonical term after C1 and C2 confirm that the referent and spans permit it.
evidence: C1 blocks automatic deletion; mark it uncertain in diagnostic output.
protect it. Repeated functionless instances may activate C6 after a distribution
audit, while C1-C5 still constrain every repair.
Read all supplied title, abstract, body sections, captions, tables, appendices,
and supplementary prose before changing anything. Identify which parts are
editable and which are evidence or protected context. Separate content
requirements from style/process instructions. For generation, treat only
supplied claims, data, citations, and explicitly marked hypotheticals as content.
Apply the semantic and terminology contracts. Build the claim/evidence ledger
with source-to-output mappings and provenance status for:
The editable draft establishes what the author currently says; it does not by
itself prove that a cited paper, result, quotation, or factual premise exists.
Mark unsupported evidence assertions as draft-only and preserve or flag them
instead of silently treating them as verified or extending the argument from them.
Build a separate terminology ledger for scientific concepts, especially newly
coined methods, modules, losses, metrics, datasets, and task names. Record:
concept_id, canonical_term, and the span that defines or first formallynames the concept;
allowed_forms, including full-name/abbreviation pairs and necessarygrammatical or bilingual mappings;
observed_variants, distinguish_from, and resolution status.Use explicit user terminology first, then formal definitions, then the first
unambiguous formal naming. Frequency alone never selects the canonical term.
Keep both ledgers internal unless the user asks for an audit trail.
variation as declared form, same-concept drift, **intentional
distinction, protected mention, or uncertain identity**.
uncertain using contrast-logic.md.
signals in the same span?
A lone word or sentence form is not enough to infer authorship or poor quality.
It can still be a local defect when it adds an unsupported claim, false relation,
or empty evaluation. Multiple weak signals in one span form one finding, not
several duplicate findings.
For multi-sentence input, map candidates by section, paragraph, sentence,
position, and rhetorical function using global-pattern-contract.md. Inspect:
Use within-document evidence and section function; never apply a universal count
or ratio. A distribution map supports findings only about the supplied editable
scope; an excerpt cannot support a whole-manuscript judgment. Optional metrics
produce a distribution map, not an authorship or quality judgment.
Classify each finding as local defect, distributional defect,
functional/protected, or uncertain. A distributional defect requires both
repetition or positional regularity and redundant rhetorical function. Several
valid ablation contrasts, method steps, reported metrics, or theorem consequences
remain protected even when their surface forms repeat.
every editable occurrence, including captions and tables. Preserve declared
abbreviations and grammatical forms; do not replace protected mentions.
preserve the text and ask or flag it outside the clean artifact.
not only X but also Y when Xand Y are supported; removing the construction must not remove either claim.
decide whether the contrast is real. Do not silently erase them.
supported proposition and relation, and vary syntax only when argument function
warrants it. Do not randomize sentence length or replace one repeated template
with another repeated template.
Never invent categories merely to make a list appear elegant.
it outside the artifact; do not strengthen it or use it to generate new claims.
Scan all editable sections together after revision. Every scientific concept
must use its canonical term or a declared allowed form. Verify that coined names
are unchanged after their formal introduction, captions and tables match the
body, bilingual mappings are declared, and distinct concepts remain distinct.
Any unresolved identity is a stop/flag result, not an automatic normalization.
Rebuild the distribution map after editing. Check that redundant clusters,
mechanical paragraph templates, uniform rhetorical peaks, and unsupported
certainty were resolved without erasing functional repetition or creating a new
dominant pattern. If the supplied scope is shorter than the claimed scope, report
the limitation and do not claim a whole-manuscript pass.
Re-read source and output side by side. The output fails if any answer is no:
scope unchanged?
declared forms and intentional distinctions remaining?
verified sources, or explicitly marked draft-only outside the artifact?
by guesswork, and tool residue?
10. Would a zero-edit result have been more accurate? If yes, restore the source.
Run metrics only as an optional residual scan. A metric never overrides this gate.
line, score, checklist, leak line, or editor preface.
local ordistributional). Each finding includes an exact source quote, rule ID,
location/distribution evidence, reason, and one of change, keep, or
uncertain.
them. Clearly separate diagnostics from text intended for the manuscript.
Stop and ask instead of guessing when:
citation, result, quotation, or factual premise whose existence or provenance
cannot be established from the source bundle;
not establish their identity, or no canonical term can be grounded;
Do not invent specifics, personal experience, citations, data, mechanisms,
baselines, or limitations to make prose sound more human. Do not casualize
academic writing merely to make it look less generated.
Take dongshuyan/academic-humanizer from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.