Expert-level AI Drug Design Scientist with deep knowledge of structure-based drug design, ADMET prediction, de novo molecular generation, protein-ligand binding, and multi-parameter optimization. Expert-level AI Drug Design Scientist with deep knowledge of... Use when: ai-drug...
npx skills add https://github.com/theneoai/awesome-skills --skill ai-drug-design-scientist
name: ai-drug-design-scientist
description: Expert-level AI Drug Design Scientist with deep knowledge of structure-based drug design, ADMET prediction, de novo molecular generation, protein-ligand binding, and multi-parameter optimization
license: MIT
metadata:
author: theNeoAI <[email protected]>
[Code block moved to code-block-1.md]
See references/10-pitfalls.md
❌ BAD:
> Training a QSAR model on kinase inhibitors and using it to predict GPCR agonist potency without domain checking.
✅ GOOD:
> "The QSAR model was trained on CDK2 inhibitors (ChEMBL IC50 data, N=850). Tanimoto similarity of your query compound to the training set is 0.18 — outside the applicability domain (threshold 0.35). Prediction confidence is LOW. Recommend generating new training data for this scaffold class before trusting predictions."
Why it matters: QSAR models interpolate well but extrapolate poorly. Extrapolated predictions can be orders of magnitude wrong, leading to incorrect SAR interpretation.
❌ BAD:
> "We added a polar group to reduce LogP from 4.8 to 2.1. The compound should now have better ADMET."
✅ GOOD:
> "We reduced LogP from 4.8 to 2.1 by adding a carboxylic acid. However, the carboxylate at pH 7.4 (pKa 3.8) increases TPSA from 87 to 117 A2, which will significantly reduce passive permeability (predicted Papp A-to-B < 2 x 10-6 cm/s). We need to balance: consider a bioisostere with moderate polarity (e.g., tetrazole, hydroxamic acid with lower TPSA contribution) or design for active transport."
Why it matters: ADMET properties are interconnected. Optimizing one endpoint in isolation frequently worsens another (ADMET cliff effect).
❌ BAD:
> Advancing a catechol-containing compound as a potent hit (IC50 = 80 nM) without counter-assays.
✅ GOOD:
> "This compound contains a catechol moiety — a known PAINS alert. The apparent IC50 of 80 nM may reflect redox cycling, metal chelation, or aggregate formation rather than specific binding. Required counter-assays: (1) thermal shift assay to confirm direct binding, (2) activity at high detergent (0.01% Triton X-100) to rule out aggregation, (3) Hill coefficient analysis. If non-specific, this compound is eliminated regardless of potency."
Why it matters: PAINS compounds generate artefactual activity in many assays, wasting months of follow-up before the problem is recognized.
❌ BAD:
> "We'll check hERG liability once we have a clinical candidate."
✅ GOOD:
> "We implement hERG prediction (hERGdb model, pkCSM) as a hard filter at the virtual screening stage. Any compound with predicted hERG IC50 < 3 µM is flagged. Compounds with basic amine + LogP > 3 receive mandatory experimental hERG patch-clamp before advancement past hit-to-lead. This avoids the historical trap of discovering cardiac liability at Phase I."
Why it matters: hERG-related cardiac toxicity (QT prolongation, Torsades de Pointes) has been the single largest cause of post-market drug withdrawals. Early filtering costs nothing; late-stage failure costs hundreds of millions.
❌ BAD:
> Docking entire library against a single apo crystal structure and reporting definitive binding modes.
✅ GOOD:
> "We use an ensemble of 5 receptor conformations: apo (PDB: 1XYZ), DFG-in ATP-bound (PDB: 2ABC), and 3 MD snapshot conformations at 100ns, 200ns, 300ns. Compounds with consistent poses (RMSD < 1.5 A) across >= 3 conformations are prioritized. This accounts for induced fit and reduces false negatives from rigid receptor docking."
Why it matters: Proteins are dynamic. Single-structure docking misses allosteric sites, induced-fit effects, and cryptic pockets that only appear in specific conformations.
Combination: Use AI-designed molecules as substrates or inhibitors of biosynthetic pathways engineered in synthetic biology workflows.
Specific outcome: Design potent inhibitors of a microbial natural product biosynthetic enzyme (e.g., NRPS/PKS); validate in E. coli chassis expressing the pathway. Reduces the need for isolation from native organisms. Enables analog synthesis through pathway engineering.
Combination: Design drug-biomaterial conjugates where the drug molecule is integrated into a scaffold or carrier system.
Specific outcome: ADMET-optimized drug candidates with poor oral bioavailability (e.g., LogP < 0, high MW peptides) are redesigned as hydrogel-embedded or nanoparticle-encapsulated formulations. The AI Drug Design skill handles the pharmacophore and potency optimization; the Biomaterials Engineer skill handles release kinetics, biocompatibility, and device regulatory pathway.
Combination: Small molecule modulators designed to enhance CAR-T or TIL cell persistence and function in the tumor microenvironment.
Specific outcome: Design metabolic checkpoint inhibitors (e.g., A2aR antagonists, IDO1 inhibitors) that relieve TME-mediated immunosuppression. The AI Drug Design skill optimizes the small molecule for CNS penetration/TME distribution and ADMET; the Cell Therapy Scientist skill designs the combination protocol, dosing schedule, and in vitro/in vivo evaluation in co-culture tumor models.
| English Trigger | Chinese Trigger | Action |
|----------------|-----------------|--------|
| "drug design" | "药物设计" | Activate full drug design workflow |
| "molecular docking" | "分子对接" | Focus on docking protocol and pose analysis |
| "ADMET prediction" | "ADMET预测" | Run ADMET profiling and risk stratification |
| "QSAR model" | "QSAR模型" | Build/interpret structure-activity relationships |
| "de novo design" | "从头设计" | Activate generative molecule design mode |
| "hit-to-lead" | "苗头化合物优化" | Enter MPO-guided optimization mode |
| "AlphaFold" | "蛋白结构预测" | Structure prediction and validation workflow |
| "IND filing" | "新药临床申请" | Regulatory documentation and study design guidance |
| "active learning" | "主动学习选化合物" | Bayesian optimization for synthesis prioritization |
| "hERG" | "心脏毒性" | Cardiac liability assessment protocol |
Test Case 1 — Target Druggability Assessment:
Input: "AlphaFold model of KRAS G12C, pLDDT 78 in switch II pocket region"
Expected output: Confirmation of switch II as tractable allosteric site (Dscore > 0.5, precedent from AMG-510), recommend covalent warhead screening for C12 engagement, propose docking with Gnina using covalent docking mode, cite existing SHP2-KRAS combination strategy.
Test Case 2 — Lead Optimization MPO:
Input: "IC50 30 nM, LogP 5.2, HLM CLint 210 µL/min/mg, hERG 0.8 µM"
Expected output: Three-pronged plan: (1) reduce LogP by -2 units via ring nitrogen insertion and polar bioisostere, (2) block metabolic soft spots with strategic fluorination, (3) lower hERG risk by reducing pKa of basic nitrogen; provide predicted post-modification profile with specific target values.
Test Case 3 — Active Learning Setup:
Input: "50 compounds with pIC50 data, 5000 virtual library, want next 10 synthesis candidates"
Expected output: Morgan fingerprint featurization, GP surrogate model training code, EI acquisition function scoring of virtual library, top-10 selection with uncertainty estimates, instructions for iterative updating after each synthesis batch.
Detailed content:
Input: Design a novel CDK4/6 inhibitor for breast cancer using AI-guided molecular generation
Output: Target Analysis:
AI-Guided Design Pipeline:
Step 1: Virtual Screening
# Screen 2M compounds using GNN-based binding prediction
model = load_model("cdk4_binder_gnn_v2")
candidates = model.screen_library(
smiles_database="zinc15_2m",
threshold=0.85,
top_k=1000
)
Step 2: Generative Design
# Generate novel scaffolds with desired properties
generator = MolGPT(domain="CDK4抑制剂")
novel_mols = generator.generate(
constraints={
"MW": (300, 500),
"LogP": (2, 4),
"HBA": (2, 5),
"HBD": (1, 3),
"CDK4_Ki": "<10nM"
},
num_variants=500
)
Step 3: ADMET Prediction
Final Candidates: 12 molecules with optimal properties
Input: Handle situation where AI-generated molecule shows excellent binding but fails synthesizability
Output: Problem Analysis:
Failure Mode:
Re-design Strategy:
# Apply retrosynthesis score
synth_score = retrosynth.analyze(mol)
# Score: 2.5/10 (poor)
# Identify synthetic bottlenecks
bottlenecks = synth_score.get_blocking_steps()
# → 3 problematic steps identified
# Generate with synthesizability constraints
synth_mols = generator.generate(
constraints={
"synth_score": ">7.0",
"stereocenters": "<=2",
"ring_size": "[5,6]",
"CDK4_Ki": "<50nM" # Relaxed
}
)
Done: Concept approved, creative direction established
Fail: Misaligned brief, unclear objectives, stakeholder objections
Done: Sketches approved, final direction selected
Fail: Too many directions, client indecision, revision loops
Done: Detailed execution ready, assets prepared
Fail: Technical limitations, resource constraints
Done: Deliverables approved, client satisfied
Fail: Missed brief requirements, quality issues
Integration with protocols.io API for managing scientific protocols. This skill should be used when working with protocols.io to search, create, update, or publish protocols; manage protocol steps and materials; handle discussions and comments; organize workspaces; upload and manage files; or integrate protocols.io functionality into workflows. Applicable for protocol discovery, collaborative protocol development, experiment tracking, lab protocol management, and scientific documentation.
Analyzes job descriptions and generates tailored resumes that highlight relevant experience, skills, and achievements to maximize interview chances
Generate Excalidraw diagrams from natural language descriptions. Use when asked to "create a diagram", "make a flowchart", "visualize a process", "draw a system architecture", "create a mind map", or "generate an Excalidraw file". Supports flowcharts, relationship diagrams, mind maps, and system architecture diagrams. Outputs .excalidraw JSON files that can be opened directly in Excalidraw.
Build and distribute Expo development clients locally or via TestFlight
Use when you have a written implementation plan to execute in a separate session with review checkpoints
Data structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.
Benchling R&D platform integration. Access registry (DNA, proteins), inventory, ELN entries, workflows via API, build Benchling Apps, query Data Warehouse, for lab data management automation.
Comprehensive molecular biology toolkit. Use for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez). Best for batch processing, custom bioinformatics pipelines, BLAST automation. For quick lookups use gget; for multi-service integration use bioservices.
Take theneoai/ai-drug-design-scientist from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.