mcpbeat Sign in

Solublempnn Agent Skill

> Inverse-fold a backbone with SolubleMPNN — ProteinMPNN retrained on a soluble-PDB subset (Dauparas et al. 2022) — for sequences biased toward cytosolic expression and reduced aggregation. Reach for this skill when designs from vanilla ProteinMPNN are aggregating or going to inclusion bodies, when redesigning a membrane-adjacent fold for soluble expression, or when an E. coli expression screen is the next step.

1k tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
859
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/xuzhougeng/wisp-science --skill solublempnn

The instruction itself

5 sections, as written by the author

SolubleMPNN

SolubleMPNN is not a separate package — it is the ProteinMPNN architecture

retrained on a soluble-PDB subset, which shifts the output distribution away

from the surface hydrophobics that the full-PDB model happily places (because

many of them are buried at crystallographic or membrane interfaces in the

training set). Reach for it when the goal is soluble yield in a heterologous

host; stick with proteinmpnn when native-like recovery matters more, since

the soluble prior trades a few points of recovery for the surface bias. Code

and weights are MIT (github.com/dauparas/ProteinMPNN, soluble_model_weights;

also exposed via github.com/dauparas/LigandMPNN). The model is small enough to

run on CPU — for a handful of sequences on one backbone that is seconds and

usually faster than dispatching; a GPU helps for batched campaigns. Either way

the repo is cloned in-job (no PyPI dist; checkpoints bundled).

Running it

pip install torch numpy   # if not already present
git clone --depth 1 https://github.com/dauparas/ProteinMPNN.git proteinmpnn
cd proteinmpnn
python protein_mpnn_run.py \
  --pdb_path backbone.pdb --pdb_path_chains "A" \
  --out_folder out --num_seq_per_target 16 \
  --sampling_temp "0.1" --use_soluble_model

The runner uses repo-relative imports, so the cd line is load-bearing —

invoking the script by absolute path from elsewhere fails with

ModuleNotFoundError. If you want threaded designed-sequence PDBs as well,

the LigandMPNN runner accepts --model_type soluble_mpnn (see ligandmpnn

for that path; it needs ProDy in addition to torch). The flag surface is

otherwise identical to proteinmpnn (or ligandmpnn for the second form),

including the string-typed temperature and the fixed-position JSONL keyed by

PDB stem — see proteinmpnn for the parsing quirks. The repo

ships soluble weights at v_48_010 and v_48_020 only; asking for

--model_name v_48_002 --use_soluble_model errors on a missing checkpoint, so

leave --model_name at its default.

Output is out/seqs/<stem>.fa with score= and seq_recovery= in each

header. Expect recovery against a native structure to drop a few points

relative to vanilla — that is the prior working, not a bug.

Wisp execution

Use python only for bounded interactive checks. For a long or GPU-backed

workload, require a selected and probed ssh:<alias> context and load

remote-compute-ssh. Put the documented invocation in a self-contained project

script, activate the remote environment explicitly, stage only small files with

input_paths, and make the command write to a known absolute remote result

path. Submit it with run_in_context and register that exact ssh:// path in

output_specs. Call monitor_run once when waiting is needed, get_run once

for a snapshot, or cancel_run to stop. Do not send a scheduler submission

through the SSH-direct runner.

Hydrophobic surface patches still recur where the fold needs them

Soluble weights shift the distribution; they do not enforce a hydrophobicity

ceiling. If a particular surface patch keeps coming back hydrophobic, that

patch is likely structurally load-bearing and the network is paying the

solubility cost to keep the fold. Layering --omit_AAs "CW" or a per-position

bias on top is fine, but check that the resulting designs still fold (via

boltz or esmfold2) before assuming the constraint was free.

"Crystallisable" training set ≠ "soluble in your host" — keep an orthogonal filter

The training set is "structures that were soluble enough to crystallise," which

correlates with but is not the same as "expresses solubly in E. coli at 37 °C."

For campaigns where expression yield is the bottleneck, rank the soluble-MPNN

output by an orthogonal sequence-based predictor before committing wet-lab

slots; treat the MPNN bias as widening the funnel, not replacing the filter.


Next: fold the designs with boltz or esmfold2 to confirm the backbone

is still recovered, then carry survivors into the expression screen.

Other skills for the same job

different authors, same section of the catalogue
Protocolsio Integration
by christophacham
×4

Integration with protocols.io API for managing scientific protocols. This skill should be used when working with protocols.io to search, create, update, or publish protocols; manage protocol steps and materials; handle discussions and comments; organize workspaces; upload and manage files; or integrate protocols.io functionality into workflows. Applicable for protocol discovery, collaborative protocol development, experiment tracking, lab protocol management, and scientific documentation.

16k tokens
Tailored Resume Generator
by frostant
×4

Analyzes job descriptions and generates tailored resumes that highlight relevant experience, skills, and achievements to maximize interview chances

3k tokens
Excalidraw Diagram Generator
by github
vendor ×3

Generate Excalidraw diagrams from natural language descriptions. Use when asked to "create a diagram", "make a flowchart", "visualize a process", "draw a system architecture", "create a mind map", or "generate an Excalidraw file". Supports flowcharts, relationship diagrams, mind maps, and system architecture diagrams. Outputs .excalidraw JSON files that can be opened directly in Excalidraw.

36k tokens scripts
Expo Dev Client
by openai
vendor ×3

Build and distribute Expo development clients locally or via TestFlight

961 tokens
Executing Plans
by ZhanlinCui
×3

Use when you have a written implementation plan to execute in a separate session with review checkpoints

542 tokens
Anndata
by christophacham
×3

Data structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.

16k tokens
Benchling Integration
by christophacham
×3

Benchling R&D platform integration. Access registry (DNA, proteins), inventory, ELN entries, workflows via API, build Benchling Apps, query Data Warehouse, for lab data management automation.

14k tokens
Biopython
by christophacham
×3

Comprehensive molecular biology toolkit. Use for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez). Best for batch processing, custom bioinformatics pipelines, BLAST automation. For quick lookups use gget; for multi-service integration use bioservices.

24k tokens

How to use it

Copy the folder

Take xuzhougeng/solublempnn from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.

Install what it needs

The instructions reference pip. Without those the skill loads but fails at the first command.