Guide for choosing and creating scientific visualizations for publications and talks. Covers chart-type selection by data structure, color theory for accessibility/print, figure composition, journal formatting (Nature, Cell, ACS), and common pitfalls. Consult when visualizing data or preparing submission figures.
npx skills add https://github.com/jaechang-hits/SciAgent-Skills --skill scientific-visualization
Effective scientific visualization communicates data clearly, honestly, and accessibly. Poor chart choices, misleading axes, or inaccessible color palettes can obscure findings or introduce bias. This guide covers the full workflow of scientific figure preparation: from selecting the right chart type for your data structure through color theory, accessibility, and journal submission formatting requirements.
Every chart type is optimized for a specific data structure. Mismatches (e.g., pie charts for continuous distributions, bar charts for time series) hide structure and distort perception.
| Data Type | Recommended Chart | Avoid |
|-----------|------------------|-------|
| Continuous distribution (1 group) | Histogram, violin plot, ridge plot | Bar chart with mean only |
| Continuous distribution (2–5 groups) | Violin + boxplot overlay, beeswarm | Grouped bar chart |
| Two continuous variables, correlation | Scatter plot, hexbin (large N) | Line chart without temporal order |
| Categorical counts / proportions | Bar chart (horizontal for long labels) | Pie chart (>4 categories) |
| Change over time (continuous) | Line chart | Bar chart |
| Change over time (sparse events) | Step chart, event raster | Connected scatter |
| Part-to-whole (≤5 parts) | Stacked bar, waffle chart | 3D pie chart |
| High-dimensional (>5 variables) | Heatmap (clustered), parallel coordinates | 3D scatter |
| Spatial data | Map, spatial heatmap | Bubble chart |
| Survival / time-to-event | Kaplan-Meier curve | Bar chart of median survival |
Color encodes information. Misused color introduces artifacts and fails readers with color vision deficiency (CVD; ~8% of males).
Sequential palettes encode ordered numeric data from low to high (e.g., expression level, concentration). Use perceptually uniform palettes: viridis, magma, cividis. These also print in grayscale.
Diverging palettes encode data with a meaningful midpoint (e.g., fold-change centered at 0, correlation from -1 to +1). Use RdBu, coolwarm, or vlag. Always ensure the midpoint maps to white/neutral.
Qualitative palettes encode unordered categories. Use Okabe-Ito (CVD-safe), tab10 (matplotlib default), or ColorBrewer qualitative palettes. Limit to ≤8 distinguishable colors; use shape or pattern as redundant encoding beyond that.
Color don'ts:
Scientific figures are typically multi-panel. Panel layout and labeling affect how readers parse information.
Major journals specify exact figure requirements for submission. Violating these causes desk-rejection delays.
| Journal/Style | Max Width | Resolution | Color Mode | Font | File Format |
|---------------|-----------|------------|------------|------|-------------|
| Nature family | 89 mm (1-col), 183 mm (2-col) | 300 dpi (photos), 600 dpi (line art) | RGB or CMYK | Arial 5–7 pt | PDF, TIFF, EPS |
| Cell/iScience | 85 mm (1-col), 170 mm (2-col) | 300 dpi raster, 600 dpi halftone | RGB | Helvetica 6–8 pt | PDF, EPS, TIFF |
| ACS journals | 3.25 in (1-col), 7 in (2-col) | 600 dpi (color), 1200 dpi (b&w line art) | RGB (screen), CMYK (print) | Arial/Helvetica 4.5–7 pt | TIFF, EPS, PDF |
| PLOS ONE | No strict width | 300 dpi (raster), 600–1200 dpi (line art) | RGB | Any | TIFF, EPS, PDF |
Use this tree to select the right visualization for your analysis goal:
What is the primary message of this figure?
|
+-- Show a distribution or spread of values
| +-- One group --> Histogram or violin plot
| +-- 2-5 groups --> Violin + jitter (show all points if N < 100)
| +-- Many groups --> Ridge plot (joy plot)
|
+-- Compare quantities between categories
| +-- Few categories (2-5) --> Bar chart with error bars + individual points
| +-- Many categories (>8) --> Lollipop chart or dot plot (horizontal)
| +-- Paired measurements --> Slopegraph or paired dot plot
|
+-- Show a relationship between two continuous variables
| +-- N < 1000 --> Scatter plot
| +-- N > 1000 --> Hexbin or 2D density plot
| +-- Time ordered --> Line chart
|
+-- Show composition or part-to-whole
| +-- 2-4 parts --> Stacked bar or waffle chart
| +-- Over time --> Stacked area chart
| +-- Avoid pie chart unless <= 3 parts and proportions are obvious
|
+-- Show high-dimensional data
| +-- Genes x samples --> Clustered heatmap (seaborn.clustermap)
| +-- Embeddings (UMAP, PCA) --> Scatter colored by metadata
| +-- Feature importance --> Horizontal bar chart (sorted)
|
+-- Show spatial or geographic data
| +-- Microscopy --> Image overlay with colorbar
| +-- Geographic --> Choropleth map
| Analysis Goal | Chart Type | Library | Key Consideration |
|---------------|-----------|---------|-------------------|
| Gene expression across groups | Violin + jitter | seaborn, plotnine | Show all points if N < 50; never bar+SEM only |
| Differential expression | Volcano plot | matplotlib | Log2FC on x-axis, -log10(p) on y-axis |
| Clustering results | UMAP scatter | scanpy, matplotlib | One plot per annotation variable |
| Correlation matrix | Clustered heatmap | seaborn.clustermap | Use diverging palette centered at 0 |
| Protein structure | Ribbon diagram | PyMOL, ChimeraX | Not covered here — use dedicated molecular graphics tools |
| Survival analysis | Kaplan-Meier | lifelines | Include confidence bands and at-risk table |
| Time course | Line chart with CI | matplotlib | Show uncertainty; connect group means, not individual points |
viridis/cividis for sequential data. Test your figure with a CVD simulator (e.g., Coblis) before submission.viridis, magma, or inferno for sequential; RdBu or coolwarm for diverging data. These are the defaults in seaborn >= 0.12.seaborn.stripplot(jitter=True)), use transparency (alpha=0.3), or switch to a hexbin / 2D density plot for large N.matplotlib.rcParams style dictionary at the top of your figure script and apply it to all panels.plt.savefig("fig1.pdf", bbox_inches="tight") in matplotlib.rcParams at the top of every figure script; use fig.set_size_inches() to enforce journal dimensions; export with dpi=300 minimum.theme_classic() or a custom theme; set ggsave(width=..., units="mm", dpi=300).matplotlib.gridspec, patchworklib (Python), or cowplot/patchwork (R) for aligned panel grids.matplotlib-figures — Python implementation of publication-quality figures with matplotlib and seaborndata-visualization — general Python plotting recipesbiostatistics — statistical test selection to accompany figure annotationsComprehensive spreadsheet creation, editing, and analysis with support for formulas, formatting, data analysis, and visualization. When Claude needs to work with spreadsheets (.xlsx, .xlsm, .csv, .tsv, etc) for: (1) Creating new spreadsheets with formulas and formatting, (2) Reading or analyzing data, (3) Modify existing spreadsheets while preserving formulas, (4) Data analysis and visualization in spreadsheets, or (5) Recalculating formulas
Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .csv, or .tsv file (e.g., adding columns, computing formulas, formatting, charting, cleaning messy data); create a new spreadsheet from scratch or from other data sources; or convert between tabular file formats. Trigger especially when the user references a spreadsheet file by name or path — even casually (like \"the xlsx in my downloads\") — and wants something done to it or produced from it. Also trigger for cleaning or restructuring messy tabular data files (malformed rows, misplaced headers, junk data) into proper spreadsheets. The deliverable must be a spreadsheet file. Do NOT trigger when the primary deliverable is a Word document, HTML report, standalone Python script, database pipeline, or Google Sheets API integration, even if tabular data is involved.
Picks random winners from lists, spreadsheets, or Google Sheets for giveaways, raffles, and contests. Ensures fair, unbiased selection with transparency.
Query openFDA API for drugs, devices, adverse events, recalls, regulatory submissions (510k, PMA), substance identification (UNII), for FDA regulatory data analysis and safety research.
MATLAB and GNU Octave numerical computing for matrix operations, data analysis, visualization, and scientific computing. Use when writing MATLAB/Octave scripts for linear algebra, signal processing, image processing, differential equations, optimization, statistics, or creating scientific visualizations. Also use when the user needs help with MATLAB syntax, functions, or wants to convert between MATLAB and Python code. Scripts can be executed with MATLAB or the open-source GNU Octave interpreter.
UMAP dimensionality reduction. Fast nonlinear manifold learning for 2D/3D visualization, clustering preprocessing (HDBSCAN), supervised/parametric UMAP, for high-dimensional data.
Creating interactive data visualisations using d3.js. This skill should be used when creating custom charts, graphs, network diagrams, geographic visualisations, or any complex SVG-based data visualisation that requires fine-grained control over visual elements, transitions, or interactions. Use this for bespoke visualisations beyond standard charting libraries, whether in React, Vue, Svelte, vanilla JavaScript, or any other environment.
Access AlphaFold 200M+ AI-predicted protein structures. Retrieve structures by UniProt ID, download PDB/mmCIF files, analyze confidence metrics (pLDDT, PAE), for drug discovery and structural biology.
Take jaechang-hits/scientific-visualization from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.