Read BEFORE optimizing any MATLAB code for speed. Without this workflow, agents commonly optimize the wrong target, fabricate speedup claims without measurement, or introduce regressions. Guides the 7-step workflow: baseline, profile, identify, optimize, measure, verify, report.
npx skills add https://github.com/matlab/matlab-agentic-toolkit --skill matlab-optimize-performance
Systematic 7-step workflow for finding and fixing performance bottlenecks in MATLAB code.
writing-matlab-perf-tests skill)Measure current performance so you have a number to improve against.
For a single function:
f = @() targetFunction(input1, input2);
baseline = timeit(f);
fprintf('Baseline: %.4f s\n', baseline);
For GPU code:
f = @() gpuFunction(gpuInput);
baseline = gputimeit(f);
For a script or multi-step workflow:
% Warmup run (first call includes JIT compilation)
myWorkflow(inputs);
% Timed run
tic;
myWorkflow(inputs);
baseline = toc;
fprintf('Baseline: %.4f s\n', baseline);
timeit is preferred because it handles warmup and runs multiple samples automatically.
Find where the time is actually spent. Do NOT guess — always profile.
profile on;
targetFunction(input1, input2);
profile off;
profile viewer;
Reading profiler results:
Tips:
Based on profiling results, identify which patterns apply. Read references/optimization-patterns.md for the full catalog.
High-impact patterns:
| Pattern | Typical Speedup | Look For |
|---------|----------------|----------|
| Vectorization | 2–200x | Loops doing element-wise math on arrays |
| Preallocation | 2–100x | Arrays growing inside loops (x = [x; newRow]) |
| Unnecessary recomputation | 2–50x | Same expensive expression computed multiple times |
| discretize/histcounts | 2–50x | Loops binning or classifying data |
| Persistent caching | 1.5–95x | Repeated load() or expensive object creation |
| Logical indexing | 1.2–5x | Using find() just to index into an array |
| arguments block | 1.1–1.8x | Functions using inputParser |
| Algebraic simplification | 1.5–3x | Redundant sqrt, abs, or matrix ops |
Before optimizing, verify the target is worth it:
filtfilt/filter instead of looping over columns).Apply the patterns identified in Step 3. See references/optimization-patterns.md for the full catalog with before/after code examples.
General principles:
discretize, cumsum, hypot) instead of hand-written equivalentsExample — move invariant work out of loops:
% Before: repeated expensive setup
for i = 1:n
opts = optimoptions('fminunc', 'Display', 'off');
result(i) = fminunc(@(x) cost(x, data(i)), x0, opts);
end
% After: setup once
opts = optimoptions('fminunc', 'Display', 'off');
for i = 1:n
result(i) = fminunc(@(x) cost(x, data(i)), x0, opts);
end
Re-measure using the same method as Step 1:
f = @() optimizedFunction(input1, input2);
optimized = timeit(f);
speedup = baseline / optimized;
fprintf('Optimized: %.4f s (%.2fx speedup)\n', optimized, speedup);
A speedup of 1.2x or more is considered significant. Below that, measurement noise makes it hard to be confident the change helped.
Every optimization must produce the same results as the original:
original = originalFunction(input1, input2);
fast = optimizedFunction(input1, input2);
% Numeric comparison (allows floating-point tolerance)
maxErr = max(abs(original(:) - fast(:)));
fprintf('Max error: %.2e\n', maxErr);
assert(maxErr < 1e-10, 'Results differ beyond tolerance!');
For non-numeric outputs:
assert(isequal(original, fast), 'Results differ!');
If results differ slightly due to floating-point reordering (e.g., summing in a different order), that's usually acceptable. Document the expected tolerance.
Summarize what was done and the improvement achieved:
fprintf('\n=== Performance Optimization Report ===\n');
fprintf('Target: %s\n', funcName);
fprintf('Baseline: %.4f s\n', baseline);
fprintf('Optimized: %.4f s\n', optimized);
fprintf('Speedup: %.2fx\n', speedup);
fprintf('Correctness: max error = %.2e\n', maxErr);
fprintf('Pattern applied: %s\n', patternName);
For multiple optimizations, report each speedup individually and the overall end-to-end improvement.
wait(gpuDevice) before and after timing GPU codetimeit — it handles warmup and averaging; avoid raw tic/toc for benchmarks| Mistake | Why It's Wrong | Do This Instead |
|---------|---------------|-----------------|
| Optimizing without profiling | You'll fix the wrong thing | Profile first (Step 2) |
| Single tic/toc without warmup | Includes JIT compilation time | Use timeit or add a warmup call |
| Timing GPU code without sync | GPU ops are async; toc fires early | wait(gpuDevice) before and after |
| Growing arrays in loops | Each append copies the entire array | Preallocate before the loop |
| Vectorizing huge arrays blindly | May exceed memory | Use chunked processing for large data |
| Reporting only subfunction speedup | Misleading if subfunction is 5% of total | Always report end-to-end timing |
| Assuming faster = correct | Bugs can make code fast (by skipping work) | Always verify results match (Step 6) |
references/optimization-patterns.md — Full catalog of optimization patterns with code examples and measured speedupsreferences/measurement-templates.md — Ready-to-use MATLAB script templates for each workflow stepCopyright 2026 The MathWorks, Inc.
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup
Comprehensive GitHub release orchestration with AI swarm coordination for automated versioning, testing, deployment, and rollback management
Migrate test files from `as` type assertions to @total-typescript/shoehorn. Use when user mentions shoehorn, wants to replace `as` in tests, or needs partial test data.
Modern JavaScript/TypeScript development with Bun runtime. Covers package management, bundling, testing, and migration from Node.js. Use when working with Bun, optimizing JS/TS development speed, or migrating from Node.js to Bun.
You are a dependency management expert specializing in safe, incremental upgrades of project dependencies. Plan and execute dependency updates with minimal risk, proper testing, and clear migration pa
Master systematic debugging techniques, profiling tools, and root cause analysis to efficiently track down bugs across any codebase or technology stack. Use when investigating bugs, performance issues, or unexpected behavior.
Opinionated backend development standards for Node.js + Express + TypeScript microservices. Covers layered architecture, BaseController pattern, dependency injection, Prisma repositories, Zod validation, unifiedConfig, Sentry error tracking, async safety, and testing discipline.
Best practices for writing JavaScript/TypeScript tests using Jest, including mocking strategies, test structure, and common patterns.
Take matlab/matlab-optimize-performance from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.