Iteratively optimizes a cuDF RapidsUDF implementation for GPU performance. Use after testing and benchmarking with udf-benchmark. Runs a loop of profiling, optimizing, testing, and benchmarking until performance converges or the iteration budget is exhausted.
npx skills add https://github.com/NVIDIA/cudf-spark --skill udf-optimize-cudf
Derive <CamelName> and <snake_name> from the UDF class name.
> Note: Commands require access to /tmp (Spark temp storage) and /dev (GPU device). If commands fail due to sandbox restrictions, re-run them unsandboxed.
cp src/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala> \
src/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala>.bak
.orig.bak exists yet, save the original unoptimized implementation:cp src/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala> \
src/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala>.orig.bak
This file is never overwritten; it preserves the pre-optimization baseline.
./run_micro_benchmark.sh --mode all --data-path data/bench_data_<rows>_rows.parquet --rows <rows>
Record the baseline GPU time and speedup. This is the number to beat.
Repeat the following steps up to N iterations (default: 10). Also stop early if no improvement is found after 3 consecutive failed attempts.
Maintain an optimization log throughout the loop: for each iteration, record what change was attempted and whether it improved, regressed, or had no effect. This prevents repeating failed approaches and feeds the final report.
Profile the current implementation to identify bottlenecks:
./run_micro_benchmark.sh --mode gpu --data-path data/bench_data_<rows>_rows.parquet --rows <rows> --profile
Summarize libcudf kernel stats:
nsys stats --report nvtx_sum --format csv -o rapidsudf results/<report>.nsys-rep
Consult references/OPTIMIZATION_PATTERNS.md for interpreting profiler output and identifying optimization opportunities.
> Tip: Profiling frequently is strongly recommended. Without profiler data, optimization changes are guesses. Try using other nsys stats commands as needed.
Based on profiling insights (or optimization patterns from the reference), make one targeted change to the RapidsUDF implementation. Isolating changes one at a time makes it possible to attribute performance impact.
mvn test -Dsuites=com.udf.CudfComparisonTest
cp src/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala>.bak \
src/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala>
Log the failure reason, then return to Step 1.
Sometimes, matching a certain edge case is impossible without a major performance tradeoff. If so, document the attempted fix, the benchmark evidence, and the exact behavior difference, then ask the user whether the performance-vs-correctness tradeoff is acceptable.
Do not comment out tests or accept a correctness difference during optimization unless the user explicitly approves that tradeoff.
./run_micro_benchmark.sh --mode all --data-path data/bench_data_<rows>_rows.parquet --rows <rows>
Compare the GPU time against the current best (from the last checkpoint).
cp src/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala> \
src/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala>.bak
Record the new best GPU time. Reset the consecutive-failure counter. Return to Step 1.
cp src/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala>.bak \
src/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala>
Increment the consecutive-failure counter. Return to Step 1.
If the user explicitly asked for the judge, a judge subagent, or a review agent, treat that as an explicit request for delegation: you MUST launch a separate subagent with model: inherit and instruct it to use the udf-judge-conversion skill. Ask it to review the UnitTest, CudfComparisonTest, optimized RapidsUDF implementation, and optimization log as a cuDF conversion.
If the user did not request a judge/review agent, mark this step as skipped and continue to Final Step 2. If a required judge subagent is blocked by tool policy, stop and tell the user that explicit permission/instruction is needed.
If you run the judge, include the judge verdict in the final report. If there are any blocking issues, fix them or report the last known-good checkpoint.
After completing all iterations (or early-stopping), review your own work to ensure the optimization did not weaken correctness, introduce hardcoded test behavior, hide CPU fallback logic, or comment out core test coverage.
After completing all iterations (or early-stopping), report:
Upon successful completion:
src/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala>src/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala>.baksrc/main/<java|scala>/com/udf/<CamelName>RapidsUDF.<java|scala>.orig.bakresults/Machine learning in Python with scikit-learn. Use when working with supervised learning (classification, regression), unsupervised learning (clustering, dimensionality reduction), model evaluation, hyperparameter tuning, preprocessing, or building ML pipelines. Provides comprehensive reference documentation for algorithms, preprocessing techniques, pipelines, and best practices.
Machine learning in Python with scikit-learn. Use when working with supervised learning (classification, regression), unsupervised learning (clustering, dimensionality reduction), model evaluation, hyperparameter tuning, preprocessing, or building ML pipelines. Provides comprehensive reference documentation for algorithms, preprocessing techniques, pipelines, and best practices.
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks
Automated scRNA-seq cell type annotation via pre-trained logistic regression. 45+ models: immune, gut, lung, brain, fetal, cancer microenvironments. Input normalized AnnData; outputs per-cell labels, majority-vote cluster labels, confidence scores. Use for fast, reference-backed annotation without manual marker inspection.
Classical ML in Python: classification, regression, clustering, dim reduction, evaluation, tuning, preprocessing pipelines. Linear models, tree ensembles, SVMs, K-Means, PCA, t-SNE. Use PyTorch/TF for deep learning; XGBoost/LightGBM for scale.
Python statistical modeling: regression (OLS, WLS, GLM), discrete (Logit, Poisson, NegBin), time series (ARIMA, SARIMAX, VAR), with rigorous inference, diagnostics, and hypothesis tests. Use scikit-learn for ML; statistical-analysis for test choice.
Take nvidia/udf-optimize-cudf from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.