Process split datasets into training, validation, and testing sets for ML model development. Use when requesting "split dataset", "train-test split", or "data partitioning". Trigger with relevant phrases based on skill purpose. '
2k tokens
context cost
the whole folder, loaded on every use
7
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
2679
stars on the repo
on the repository, not the skill itself
Install
one command, takes just this skill from the repository
Bashruns shell commands — read the instruction before connecting
The instruction itself
14 sections, as written by the author
Dataset Splitter
Split datasets into training, validation, and testing sets with configurable ratios and stratification options.
Overview
This skill automates the process of dividing a dataset into subsets for training, validating, and testing machine learning models. It ensures proper data preparation and facilitates robust model evaluation.
How It Works
Analyze Request: The skill analyzes the user's request to determine the dataset to be split and the desired proportions for each subset.
Generate Code: Based on the request, the skill generates Python code utilizing standard ML libraries to perform the data splitting.
Execute Splitting: The code is executed to split the dataset into training, validation, and testing sets according to the specified ratios.
When to Use This Skill
This skill activates when you need to:
Prepare a dataset for machine learning model training.
Create training, validation, and testing sets.
Partition data to evaluate model performance.
Examples
Example 1: Splitting a CSV file
User request: "Split the data in 'my_data.csv' into 70% training, 15% validation, and 15% testing sets."
The skill will:
Generate Python code to read the 'my_data.csv' file.
Execute the code to split the data according to the specified proportions, creating 'train.csv', 'validation.csv', and 'test.csv' files.
Example 2: Creating a Train-Test Split
User request: "Create a train-test split of 'large_dataset.csv' with an 80/20 ratio."
The skill will:
Generate Python code to load 'large_dataset.csv'.
Execute the code to split the dataset into 80% training and 20% testing sets, saving them as 'train.csv' and 'test.csv'.
Best Practices
Data Integrity: Verify that the splitting process maintains the integrity of the data, ensuring no data loss or corruption.
Stratification: Consider stratification when splitting imbalanced datasets to maintain class distributions in each subset.
Randomization: Ensure the splitting process is randomized to avoid bias in the resulting datasets.
Integration
This skill can be integrated with other data processing and model training tools within the Claude Code ecosystem to create a complete machine learning workflow.
Prerequisites
Appropriate file access permissions
Required dependencies installed
Instructions
Invoke this skill when the trigger conditions are met
Provide necessary context and parameters
Review the generated output
Apply modifications as needed
Output
The skill produces structured output relevant to the task.
Error Handling
Invalid input: Prompts for correction
Missing dependencies: Lists required components
Permission errors: Suggests remediation steps
Resources
Project documentation
Related skills and commands
How to use it
Copy the folder
Take jeremylongshore/splitting-datasets from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
Check the name does not clash
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.