seb1n/data-labeling
Set up and manage data labeling workflows using manual annotation tools, semi-automated pipelines, active learning, and programmatic weak supervision.
npx skills add https://github.com/seb1n/awesome-ai-agent-skills --skill data-labeling
This skill enables an AI agent to design and execute data labeling workflows for machine learning projects. It covers manual annotation with tools like Label Studio, semi-automated labeling with model-assisted pre-annotation, active learning loops that prioritize the most informative samples, and programmatic weak supervision using labeling functions. The agent handles label schema design, annotator guidelines, quality control through inter-annotator agreement, and export to ML-ready formats.
Provide the agent with the raw dataset, the task type (classification, NER, object detection, etc.), and the label categories. Optionally specify the labeling tool preference and quality requirements (minimum inter-annotator agreement). The agent will configure the labeling environment, set up quality control, and manage the annotation workflow.
Label Studio labeling interface configuration (config.xml):
<View>
<Header value="Classify the customer review sentiment:" />
<Text name="text" value="$text" />
<Choices name="sentiment" toName="text" choice="single-column" showInline="true">
<Choice value="positive" />
<Choice value="negative" />
<Choice value="neutral" />
</Choices>
<Textarea name="notes" toName="text" placeholder="Optional: explain ambiguous cases"
maxSubmissions="1" editable="true" />
</View>
Python script to set up the project and import data:
from label_studio_sdk import Client
ls = Client(url="http://localhost:8080", api_key="your-api-key")
project = ls.start_project(
title="Customer Review Sentiment",
label_config=open("config.xml").read(),
description="Label customer reviews as positive, negative, or neutral.",
)
# Import tasks from a CSV file
import csv
tasks = []
with open("reviews.csv") as f:
for row in csv.DictReader(f):
tasks.append({"data": {"text": row["review_text"]}, "meta": {"source_id": row["id"]}})
project.import_tasks(tasks)
# Configure inter-annotator overlap: each task gets 2 annotators
project.set_params(maximum_annotations=2, overlap_cohort_percentage=100)
print(f"Created project with {len(tasks)} tasks, 2 annotators per task")
# After annotation, export results
annotations = project.export_tasks(export_type="JSON")
# Compute agreement
from sklearn.metrics import cohen_kappa_score
labels_a1 = [a["annotations"][0]["result"][0]["value"]["choices"][0] for a in annotations if len(a["annotations"]) >= 2]
labels_a2 = [a["annotations"][1]["result"][0]["value"]["choices"][0] for a in annotations if len(a["annotations"]) >= 2]
print(f"Cohen's kappa: {cohen_kappa_score(labels_a1, labels_a2):.3f}")
import pandas as pd
import numpy as np
from snorkel.labeling import labeling_function, PandasLFApplier, LFAnalysis
from snorkel.labeling.model import LabelModel
SPAM = 1
HAM = 0
ABSTAIN = -1
df = pd.DataFrame({
"text": [
"Congratulations! You've won a free iPhone!", "Meeting at 3pm tomorrow",
"URGENT: claim your prize now!!!", "Can you review the Q3 report?",
"Buy cheap meds online fast", "Lunch plans for Thursday?",
"Click here for a free vacation", "Project deadline is next Friday",
]
})
@labeling_function()
def lf_contains_free(x):
return SPAM if "free" in x.text.lower() else ABSTAIN
@labeling_function()
def lf_contains_urgent(x):
return SPAM if "urgent" in x.text.lower() else ABSTAIN
@labeling_function()
def lf_contains_click(x):
return SPAM if "click" in x.text.lower() else ABSTAIN
@labeling_function()
def lf_excessive_punctuation(x):
return SPAM if x.text.count("!") >= 3 else ABSTAIN
@labeling_function()
def lf_contains_meeting(x):
return HAM if any(w in x.text.lower() for w in ["meeting", "project", "report", "deadline"]) else ABSTAIN
@labeling_function()
def lf_short_and_casual(x):
return HAM if len(x.text.split()) < 8 and "?" in x.text else ABSTAIN
lfs = [lf_contains_free, lf_contains_urgent, lf_contains_click,
lf_excessive_punctuation, lf_contains_meeting, lf_short_and_casual]
applier = PandasLFApplier(lfs=lfs)
L_train = applier.apply(df=df)
print(LFAnalysis(L=L_train, lfs=lfs).lf_summary())
# Train the label model to combine noisy labeling functions
label_model = LabelModel(cardinality=2, verbose=True)
label_model.fit(L_train=L_train, n_epochs=500, log_freq=100, seed=42)
# Get probabilistic labels
probs = label_model.predict_proba(L=L_train)
df["label"] = label_model.predict(L=L_train)
df["confidence"] = np.max(probs, axis=1)
# Filter out low-confidence samples for manual review
confident = df[df["confidence"] > 0.8]
needs_review = df[df["confidence"] <= 0.8]
print(f"Confidently labeled: {len(confident)}, needs manual review: {len(needs_review)}")
Take seb1n/data-labeling from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.