mcpbeat Sign in

Product Data Cleanup Agent Skill

Clean a local FF&E CSV schedule by normalizing casing, dimensions, units, language, materials, and formatting. Use when asked to clean, fix, or standardize product data.

3k tokens
context cost
the whole folder, loaded on every use
2
files
instructions only
0
copies elsewhere
how many repositories repackaged it
302
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/AlpacaLabsLLC/skills-for-architects --skill product-data-cleanup

What comes with it

3 312 bytes besides the instruction
README.md

The instruction itself

19 sections, as written by the author

/as:product-data-cleanup — Product Data Normalizer

Takes a messy FF&E schedule and normalizes everything: casing, dimensions, units, language, materials vocabulary, currency formatting, and duplicates. Outputs a clean, consistent, spec-ready schedule.

Persistent cleanup operates on the nearest project's product-library.csv. Pasted tables may be previewed in Markdown but are not another persistent format.

Input

The user provides a schedule in one of these ways:

  • Project library — the nearest product-library.csv under an ancestor containing PROJECT.md.
  • CSV file path — import only after its exact 33-column header is validated.
  • Pasted table — preview a proposed canonical mapping before any persistence.

If the input format is unclear, ask.

Cleanup Rules

1. Casing

| Field | Rule | Example |

|-------|------|---------|

| Product Name | Title Case | eames lounge chairEames Lounge Chair |

| Brand | Title Case, preserve known abbreviations | HERMAN MILLERHerman Miller, HAYHAY |

| Collection | Title Case | cosmCosm |

| Category | Title Case, singular | chairsChair, TABLESTable |

| Materials | Sentence case, lowercase after first word | MOLDED PLYWOOD, FULL GRAIN LEATHERMolded plywood, full grain leather |

| Colors/Finishes | Title Case per item | walnut/black leatherWalnut / Black Leather |

Known brand abbreviations to preserve: HAY, USM, B&B, DWR, CB2, HBF, OFS, SitOnIt, 3form, ICF

2. Category Normalization

Map free-text categories to the canonical vocabulary and alias table defined in ../../schema/product-schema.md. Read that file for the full mapping of variations (English, Spanish, legacy terms) to canonical category names.

If a category is ambiguous, keep the closest match and add a [?] flag for the user to review.

3. Dimensions

Splitting combined dimensions:

| Input | → W | → D | → H | → Unit |

|-------|-----|-----|-----|--------|

| 32 x 24 x 30 in | 32 | 24 | 30 | in |

| 80 × 60 × 75 cm | 80 | 60 | 75 | cm |

| W32 D24 H30 | 32 | 24 | 30 | (infer) |

| 32"W x 24"D x 30"H | 32 | 24 | 30 | in |

| Ancho: 80, Prof: 60, Alto: 75 cm | 80 | 60 | 75 | cm |

Dimension rules:

  • Always store as separate W, D, H columns with a Unit column
  • If dimensions are already split, validate they're numeric (strip any unit text from the number)
  • Interpret " as inches, ' as feet (convert to inches: 2'6"30)
  • Accept ×, x, X, by, por as separators
  • Convention: W × D × H (width × depth × height). If only 2 values, ask which is missing.
  • Round to 2 decimal places max
  • If unit is missing but values suggest inches (all < 100 for furniture), assume in. If values suggest cm (> 100 or explicit), use cm. If truly ambiguous, flag with [?].

Do NOT convert units. Keep the original unit. Designers need the manufacturer's spec for ordering.

4. Language Normalization

Detect the language of each field value and normalize to English unless the user specifies otherwise.

| Spanish (common in UY sources) | → English |

|-------------------------------|-----------|

| Silla | Chair (category) |

| Mesa | Table (category) |

| Escritorio | Desk (category) |

| Madera | Wood (material) |

| Cuero | Leather (material) |

| Acero | Steel (material) |

| Vidrio | Glass (material) |

| Tela | Fabric (material) |

| Mármol | Marble (material) |

| Roble | Oak (material) |

| Nogal | Walnut (material) |

| Blanco | White (color) |

| Negro | Black (color) |

| Natural | Natural (keep as-is) |

| Cromado | Chrome (finish) |

Rule: Translate category, material, and color/finish fields. Leave Product Name and Brand as-is (proper nouns).

If the user says "keep in Spanish" or specifies a target language, respect that.

5. Materials & Finishes Vocabulary

Standardize common material terms:

| Variations | → Standard |

|-----------|------------|

| SS, Stainless, S/S | Stainless steel |

| Ply, Plywood, Mold ply | Molded plywood |

| MDF, Medium density | MDF |

| HPL, High pressure laminate | HPL |

| Lam, Laminate | Laminate |

| Fab, Textile | Fabric |

| COM, C.O.M. | COM (Customer's Own Material) |

| COL, C.O.L. | COL (Customer's Own Leather) |

| Powder coat, PC, Pwdr | Powder-coated |

| Chrm, Chrome plated | Chrome |

| Anodized alum, Anod. | Anodized aluminum |

| Ven, Veneer | Veneer |

| Sol. wood, Solid | Solid wood |

6. Price & Currency

  • Strip currency symbols ($, , £, ¥) — store symbol as currency code in separate column
  • Remove thousands separators (both . and , — detect locale: 1.234,56 is EU format, 1,234.56 is US)
  • Store as plain decimal number: 5695.00
  • If price says "Contact", "Quote", "Trade", "A consultar", "Consultar" → set to empty
  • Currency detection: $ alone defaults to USD unless context suggests otherwise (UY site → UYU, EU site → EUR)
  • If a schedule mixes currencies, keep each row's original currency. Add a note at the top.

7. Duplicate Detection

  • Flag rows with identical Product Name + Brand as potential duplicates
  • Flag rows with identical URL as definite duplicates
  • Don't auto-delete — present duplicates to the user and ask what to keep

8. Whitespace & Formatting

  • Trim leading/trailing whitespace from all fields
  • Collapse multiple spaces to single space
  • Remove line breaks within field values
  • Normalize list separators: wood / metal / glassWood, Metal, Glass (comma-separated)
  • Remove trailing commas or semicolons

Workflow

Step 1: Load the schedule

Read the input. Report: "Loaded N rows with M columns."

Map input columns to the canonical schema. If column mapping is ambiguous (e.g., a column called "Size" could be combined dimensions), ask the user.

Step 2: Analyze issues

Scan all rows and produce a summary:

## Cleanup Preview

- **Casing**: X product names need Title Case
- **Categories**: Y rows have non-standard categories (mapping: "chairs" → Seating, etc.)
- **Dimensions**: Z rows have combined dimensions to split
- **Language**: W rows have Spanish-language fields to translate
- **Materials**: V rows have non-standard material terms
- **Prices**: U rows need currency formatting cleanup
- **Duplicates**: T potential duplicate rows found
- **Empty fields**: S rows missing dimensions, R rows missing price

Step 3: Confirm scope

The issue summary is the change preview. Present selectable cleanup groups directly through the single confirmation gate; do not ask the same question first in prose.

Step 4: Apply fixes

Process every row through the active cleanup rules. Track every change made.

Step 5: Present results

Show a before/after diff for a sample of changed rows (up to 5 examples). Then show the full cleaned table.

Report:

## Cleanup Complete

- Rows processed: N
- Changes made: X
- Flagged for review: Y (marked with [?])

Step 6: Save

Read ../../schema/product-schema.md and ../../schema/csv-conventions.md. For multiple changed rows, materialize the complete proposed 33-column CSV as a temporary or user-visible review file, validate that candidate, preview the whole change once, and use the single confirmation gate. After approval, invoke python3 "${CLAUDE_PLUGIN_ROOT}/skills/master-schedule/scripts/csv-library.py" import product --project <project-root> --source <review.csv> exactly once for one atomic replacement; never loop update. A genuinely single-record edit may instead invoke the same plugin-root helper's update command once with one uniquely matching stable field. Never overwrite an arbitrary input or hand-edit product-library.csv.

Edge Cases

  • Mixed-language schedule: Detect dominant language per column, normalize to one language
  • Merged cells or irregular formatting: Flag and ask user how to handle
  • Extra columns not in schema: Reject persistence and preview how they would map to canonical fields
  • Empty rows: Remove silently
  • Header detection: Auto-detect header row (first row with text that matches known field names). If uncertain, ask.

Other skills for the same job

different authors, same section of the catalogue
Startup Analyst
by ComeOnOliver
×2

Expert startup business analyst specializing in market sizing, financial modeling, competitive analysis, and strategic planning for early-stage companies. Use PROACTIVELY when the user asks about market opportunity, TAM/SAM/SOM, financial projections, unit economics, competitive landscape, team planning, startup metrics, or business strategy for pre-seed through Series A startups.

5k tokens
Team Composition Analysis
by ComeOnOliver
×2

This skill should be used when the user asks to "plan team structure", "determine hiring needs", "design org chart", "calculate compensation", "plan equity allocation", or requests organizational design and headcount planning for a startup.

5k tokens
Bulk Rnaseq
by K-Dense-AI
×1

End-to-end bulk RNA-seq orchestrator — takes raw FASTQ reads through QC and trimming (FastQC, fastp/Trim Galore), alignment and quantification (STAR, Salmon, featureCounts), assembles a gene-level counts matrix, then hands off to differential expression (pydeseq2), pathway/GSEA enrichment (pathway-enrichment), and publication figures (scientific-visualization). Use whenever the user has bulk RNA-seq reads or quant output and wants a complete, reproducible differential-expression workflow — e.g. "analyze my RNA-seq", "FASTQ to DESeq2", "run nf-core/rnaseq", "STAR/Salmon quantification", "build a counts matrix for DESeq2", or "go from reads to differentially expressed genes and enriched pathways". Routes between an nf-core/rnaseq (Nextflow) path and a standalone STAR/Salmon path, and covers experimental design, strandedness, and QC gates. For single-cell RNA-seq use the scanpy skill instead.

13k tokens scripts
Generate Status Report
by openai
vendor ×1

Generate project status reports from Jira issues and publish to Confluence. When an agent needs to: (1) Create a status report for a project, (2) Summarize project progress or updates, (3) Generate weekly/daily reports from Jira, (4) Publish status summaries to Confluence, or (5) Analyze project blockers and completion. Queries Jira issues, categorizes by status/priority, and creates formatted reports for delivery managers and executives.

6k tokens scripts
Us Market Bubble Detector
by nicepkg
×1

Evaluates market bubble risk through quantitative data-driven analysis using the revised Minsky/Kindleberger framework v2.1. Prioritizes objective metrics (Put/Call, VIX, margin debt, breadth, IPO data) over subjective impressions. Features strict qualitative adjustment criteria with confirmation bias prevention. Supports practical investment decisions with mandatory data collection and mechanical scoring. Use when user asks about bubble risk, valuation concerns, or profit-taking timing.

22k tokens scripts
Gws Workflow Standup Report
by googleworkspace
vendor

Google Workflow: Today's meetings + open tasks as a standup summary.

278 tokens
Recipe Create Events From Sheet
by googleworkspace
vendor

Read event data from a Google Sheets spreadsheet and create Google Calendar entries for each row.

223 tokens
Baoyu Diagram
by JimLiu

Create professional, dark-themed SVG diagrams of any type — architecture diagrams, flowcharts, sequence diagrams, structural diagrams, mind maps, timelines, illustrative/conceptual diagrams, and more. Use this skill whenever the user asks for any kind of technical or conceptual diagram, visualization of a system, process flow, data flow, component relationship, network topology, decision tree, org chart, state machine, or any visual representation of structure/logic/process. Also trigger when the user says "画个图" "画一个架构图" "diagram" "flowchart" "sequence diagram" "draw me a ..." or uploads content and asks to visualize it. Output is always a standalone .svg file.

7k tokens scripts

How to use it

Copy the folder

Take alpacalabsllc/product-data-cleanup from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.