Semantic search in DDC CWICR construction database using vector embeddings. Find similar work items and resources for cost estimation.
npx skills add https://github.com/datadrivenconstruction/DDC_Skills_for_AI_Agents_in_Construction --skill semantic-search-cwicr
Construction cost estimation requires finding relevant work items from large databases. Traditional keyword search fails when:
DDC CWICR database provides pre-computed embeddings (OpenAI text-embedding-3-large, 3072 dimensions) enabling semantic similarity search across 55,719 work items in 9 languages.
pip install qdrant-client openai pandas
# Download Qdrant snapshot
wget https://github.com/datadrivenconstruction/OpenConstructionEstimate-DDC-CWICR/releases/download/v0.1.0/qdrant_snapshot_en.tar.gz
# Start Qdrant with Docker
docker run -p 6333:6333 -v $(pwd)/qdrant_storage:/qdrant/storage qdrant/qdrant
import pandas as pd
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams
import openai
class CWICRSemanticSearch:
def __init__(self, qdrant_host: str = "localhost", port: int = 6333):
self.client = QdrantClient(host=qdrant_host, port=port)
self.collection_name = "ddc_cwicr_en"
self.embedding_model = "text-embedding-3-large"
self.embedding_dim = 3072
def get_embedding(self, text: str) -> list:
"""Generate embedding for search query."""
response = openai.embeddings.create(
model=self.embedding_model,
input=text
)
return response.data[0].embedding
def search_work_items(self, query: str, limit: int = 10,
min_score: float = 0.7) -> pd.DataFrame:
"""Search for similar work items."""
query_vector = self.get_embedding(query)
results = self.client.search(
collection_name=self.collection_name,
query_vector=query_vector,
limit=limit,
score_threshold=min_score
)
items = []
for result in results:
item = result.payload
item['similarity_score'] = result.score
items.append(item)
return pd.DataFrame(items)
def search_by_category(self, query: str, category: str,
limit: int = 10) -> pd.DataFrame:
"""Search within specific category."""
query_vector = self.get_embedding(query)
results = self.client.search(
collection_name=self.collection_name,
query_vector=query_vector,
query_filter={
"must": [{"key": "category", "match": {"value": category}}]
},
limit=limit
)
return pd.DataFrame([{**r.payload, 'score': r.score} for r in results])
def estimate_cost(self, work_items: pd.DataFrame,
quantities: dict) -> dict:
"""Calculate cost from matched work items."""
total_cost = 0
breakdown = []
for _, item in work_items.iterrows():
if item['work_item_code'] in quantities:
qty = quantities[item['work_item_code']]
cost = qty * item.get('unit_price', 0)
total_cost += cost
breakdown.append({
'item': item['description'],
'quantity': qty,
'unit_price': item.get('unit_price', 0),
'total': cost
})
return {
'total_cost': total_cost,
'breakdown': breakdown,
'currency': 'Regional default'
}
search = CWICRSemanticSearch()
# Natural language query
results = search.search_work_items("brick masonry wall construction")
print(results[['description', 'unit', 'unit_price', 'similarity_score']])
# Find work items for foundation work
foundation_items = search.search_work_items(
"reinforced concrete foundation excavation and pouring",
limit=20
)
# Estimate with quantities
quantities = {
'CONC-001': 150, # cubic meters
'EXCV-002': 200, # cubic meters
}
estimate = search.estimate_cost(foundation_items, quantities)
print(f"Estimated Cost: ${estimate['total_cost']:,.2f}")
| Field | Type | Description |
|-------|------|-------------|
| work_item_code | string | Unique identifier |
| description | string | Work item description |
| unit | string | Measurement unit |
| labor_norm | float | Labor hours per unit |
| material_cost | float | Material cost per unit |
| equipment_cost | float | Equipment cost per unit |
| unit_price | float | Total price per unit |
| category | string | Work category |
| embedding | vector[3072] | Pre-computed embedding |
Access NCBI GEO for gene expression/genomics data. Search/download microarray and RNA-seq datasets (GSE, GSM, GPL), retrieve SOFT/Matrix files, for transcriptomics and expression analysis.
Implement efficient similarity search with vector databases. Use when building semantic search, implementing nearest neighbor queries, or optimizing retrieval performance.
Implement efficient similarity search with vector databases. Use when building semantic search, implementing nearest neighbor queries, or optimizing retrieval performance.
Systematic database and table profiling for DBX Studio. Use when a user wants to understand their data, explore schema structure, or profile a dataset.
Systematic database and table profiling for DBX Studio. Use when a user wants to understand their data, explore schema structure, or profile a dataset.
Turn JSON or PostgreSQL jsonb payloads into compact readable context for LLMs. Use when a user wants to compress JSON, reduce token usage, summarize API responses, or convert structured data into model-friendly text without dumping raw paths.
Implement ReasoningBank adaptive learning with AgentDB's 150x faster vector database. Includes trajectory tracking, verdict judgment, memory distillation, and pattern recognition. Use when building self-learning agents, optimizing decision-making, or implementing experience replay systems.
Access NCBI GEO for gene expression/genomics data. Search/download microarray and RNA-seq datasets (GSE, GSM, GPL), retrieve SOFT/Matrix files, for transcriptomics and expression analysis.
Take datadrivenconstruction/semantic-search-cwicr from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip, docker.
Without those the skill loads but fails at the first command.