mcpbeat Sign in

Lexical Kb Skill for Claude

Query a portable, embedding-free lexical knowledgebase bundled as a `.skill`. Use when the user references this KB or asks a question whose answer is in its corpus ({{SOURCE}}). Retrieval is BM25 over a precomputed inverted index — there is no embedding model, so YOU expand the query into search terms before searching. Bundle holds index.json + chunks.jsonl + search.js + search.py; pure stdlib, no install, no network.

1k tokens
context cost
the whole folder, loaded on every use
1
files
ships runnable scripts
0
copies elsewhere
how many repositories repackaged it
137
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/oaustegard/claude-skills --skill lexical-kb

The instruction itself

7 sections, as written by the author

lexical-kb — query an embedding-free knowledgebase

This KB has no semantic search and no embedding model. Retrieval is pure

lexical BM25 over a precomputed inverted index. That design moves one job onto

you: bridging the gap between how the user phrases a question and how the corpus

phrases the answer. An embedding model would do this with a vector; here **you

are the semantic layer** — you expand the query into terms before searching.

Corpus: {{SOURCE}} ({{CHUNK_COUNT}} chunks).

The retrieval protocol — follow every step

A raw user question fed straight to BM25 underperforms: it matches only the

exact words the user happened to use. The expansion step is what makes lexical

retrieval competitive with embeddings. Do not skip it.

  • Read the question. Extract core terms — the essential nouns, proper

nouns, and identifiers the answer MUST contain. These carry full weight.

  • Generate expand terms — synonyms, morphological variants (plural/verb

forms), acronym expansions and contractions, and adjacent concepts. These

carry lower weight. This is the work the missing embedding model would have

done. Be generous: 5–15 expansion terms is normal.

  • Run the searcher. It ships in this bundle in two equivalent runtimes —

node search.js or python3 search.py, identical flags and identical

results. Use whichever your environment has. Pass the user's original

question via --query AND your term groups — expansion is additive, it

never replaces the user's words:

   node search.js \
     --query "how does centered simhash differ from random projection?" \
     --core "simhash" --core "centered" \
     --expand "random projection" --expand "hyperplane" --expand "LSH" \
     --expand "binary quantization" --expand "hamming distance" \
     --k 5

--core/--expand are repeatable; pass phrases, the searcher tokenizes

them. The --query terms contribute at a low floor weight so a curated

synonym can lift a result but can never drop a doc the literal question would

have matched. Defaults: core 1.0, expand 0.4, query-floor 0.25, top-k 5.

Keep expansion targeted — terms too generic ("system", "process") leak into

unrelated chunks and blur the ranking. A precise word can mislead too if it

is polysemous: prefer the disambiguating phrase as one --core term (e.g.

--core "centered simhash") over a bare ambiguous word (--core "centered",

which also matches "centered around …" in unrelated chunks). The passage

extractor is lexical too, so it will highlight the wrong sense rather than

correct it.

  • Read the returned chunks. Answer from them, and cite chunk ids inline.

The chunks are the source of truth the user installed. When a chunk

contradicts your prior knowledge, the chunk wins — say so. When the chunks

do not contain the answer, say that plainly rather than filling the gap from

memory.

When you cannot expand — RM3 fallback

If the query is outside any domain you can expand confidently, pass it raw with

pseudo-relevance feedback. The searcher harvests expansion terms from the

corpus's own top hits — model-free, weaker than your expansion, and prone to

drift when the first pass is off-topic, so prefer real expansion when you can:

node search.js --query "the user's raw question" --rm3 --k 5

Metadata filtering

Each chunk carries structured meta (e.g. title, source_path, section).

Filter on it with --filter (repeatable). Filtering narrows by attribute; it

does not rank — combine it with term search.

node search.js --core "factions" --filter "section=blog" --filter "date>=2025" --k 5

Operators: =, !=, ~ (substring), >, >=, <, <= (numeric when both

sides parse, else lexicographic — ISO dates sort correctly).

Passage vs. full document

Ranking uses the whole chunk (best recall), but each hit's text is by default

the query-densest passage of that chunk (~1200 chars), not the entire chunk —

so your reasoning context is signal, not the surrounding noise of a long document.

Each matched sentence keeps its neighbours (--context, default 1 each side) so

it reads in context rather than as an orphaned fragment, and nearby matches merge

into contiguous passages; ' … ' marks elisions between them. When a hit is a

passage, the result carries full_chars (the chunk's full length). If you need

the complete document for a hit — broader context, a quote in a section the

passage elided — re-run with --snippet 0:

node search.js --query "…" --core "…" --snippet 0 --k 3

Tune with --snippet 2000/600 (budget) and --context 2/0 (neighbours).

Output

The searcher prints JSON: {"hits": [{id, score, text, meta, full_chars?}, ...]},

sorted by descending BM25 score. text is the focused passage (or the full chunk

if it was already short / --snippet 0); full_chars appears only when text

is a passage. Surface the top hits to the user with their ids, then answer using

them as authoritative context.

Mechanics

  • Pure stdlib — Node or Python. No npm install / pip install, no model

download, no network.

  • The bundle is self-contained: search.js + search.py (pick one),

index.json, chunks.jsonl. Run the searcher from inside the bundle

directory (it defaults --index to its own location) or pass

--index /path/to/bundle.

Other skills for the same job

different authors, same section of the catalogue
Skill Creator
by anthropics
vendor ×10

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

56k tokens scripts
Geo Database
by christophacham
×4

Access NCBI GEO for gene expression/genomics data. Search/download microarray and RNA-seq datasets (GSE, GSM, GPL), retrieve SOFT/Matrix files, for transcriptomics and expression analysis.

12k tokens
Pymc Bayesian Modeling
by christophacham
×4

Bayesian modeling with PyMC. Build hierarchical models, MCMC (NUTS), variational inference, LOO/WAIC comparison, posterior checks, for probabilistic programming and inference.

24k tokens scripts
Pymoo
by christophacham
×4

Multi-objective optimization framework. NSGA-II, NSGA-III, MOEA/D, Pareto fronts, constraint handling, benchmarks (ZDT, DTLZ), for engineering design and optimization problems.

19k tokens scripts
Statsmodels
by ComeOnOliver
×4

Statistical modeling toolkit. OLS, GLM, logistic, ARIMA, time series, hypothesis tests, diagnostics, AIC/BIC, for rigorous statistical inference and econometric analysis.

41k tokens
Add Uint Support
by pytorch
vendor ×3

Add unsigned integer (uint) type support to PyTorch operators by updating AT_DISPATCH macros. Use when adding support for uint16, uint32, uint64 types to operators, kernels, or when user mentions enabling unsigned types, barebones unsigned types, or uint support.

2k tokens
At Dispatch V2
by pytorch
vendor ×3

Convert PyTorch AT_DISPATCH macros to AT_DISPATCH_V2 format in ATen C++ code. Use when porting AT_DISPATCH_ALL_TYPES_AND*, AT_DISPATCH_FLOATING_TYPES*, or other dispatch macros to the new v2 API. For ATen kernel files, CUDA kernels, and native operator implementations.

2k tokens
Docstring
by pytorch
vendor ×3

Write docstrings for PyTorch functions and methods following PyTorch conventions. Use when writing or updating docstrings in PyTorch code.

3k tokens

How to use it

Copy the folder

Take oaustegard/lexical-kb from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.

Install what it needs

The instructions reference pip, npm. Without those the skill loads but fails at the first command.