MCP Local RAG runs on your own machine — the client starts it, so there is no endpoint to ping. 105 installs a week from npm. Last commit 23 Jul 2026.
Semantic code & doc search with keyword boost. AST code nav, auto HF mirror, local, privacy-first.
Today is the operative word: we check MCP Local RAG every 15 minutes and re-read its code on every release. Watch it and you find out the day that stops being true.
This server runs on your own machine — install it with the package manager and the client starts it for you. Package name taken from the official registry entry.
claude mcp add mcp-local-rag -- npx -y @damoqiongqiu/mcp-local-rag
{
"mcpServers": {
"mcp-local-rag": {
"args": [
"-y",
"@damoqiongqiu/mcp-local-rag"
],
"command": "npx"
}
}
}
[mcp_servers.mcp-local-rag]
command = "npx"
args = ["-y", "@damoqiongqiu/mcp-local-rag"]
{
"mcpServers": {
"mcp-local-rag": {
"args": [
"-y",
"@damoqiongqiu/mcp-local-rag"
],
"command": "npx"
}
}
}
{
"mcpServers": {
"mcp-local-rag": {
"args": [
"-y",
"@damoqiongqiu/mcp-local-rag"
],
"command": "npx"
}
}
}
This one needs environment variables set before it will start:
BASE_DIR (Base directory for document storage (defaults to current working directory). Ignored when BASE_DIRS is set.), BASE_DIRS (JSON array of base directories (e.g. '["/a","/b"]'). Takes precedence over BASE_DIR.), DB_PATH (Path to LanceDB database directory (defaults to ./lancedb/)), CACHE_DIR (Directory where Transformers.js models are cached (defaults to ./models/)), MODEL_NAME (Embedding model name (defaults to Xenova/all-MiniLM-L6-v2)), MAX_FILE_SIZE (Maximum file size in bytes (defaults to 104857600 / 100MB)), RAG_MAX_DISTANCE (Maximum distance threshold for filtering search results. Results with distance greater than this value will be excluded. Lower values mean stricter filtering (e.g., 0.5 for high relevance only)), RAG_GROUPING (Grouping mode for quality filtering. 'similar' returns only the most similar group (stops at first distance jump). 'related' includes related groups (stops at second distance jump). Unset means no grouping filter), RAG_MAX_FILES (Maximum number of files to keep in search results. Results are filtered to include only chunks from the top N best-scoring files. For example, 1 returns only the single best-matching file's chunks. Unset means no file filtering.), CHUNK_MIN_LENGTH (Minimum chunk length in characters (1-10000, defaults to 50). Chunks shorter than this threshold are filtered out during ingestion.), RAG_DEVICE (Execution device for the embedder (defaults to cpu). Passed straight to ONNX Runtime; see the Transformers.js device source for the supported backend names. If the requested device fails to initialize, the server throws an error.), RAG_DTYPE (Embedding quantization dtype for the embedder (defaults to fp32). Opt-in and pass-through; accepts any dtype the chosen model provides (fp32, fp16, q8, int8, ...). If the model has no variant for the requested dtype, the server throws an error. Changing this changes the embedding space — re-ingest existing data.), RAG_HYBRID_WEIGHT (Keyword boost factor for hybrid search (0.0-1.0, defaults to 0.6). 0 means semantic similarity only; higher values increase the keyword-match contribution to the final score.), RAG_WATCH (Enable automatic file watcher to detect changes and re-index modified files (true/1 to enable). Useful during active development.), HF_ENDPOINT (Explicit HuggingFace endpoint URL to use for model downloads. When set, auto-mirror detection is skipped and this URL is used directly. Useful when you have a full-featured mirror.), HF_AUTO_MIRROR (Enable automatic mirror detection for HuggingFace downloads (defaults to true). Set to "false" or "0" to disable and use huggingface.co directly. The mirror chain is: huggingface.co → hf-mirror.com → modelscope.cn.), HTTPS_PROXY (HTTPS proxy URL for downloading embedding models (e.g. http://proxy:8080). Required only if behind a corporate proxy.), HTTP_PROXY (HTTP proxy URL for downloading embedding models. Fallback if HTTPS_PROXY is not set.).
The author declared them in the registry entry; get the values from the project itself.
Index source code into a local knowledge base, search with keyword + semantic + hybrid modes.
Local semantic code search with Git history. Works offline, no API key needed.
Local file search for AI agents — semantic + keyword indexing over your codebase.
Local-first semantic code search across all your repos. Private context for AI coding assistants.
Semantic codebase search + persistent working memory for AI code editors. Local, no API key.
Local, privacy-focused RAG service for code search via MCP. https://linggen.dev
Local repo context for coding assistants with semantic search and graph relationships.
Productivity-boosting RAG engine for codebases with multi-provider AI support and semantic search.
Answers built from our own checks of this server.