>- Builds a content-addressed semantic-unit index from source code for structural context. Use when a pinned or live codebase is available and structural cross-reference data would improve research quality. Don't use for findings analysis, patching, or reporting.
npx skills add https://github.com/google/mantis --skill mantis-structural-index
This is an optional first-class stage in the Pass Lifecycle Contract. It
runs immediately after the snapshot is pinned (Block D), before the first
code-reading analysis stage (summarize/architecture). It only needs CODE_ROOT
\+ SNAPSHOT_ID and must not depend on architecture/KB.
Structural Code Index Builder. Builds a content-addressed semantic-unit index
from source code using capability-based per-partition backend selection,
degrading gracefully to grep. Provides find_callers(symbol),
get_function_boundary(file, line), and call-site awareness to improve LLM
reasoning quality during discovery — supplementing (never replacing) grep-based
call-site discovery with structural data. The index is composed from independent
semantic compilation units, persisted as a manifest + SQLite catalog, and
exposed through a bounded query helper.
/mantis-structural-indexcode under CODE_ROOT.
--snapshot_root/--snapshot_id/--state_root. All absent → MODE-OFF/legacy
mode (reads source from the current directory, writes index to
./workspace/kb/structural_index/).
workspace/.mantis_state.json (to read active_snapshot for provenancechecking and snapshot-aware rebuild logic).
workspace/kb/structural_index/manifest.json (to check snapshot_id forreuse-on-match idempotency — primary check).
workspace/kb/structural_index.jsonl (backward-compat provenance check ifmanifest absent).
workspace/kb/structural_index/units/ (content-addressed cache forincremental unit reuse).
workspace/kb/structural_index/native/ and sidecar provenance.json files(prebuilt index attachments and metadata manifests).
source files under CODE_ROOT).
workspace/kb/structural_index/manifest.json (STATE-RELATIVE — atomiccommit point, written LAST).
workspace/kb/structural_index/catalog.sqlite (STATE-RELATIVE —query-optimized serving store).
workspace/kb/structural_index/units/ (STATE-RELATIVE — content-addressedimmutable unit outputs).
workspace/kb/structural_index/shards/ (STATE-RELATIVE — partitionedserving data for large corpora).
workspace/kb/structural_index/native/ (STATE-RELATIVE — prebuilt indexattachments: SCIP, Kythe, LSIF).
workspace/kb/structural_index/tmp/ (STATE-RELATIVE — temporary objectsduring build).
workspace/kb/structural_index.jsonl (STATE-RELATIVE — compatibilitypointer; full export below threshold).
workspace/helpers/build_structural_index.py (STATE-RELATIVE — the builderhelper script).
workspace/helpers/query_structural_index.py (STATE-RELATIVE — the queryhelper script).
CODE_ROOT. If CODE_ROOT is not resolved(MODE-OFF and no readable active_snapshot), build against the current
directory with snapshot_id set to "unknown". Do NOT skip — this is the
standalone-efficiency case.
invokes it (the harness, mantis-plan, or mantis-researcher). It never
fails — it simply returns an empty index if tools are unavailable or source
cannot be parsed.
CODE_ROOT. Writes only to STATE-RELATIVE paths. Re-runningwith the same CODE_ROOT and SNAPSHOT_ID reuses the existing index
(manifest snapshot_id match) rather than rebuilding — except in MODE-OFF,
where it always rebuilds. Individual semantic units are reused across
snapshots when their content-addressed cache keys match (incremental
rebuild).
LOCATOR RESOLUTION (before reading ANY target code or artifact):
0. ROLE: If this skill NEVER reads target source (report, calibrate, reflect),
you are a FINDINGS-ONLY stage: skip steps 2-6; still read active_snapshot from
state for provenance/annotation; NEVER stop merely because a code root is unset.
1. Determine CODE_ROOT, in this priority order:
a. If --target_root is passed on THIS invocation, CODE_ROOT = --target_root.
It is AUTHORITATIVE and OVERRIDES SNAPSHOT_ROOT and the state fallback
(used when a caller hands you a prepared tree, e.g. a patched shadow).
b. Else if --snapshot_root (or SNAPSHOT_ROOT) is passed, use it.
c. Else read state_root/workspace/.mantis_state.json (state_root from
--state_root if passed, else ./workspace/... relative to the current dir)
-> active_snapshot.root / .snapshot_id / .snapshot_pinned.
d. Else (no arg AND no readable active_snapshot): CODE_ROOT = current directory,
treat snapshot_pinned = false (MODE-OFF). Do NOT stop.
2. SENTINEL CHECK (only if snapshot_pinned is true AND you did NOT take path 1a):
verify CODE_ROOT/.mantis_snapshot_id exists and equals SNAPSHOT_ID. If missing
or different -> STOP "snapshot sentinel mismatch". (A --target_root tree (1a) is
deliberately mutated and is sentinel-EXEMPT.)
3. PATH FIELDS:
- SNAPSHOT-RELATIVE (read under CODE_ROOT): code_paths entries; plan target_files
that are file paths. Strip ONLY a trailing ":<digits>". A code_paths entry
containing "://" is a URL/endpoint, NOT a file read. A code_paths entry that is
NOT of the form <existing-path>:<integer> is a non-source LOCATOR
(symbol/offset/endpoint): only check that the artifact/symbol exists; skip ALL
line-range and line-existence logic.
- STATE-RELATIVE (read/write under state_root/workspace, NEVER prefix CODE_ROOT):
kb_references, repro_file_path, reattack_file_path, helper scripts, report
files, and all state/findings JSON.
4. Never WRITE under CODE_ROOT when snapshot_pinned is true. Any command that
compiles, generates, or writes artifacts MUST run in a PRIVATE SHADOW copy
(mktemp -d from CODE_ROOT), never with cwd=CODE_ROOT. Read-only inspection may
cd into CODE_ROOT.
5. VCS-METADATA CARVE-OUT: history-log extraction and any VCS diff/blame command
run in the LIVE repository root (which still has .git/.hg/.repo), NOT CODE_ROOT
(the snapshot copy strips VCS metadata). Do NOT stop merely because CODE_ROOT
lacks .git/.hg/.repo.
6. Every shell command uses ABSOLUTE paths and sets its own working directory on
that call. Do NOT assume the working directory persists between calls.
This is a CODE-READING stage — it reads target source files under CODE_ROOT via
the helper script. Block A step 0's findings-only skip does NOT apply.
MANDATORY FIRST STEP: Before writing any build scripts, probing backends, or
extracting symbols, check for MODE-OFF or an existing manifest:
SNAPSHOT_ID is "unknown" orabsent) → always rebuild. Do NOT reuse a previous "unknown" index,
because the live tree is mutable and "unknown" is a constant (not a
freshness signal). Skip directly to Step 2 (individual units whose content
has not changed may still hit the content-addressed cache).
workspace/kb/structural_index/manifest.json exists, read itssnapshot_id field.
snapshot_id matches the current SNAPSHOT_ID (and SNAPSHOT_ID is NOT"unknown") → reuse the index immediately and STOP. Do NOT invoke Step
2, probe backends, or write build scripts. This bounds cost across
retries/crash-resume.
manifest.json is absent but workspace/kb/structural_index.jsonl exists(backward compat), read its provenance header (_provenance, snapshot_id
keys). If snapshot_id matches (and is NOT "unknown") → **reuse and
STOP**. Otherwise proceed to rebuild.
SNAPSHOT_ID differs → proceed to rebuild (Steps 2–5). For incrementalreuse: before rebuilding a semantic unit, check units/ for a
content-addressed cache hit (see Content-Addressed Cache Key below). A cache
hit reuses the unit output without re-extraction.
Partition the codebase into semantic units (see Per-Language Semantic Units
below). For each unit, select the most precise backend available in this
environment using a capability-based, per-partition decision — not a single
global ladder. The following are examples, not an exhaustive enum. The
helper probes each tier per partition and selects the highest available:
(precision: semantic) — if a pre-built index matching the current
snapshot_id or root_fingerprint is available. Most precise: full
type-aware cross-reference, call hierarchy, and hover/signature data.
typecheck) — ifcompile_commands.json, build context, or typechecker is available.
Type-accurate symbol resolution and call edges.
ast) — if tree-sitter,ast-grep, or a language-specific parser is available. Full AST parsing:
function boundaries, call expressions, signatures.
symbol-only) — if ctags orequivalent is on PATH. Symbol table only (function definitions, locations —
no call graph). Call-site extraction uses a lightweight regex pass within
known function boundaries.
heuristic) — Python stdlib regex passover source files. Identifies function definitions and call patterns using
language-agnostic heuristics. Less precise but zero-dependency.
coverage-only) —grep; no structural index is written. Manifest records status: "empty".
Consumers fall back to grep-based discovery (today's behavior byte-for-byte).
Native Index Probing & Resolution Rules:
workspace/kb/structural_index/native/ and subdirectories
native/{scip,lsif,kythe}/ for prebuilt index files.
Kythe) do not embed snapshot identity directly in their binary payload,
prebuilt indexes MUST declare their target snapshot using a sidecar
provenance.json manifest located at
workspace/kb/structural_index/native/provenance.json or
native/<kind>/provenance.json. The manifest contains an array of
attachments:
[{"kind": "scip|lsif|kythe", "path": "...", "snapshot_id": "...", "root_fingerprint": "...", "language": "...", "indexer": "...", "precision": "semantic", "files": [...]}].
A native index is matched if its declared snapshot_id equals SNAPSHOT_ID
(when SNAPSHOT_ID != "unknown") or its root_fingerprint matches the
workspace's calculated root fingerprint. If the files array is absent or
empty, treat the native index as covering no individual files directly (record
in manifest.native_indexes but do not update coverage rows; fall through
to lower tiers for all files).
SCIP protobuf) in pure Python without dependencies is costly and complex.
Option A builder scripts MUST detect matching prebuilt native indexes, record
their entries in manifest.native_indexes, and set the coverage table
backend (e.g., "scip" or "scip-clangd") for all files listed in the
provenance manifest. If catalog.sqlite is NOT populated with symbols from
the native index (raw binary deferred to harness/MCP readers), set
coverage.status = "deferred" and precision = "deferred" (with
indexed_files = 0). This prevents the query helper from claiming an
un-ingested partition is "authoritative empty" at semantic precision,
ensuring consumers run the mandatory grep fallback. When catalog.sqlite IS
populated (e.g., via Option B pre-ingestion or scip-to-sqlite), set
precision = "semantic" and status = "indexed".
LSP is NOT equivalent to SCIP / LSIF. LSP is an interactive protocol whose
workspace state may be partial or mutable. Use it only when the server can
demonstrate snapshot identity AND complete workspace coverage. A running
language server does not automatically qualify as a semantic backend.
SCIP merging. SCIP explicitly permits merging complementary information from
indexers with different precision levels. Its format records indexer version and
per-document language metadata. The builder MAY merge results from multiple
indexers (e.g., a SCIP index for Go + tree-sitter for Python) within a single
catalog.
The determinism lives in a runtime-generated versioned helper
(build_structural_index.py, # MANTIS_HELPER_VERSION = 5, grep-and-regenerate
on reuse) that probes and selects backends per partition. No shipped binaries;
air-gapped-safe.
Deterministic partial coverage. On very large source trees, a full rebuild
can dominate stage-0 wall-clock. Replace discovery-order truncation with a
deterministic priority queue:
plan.json, if available).reverse-dependency edges in the existing catalog).
Apply deterministic max_units (default: 10000) or max_source_bytes (default:
500MB) bounds to that queue. Persist deferred units in the manifest so another
invocation can resume. Statuses: complete, partial, empty, failed.
A consumer should only interpret "no callers" strongly when the relevant
partition is complete, snapshot-matched, covered, and semantically precise.
Otherwise the answer is "no indexed callers", followed by the existing lexical
fallback. The index is HINT-only, so partial coverage is safe (grep remains
authoritative).
Two runtime-generated helpers are used. Both follow the grep-and-regenerate
pattern: before reuse, grep the first line for the version marker; if absent or
a different integer, REGENERATE.
build_structural_index.pyworkspace/helpers/build_structural_index.py. The FIRSTLINE MUST be exactly # MANTIS_HELPER_VERSION = 5. Before reusing an
existing helper, grep its first lines for MANTIS_HELPER_VERSION = 5; if
that marker is absent or a different integer, REGENERATE the helper.
Semantic Units), computes content-addressed cache keys, checks units/ for
reuse, selects a backend per partition (Step 2), extracts symbols + edges,
and writes results to catalog.sqlite.
symbol_id, name, qualified_name, namespace,language, file_path, start_line, end_line, kind, signature,
backend, precision.
caller_id, callee_id, callee_name, file_path,line, edge_kind.
symbol_id, file_path, start_line,end_line, signature, language.
file_path, indexed, backend, precision,unit_cache_key, status.
A step 6). It MUST NOT write anything under CODE_ROOT when
snapshot_pinned is true (Block A step 4). Backends that produce sidecar
files (ctags tags, cscope cscope.out, clangd cache) MUST be redirected to
STATE-RELATIVE paths: ctags -f <state>/helpers/tags,
cscope -f <state>/helpers/cscope.out,
CLANGD_INDEX_STORAGE=<state>/helpers/. Read-only LSP/SCIP queries to a
running server need no redirect.
query_structural_index.pyworkspace/helpers/query_structural_index.py. TheFIRST LINE MUST be exactly # MANTIS_HELPER_VERSION = 5. Before reusing an
existing helper, grep its first lines for MANTIS_HELPER_VERSION = 5; if
absent or a different integer, REGENERATE.
catalog.sqlite (or a remote endpoint — identical API). It IS the
consumption contract (see Query Interface below).
catalog.sqlite is absent but structural_index.jsonl exists, the queryhelper falls back to linear scanning of the JSONL file (slower but
functional). If both are absent, it returns empty results with
coverage.partition_status = "empty".
workspace/kb/structural_index/tmp/manifest.jsonfirst. Then atomically rename it to
workspace/kb/structural_index/manifest.json. This is the atomic commit
point — the manifest is written LAST, after all units, catalog, and coverage
data are written.
workspace/kb/structural_index.jsonl(STATE-RELATIVE — NEVER under CODE_ROOT). Below a configurable threshold
(default: 10K records), emit a complete JSONL export (provenance header + all
records). Above the threshold, emit only a provenance header with
compat_pointer.full_export = false and compat_pointer.symbol_count set.
Large consumers MUST use the query interface.
provider.backend_versions field records which backend was usedper language (e.g.,
{"go": {"backend_name": "scip-clangd", "precision": "semantic"}, "python": {"backend_name": "tree-sitter", "precision": "ast"}}),
so consumers know the precision level.
tmp/ withoutcorrupting the last published index. On resume, check manifest.json status
and snapshot_id (Step 1).
manifest.json, catalog.sqlite, and the query helper.Include a summary (unit count, symbol count, call-edge count, backends used,
coverage status).
the caller: "Structural index is empty — structural context unavailable."
harness, planner, or researcher.
workspace/kb/structural_index/
├── manifest.json # Atomic commit point — written LAST
├── catalog.sqlite # Query-optimized serving store (both directions indexed)
├── units/ # Content-addressed immutable unit outputs
│ └── ab/cd/abcdef... # sha256 prefix sharding (2+2 hex dirs)
├── shards/ # Partitioned serving data (large corpora)
│ └── shard_0000.sqlite
├── native/ # Prebuilt index attachments (SCIP, Kythe, LSIF)
│ ├── provenance.json # Prebuilt index provenance manifest
│ ├── scip/
│ └── kythe/
└── tmp/ # Temporary objects during build
workspace/helpers/
├── build_structural_index.py # Builder (MANTIS_HELPER_VERSION = 5)
└── query_structural_index.py # Query helper (MANTIS_HELPER_VERSION = 5)
workspace/kb/structural_index.jsonl # Compatibility pointer
manifest.json){
"schema_version": 1,
"snapshot_id": "<SNAPSHOT_ID or 'unknown'>",
"root_fingerprint": "<sha256 of sorted (path, content_sha256) for all source files>",
"status": "complete|partial|empty|failed",
"provider": {
"kind": "local-build|baseline+overlay|remote",
"catalog": "catalog.sqlite",
"backend_versions": {
"cpp": {"backend_name": "tree-sitter", "backend_version": "0.20.8", "precision": "ast"},
"go": {"backend_name": "scip-clangd", "backend_version": "0.2.3", "precision": "semantic"}
}
},
"units": {"total": 0, "reused": 0, "rebuild": 0, "failed": 0},
"coverage": {"total_files": 0, "indexed_files": 0, "failed_files": 0, "deferred_files": 0},
"shards": [{"id": "", "path": "", "checksum": "", "partition_key": "", "symbol_count": 0, "edge_count": 0}],
"deferred_units": [{"unit_id": "", "language": "", "files": [], "priority": 4, "reason": ""}],
"native_indexes": [{"kind": "scip", "path": "", "snapshot_id": "", "root_fingerprint": "", "language": "", "indexer": "", "precision": ""}],
"baseline": {"source": "ci|local|none", "snapshot_id": "", "manifest_path": ""},
"overlay": {"units_added": 0, "units_modified": 0, "files": []},
"compat_pointer": {"path": "structural_index.jsonl", "full_export": true, "symbol_count": 0},
"created_at": "<ISO 8601>",
"build_duration_ms": 0
}
Atomic commit: The manifest is written LAST (atomic rename from tmp/).
Interrupted builds leave unreferenced temp objects without corrupting the last
published index.
catalog.sqlite)CREATE TABLE IF NOT EXISTS schema_meta (
key TEXT PRIMARY KEY, value TEXT NOT NULL
);
CREATE TABLE IF NOT EXISTS symbols (
symbol_id TEXT PRIMARY KEY,
name TEXT NOT NULL,
qualified_name TEXT NOT NULL,
namespace TEXT,
language TEXT NOT NULL,
file_path TEXT NOT NULL,
start_line INTEGER NOT NULL,
end_line INTEGER,
kind TEXT NOT NULL,
signature TEXT,
backend TEXT NOT NULL,
precision TEXT NOT NULL CHECK (precision IN ('semantic','typecheck','ast','symbol-only','heuristic','deferred','coverage-only')),
corpus TEXT DEFAULT 'default',
partition_key TEXT,
unit_cache_key TEXT,
source_layer TEXT NOT NULL DEFAULT 'baseline'
);
CREATE TABLE IF NOT EXISTS call_edges (
edge_id INTEGER PRIMARY KEY AUTOINCREMENT,
caller_id TEXT NOT NULL,
callee_id TEXT,
callee_name TEXT NOT NULL,
file_path TEXT NOT NULL,
line INTEGER NOT NULL,
edge_kind TEXT NOT NULL CHECK (edge_kind IN ('direct','indirect','virtual','macro','unresolved')),
corpus TEXT DEFAULT 'default',
partition_key TEXT,
source_layer TEXT NOT NULL DEFAULT 'baseline',
FOREIGN KEY (caller_id) REFERENCES symbols(symbol_id),
FOREIGN KEY (callee_id) REFERENCES symbols(symbol_id)
);
CREATE TABLE IF NOT EXISTS function_boundaries (
symbol_id TEXT PRIMARY KEY,
file_path TEXT NOT NULL,
start_line INTEGER NOT NULL,
end_line INTEGER NOT NULL,
signature TEXT,
language TEXT NOT NULL,
FOREIGN KEY (symbol_id) REFERENCES symbols(symbol_id)
);
CREATE TABLE IF NOT EXISTS coverage (
file_path TEXT NOT NULL,
indexed INTEGER NOT NULL DEFAULT 0,
backend TEXT,
precision TEXT,
unit_cache_key TEXT,
status TEXT NOT NULL DEFAULT 'pending',
corpus TEXT DEFAULT 'default',
partition_key TEXT,
PRIMARY KEY (file_path, corpus)
);
CREATE TABLE IF NOT EXISTS unit_cache (
cache_key TEXT PRIMARY KEY,
unit_id TEXT NOT NULL,
language TEXT NOT NULL,
extractor_name TEXT NOT NULL,
extractor_version TEXT NOT NULL,
file_count INTEGER NOT NULL,
symbol_count INTEGER NOT NULL,
edge_count INTEGER NOT NULL,
source_bytes INTEGER NOT NULL,
created_at TEXT NOT NULL,
snapshot_id TEXT NOT NULL
);
CREATE TABLE IF NOT EXISTS deferred_units (
unit_id TEXT PRIMARY KEY,
language TEXT NOT NULL,
files TEXT NOT NULL,
priority INTEGER NOT NULL,
reason TEXT NOT NULL,
cache_key TEXT,
created_at TEXT NOT NULL
);
-- Indexes: both directions!
CREATE INDEX IF NOT EXISTS idx_symbols_name ON symbols(name);
CREATE INDEX IF NOT EXISTS idx_symbols_qualified ON symbols(qualified_name);
CREATE INDEX IF NOT EXISTS idx_symbols_file ON symbols(file_path, start_line);
CREATE INDEX IF NOT EXISTS idx_symbols_lang ON symbols(language);
CREATE INDEX IF NOT EXISTS idx_edges_caller ON call_edges(caller_id);
CREATE INDEX IF NOT EXISTS idx_edges_callee ON call_edges(callee_id);
CREATE INDEX IF NOT EXISTS idx_edges_callee_name ON call_edges(callee_name);
CREATE INDEX IF NOT EXISTS idx_edges_file ON call_edges(file_path, line);
CREATE INDEX IF NOT EXISTS idx_boundaries_file ON function_boundaries(file_path, start_line, end_line);
CREATE INDEX IF NOT EXISTS idx_coverage_status ON coverage(status);
CREATE INDEX IF NOT EXISTS idx_deferred_priority ON deferred_units(priority);
Partitioning: When symbols exceed 500K or edges exceed 2M, partition into
shard SQLite files under shards/ by lang:{language}:bucket:{NN} (stable
hash). Each shard has the same schema. The query helper routes queries to the
correct shard(s) based on the manifest.
cache_key = sha256(
schema_version +
extractor_name@version +
language +
compile_context_digest +
ordered_input_content_digests +
dependency_or_interface_digest
)
Critical: snapshot_id is NOT in the cache key. It goes in provenance only.
This permits reuse across commits, branches, and snapshots with identical units.
Native IDs (when semantic backend is available):
scip:{symbol}kythe:{uri}clangd:{usr}Fallback IDs (when no semantic backend):
fallback:{language}:{file_path}:{sha256(qualified_name|start_line|signature)[:16]}
This distinguishes namespaces, overloads, methods, local functions, and
duplicate names across languages or repositories.
Edge kinds are kept explicitly separate — never silently merged:
direct — statically resolved call to a known symbolindirect — function pointer, closure, callbackvirtual — virtual method dispatch (runtime-resolved)macro — call introduced by macro expansionunresolved — callee name found but no symbol_id resolved| Language | Unit Type | Compile Context | Dependency Digest |
| ----------- | ---------------- | --------------------------- | ------------------------------------------- |
| C/C++ | Compilation unit | compile_commands.json entry | sha256 of transitive header interfaces |
| Go | Package | go.mod + build tags | sha256 of imported packages' exported API |
| Rust | Crate | Cargo.toml + features | sha256 of extern crate signatures |
| Java/Kotlin | Compiler batch | module path + classpath | sha256 of imported class signatures |
| TypeScript | Project | tsconfig.json | sha256 of imported module type declarations |
| Fallback | Individual file | None | Empty string |
Fallback Granularity Rule: For the fallback tier (regex, AST parser, or
lightweight crawler without a language build system), **each source file MUST be
its own independent semantic unit**. Never bundle multiple source files into a
single fallback unit. Per-file unit isolation is essential for content-addressed
cache efficiency — when one file changes in a future snapshot, only that single
file's unit invalidates while all unchanged files hit the cache
(reused = N-1).
baseline.
At query time, the query helper merges baseline + overlay results. Overlay
symbols supersede baseline symbols for the same file. The manifest records
provider.kind = "baseline+overlay" with baseline manifest reference and
overlay unit count.
The old structural_index.jsonl format is retained as a compatibility export
only — it is NOT the query contract. Below the threshold (default 10K records),
a complete JSONL export is emitted. Above the threshold, only a provenance
pointer is written. Large consumers MUST use the query interface.
Line 1 — provenance header:
{"_provenance": true, "snapshot_id": "abc123", "tool": "tree-sitter"}
Lines 2+ — structural records (one per line, _type discriminator):
{"_type": "function", "key": "src/parser.c:parse_input", "start_line": 45, "end_line": 120, "signature": "int parse_input(char *buf, size_t len)", "calls": ["malloc", "validate_input", "memcpy"]}
{"_type": "call_edge", "caller": "parse_input", "callee": "malloc", "file": "src/parser.c", "line": 78}
{"_type": "call_edge", "caller": "main", "callee": "parse_input", "file": "src/main.c", "line": 203}
Provenance header fields:
| Field | Type | Description |
| ------------- | ------ | ----------------------------------------------------------------------------------- |
| _provenance | bool | Always true — marks this as the provenance header line |
| snapshot_id | string | SNAPSHOT_ID the index was built against |
| tool | string | Backend used (e.g. "lsp-clangd", "tree-sitter", "ctags", "regex", "grep") |
Record types:
| _type | Description | Key fields |
| ----------- | ----------------------------------------------- | ----------------------------------------------------- |
| function | Function definition with boundary and signature | key, start_line, end_line, signature, calls |
| call_edge | A call from caller to callee at file:line | caller, callee, file, line |
CODE_ROOT, not live tree (when pinned). When the snapshot ispinned, the structural index is built from the pinned CODE_ROOT, ensuring
it reflects the exact bytes the pipeline is analyzing.
manifest.json already carries the currentSNAPSHOT_ID (and SNAPSHOT_ID is NOT "unknown" — in MODE-OFF, always
rebuild), reuse it — do not rebuild. Rebuild when SNAPSHOT_ID differs or is
"unknown". Individual units may still be reused from the content-addressed
cache. This bounds cost across retries/crash-resume.
tmp/. Interrupted builds leave unreferenced temp objects without corrupting
the last published index. On resume, check manifest.json status and
snapshot_id.
snapshot_pinned is false (HALT), theindex may be built from the unpinned CODE_ROOT but is marked as potentially
stale. Consumers treat structural hints as advisory.
active_snapshot is absent(MODE-OFF), build against the current directory (cwd) with provenance
snapshot_id set to "unknown". This is the standalone-efficiency case —
the index is still useful for the researcher/planner even without snapshot
pinning. Do NOT return an empty index merely because the snapshot is absent.
These are runtime instructions for callers — they define how consumers use the
structural index. The structural index is a HINT-only enhancement; skills that
do not use it behave exactly as they do today.
The query helper IS the contract. All consumers use
query_structural_index.py (or a compatible remote endpoint). The JSONL file is
NOT the query contract — it is a compatibility export only.
Operations:
resolve_symbol(name, language?, file?, namespace?) →{results, total, ambiguous, coverage}
find_callers / find_callees.find_callers(symbol_id, limit=100, offset=0) →{results, total, has_more, coverage}
precision, backend, edge_kind.coverage.partition_status.find_callees(symbol_id, limit=100, offset=0) →{results, total, has_more, coverage}
find_callers.get_function_boundary(file, line) →{symbol_id, start_line, end_line, signature, precision, backend}
get_coverage(file?) →{total_files, indexed_files, failed, deferred, partition_status, backends_used}
Key properties:
ambiguous symbols).
MANTIS_STRUCTURAL_INDEX_URL env var or manifest provider.kind = "remote").
Coverage on empty results:
| partition_status | Meaning | Consumer action |
| ------------------ | ----------------------------- | -------------------------------------------------------- |
| complete | All files indexed | "No indexed callers" (still run grep per HINT-only rule) |
| partial | Some files deferred or failed | "Not fully indexed" — MUST run grep fallback |
| empty | No backend available | "Not indexed" — MUST run grep fallback |
| failed | Backend attempted but failed | "Index failed" — MUST run grep fallback |
When planning investigations, call resolve_symbol() then find_callers() to
identify all functions that call into a target — this broadens the audit set
beyond single-file analysis.
first based on call-graph centrality), never MEMBERSHIP. It may broaden the
audit set (safe over-reporting), but must never REMOVE or drop a file; the
planner's existing logic remains the membership floor.
the exhaustive set of candidate call-sites — this is the mandatory floor. Then
use the structural index query helper (resolve_symbol then find_callers)
to RANK and prioritize which call-sites to audit first (the index
distinguishes actual calls from comments/strings/variable names). Audit the
union of both result sets — the structural index may miss macro-based calls,
function pointers, and dynamic dispatch, so grep remains the floor.
get_function_boundary(file, line) to start withthe enclosing function, expanding to callers/callees/file as needed for
cross-function context — this saves context while preserving coverage.
(today's behavior). The structural index is a coverage HINT only — it improves
audit quality but is never required.
Non-negotiable invariants:
empty index → grep fallback = today's behavior byte-for-byte. Optional in the
Pass Lifecycle Contract; a non-conformant harness simply skips it.
grep; a symbol the index misses must still be reachable by the exhaustive
sweep. The structural index decides ORDER, never MEMBERSHIP. The query
contract enforces this: find_callers returns HINTs, never authoritative
membership — consumers MUST union with grep.
drop a finding.
reachability ("is it compiled into production") is explicitly out of scope —
a dev customization, not part of this skill, because absence-from-a-build can
hide a real finding (INV-2). Do not fold build-derived reachability into this
skill.
A reference blueprint is available at
mantis-pipeline-adapter/references/mantis-structural-index.md.
It is a stub that points to this SKILL.md as the single source of truth — do not
duplicate spec content there.
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
Automatically creates user-facing changelogs from git commits by analyzing commit history, categorizing changes, and transforming technical commits into clear, customer-friendly release notes. Turns hours of manual changelog writing into minutes of automated generation.
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
React Native and Expo best practices for building performant mobile apps. Use when building React Native components, optimizing list performance, implementing animations, or working with native modules. Triggers on tasks involving React Native, Expo, mobile performance, or native platform APIs.
React and Next.js performance optimization guidelines from Vercel Engineering. This skill should be used when writing, reviewing, or refactoring React/Next.js code to ensure optimal performance patterns. Triggers on tasks involving React components, Next.js pages, data fetching, bundle optimization, or performance improvements.
Next.js best practices - file conventions, RSC boundaries, data patterns, async APIs, metadata, error handling, route handlers, image/font optimization, bundling
Use when starting feature work that needs isolation from current workspace or before executing implementation plans - creates isolated git worktrees with smart directory selection and safety verification
Take google/mantis-structural-index from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.