trailofbits/audit-augmentation
> Augments Trailmark code graphs with external audit findings from SARIF static analysis results, weAudit annotation files, and version-gated Trailmark 0.4.x binary-analysis graph exports. Maps findings to graph nodes by file and line overlap, creates severity-based subgraphs, and enables cross-referencing findings with pre-analysis data (blast radius, taint, etc.). Use when projecting SARIF results onto a code graph, overlaying weAudit annotations, importing binary graph findings, cross-referencing Semgrep, CodeQL, or binary-analysis findings with call graph data, or visualizing audit findings in the context of code structure.
npx skills add https://github.com/trailofbits/skills --skill audit-augmentation
Projects findings from external tools (SARIF) and human auditors (weAudit)
onto Trailmark code graphs as annotations and subgraphs. Trailmark 0.4.0+ can
also import an external binary-analysis graph JSON export via
engine.augment_binary().
trailmark-finding-triagetrailmark skill)diagramming-code skill after augmenting)| Rationalization | Why It's Wrong | Required Action |
|-----------------|----------------|-----------------|
| "The user only asked about SARIF, skip pre-analysis" | Without pre-analysis, you can't cross-reference findings with blast radius or taint | Always run engine.preanalysis() before augmenting |
| "Unmatched findings don't matter" | Unmatched findings may indicate parsing gaps or out-of-scope files | Report unmatched count and investigate if high |
| "One severity subgraph is enough" | Different severities need different triage workflows | Query all severity subgraphs, not just error |
| "SARIF results speak for themselves" | Findings without graph context lack blast radius and taint reachability | Cross-reference with pre-analysis subgraphs |
| "weAudit and SARIF overlap, pick one" | Human auditors and tools find different things | Import both when available |
| "Tool isn't installed, I'll do it manually" | Manual analysis misses what tooling catches | Install trailmark first |
MANDATORY: If uv run trailmark fails, install trailmark first:
uv pip install trailmark
SARIF and weAudit augmentation are v0.2-safe. Binary graph augmentation is
Trailmark 0.4.0+ only. Before calling engine.augment_binary(), check:
if not hasattr(engine, "augment_binary"):
raise RuntimeError("Binary augmentation requires Trailmark >= 0.4.0")
On Trailmark 0.5.0+, known links between source functions and imported binary
or external endpoints can also be declared once in .trailmark/links.toml
(see the main trailmark skill's Repository Links section) instead of being
re-derived per session. Declared external endpoints materialize as
proxy.external:<symbol> nodes on every parse.
# Augment with SARIF
uv run trailmark augment {targetDir} --sarif results.sarif
# Augment with weAudit
uv run trailmark augment {targetDir} --weaudit .vscode/alice.weaudit
# Both at once, output JSON
uv run trailmark augment {targetDir} \
--sarif results.sarif \
--weaudit .vscode/alice.weaudit \
--json
Binary graph augmentation is programmatic in Trailmark 0.4.0+; do not invent a
CLI flag if trailmark augment --help does not show one.
from trailmark.query.api import QueryEngine
engine = QueryEngine.from_directory("{targetDir}", language="auto")
# Run pre-analysis first for cross-referencing
engine.preanalysis()
# Augment with SARIF
result = engine.augment_sarif("results.sarif")
# result: {matched_findings: 12, unmatched_findings: 3, subgraphs_created: [...]}
# Augment with weAudit
result = engine.augment_weaudit(".vscode/alice.weaudit")
# Augment with an external binary graph export (v0.4+)
if hasattr(engine, "augment_binary"):
result = engine.augment_binary("binary_graph.json")
# Query findings
engine.findings() # All findings
engine.subgraph("sarif:error") # High-severity SARIF
engine.subgraph("weaudit:high") # High-severity weAudit
engine.subgraph("sarif:semgrep") # By tool name
engine.annotations_of("function_name") # Per-node lookup
If auto-detection is wrong for the target, rerun with an explicit language or
comma-separated list such as python,rust.
Augmentation Progress:
- [ ] Step 1: Build graph and run pre-analysis
- [ ] Step 2: Locate SARIF/weAudit/binary graph files
- [ ] Step 3: Run augmentation
- [ ] Step 4: Inspect results and subgraphs
- [ ] Step 5: Cross-reference with pre-analysis
Step 1: Build the graph and run pre-analysis for blast radius and taint
context:
engine = QueryEngine.from_directory("{targetDir}", language="auto")
engine.preanalysis()
If auto-detection is wrong for the target, rerun with an explicit language or
comma-separated list such as python,rust.
Step 2: Locate input files:
semgrep --sarif -o results.sarifor codeql database analyze --format=sarif-latest
.vscode/<username>.weaudit within the workspaceartifact, functions, andcalls fields. Trailmark imports this graph; it does not disassemble
binaries itself.
Step 3: Run augmentation via engine.augment_sarif() or
engine.augment_weaudit(). For binary graphs, run engine.augment_binary()
only after the Version Gate succeeds. Check unmatched_findings in SARIF and
weAudit results — these are findings whose file/line locations didn't overlap
any parsed code unit.
Step 4: Query findings and subgraphs. Use engine.findings() to list all
annotated nodes. Use engine.subgraph_names() to see available subgraphs.
Step 5: Cross-reference with pre-analysis data to prioritize:
sarif:error with tainted subgraphhigh_blast_radiusprivilege_boundaryFor one candidate finding that needs a reachability verdict or PoC handoff,
continue with trailmark-finding-triage and use the augmented node as the
bound candidate.
Findings are stored as standard Trailmark annotations:
finding (tool-generated) or audit_note (human notes)sarif:<tool_name> or weaudit:<author>[SEVERITY] rule-id: message (tool)
| Subgraph | Contents |
|----------|----------|
| sarif:error | Nodes with SARIF error-level findings |
| sarif:warning | Nodes with SARIF warning-level findings |
| sarif:note | Nodes with SARIF note-level findings |
| sarif:<tool> | Nodes flagged by a specific tool |
| weaudit:high | Nodes with high-severity weAudit findings |
| weaudit:medium | Nodes with medium-severity weAudit findings |
| weaudit:low | Nodes with low-severity weAudit findings |
| weaudit:findings | All weAudit findings (entryType=0) |
| weaudit:notes | All weAudit notes (entryType=1) |
| binary:<artifact> | Binary function nodes imported from a v0.4+ binary graph |
Findings are matched to graph nodes by file path and line range overlap:
root_pathlocation.file_path matches AND whose line range overlaps areselected
SARIF paths may be relative, absolute, or file:// URIs — all are handled.
weAudit uses 0-indexed lines which are converted to 1-indexed automatically.
Binary graph imports create origin=binary function nodes, origin=proxy
external proxy nodes for unresolved binary calls, and inferred
corresponds_to edges when a binary function maps back to a source node. The
expected JSON shape is intentionally small:
{
"artifact": {"name": "libexample", "architecture": "x86_64", "sha256": "..."},
"functions": [
{"symbol": "parse_packet", "address": "0x401000",
"source": {"file": "src/parser.c", "line": 42}}
],
"calls": [
{"source": "parse_packet", "target": "malloc", "confidence": "inferred"}
]
}
weAudit file format field reference
Take trailofbits/audit-augmentation from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference pip, uv.
Without those the skill loads but fails at the first command.