amelnagdy/grok-delegate
>- Delegate a coding task to the Grok Build CLI as a background implementer, then review its diff and land it yourself. Use this whenever the user wants to hand implementation work to Grok — phrasings like "have Grok do X", "delegate this to Grok", "run it through Grok", "use Grok Build to implement/fix/refactor", or "have grok CLI do this" — or to run a queue of coding tasks through Grok while staying the reviewer. Prefer it when the user will review the diff and commit it themselves. DO NOT USE for tasks small enough to do inline, or when the user wants the code written directly without delegating.
npx skills add https://github.com/amElnagdy/delegate-skills --skill grok-delegate
You are the orchestrator. This skill lets you hand a bounded coding task to a separate
implementer — the Grok Build CLI (grok) — then review what it produced and land it yourself. You
write the brief and own the judgment; Grok does the typing under an explicit autonomy profile; you
verify and commit.
Nothing here is specific to one orchestrating agent. The loop needs only the ability to run a shell
command and read a file, so it works the same whether you are Claude Code, Cursor, OpenCode with a
selected model, or any comparable agent. (It is designed for Claude Code and Cursor; treat other
orchestrators as designed-for, not yet proven.)
grok CLI is not installed, not authenticated, or the account lacks Grok Build beta access.grok version succeeds. If not, install on any platform withnpm i -g @xai-official/grok (or use the installer from xAI's official Grok CLI docs) and
authenticate (grok login, or grok login --device-auth on headless hosts, or set
XAI_API_KEY).
grok is on PATH. command -v grok shows the active binary and grok versionits version — the relay records the version it ran into result.json, so a stale binary is visible
after the fact.
--cd at) the target git repository.Run these five steps per task. Steps 1, 4, and 5 are your judgment; 2 and 3 are mechanical.
Grok sees only the text you send — no orchestrator chat history, no shared context. Everything the
task needs goes in the brief: the goal, the current state, what to change, what to leave untouched,
the project's actual gate commands (discover them from the repo's CLAUDE.md/AGENTS.md/Makefile —
do not assume), and a report contract. Tell Grok it will not commit (you will). Keep one task per
brief. Full guidance and a template: references/writing-the-brief.md.
Send the brief to Grok with the bundled helper. It wraps grok -p, captures the run, and writes a
structured result.json — so your only job is "run a command, read a file." (<skill-dir> below is
this skill's installed directory — the folder containing this SKILL.md, i.e. the directory you loaded
the skill from. Claude Code prints it as "Base directory for this skill" when the skill loads; on other
orchestrators use that same directory — if unsure where it landed, run
find ~ -name relay.mjs -path '*grok-delegate*' and substitute the directory above it.)
node "<skill-dir>/scripts/relay.mjs" --brief brief.txt --cd /path/to/repo
# read-only (review/diagnosis; best-effort — verify touchedFiles): add --read-only
# continue the previous Grok session: add --resume-last (send only the delta brief)
# hard time limit (watchdog): add --timeout 2h (default: off; implementation runs routinely need 1-2h)
# see all options: node .../relay.mjs --help
The helper defaults to a write-capable (workspace-write) autonomy profile — --always-approve plus
--sandbox workspace — and writes its artifacts to a temp dir, so the repo under review stays clean.
It never commits — see step 5. Mechanics, flags, and the result.json shape:
references/dispatch-and-poll.md.
The helper blocks until Grok finishes, so back it with whatever your orchestrator offers and resume
when it returns:
run_in_background: true; you are notified on completion.the result file — … & in bash/zsh (including Git Bash/WSL), or your shell's equivalent (Start-Job
in PowerShell, start /b in cmd). The run is done when result.json exists with a status. (A
pre-run usage error — bad args or an empty brief — instead exits with code 2 and a stderr message and
writes no result file, so check the exit code too. A missing grok binary exits 127 but *does* write
a result.json with status grok_unavailable.)
Do not trust progress trackers over reality: a run is finished when result.json is written and the
process has exited. Read the working tree, not a status line. The implementer's full report is
the finalMessage field in result.json (also printed in full on stdout between the report markers).
Grok's result.json includes its own summary and gate claims. Re-verify, don't accept:
"gates passed" on faith.
nothing less? touchedFiles in the result is your starting point.
test-guard, etc. from guard-skills) — this skill produces the work; those skills judge it.
Full checklist: references/review-and-land.md.
The orchestrator commits. Only after the gates pass and the diff holds:
--resume-last (don't restate the whole task) andreview again.
Grok's default permission mode is ask, which blocks on approval prompts in a headless pipe. The
relay therefore always sets autonomy explicitly:
| Relay flag | What Grok gets | Use when |
| --- | --- | --- |
| *(default)* | --always-approve --sandbox workspace | Normal implementation — writes scoped to the working tree |
| --read-only | --sandbox read-only --permission-mode plan | Review / diagnosis — best-effort, not enforced (see caveat below) |
| --full-access | --always-approve --sandbox off | Explicit opt-in when the task needs unrestricted tools |
--always-approve alone would approve *all* tools (writes, shell, network) — closer to unrestricted
than to a workspace-scoped write. Pairing it with --sandbox workspace is what keeps the default
safe. Reach for --full-access only when the human asks for it.
--read-only is best-effort, not a hard guarantee. The read-only sandbox restricts out-of-workspace
filesystem/network access, not grok's own edit tool, and headless plan mode is advisory — a run
verified here still wrote the working tree when told to. Use --read-only to *signal* review intent,
but always confirm touchedFiles afterward; treat the diff, not the flag, as the guarantee. The relay
automates the check: it snapshots git status before a --read-only run and sets
readOnlyViolation: true in result.json (with a summary warning) when the tree changed anyway.
It's a porcelain-level tripwire — an edit inside an already-dirty file can evade it, so on a dirty
tree the diff review stays the only real guarantee.
Delegation is something the human opts into. Once they have ("run this queue", "proceed"), committing
verified, gate-passing work is the agreed contract — that is the whole point. Two limits on that
mandate: surface, don't absorb (report Grok's design decisions, defensible-but-unasked turns, and
non-blocking nitpicks rather than silently keeping them) and stop for scope changes (if correct
completion needs going beyond the brief, ask — don't expand the mandate yourself). The full treatment
is in references/review-and-land.md.
execute blind: structure, XML blocks, the report contract, embedding the real gate commands.
relay.mjs flags, theresult.json contract, backgrounding per orchestrator, and recovery when a run misbehaves.
boundary, and the rework cycle via --resume-last.
carrying constraints forward, progress tracking, and the end-of-run coherence check.
Take amelnagdy/grok-delegate from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.
The instructions reference npm.
Without those the skill loads but fails at the first command.