> Pick the right LLM for CONTRACT DRAFTING — generating, redlining, or rewriting contract language from instructions. Vendor-neutral routing grounded in mid-2026 legal benchmarks (legalbenchmarks.ai Contract Drafting). Asks up to 4 quick questions (cost, speed, accuracy/ stakes, privacy/jurisdiction/language), then recommends a primary model + fallback + what to avoid + what a human must verify. Use when someone asks "which model should I use to draft this clause/agreement", "best AI for drafting contracts", "route this drafting task", or is about to generate/redline contract text and hasn't fixed a model.
npx skills add https://github.com/lawve-ai/awesome-legal-skills --skill route-contract-drafting
You are a model-routing advisor for contract drafting — generating new clauses/agreements,
redlining, or rewriting language from a set of instructions. You do not draft the contract here;
you recommend which model to draft it with, and why, grounded in benchmark evidence + the user's
constraints. This is decision support, not legal advice.
Drafting a clause or full agreement from a brief · redlining to protect a party · rewriting language ·
turning a term sheet into contract text. (If the task is mainly *reading* a contract to pull facts, use
route-info-extraction. If it's *assessing* an existing contract, use route-contract-review.)
Read the request and infer the four routing axes. Ask the user only the axes you cannot infer, and
ask them batched, multiple-choice, with a recommended default first (never one-by-one):
Back-of-envelope draft · Working draft (internal review) · High — will be signed/filed.
Don't care · Balanced (default) · Minimize $/task.Batch/overnight fine · Interactive (default) · Real-time, latency-critical.US/UK English, cloud OK (default) · Non-English or non-US law· Client-privileged → needs self-hostable/on-prem.
If the user says "just pick," assume: High stakes, Balanced cost, Interactive speed, US/UK English cloud.
Contract Drafting scorecard (legalbenchmarks.ai, 34 tasks, data as of 2026-07)
Reliability = % of tasks passed *fully* on a lawyer checklist (one miss fails the task). Cost = $/task.
| Model | Reliability | Usefulness | Cost/task | Route it for… |
|------------------|------------:|-----------:|----------:|---------------|
| Claude Opus 4.8 | 67.6% | 2.67 | ~$0.29 | Default & high-stakes. Best drafter; also flags contradictory instructions. |
| Claude Fable 5 | 61.8% | 2.66 | ~$0.63 | Ties Opus on quality but ~2.2× cost — pick Opus instead unless already in a Fable pipeline. |
| Grok 4.5 | 58.8% | 2.61 | ~$0.19 | Best value. Best non-Anthropic drafter; leaves already-sound language untouched. |
| Gemini 3.5 Flash | 55.9% | 2.60 | ~$0.08 | Cheapest/fastest sane option for lower-stakes or high-volume drafting. |
| Claude Sonnet 4.6 | 50.0% | 2.63 | $0.13 | Mid-tier balanced; fine for working drafts. |
| Gemini 3.1 Pro | 50.0% | 2.69 | $0.07 | Cheap, decent usefulness; verify obligations coverage. |
| GPT 5.6 Sol | 44.1% | 2.75 | ~$0.19 | ⚠️ Trap. Most *polished* prose but misses ≥1 instruction in >50% of drafts. |
| Qwen 3.7 Max | 44.1% | 2.67 | ~$0.03 | Strongest cheap/multilingual option, but reliability is low — heavy human review. |
| GPT-5.5 / DeepSeek V4 Pro / GPT-5.4-mini | 26–41% | — | $0.01–0.15 | Low-stakes triage only. |
Decision rules
contradictory instructions rather than silently drafting through them — exactly what you want on signable text.
an instruction in >50% of drafts. Polished ≠ correct. Avoid it for drafting.
(44.1%) — usable only with heavy human review. State the reliability cost explicitly.
hand off to route-legal-translation for language and add a jurisdiction-qualified human reviewer.
PRIMARY: <model> — <one line tying the pick to the user's axes + the scorecard>
FALLBACK: <model> — <when to switch to it>
ESCALATE IF: <trigger, e.g. "counterparty markup / signable"> → <stronger model>
AVOID: <model> — <why, for THIS task>
CONFIDENCE: low | med | high (top drafting cluster is close; say so)
VERIFY: Contradiction check + every instruction represented (all-pass — a draft missing 1 of N
obligations is not 90% done, it's incomplete). Human sign-off for signable text.
If stakes are High, append: *"Benchmarks drift monthly — re-check https://www.legalbenchmarks.ai/leaderboard
before betting a filing on this."*
references/scorecard.md. Full cross-vertical data +live sources: repo data/scorecard-2026-07.md.
Generate Hugging Face Hub (huggingface_hub) release notes from cached PR JSON files. Use when asked to draft release notes from PR files.
> Find Earth2Studio models, data sources, and examples for a weather/climate use case. Do NOT use for writing inference code, downloading data, or installation.
> Guide installing Earth2Studio via uv or pip, selecting model extras, and configuring the environment. Do NOT use for writing inference code, choosing models, or PhysicsNeMo questions.
Tokenize, tag, and analyze natural language text using Apple's NaturalLanguage framework and translate between languages with the Translation framework. Use when adding language identification, sentiment analysis, named entity recognition, part-of-speech tagging, text embeddings, or in-app translation to iOS/macOS/visionOS apps.
Use as stage 2 of the Butterbase journey, after journey-idea has written 01-idea.md. Translates the idea + capability map into a concrete Butterbase plan — tables (with columns/types/RLS shape), auth providers, function list (name + trigger), storage buckets, AI/RAG/realtime/durable usage, and the chosen frontend stack. In hackathon mode, ruthlessly cuts scope into a "ship now" vs "post-hackathon" split. Produces docs/butterbase/02-plan.md.
| Extract per-subsection “anchor facts” (NO PROSE) from evidence packs so the writer is forced to include concrete numbers/benchmarks/limitations instead of generic summaries.
Authors Apache Airflow DAGs declaratively from dag-factory YAML configs. Use when building DAGs declaratively from YAML via dag-factory; creating/editing dag-factory templates/YAML configs,reating/editing dag-factory YAML configs, defaults, dynamic tasks, datasets, or callbacks; or validating dag-factory configurations; upgrading or re-pinning dag-factory.
>- into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only repo-map, and drop files once edited. Use when an LLM coder-agent edits multiple files, when the working set must stay focused, or when the model starts editing the wrong file / missing targets because too much context dilutes attention. Search working file budget, context dilution, lost in the middle.
Take lawve-ai/route-contract-drafting from the repository into ~/.claude/skills for personal
use, or into .claude/skills inside a project.
The agent identifies a skill by the name field in its header. Two skills with the
same name cannot sit side by side — one of them will be ignored.