mcpbeat Sign in

Define Goal Skill for Codex

Help the user define a concrete, measurable goal before starting work, especially when they ask to use the goal tool, create a goal, set an objective, clarify success criteria, or turn a fuzzy intention into a quantitative outcome. Use this skill for goal creation and goal refinement only; it does not manage durable snapshots, decision logs, or long-running execution artifacts.

4k tokens
context cost
the whole folder, loaded on every use
3
files
instructions only
0
copies elsewhere
how many repositories repackaged it
24488
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/openai/skills --skill define-goal

What comes with it

10 984 bytes besides the instruction
LICENSE.txt
agents/openai.yaml

The instruction itself

6 sections, as written by the author

Define Goal

Overview

Shape the user's intent into an objective an agent can pursue honestly. Prefer measurable outcomes, explicit evidence, and bounded scope over activity descriptions.

This skill covers goal definition and goal-tool creation only. Do not create intermediate planning artifacts, durable snapshots, ledgers, decision logs, or resume files from this skill.

Workflow

  • Confirm that goal definition is actually needed.
  • Use this skill when the user asks for $define-goal, asks to create or set a goal, asks for the goal tool, or wants help turning an intention into a clear objective.
  • If the user only asks for ordinary implementation work, do the work directly instead of forcing goal creation.
  • Restate the likely goal in concrete terms.

A usable goal names:

  • the specific outcome that will be true
  • the main artifact, system, repo, environment, or user-facing behavior involved
  • how completion will be verified
  • what is in scope
  • what is out of scope when ambiguity would matter
  • the stop condition for asking the user instead of grinding
  • Make it quantitative when the domain supports it.

Prefer numbers that represent real success, not decorative precision:

  • pass/fail validators: exact tests, checks, CI jobs, evals, commands, or acceptance criteria
  • quality thresholds: latency, error rate, cost, accuracy, recall, precision, coverage, flake rate, bundle size, memory, uptime, completion rate, or manual review criteria
  • artifact constraints: file paths, affected modules, allowed commands, output formats, target environments, deadlines, or maximum blast radius
  • evidence counts: number of reproduced failures, successful reruns, reviewed examples, migrated records, addressed comments, or verified cases
  • Repair weak goals before setting them.
  • Rewrite vague goals into measurable objectives when local context makes the rewrite safe.
  • Ask one concise clarification question when the missing detail changes the intended outcome or validation.
  • Reject pure activity goals such as "make progress," "keep investigating," "improve things," or "work on X" unless they are sharpened into a verifiable outcome.
  • Check active goal state before creating a goal.
  • Call get_goal.
  • If there is no active goal and the objective meets the quality bar, call create_goal.
  • If there is an active goal that still matches the user's intent, continue using it instead of creating a duplicate.
  • If there is an active goal that conflicts with the new request, ask whether to finish the current goal, mark it complete if done, or start a separate goal-backed thread.
  • Create the goal only after it passes the quality bar.
  • Use a single concise objective string.
  • Include the verification evidence in the objective itself.
  • Include scope bounds when they constrain the work.
  • Include a token budget only when the user explicitly requested one.
  • Do not call create_goal for an ordinary multi-step task unless the user explicitly asked for goal-backed work.

Goal Quality Bar

Before create_goal, the objective should answer:

  • What concrete thing will be true when this is done?
  • What evidence will prove it?
  • What quantitative or binary threshold defines success?
  • What scope boundaries matter?
  • What should cause the agent to stop and ask?

Good:

> Reduce checkout API p95 latency below 250 ms for the documented slow path by making the smallest safe server-side change, then verify with npm run test:checkout and the existing local latency benchmark showing p95 under 250 ms across 3 consecutive runs.

Good:

> Resolve the open review comments on PR 123 that request code changes, update only the affected auth files and tests, and verify with the targeted auth test command plus gh pr view 123 showing no unresolved change-request threads.

Weak:

> Make checkout faster.

Weak:

> Keep investigating the PR comments.

Quantification Heuristics

  • For bugs, define success as reproduction first, fix second, and a failing-then-passing validator when possible.
  • For tests, name the exact command and required pass condition.
  • For performance, name the metric, target threshold, measurement method, and number of runs.
  • For quality work, define an observable acceptance bar such as reviewed examples, lint/typecheck/test pass, or user-approved artifact.
  • For research, define the decision the research must enable, the sources or systems in scope, and the evidence standard.
  • For operations, define healthy state, monitoring window, failure threshold, and rollback or escalation trigger.

Clarifying Questions

Ask only when a reasonable rewrite would risk pursuing the wrong outcome. Keep the question short and oriented around the missing validator or scope boundary.

Useful question shapes:

  • "What metric should define success here: latency, cost, accuracy, or user-visible behavior?"
  • "Which environment should I verify against: local, staging, or production?"
  • "What is the minimum evidence you want before I mark this goal complete?"

If the user cannot provide a metric, propose the most honest binary validator available and ask for confirmation.

Other skills for the same job

different authors, same section of the catalogue
Doc Coauthoring
by anthropics
vendor ×10

Guide users through a structured workflow for co-authoring documentation. Use when user wants to write documentation, proposals, technical specs, decision docs, or similar structured content. This workflow helps users efficiently transfer context, refine content through iteration, and verify the doc works for readers. Trigger when user mentions writing docs, creating proposals, drafting specs, or similar documentation tasks.

4k tokens
File Organizer
by frostant
×10

Intelligently organizes your files and folders across your computer by understanding context, finding duplicates, suggesting better structures, and automating cleanup tasks. Reduces cognitive load and keeps your digital workspace tidy without manual effort.

3k tokens
Domain Name Brainstormer
by frostant
×8

Generates creative domain name ideas for your project and checks availability across multiple TLDs (.com, .io, .dev, .ai, etc.). Saves hours of brainstorming and manual checking.

1k tokens
Brainstorming
by ZhanlinCui
×4

You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.

626 tokens
Planning With Files
by ZhanlinCui
×3

Implements Manus-style file-based planning for complex tasks. Creates task_plan.md, findings.md, and progress.md. Use when starting complex multi-step tasks, research projects, or any task requiring >5 tool calls.

9k tokens scripts
Scientific Brainstorming
by christophacham
×3

Creative research ideation and exploration. Use for open-ended brainstorming sessions, exploring interdisciplinary connections, challenging assumptions, or identifying research gaps. Best for early-stage research planning when you do not have specific observations yet. For formulating testable hypotheses from data use hypothesis-generation.

5k tokens
GitHub Project Management
by ComeOnOliver
×3

Comprehensive GitHub project management with swarm-coordinated issue tracking, project board automation, and sprint planning

14k tokens
Grill Me
by ComeOnOliver
×3

Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".

3k tokens

How to use it

Copy the folder

Take openai/define-goal from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.