mcpbeat Sign in

Generate Image Agent Skill

>- Generate images using AI. Use when asked to generate, create, or make images, textures, icons, sprites, artwork, visual assets, or mockups. Supports OpenAI (gpt-image-2) and Google Gemini (Nano Banana). Requires an API key for the chosen provider.

962 tokens
context cost
the whole folder, loaded on every use
1
files
instructions only
0
copies elsewhere
how many repositories repackaged it
37394
stars on the repo
on the repository, not the skill itself

Install

one command, takes just this skill from the repository
npx skills add https://github.com/github/awesome-copilot --skill generate-image

The instruction itself

6 sections, as written by the author

Generate Image

You are an image generation assistant. When invoked, follow the workflow below.

Workflow

  • Check for API keys — check whether SKILL_IMAGE_GEN_OPENAI_KEY and/or SKILL_IMAGE_GEN_GEMINI_KEY are set in the environment.
  • If one key is set — use that provider. No need to ask.
  • If both are set — pick based on context (OpenAI for polish, Gemini for speed), or ask if the user has a preference.
  • If no keys are set — run the Onboarding section.
  • Generate the image using the appropriate API reference.
  • Tell the user where the image was saved.

Onboarding

Only run this if no keys are set. Guide the user conversationally.

  • Ask which provider they'd like to use:
  • OpenAI (gpt-image-2) — High quality, excellent text rendering, paid per image
  • Google Gemini (Nano Banana) — Fast, free tier available, great for iteration
  • Direct them to get an API key:
  • OpenAI → https://platform.openai.com/api-keys
  • Gemini → https://aistudio.google.com/apikey
  • Once they provide the key, set SKILL_IMAGE_GEN_OPENAI_KEY or SKILL_IMAGE_GEN_GEMINI_KEY in the current session and persist it to the appropriate shell profile.
  • Proceed to generate the image they originally asked for.

API Reference: OpenAI

Method: POST

URL: https://api.openai.com/v1/images/generations

Headers:

  • Authorization: Bearer <SKILL_IMAGE_GEN_OPENAI_KEY>
  • Content-Type: application/json

Body (JSON):

{
  "model": "gpt-image-2",
  "prompt": "<user prompt>",
  "n": 1,
  "size": "1024x1024",
  "quality": "medium"
}

| Field | Default | Options |

|---|---|---|

| model | gpt-image-2 | gpt-image-2, gpt-image-1 |

| size | 1024x1024 | 1024x1024, 1024x1536, 1536x1024, auto |

| quality | medium | low, medium, high |

Response: data[0].b64_json contains the base64-encoded image. Decode it and save to the output path. If data[0].url is present instead, download the image from that URL.

API Reference: Google Gemini (Nano Banana)

Method: POST

URL: https://generativelanguage.googleapis.com/v1beta/models/<model>:generateContent

Headers:

  • x-goog-api-key: <SKILL_IMAGE_GEN_GEMINI_KEY>
  • Content-Type: application/json

Body (JSON):

{
  "contents": [{"parts": [{"text": "Generate an image: <user prompt>"}]}],
  "generationConfig": {"responseModalities": ["TEXT", "IMAGE"]}
}

| Field | Default | Options |

|---|---|---|

| model (in URL) | gemini-2.0-flash-exp | gemini-2.0-flash-exp, gemini-2.5-flash-image |

Response: Find candidates[0].content.parts[] — look for a part with inlineData.data (base64 image) and inlineData.mimeType. Decode and save.

Error cases: error key (API error), promptFeedback.blockReason (safety block), finishReason: "SAFETY" (filtered).

Agent Guidelines

  • Choose the output path intelligently — save to the project's relevant directory (e.g., assets/, images/, or the current directory).
  • For game textures, enrich prompts with "seamless", "tileable", "game asset".
  • For batch generation, make multiple API calls in parallel.
  • If the user asks to switch providers or what options are available, explain both and help them set up.
  • Always create the output directory before saving.
  • Ensure special characters in the user's prompt are properly escaped in the JSON body.

Other skills for the same job

different authors, same section of the catalogue
Canvas Design
by anthropics
vendor ×13

Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.

1388k tokens
Algorithmic Art
by anthropics
vendor ×10

Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.

15k tokens scripts
Image Enhancer
by frostant
×6

Improves the quality of images, especially screenshots, by enhancing resolution, sharpness, and clarity. Perfect for preparing images for presentations, documentation, or social media posts.

635 tokens
Video Downloader
by CommandCodeAI
×4

Downloads videos from YouTube and other platforms for offline viewing, editing, or archival. Handles various formats and quality options.

671 tokens
Histolab
by christophacham
×3

Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatial proteomics, multiplexed imaging, or deep learning pipelines use pathml.

18k tokens
Omero Integration
by christophacham
×3

Microscopy data management platform. Access images via Python, retrieve datasets, analyze pixels, manage ROIs/annotations, batch processing, for high-content screening and microscopy workflows.

32k tokens
Pydicom
by christophacham
×3

Python library for working with DICOM (Digital Imaging and Communications in Medicine) files. Use this skill when reading, writing, or modifying medical imaging data in DICOM format, extracting pixel data from medical images (CT, MRI, X-ray, ultrasound), anonymizing DICOM files, working with DICOM metadata and tags, converting DICOM images to other formats, handling compressed DICOM data, or processing medical imaging datasets. Applies to tasks involving medical image analysis, PACS systems, radiology workflows, and healthcare imaging applications.

13k tokens scripts
Transformers
by christophacham
×3

This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.

13k tokens

How to use it

Copy the folder

Take github/generate-image from the repository into ~/.claude/skills for personal use, or into .claude/skills inside a project.

Check the name does not clash

The agent identifies a skill by the name field in its header. Two skills with the same name cannot sit side by side — one of them will be ignored.