agentic-ai-engineering/skills/individual/planf3/workflows/image-generation.md

3.2 KiB

Image Generation

Fill or update the embedded images in an existing plan .html file. Pick the sub-workflow based on the incoming USER_PROMPT:

Sub-workflow When to call it
Create The prompt asks to generate, fill, or add the plan's images from scratch (empty {{...IMAGE slots)
Update The prompt asks to change, refine, regenerate, or replace images that already exist in the plan

Script selection (two providers, same CLI signature)

Scripts run with uv run and need an API key. Two interchangeable providers — pick whichever has a working key, or fall back to commented placeholders if neither is available:

Provider Script Env var needed Notes
OpenAI (original) scripts/generate_gpt_image.py OPENAI_API_KEY (real sk-... key) Uses gpt-image-2. Paid per image.
OpenRouter (adapter) scripts/generate_or_image.py OPENROUTER_API_KEY Uses Gemini 3 image models. Reuses an existing OpenRouter key — but image models require credits (not free-tier); requests fail with HTTP 402 if the account can't afford the output tokens.

Invoke (either create script, same args):

  • uv run scripts/generate_or_image.py "<prompt>" <output.png> --size 1536x1024 --quality high
  • uv run scripts/generate_gpt_image.py "<prompt>" <output.png> --size 1536x1024 --quality high
  • Edit (OpenAI only): uv run scripts/edit_gpt_image.py "<instruction>" <output.png> <input.png> --size 1536x1024 --quality high

If neither provider is usable (no key, or insufficient credits), leave the image slots as commented placeholders (<!-- {{...IMAGE: subject}} -->) and note in the plan that images are pending a key. The plan is still complete and usable without them — images aid comprehension but are not load-bearing.

Shared rules for every image prompt:

  • always generate in wide format (--size 1536x1024) at high quality (--quality high)
  • convey the one or two core ideas of that section for a professional software engineer
  • match the plan's synced visual identity (professional, focused, minimal)
  • keep total words shown in the image under 10
  • save images to IMAGES_OUTPUT_DIR (create it if missing)

Create

  1. Find slots - Grep the plan for {{...IMAGE placeholders (hero + per-phase). Each comment names the intended subject.
  2. Write prompts - For each slot, write a prompt following the shared rules above.
  3. Generate - Run generate_gpt_image.py once per slot, writing to IMAGES_OUTPUT_DIR.
  4. Embed - Replace each <!-- {{...IMAGE: ...}} --> placeholder with <img src="<plan-name>/<file>.png" alt="...">, keeping the existing <figure>/<figcaption>.
  5. Report - List the images generated and the slots filled.

Update

  1. Identify targets - From the USER_PROMPT, determine which embedded <img> images to change.
  2. Write instruction - Write an edit instruction describing the change, following the shared rules above.
  3. Edit - Run edit_gpt_image.py with the existing PNG as input, overwriting it (the script backs up the original first).
  4. Verify embed - Confirm the <img> still points at the updated file; update src/alt/<figcaption> if the change warrants it.
  5. Report - List the images updated and what changed.