agentic-ai-engineering/skills/individual/planf3/workflows/image-generation.md

48 lines
3.2 KiB
Markdown

# Image Generation
Fill or update the embedded images in an existing plan `.html` file. Pick the sub-workflow based on the incoming `USER_PROMPT`:
| Sub-workflow | When to call it |
| --- | --- |
| Create | The prompt asks to generate, fill, or add the plan's images from scratch (empty `{{...IMAGE` slots) |
| Update | The prompt asks to change, refine, regenerate, or replace images that already exist in the plan |
## Script selection (two providers, same CLI signature)
Scripts run with `uv run` and need an API key. **Two interchangeable providers** — pick whichever has a working key, or fall back to commented placeholders if neither is available:
| Provider | Script | Env var needed | Notes |
| --- | --- | --- | --- |
| OpenAI (original) | `scripts/generate_gpt_image.py` | `OPENAI_API_KEY` (real `sk-...` key) | Uses gpt-image-2. Paid per image. |
| OpenRouter (adapter) | `scripts/generate_or_image.py` | `OPENROUTER_API_KEY` | Uses Gemini 3 image models. Reuses an existing OpenRouter key — but **image models require credits** (not free-tier); requests fail with HTTP 402 if the account can't afford the output tokens. |
Invoke (either create script, same args):
- `uv run scripts/generate_or_image.py "<prompt>" <output.png> --size 1536x1024 --quality high`
- `uv run scripts/generate_gpt_image.py "<prompt>" <output.png> --size 1536x1024 --quality high`
- Edit (OpenAI only): `uv run scripts/edit_gpt_image.py "<instruction>" <output.png> <input.png> --size 1536x1024 --quality high`
**If neither provider is usable** (no key, or insufficient credits), leave the image slots as commented placeholders (`<!-- {{...IMAGE: subject}} -->`) and note in the plan that images are pending a key. The plan is still complete and usable without them — images aid comprehension but are not load-bearing.
Shared rules for every image prompt:
- always generate in wide format (`--size 1536x1024`) at high quality (`--quality high`)
- convey the one or two core ideas of that section for a professional software engineer
- match the plan's synced visual identity (professional, focused, minimal)
- keep total words shown in the image under 10
- save images to `IMAGES_OUTPUT_DIR` (create it if missing)
## Create
1. Find slots - Grep the plan for `{{...IMAGE` placeholders (hero + per-phase). Each comment names the intended subject.
2. Write prompts - For each slot, write a prompt following the shared rules above.
3. Generate - Run `generate_gpt_image.py` once per slot, writing to `IMAGES_OUTPUT_DIR`.
4. Embed - Replace each `<!-- {{...IMAGE: ...}} -->` placeholder with `<img src="<plan-name>/<file>.png" alt="...">`, keeping the existing `<figure>`/`<figcaption>`.
5. Report - List the images generated and the slots filled.
## Update
1. Identify targets - From the `USER_PROMPT`, determine which embedded `<img>` images to change.
2. Write instruction - Write an edit instruction describing the change, following the shared rules above.
3. Edit - Run `edit_gpt_image.py` with the existing PNG as input, overwriting it (the script backs up the original first).
4. Verify embed - Confirm the `<img>` still points at the updated file; update `src`/`alt`/`<figcaption>` if the change warrants it.
5. Report - List the images updated and what changed.