3.2 KiB
3.2 KiB
Image Generation
Fill or update the embedded images in an existing plan .html file. Pick the sub-workflow based on the incoming USER_PROMPT:
| Sub-workflow | When to call it |
|---|---|
| Create | The prompt asks to generate, fill, or add the plan's images from scratch (empty {{...IMAGE slots) |
| Update | The prompt asks to change, refine, regenerate, or replace images that already exist in the plan |
Script selection (two providers, same CLI signature)
Scripts run with uv run and need an API key. Two interchangeable providers — pick whichever has a working key, or fall back to commented placeholders if neither is available:
| Provider | Script | Env var needed | Notes |
|---|---|---|---|
| OpenAI (original) | scripts/generate_gpt_image.py |
OPENAI_API_KEY (real sk-... key) |
Uses gpt-image-2. Paid per image. |
| OpenRouter (adapter) | scripts/generate_or_image.py |
OPENROUTER_API_KEY |
Uses Gemini 3 image models. Reuses an existing OpenRouter key — but image models require credits (not free-tier); requests fail with HTTP 402 if the account can't afford the output tokens. |
Invoke (either create script, same args):
uv run scripts/generate_or_image.py "<prompt>" <output.png> --size 1536x1024 --quality highuv run scripts/generate_gpt_image.py "<prompt>" <output.png> --size 1536x1024 --quality high- Edit (OpenAI only):
uv run scripts/edit_gpt_image.py "<instruction>" <output.png> <input.png> --size 1536x1024 --quality high
If neither provider is usable (no key, or insufficient credits), leave the image slots as commented placeholders (<!-- {{...IMAGE: subject}} -->) and note in the plan that images are pending a key. The plan is still complete and usable without them — images aid comprehension but are not load-bearing.
Shared rules for every image prompt:
- always generate in wide format (
--size 1536x1024) at high quality (--quality high) - convey the one or two core ideas of that section for a professional software engineer
- match the plan's synced visual identity (professional, focused, minimal)
- keep total words shown in the image under 10
- save images to
IMAGES_OUTPUT_DIR(create it if missing)
Create
- Find slots - Grep the plan for
{{...IMAGEplaceholders (hero + per-phase). Each comment names the intended subject. - Write prompts - For each slot, write a prompt following the shared rules above.
- Generate - Run
generate_gpt_image.pyonce per slot, writing toIMAGES_OUTPUT_DIR. - Embed - Replace each
<!-- {{...IMAGE: ...}} -->placeholder with<img src="<plan-name>/<file>.png" alt="...">, keeping the existing<figure>/<figcaption>. - Report - List the images generated and the slots filled.
Update
- Identify targets - From the
USER_PROMPT, determine which embedded<img>images to change. - Write instruction - Write an edit instruction describing the change, following the shared rules above.
- Edit - Run
edit_gpt_image.pywith the existing PNG as input, overwriting it (the script backs up the original first). - Verify embed - Confirm the
<img>still points at the updated file; updatesrc/alt/<figcaption>if the change warrants it. - Report - List the images updated and what changed.