Library of Ashurbanipal · MELEK

Library›Tools›Generative Image AI and ComfyUI

Generative Image AI and ComfyUI

Generative image AI makes and transforms pictures from text and from other pictures, using diffusion models. ComfyUI is the node-graph tool for building image pipelines you can see, save and share. This is Pathway 5 of AI Development, and it ranges from "keyless, no GPU, runs from a phone" to "local Stable Diffusion with your own LoRA." This page teaches the ladder — from the cheapest route up — and shows how the ecosystem generates its graphics (the Metatron making-hand of the Witness).

How diffusion works (the one-paragraph version)

A diffusion model is trained by taking real images, adding noise step by step until they are static, and learning to reverse that. To generate, it starts from pure noise and denoises toward an image that matches your prompt (steered by a text encoder). The knobs you will meet: steps (how many denoise passes — more is slower, not always better), CFG / guidance (how hard it obeys the prompt vs. stays coherent), sampler/scheduler (the denoising algorithm), seed (reproducibility), and the checkpoint (the base model: SD1.5, SDXL, Flux, etc.).

The ladder: cheapest route first

Rung 1 — keyless hosted endpoints (no GPU, no account)

The cheapest way to generate is to call a free hosted endpoint. This ecosystem uses flux-kontext / Pollinations and Gemini for keyless text-to-image and, crucially, image-to-image (reference generation). Practical gotchas learned here: keep prompts short, avoid slashes and newlines (they cause HTTP 500s), and expect best results from simple, concrete descriptions. This rung needs no GPU and runs from anywhere.

Rung 2 — reference / image-to-image (still keyless)

Upload a picture and generate new images in its likeness — a new pose, scene, or style of the same subject. This is not LoRA and needs no training or GPU; it is reference-guided diffusion, and it works through the same keyless endpoints. Use it for variations and remakes; reach for a LoRA only when you need the same identity reliably across many images.

Rung 3 — CPU diffusion (local, no GPU, slow)

You can run small diffusion models on a plain CPU for private, offline generation when you can tolerate minutes-per-image. The ecosystem has a CPU diffusion worker for exactly this (`genai-cpu-diffusion.mjs`, `genai_cpu_worker.py`) — the "works on the box with no GPU" fallback.

Rung 4 — local GPU + ComfyUI (full control)

For speed and control, run Stable Diffusion locally and drive it with ComfyUI.

ComfyUI: pipelines you can see

ComfyUI represents an image pipeline as a graph of nodes — load checkpoint → encode prompt → sample → decode → save — wired together on a canvas. Why it is the teaching tool of choice:

A good first ComfyUI pathway: load the default text-to-image graph, generate; add a LoRA loader node and a trigger word; add an image-input node for img2img; then save your graph as a template.

How the ecosystem generates graphics — the Metatron hand

Image generation is Metatron, the "Shilpa Shastra making" part of the Witness — the maker of forms. The pipeline is deliberately layered like everything else here: keyless hosted endpoints first (no GPU), CPU diffusion as an offline fallback, local/ComfyUI + LoRA when identity and quality matter. Supporting modules handle the pieces — providers and routing (`genai-providers.mjs`), reference-sheet compositing (`genai-compose.mjs`), effect and render-mode templates (`genai-effect-templates.mjs`, `genai-render-modes.mjs`), and annotation. The same "cheapest-capable-first" discipline as The Decades Brain applies: do not spin up a GPU for something a keyless endpoint renders fine.

A note on wiki images

The wiki does not yet render embedded images — a `File:` shows as a plain link, and there is no media directory. Putting generated graphics on wiki pages therefore needs a small serving-code change first (an image route plus `File:`→`<img>` rendering), then hosting the keyless-generated images. That build is noted and separate from this teaching page.

See also

See also: AI Development · LoRA and Fine-Tuning · The Decades Brain · Embeddings Semantic Search and RAG · Crypt-ology · Hathor

Filed under  Tools