I started by asking how to turn a still image into a short video with text guiding the motion. That sounded like one creative task, but ComfyUI made me confront the pieces below it. I narrowed the goal to a basic still-image workflow first. Even finding the right CLIP text-encoding node, connecting a checkpoint, and deciding what to put in each prompt box were separate problems when I was learning the graph.
The assistant proposed a node path and trial settings. As I tried to follow it, missing-node and checkpoint-loading suggestions appeared, but the later diagnosis was still only a suggestion in the chat. There is no confirmed render in this record. What it actually shows is a more useful first step than the original reel idea: learn which node supplies the model, where text enters, and how to tell a connection problem from a generation problem.
The first video target
The suggested initial settings were 512 × 768 pixels, 60 frames, 12 frames per second, 16 sampling steps, and CFG 5. The proposed graph loaded a diffusion model, text encoder, base image, video latent, sampler, decoder, and video output node.
I asked for explicit guidance on the canvas because node names alone were not enough. Finding the correct box, its input socket, and its output connection was part of the learning task.
Afterward, I reported that the face was going “a bit wonky.” The assistant treated that as possible identity drift and suggested narrowing the motion. I also asked about other models, including SVD XT 1.1. The conversation did not settle that comparison with a finished, verified reel.
Returning to a still image
Eventually I changed the immediate goal to basic text-to-image. The smaller graph made the separate responsibilities easier to see:
| Node | Job in the proposed graph |
|---|---|
| Load Checkpoint | Supply the model, CLIP encoder, and VAE |
| Two CLIP Text Encode nodes | Encode positive and negative prompts |
| Empty Latent Image | Set image dimensions and batch size |
| KSampler | Generate the latent result from model and conditioning |
| VAE Decode | Turn the latent into a visible image |
| Save Image | Save the decoded output |
The connections were the important content. The checkpoint’s CLIP output fed both prompt encoders. Their conditioning outputs went to the sampler’s positive and negative inputs. The empty latent went to the sampler; the sampled latent then went through VAE Decode before saving.
The assistant used a golden retriever wearing sunglasses as a simple test subject. Its proposed baseline was a 512 × 512 image, batch size 1, a fixed seed, 20 steps, CFG 7, Euler sampling, and denoise 1.0.
CLIP was a connection problem before it was a model problem
I asked where CLIP was, then said I could not find CLIP Text Encode. The clarification distinguished the checkpoint’s CLIP output from the separate prompt-encoding node. Dragging from that output to empty canvas was suggested as a way to show compatible nodes.
Later, the assistant read an attached report differently: the prompt nodes were present, while loading the SD 1.5 checkpoint was failing with ModelMMAP allocation failed. It also identified a filename mismatch involving an SDXL checkpoint whose actual filename had a parenthesized suffix.
That final diagnosis was the assistant’s interpretation of the supplied report. It suggested using the exact filename offered in the checkpoint dropdown and separating file-loading trouble from graph wiring. I have not turned its claims about successful generations into my own independently checked result.
I had backed away from the short reel to a still-image graph I could examine one connection at a time. The next honest milestone was a confirmed render. The node and checkpoint suggestions in the chat had not yet produced one here.