Checkpoints & Debugging — finding the last provably correct stage
A large workflow cannot be diagnosed from the final image alone. Reliable debugging is a ladder of checkpoints: after each meaningful module there should be an observable result that tells you exactly where the pipeline stopped being correct.
The last correct preview determines the next check
If stage N is correct, do not return to the prompt; inspect the connection and parameters of stage N+1.
If a module cannot be observed independently, it is difficult to debug
PreviewImage, MaskPreview, comparer, text/json preview and save probes are not decorative. They create observability: the ability to inspect data before it becomes mixed with the next system.
The longer the chain without a checkpoint, the more possible causes can produce the same final failure.
The key debugging question: where is the last correct result?
Do not begin with “why is the final image wrong?” Start at the end and move upstream until you reach the last checkpoint that looks correct and behaves correctly in the graph.
The next stage becomes the first suspect. This reduces the search space from hundreds of nodes to one branch.
Different data types need different probes
| Data | Checkpoint | What to verify |
|---|---|---|
| IMAGE | PreviewImage / comparer | pixels, composition, color, geometry |
| MASK | MaskPreview | silhouette, polarity, holes, bounds |
| BBOX / JSON | text/json preview | coordinates, count, labels |
| STRING | text preview | assembled prompt / control text |
| LATENT | usually a decode-only debug path | visual result after VAE Decode |
| OUTPUT FILE | SaveImage + actual file check | whether the delivery path really worked |
Debug in dependency order
Enable the smallest set of modules that reproduces the problem
BASE CONFIG should allow everything unrelated to the failure to be disabled. If you are testing a SAM2 mask, there is no reason to run main FLUX and upscale at the same time.
A minimal runtime speeds up iteration and reduces the chance that a side branch hides the real source of the failure.
Classify the failure before changing parameters
| Class | Examples | First action |
|---|---|---|
| Dependency | missing node / model / loader | Check manifest and paths |
| Type contract | IMAGE vs LATENT / missing CONDITIONING | Check socket types |
| Spatial contract | mask shift / bbox mismatch | Check W×H and coordinate space |
| Batch contract | index out of bounds / cardinality mismatch | Check batch counts |
| Routing | wrong source / branch has no effect | Check selector + bypass |
| Generation quality | anatomy / material / prompt mismatch | Tune AI parameters only after the technical preflight passes |
Do not treat topology problems with generation parameters
If a mask is shifted because of a resize mismatch, changing denoise will not help. If a selector chose the wrong source, the prompt cannot repair the route. If CLIP input is missing, CFG is irrelevant.
Production debugging starts with architecture and data contracts, then moves to generation tuning.
Record the test state so debugging can be repeated
- Which input was used.
- Which modules are enabled in BASE CONFIG.
- Effective selector values.
- Seed and sampling parameters.
- Dimensions and batch at the failing stage.
- Last correct checkpoint.
- Exact error message / node ID.
Practice: break a branch deliberately and localize the fault
Take a small standalone workflow with three checkpoints. Deliberately break one contract — for example the selector source or mask size. Do not fix it immediately. Walk the debug ladder and record the first checkpoint where the result diverges.
Readiness criterion
A learner has internalized debugging when they can name the last correct stage and the class of failure before changing the prompt or sampler.