Resolution, Coordinate Space & Batch Semantics — the hidden geometry of a graph
Many difficult-to-explain failures in production ComfyUI are caused not by the model, but by mismatched dimensions, coordinate spaces or batch cardinality. This chapter teaches you to verify data geometry before running expensive generative stages.
Before connecting branches, verify dimensions, coordinates and cardinality
The same visual frame can exist at several resolutions
The base image, latent canvas, detection image, external depth map, mask atlas and crop can all have different resolutions. That is acceptable while they are processed independently. The problem starts when coordinates or a mask from one canvas are applied to another.
Resize is therefore not a cosmetic operation. It changes the downstream coordinate system.
| Object | What to verify |
|---|---|
| IMAGE | width × height and aspect ratio |
| MASK | same canvas or an explicit resize |
| LATENT | expected model canvas / divisibility constraints |
| BBOX / points | which image coordinate system produced them |
| CROP | origin + width/height relative to the source canvas |
A BBOX only has meaning together with the image on which it was calculated
A bounding box is not an abstract rectangle. Its x/y coordinates are tied to a specific width and height. If detection ran after a resize, that bbox cannot be applied directly to an original image of another size without a coordinate transform.
The same rule applies to placement masks, crops and segmentation points.
A mask must match the canvas it controls
Even a semantically correct mask becomes wrong when it was created on a different resize or crop. Before composite/inpaint, verify the mask dimensions and its mapping to the target image.
For local-crop workflows, it is useful to distinguish explicitly between a local mask and a full-scene mask — they are different contracts.
Batch means element count, not simply “several images”
IMAGE may contain a batch of N elements. MASK may contain a different batch. A detection result may contain a list of objects. The downstream node must define how those cardinalities are paired.
Some nodes support broadcasting while others expect a strict 1:1 relationship. Never assume the behavior without checking the specific node implementation.
| Scenario | Risk |
|---|---|
| 1 image + 1 mask | Basic and usually safe |
| N images + N masks | The pairing order must be confirmed |
| N images + 1 mask | Requires broadcasting support or explicit repetition |
| 1 image + N object masks | May use object-batch / combined-mask behavior |
| N images + M bboxes | Unsafe if the node indexes bbox by image index |
Why segmentation can fail even when the models are correct
In our Florence2 → SAM2 test, Florence had already detected the objects successfully, but downstream segmentation received inconsistent cardinality: the image batch contained more elements than the bbox list. The node indexed bbox by image index and ran past the end of the array.
This demonstrates an important principle: an error after AI-model inference does not necessarily mean the model is at fault. Verify the data contract first.
Prove the single-image route before increasing batch size
For a complex detection/segmentation branch, a reliable benchmark starts with batch=1. Once the single-image route succeeds, test multiple objects separately, and only then move to a multi-image batch.
Resize should have an owner and an explicit place in the pipeline
If every branch resizes its source independently, coordinate mapping quickly becomes opaque. Prefer explicit working canvases and document which stage owns each size transformation.
Engineering rule
Keep spatially coupled operations — image/mask/bbox/composite — on one named working canvas whenever possible, or document the transform between canvases.
Practice: build a spatial ledger for one branch
If even one row is unknown, the branch is not yet ready for reliable integration into the Master Workflow.
| Stage | W×H | Batch | Coordinate owner |
|---|---|---|---|
| Input | fill in | fill in | BASE |
| Resize | fill in | fill in | WORKING CANVAS |
| Detection | fill in | fill in | DETECTION SPACE |
| Mask | fill in | fill in | MASK SPACE |
| Composite | fill in | fill in | DESTINATION SPACE |