AV ComfyUI Manual
15 / 64
Files
15
PART I · WORKFLOW ENGINEERING FOR COMFYUI

Reproducibility & One-Variable Testing — turning experiments into an engineering process

If two runs differ simultaneously in seed, denoise, prompt, model and resolution, the cause of the result is impossible to isolate. A professional workflow should support repeatable tests and one-variable changes per iteration.

CONFIRMEDThe reproducible-benchmark principle is directly connected to shared seed, linked controls, sampler settings and the Golden Run concept already present in the technical manual.
LOCAL CHECKLIST

0 of 11 checks

Workflow readiness0%
x

State is stored in this browser's localStorage. The ComfyUI workflow and files are not modified.

01 · WHY

Repeatability is not academic overhead — it makes decisions faster

When a result can be reproduced, an improvement or regression can be tied to a specific change. Without that, you are judging a bundle of random variation rather than building a reliable production recipe.

This is especially important in ArchViz, where we need to separate material improvement from changes in geometry, lighting, seed or a local mask.

02 · TEST STATE

What to record for a reproducible run

CategoryWhat to save
Workflowexact JSON version / commit / hash
Inputsthe same images, masks and references
Modelsfilenames / versions / relevant custom nodes
Controlseffective values, including linked inputs
Samplingseed, steps, sampler, scheduler, CFG/guidance, denoise
Canvasworking resolution / resize mode / batch
Runtime profilewhich groups are enabled/bypassed
Resultcheckpoint previews + final output
03 · ONE VARIABLE

Change one variable per test

If the goal is to measure denoise, then seed, prompt and ControlNet strength should remain unchanged. If Canny is being tested, Depth and the sampler should not change between variants.

That makes the A/B comparison interpretable.

04 · SEED

Seed is part of the experiment configuration, not a randomness button

When comparing architectural settings, keep the seed fixed. Otherwise changes in composition, people or materials may come from a different noise pattern rather than the parameter under test.

In a workflow with a shared seed, verify every consumer: one source may influence several samplers / noise nodes at once.

05 · EFFECTIVE VALUES

Record effective values, not only what the widget displays

If a sampler receives steps or denoise through a linked input, the locally stored widget value is not authoritative. A benchmark log should record the effective upstream value.

The same applies to selectors: a stored value and the currently selected source can diverge.

06 · GOLDEN RUN

Golden Run — a reference state against which changes are measured

Once the workflow is stable, capture one validated run: inputs, model manifest, controls, seed, checkpoints and final output. This becomes the baseline.

Updates to custom nodes, model weights or topology can then be checked against the Golden Run to reveal regressions quickly.

07 · BENCHMARK MATRIX

Evaluate against predefined criteria, not simply “looks better”

CriterionWhat to observe in ArchViz
Geometry preservationcamera, proportions, openings, facade rhythm
Material realismmicrodetail, roughness cues, texture stability
Lighting coherencedirection, exposure, local integration
Artifact rateAI chaos, halos, duplicated details, broken people
Localitychanges occur only where they are allowed
Runtime costVRAM, time, number of active models
08 · VERSION DRIFT

Updating a model or custom node changes the system

Even when the JSON is unchanged, a new custom-node version can alter inputs, defaults or execution behavior. Environment versioning is therefore part of the reproducibility contract.

This is exactly why separating production STABLE from experimental LAB is useful: experiments should not silently change the baseline production workflow.

09 · PRACTICE

Practice: run one real A/B benchmark

  • Choose one stable input.
  • Lock seed and all controls.
  • Choose exactly one variable.
  • Produce A and B with no other changes.
  • Save identical checkpoints.
  • Evaluate against criteria chosen in advance, not general impression.
  • Record the conclusion and promote the better variant to baseline only after a repeat check.