Reproducibility & One-Variable Testing — turning experiments into an engineering process
If two runs differ simultaneously in seed, denoise, prompt, model and resolution, the cause of the result is impossible to isolate. A professional workflow should support repeatable tests and one-variable changes per iteration.
0 of 11 checks
State is stored in this browser's localStorage. The ComfyUI workflow and files are not modified.
Repeatability is not academic overhead — it makes decisions faster
When a result can be reproduced, an improvement or regression can be tied to a specific change. Without that, you are judging a bundle of random variation rather than building a reliable production recipe.
This is especially important in ArchViz, where we need to separate material improvement from changes in geometry, lighting, seed or a local mask.
What to record for a reproducible run
| Category | What to save |
|---|---|
| Workflow | exact JSON version / commit / hash |
| Inputs | the same images, masks and references |
| Models | filenames / versions / relevant custom nodes |
| Controls | effective values, including linked inputs |
| Sampling | seed, steps, sampler, scheduler, CFG/guidance, denoise |
| Canvas | working resolution / resize mode / batch |
| Runtime profile | which groups are enabled/bypassed |
| Result | checkpoint previews + final output |
Change one variable per test
If the goal is to measure denoise, then seed, prompt and ControlNet strength should remain unchanged. If Canny is being tested, Depth and the sampler should not change between variants.
That makes the A/B comparison interpretable.
Seed is part of the experiment configuration, not a randomness button
When comparing architectural settings, keep the seed fixed. Otherwise changes in composition, people or materials may come from a different noise pattern rather than the parameter under test.
In a workflow with a shared seed, verify every consumer: one source may influence several samplers / noise nodes at once.
Record effective values, not only what the widget displays
If a sampler receives steps or denoise through a linked input, the locally stored widget value is not authoritative. A benchmark log should record the effective upstream value.
The same applies to selectors: a stored value and the currently selected source can diverge.
Golden Run — a reference state against which changes are measured
Once the workflow is stable, capture one validated run: inputs, model manifest, controls, seed, checkpoints and final output. This becomes the baseline.
Updates to custom nodes, model weights or topology can then be checked against the Golden Run to reveal regressions quickly.
Evaluate against predefined criteria, not simply “looks better”
| Criterion | What to observe in ArchViz |
|---|---|
| Geometry preservation | camera, proportions, openings, facade rhythm |
| Material realism | microdetail, roughness cues, texture stability |
| Lighting coherence | direction, exposure, local integration |
| Artifact rate | AI chaos, halos, duplicated details, broken people |
| Locality | changes occur only where they are allowed |
| Runtime cost | VRAM, time, number of active models |
Updating a model or custom node changes the system
Even when the JSON is unchanged, a new custom-node version can alter inputs, defaults or execution behavior. Environment versioning is therefore part of the reproducibility contract.
This is exactly why separating production STABLE from experimental LAB is useful: experiments should not silently change the baseline production workflow.
Practice: run one real A/B benchmark
- Choose one stable input.
- Lock seed and all controls.
- Choose exactly one variable.
- Produce A and B with no other changes.
- Save identical checkpoints.
- Evaluate against criteria chosen in advance, not general impression.
- Record the conclusion and promote the better variant to baseline only after a repeat check.