Improve and replace existing people with Florence2 + SAM2 + FLUX
This mode does not decide where to place a person. Position, scale, perspective and approximate pose are already defined by the source 3D/SDXL scene; AI isolates the existing figure, regenerates it locally and returns it to the same spatial slot.
Two routes converge at selector 459
3D owns placement; AI owns quality
The source frame already contains a person as a spatial anchor: position, height, scale, perspective and approximate pose are defined before generation. This can be a 3D proxy, library character, cutout or a person already present in the base image.
AI does not redesign the scene composition or choose a new location for the character. Its role is to improve or replace the local figure while preserving the architectural context around it.
Placement contract
In the showcase, people are already present in the base image before the PPL replacement pass.
ArchViz principle
For production, this separates responsibilities: 3D fixes spatial truth; the generator handles visual realism.
Complete Variant 2 route
Florence2 finds the person; SAM2 builds the precise mask
Florence2 performs semantic detection / phrase grounding and returns the person region as a bbox or coordinates. At this stage, the model answers “where is the person?”
SAM2 receives the spatial cue from Florence2 and builds a pixel-level mask. The task is no longer class recognition, but clean separation of the figure from architecture and environment.
- Florence2 and SAM2 must receive the same source image and coordinate space.
- Do not continue downstream if bbox or mask is already wrong.
- Image batch size and bbox count must be compatible; otherwise Sam2Segmentation may fail during indexing.
| Stage | Input | Output | QA |
|---|---|---|---|
| Florence2 | Base image + person/people prompt | BBOX / coordinates | BBox should cover the intended figure, not the whole frame |
| SAM2 | Same image + Florence spatial cue | Person mask | Head, arms, legs and accessories without facade capture |
The generator works on a crop, not the full architectural frame
A local region around the person is cut from the frame using the mask or bbox. The crop is resized to the working resolution and passed to a separate FLUX branch.
Architecture outside the crop is therefore not part of generation: the model can improve face, clothing, materials, hair and photographic response without gaining freedom to rebuild the facade or camera.
Local scope
The showcase uses a separate PPL FLUX Generate section between segmentation and composite.
Production value
A local crop reduces the risk area and improves repeatability compared with full-frame img2img.
The generated person becomes a compositing element
After generation, the result is separated from its local background using operations such as Remove Background / Cut By Mask. The required output is an RGB figure with a correct alpha/mask representation.
Before returning to the scene, Color Match is used so the person does not look pasted in from a different exposure or camera response.
| Operation | Purpose | Typical failure |
|---|---|---|
| Remove Background / Cut By Mask | Separate generated person from crop background | Halo, missing limbs, residual old background |
| Color Match | Align tone / contrast / temperature with base render | Skin and clothing look like they came from another photograph |
| Mask cleanup | Stabilize contour before paste | Hard edge, dirty alpha, floating silhouette |
Paste By Mask returns the figure to the original spatial slot
The prepared cutout is pasted back into the base scene using the mask and positioning data. The goal is to replace visual appearance without changing the approved composition.
After paste, verify scale, foot position, surface contact, occlusion and lighting match.
- Check the paste before downstream selectors or upscale.
- Compare the generated-person silhouette with the original proxy: center or scale drift is an error.
- Do not use Color Match as a substitute for a correct lighting prompt; it helps integration but cannot fix physically wrong light.
The final local pass fixes integration; it does not create a new scene
After composite, a local detail/inpaint pass can repair hands, hair, clothing edges, foot contact, small intersections and local shadows.
This stage should remain local. If the mask expands across a large part of the architecture, the workflow loses the main advantage of controlled replacement.
Where Variant 2 is especially useful
- Hero people in foreground and midground.
- Cyclists and characters with predefined poses.
- People near entrances, storefronts or facades.
- Visitors in exhibitions and public spaces where placement is already approved.
- Replacing a library 3D character without changing camera / architecture.
Not crowd placement
Variant 2 is not an automatic people-placement algorithm for an empty scene.
Workflow 01 documented separately
Generate & Place New People by Mask has its own chapter. Workflows 01 and 02 remain separate because they use different spatial contracts.
Variant 2 pre-final checklist
- The source person is already in the correct location and scale before the AI pass.
- Florence2 bbox covers only the intended person.
- SAM2 mask is clean and does not capture architecture.
- The crop contains the whole figure and enough context for generation.
- FLUX preserves the intended spatial role of the person.
- Background removal leaves no halo or holes in limbs.
- Color Match aligns the person with overall frame exposure.
- Paste preserves original scale / position.
- Feet, contact shadow and occlusion look physically plausible.
- Final local inpaint does not alter architectural geometry.