Abstract
Full-frame regeneration is a blunt instrument. When a professional creator needs to swap a background, fix a lighting artifact, or add a prop to a specific region, regenerating the entire image risks destabilizing everything else — identity drift, color shifts, composition changes. The result is a frustrating cycle of regeneration and compromise.
We present a regional editing paradigm where creators paint a spatial mask over the area they want to change, describe the targeted edit in a prompt, and the system regenerates only the selected region while preserving surrounding context. This brush-select-then-prompt workflow enables surgical edits that full-frame approaches cannot achieve.
The Full-Frame Regeneration Problem
Instruction-based full-frame editing (e.g., "make the sky more dramatic") applies changes globally. While effective for tonal adjustments, it introduces three failure modes when used for spatially targeted intent:
- Collateral mutation — changing the sky also shifts skin tones, clothing color, and surface reflections
- Identity drift — repeated full-frame edits compound small variations, gradually shifting the subject away from the original
- Composition instability — the model may reinterpret framing, crop, or spatial relationships with each regeneration
Supporting research:
"While InstructPix2Pix can follow editing instructions, it often produces unintended changes to regions that should remain unmodified — a fundamental limitation when the editing intent is spatially localized."
— Brooks et al., "InstructPix2Pix" (CVPR 2023)
The Spatial Mask Paradigm
Regional editing solves the localization problem by introducing an explicit spatial contract between creator and system:
Full-frame editing:
Prompt → Entire image regenerated
No spatial boundaries. Side effects everywhere.
Regional editing:
Brush → Mask → Prompt → Region only
Explicit boundaries. Surgical precision.
The mask serves as a spatial contract: pixels outside the mask are frozen, pixels inside are regenerated conditioned on both the prompt and the surrounding context. This guarantee is what makes regional editing viable for professional workflows where preserving approved elements is non-negotiable.
Brush-Select-Then-Prompt Workflow
The AIM regional editing workflow follows four steps designed to mirror natural creative thinking — point at what you want to change, then describe the change:
- 1. Select the brush tool — Enter brush mode from the frame edit interface. Adjustable brush size accommodates both broad regions and fine details.
- 2. Paint the target region — Brush over the area you want to modify. The mask is visualized as a semi-transparent overlay so you can see exactly what's selected.
- 3. Describe the edit — Write a natural-language prompt describing what should appear in the masked region: "replace with autumn foliage," "add dramatic storm clouds," "change to brick wall."
- 4. Generate — The system regenerates only the masked pixels, blending the new content seamlessly with the preserved surroundings.
Paint region → describe the change
↓
Masked area regenerated from your prompt
↓
Region-only regeneration with boundary blending
↓
Composite: Preserved pixels + Generated region → Final frame
Technical Foundation
Latent-Space Inpainting
Regional editing builds on latent diffusion inpainting, where the masked region is denoised in latent space while conditioned on the encoded surrounding context. This ensures that generated content inherits the lighting, color temperature, and perspective of the existing frame rather than producing a visually disconnected patch.
"By concatenating the masked image and mask as additional input channels... the model learns to generate content that is both semantically appropriate and visually coherent with the surrounding unmasked regions."
— Rombach et al., "High-Resolution Image Synthesis with Latent Diffusion Models" (CVPR 2022)
Boundary Coherence
The critical challenge in regional editing is the boundary — where generated content meets preserved content. Hard mask edges produce visible seams. AIM addresses this through mask feathering at the boundary, where a soft gradient transitions between preserved and generated regions, keeping the seam imperceptible.
Context-Aware Generation
The inpainting model doesn't generate in isolation — it analyzes the unmasked surrounding pixels to match:
Shadow angles and highlight positions in surrounding content dictate light direction for the generated region.
The model samples the chromatic range of adjacent pixels to ensure the generated content shares the same white balance and color grading.
Vanishing lines and depth cues from the preserved frame inform spatial consistency — a replaced background matches the scene's perspective geometry.
Professional Use Cases
Regional editing addresses specific professional needs that full-frame editing cannot:
| Use Case | Mask Region | Edit Prompt |
|---|---|---|
| Background swap | Sky / environment | "Dramatic thunderstorm sky" |
| Prop addition | Empty area | "Vintage leather briefcase" |
| Material change | Surface / texture | "Weathered concrete wall" |
| Artifact removal | Defect area | "Clean continuation of surface" |
| Wardrobe change | Clothing region | "Navy tailored blazer" |
Regional vs. Full-Frame Editing
| Dimension | Full-Frame Edit | Regional Edit |
|---|---|---|
| Scope | Entire image | Masked region only |
| Identity preservation | At risk | Guaranteed (outside mask) |
| Iteration safety | Compounds drift | Isolated changes |
| Best for | Global tonal shifts | Spatial modifications |
| Wasted generations | High (side effects) | Low (predictable scope) |
Preliminary Observations
From our ongoing AIM pilot study (n=20), regional editing addresses a clear gap in professional creative workflows:
Reported "collateral damage" frustration with full-frame edits
Reported edit iterations for targeted changes
Found brush selection "intuitive" from photo editing experience
* Preliminary findings from pilot study. Expanded validation in progress through 2026.
Design Decisions
Why Brush Over Bounding Box?
We chose freeform brush selection over rectangular bounding boxes for two reasons: real-world edit regions are rarely rectangular (clothing, hair, architectural features follow organic contours), and brush selection transfers naturally from existing photo-editing muscle memory (Photoshop, Lightroom masking). The cognitive overhead of learning a new selection paradigm is near zero.
Prompt After Mask, Not Before
The workflow deliberately requires spatial selection before prompt entry. This mirrors natural creative thinking: you look at a frame, notice something to change, point at it, then describe the change. Reversing this order (describe first, then select) would force premature verbalization of intent before spatial context is established.
Non-Destructive by Default
Every regional edit creates a new version. The original frame is preserved in the version history, allowing creators to compare, revert, or branch from any edit in the chain. This is essential for professional workflows where client changes may be reversed or iterated.
Conclusion
Regional editing transforms frame refinement from a full-image gamble into a targeted, predictable operation. By introducing an explicit spatial mask as the boundary contract between creator intent and AI generation, we eliminate the collateral mutations that make full-frame editing unreliable for professional work.
The brush-select-then-prompt workflow respects how professionals actually think about edits: see a problem, point at it, describe the fix. No spatial reasoning encoded in language. No praying that the model leaves the rest alone. What you mask is what changes — everything else is guaranteed.