Abstract
Traditional text-to-image AI generation suffers from a fundamental problem we term "The Angle Lottery" — creators describe a desired camera angle or object position in text, then repeatedly generate hoping to hit their intended composition by chance.
We introduce CineForge — a 3D scene composition system that inverts this paradigm. Instead of describing spatial intent through language and praying for the right output, professionals compose their scene directly in 3D space, then render exactly what they see. This eliminates the translation gap between spatial intent and linguistic description.
The Angle Lottery Problem
The core pain point this paper addresses is the fundamental unreliability of text-based spatial description in AI image generation:
- Text prompts like "dramatic low angle of a sports car" produce wildly inconsistent results
- Professional workflows require reproducible, precise compositions
- Multiple regeneration attempts waste credits and break creative flow
- The AI interprets spatial language differently each time (semantic ambiguity)
Supporting research:
"Image generation today can produce somewhat realistic images from text prompts. However, if one asks the generator to synthesize a specific camera setting... the generator will not be able to interpret and generate scene-consistent images."
— Yuan et al., "Generative Photography" (CVPR 2025)
Why Text Fails for Spatial Description
Academic research reveals fundamental reasons why spatial concepts resist linguistic encoding:
"Low angle" has no standardized definition across the 3 billion+ LAION training images. Each training example had its own interpretation.
"Left of frame" depends on framing, which depends on focal length, which depends on camera distance — a chain of unstated dependencies.
Combining multiple spatial constraints in a single prompt compounds failure rates exponentially as ambiguities multiply.
"Existing approaches for layout control are limited to 2D layouts... and fail to preserve generated images under layout changes. This makes these approaches unsuitable for applications that require 3D object-wise control."
— Eldesokey & Wonka, "Build-A-Scene" (2024)
The Scene-First Paradigm
CineForge inverts the traditional prompt-first workflow:
Traditional workflow:
Describe → Generate → Hope → Retry
CineForge workflow:
Stage → Position → Capture → Forge
The key workflow steps:
- 1. Generate or upload 3D objects — Text-to-3D, image-to-3D, custom .glb/.gltf uploads, or pull from Sketchfab
- 2. Position objects in 3D space — Translate, rotate, scale with professional transform controls plus snap-to-grid, edge magnets, drop-to-floor, group, and lock
- 3. Apply PBR materials — Per-mesh albedo, normal, roughness, and metalness maps with triplanar projection. Curated starter library plus AI-generated textures and a smart-thumbnail user library
- 4. Light the set — 3-point cinematic rig (key/fill/rim) with Kelvin temperature control, sun, ambient, and exponential atmospheric fog
- 5. Block the camera — Orbital controls, focal length selection (15mm–200mm), height adjustment
- 6. Add composition layers — Environment, atmosphere, lighting, effects as creative direction
- 7. Render — Choose your tier: real-time preview, AI-forged cinematic frame with mandatory spatial preservation, or pure path-traced Cinematic+ for ray-traced caustics and GI
Technical Implementation
3D Scene Composition
- React-Three-Fiber powered viewport with transform gizmos
- Automatic model normalization (centering and scaling)
- Real 35mm equivalent focal length to FOV mapping
- Composition layers compiled bottom-to-top into structured prompts
PBR Material System
- Per-mesh material targeting — paint individual sub-meshes within a model
- Full PBR channel set: albedo, normal, roughness, metalness with independent tiling
- Triplanar projection eliminates seams on uvless or procedurally-generated geometry
- Curated starter library (Oak, Brick, Concrete, Steel, Cobblestone, Canvas) plus AI-generated textures and a personal preset library with viewport-snapshot thumbnails
Cinematic Lighting Rig
- Three-point fixture system (key, fill, rim) with spherical positioning
- Kelvin temperature control (1500K–10000K) translated to physically-accurate color
- Soft fixtures use point-light falloff; hard fixtures use spotlight with penumbra control
- Directional sun lamp plus ambient and exponential atmospheric fog
Render Tiers
- Real-time: rasterized preview for staging and lookdev
- Cinematic: AI-forged frame preserving exact camera and composition
- Cinematic+: path-traced render with ray-traced caustics, global illumination, and accurate soft shadows — graceful fallback to Cinematic when unsupported
Composition Toolbar
DCC-grade staging shortcuts: snap-to-grid (¼/½/1 unit), edge magnets to neighbouring objects, drop-to-floor, snap-rotation to 15°, group/ungroup, and per-object lock. Built for set-builders, not just hobbyists.
The Forge Pipeline
3D viewport capture → encoded reference → vision model
↓
Structured Prompt with CRITICAL spatial preservation instructions
↓
AI Interprets viewport as "ground truth" for angle/position
↓
Photorealistic output with EXACT camera preservation
Spatial Fidelity
The render stage is constrained to treat your composed viewport as the ground truth for camera angle, object position, and framing — so the photorealistic output matches what you blocked, rather than reinterpreting the scene.
Composition Layers: AI-Assisted Scene Direction
CineForge implements a 5-layer composition system that compiles into structured AI prompts:
| Layer | Purpose | Example |
|---|---|---|
| Environment | Background setting | "Desert highway stretching to horizon" |
| Atmosphere | Weather & air | "Morning fog rolling through" |
| Lighting | Light sources | "Golden hour sunset glow" |
| Effects | Camera effects | "Shallow depth of field" |
| Props | Additional objects | "Distant city skyline" |
The system also provides AI layer suggestions — analyzing 3D objects on stage (names and prompts) to recommend contextual layers, providing professional creative direction based on scene context.
Comparison: CineForge vs. Alternative Approaches
| Approach | Camera Control | Object Positioning | Reproducibility |
|---|---|---|---|
| Text Prompts | Semantic lottery | Approximate | No |
| ControlNet Pose | Limited | Fixed reference | Partial |
| 2D Bounding Boxes | None | 2D only | Partial |
| CineForge | Full 3D orbital | Full 3D transform | Yes |
Preliminary Observations
From our ongoing AIM pilot study (n=20), scene composition workflows in CineForge align with mental models from set design and cinematography:
Reported "angle frustration" with prompt-based tools
Reported iteration cycles vs. prompt-only workflows
Found staging "natural" from set design experience
* Preliminary findings from pilot study. Expanded validation in progress through 2026.
Conclusion
CineForge represents a paradigm shift from "describing intent" to "staging intent."By providing a 3D composition environment where spatial decisions are made visually — not linguistically — we eliminate the angle lottery that plagues prompt-based generation.
Professional creators stage exactly what they want, then forge it into photorealistic output. The viewport becomes the contract — what you see is what you get.