Recently tried to generating some locally and its amazing what you can do now. Haven't done so since the original Stable Diffusion that struggled with any scene more complex than a common pose. Inpainting and controlnet steps was too finicky for my casual use.
The Krea 2 model is really good for complex scenes by having nuance prompting with natural text with some story hints rather than sledgehammer global booru style tag prompting. Specifically using fine-tuned model
Ultra, it recognizes the complex scene and style prompts better than the base model in general.
Here's a couple random assortment I made using comfyui and a messy reference workflow I've currently settled on. Would like to be able to one shot generate a delmo without getting into LoRA or fine tuning, but so far I can only reinforce the outfit details with a reference image with if it got trained on it.
Update:
Interesting, scaling up the resolution lowers fixed a lot of distortions so no real need for a 2nd upscale pass to maybe correct issues. At 1440p, I'm able to get mostly clean results that were impossible to get right on 720p.
prompt for reference (same flow without the outfit embedding):
Scene in CGI movie taking place in World War 2 warfare, winter.
A high angle, wide, camera shot overlooking a deep circular dugout used for cover.
Shadoz a villian organization consisting of mature, cute female agents. Agents wear a uniform, skirt, plain white thigh-high smooth boots, rifle strapped on their right side back.
Story: Team of 6 agents are to assault, leaving the dugout and towards the snowy plains.
In the far background, blurred by a camera depth of field, is dotted with dead fellow Shadoz agents.
Some dead agents are sprawled out along the dugout wall and floor.
Agents attempt to climb up the mud wall towards top-right corner of the camera shot.
Update 2:
Cleaned up the workflow and added to this post. Removed the useless composition prompt block that I'm going to try again later. Since I'm not getting a precise outfit and the outfit block bleeds style (anime outfit causes anime output), I added a reinterpretation pre-processing block to reinterpret the original reference. With the style more aligned, the uniform signal (multiplier block) can be set higher to better inject the outfit in, not perfect though.
Update 4:
Found out the CLIP model is where prompt cross-referencing ("attention") occurs, when I tried to use append clip vectors it just had this free floating idea of the prompt, which did happen to work when all characters wore the same uniform but stopped working when I had a unique character.
Was able to insert both a character and zako uniform prompt but definitely tricky and a hack. Maybe I'm missing some magic sauce system prompt, but Qwen VL and Krea 2 was really struggle to reconcile my inputs (probably my fault).
Here's an example (not the best) and workflow. It did do much better in some generations but had anatomy issues.
Golden delmo takes care of their government agent problem directly: