Tag: image-generation
41 discussions across 10 posts tagged "image-generation".
AI Signal - August 18, 2026
-
Workflow for using MiniMax H3 model to generate consistent character sheets by leveraging multi-image reference (up to 9 images). The approach uses "less than ideal" Google Images to build consistent characters and output 360-degree reference sheets for future generations. This demonstrates practical application of video models for character consistency in creative workflows beyond pure video generation.
- Follow-up to my 6-minute TNG video — I changed the workflow a lot for the second one r/StableDiffusion Score: 292
Refined workflow for creating longer-form video content with MiniMax H3 in ComfyUI, with major improvements in character and scene consistency. The creator evolved from independent shot generation to more integrated continuity management, demonstrating iterative learning about video generation pipelines. This shows maturation of local video generation techniques as users develop production-ready workflows.
- If you are generating MMH3 video with Sage Attention. I highly reccomend trying ComfyKitchen instead. r/StableDiffusion Score: 224
Detailed comparison of attention mechanisms for MiniMax H3 video generation, finding that ComfyKitchen provides better quality and prompt adherence than Sage Attention alternatives, though with different performance tradeoffs. This kind of systematic component comparison helps the community optimize local video generation pipelines for both quality and efficiency.
-
Configuration details for running MiniMax H3 video generation on consumer hardware (RTX 5060 Ti 16GB), including specific model variants, VAE settings, and optimization patches. The cherry-picked results demonstrate that high-quality video generation is achievable on mid-tier hardware with proper configuration. This democratizes access to video synthesis capabilities.
AI Signal - August 11, 2026
-
A creative demonstration of MiniMax video generation creating a "what if Isildur destroyed the Ring" alternate scene, showcasing the model's ability to generate coherent narrative video. While entertainment-focused, it demonstrates progress in controllable video generation.
-
ComfyUI merged a new attention mechanism from comfy-kitchen package that claims better speed and visual quality than SageAttention. Early community testing is underway to validate performance claims.
- I trained an open-source realism LoRA for MiniMax H3 - it makes generated people actually look real (weights inside) r/StableDiffusion Score: 481
A LoRA for MiniMax H3 video model that significantly improves human realism, addressing the "uncanny valley" in generated people. The open-source release enables higher-quality human-centric video generation for the community.
- AMA: MiniMax H3 Team — Ask us anything about our open video generation model, training, and future plans r/StableDiffusion Score: 951
The MiniMax H3 team conducted a comprehensive AMA addressing 400+ questions about their open video generation model, training methodology, and roadmap including planned 2K resolution support, audio improvements, and LoRA training capabilities.
-
A 6-minute Star Trek: TNG fan scene built by chaining MiniMax H3 generations, demonstrating the model's capability for extended narrative video when properly managed. Character and environment consistency required reference images and careful editing.
-
Release of an 8-step turbo LoRA for MiniMax H3, with an even faster 4-step version at 768p. Turbo models trade some quality for dramatic speed improvements, making iteration and experimentation more practical.
AI Signal - August 04, 2026
-
MiniMax releases H3, an omni-modal generative system supporting text, images, video, and audio input with native video generation up to 2K resolution and 15-second clips with stereo audio. Multiple workflow optimizations and acceleration nodes are already emerging from the community. Users report ~7 minute renders for 10-second clips on 3090s.
-
New Spectrum acceleration node for MiniMax H3 reduces Euler sampling time by 34% and RES time by 30%. Part of growing collection of model-specific optimizations available through ComfyUI-Manager. Community optimization work is accelerating video generation accessibility.
-
Demonstration of MiniMax H3's video generation capabilities showing physical consistency—a table shifts weight realistically during a conversation scene. The attention to physics and temporal coherence represents significant progress in video generation quality.
-
User leverages Claude to optimize MiniMax H3 ComfyUI workflow for RTX 3090 hardware. Identifies comfy-kitchen version incompatibility breaking ConvRot optimization path. Achieves 7:39 render time for 10-second 864x480 video. Shows AI assisting with AI infrastructure optimization.
AI Signal - July 14, 2026
-
Swift-mlx port of Hunyuan3D enabling image-to-3D generation on Apple Silicon in under 20 seconds using less than 2GB RAM, even running on iPhones. Represents significant progress in making 3D generation accessible on consumer devices.
- I spent weeks optimizing Krea 2 & LTX 2.3 workflows—here they are for free r/StableDiffusion Score: 653
Community member shared optimized workflows for Krea 2 and LTX 2.3 image/video generation, providing free access to weeks of experimentation. Demonstrates the collaborative knowledge-sharing culture around open-source generative models.
- I benchmarked every Krea 2 Turbo checkpoint format in ComfyUI - BF16 vs FP8 vs INT8 ConvRot vs MXFP8 vs NVFP4 (150 matched images) r/StableDiffusion Score: 266
Comprehensive benchmark of Krea 2 quantization formats showing INT8 ConvRot provides the best quality/speed tradeoff on consumer GPUs, outperforming both NVIDIA's NVFP4 and higher-precision formats. Rigorous methodology with 150 matched images across perceptual, semantic, and latent measurements.
AI Signal - July 07, 2026
- I created a node for Krea2 that adds Multi-LORA support with no identity bleeding and per region bounding box control like Ideogram 4 r/StableDiffusion Score: 215
A custom ComfyUI node for Krea2 enables multiple character LoRAs in a single image with bounding-box control, preventing identity bleeding. This brings Ideogram 4-style regional prompting to Krea2.
- SesquiLSR: tiny 1-2x learned latent upscaler for Flux2, Anima, SDXL and more r/StableDiffusion Score: 236
A tiny, fast latent upscaler offering arbitrary scale upscaling as an alternative to bilinear/bicubic for multiple model architectures. The ComfyUI implementation targets improved quality over traditional upscaling methods.
-
New LTX Face ID LoRA from alissonerdx retains identity from close-up reference images, addressing one of the key challenges in consistent character generation across video models.
-
After additional testing, user completely reverses earlier position and finds Krea2 matches or exceeds Ideogram and Z Image for character accuracy with proper LoRA training. Shares training settings that work well.
AI Signal - June 30, 2026
-
Complete rebuild of VNCCS, a ComfyUI extension, with so many changes it's effectively a new project. Represents continued innovation in the Stable Diffusion ecosystem, making complex workflows more accessible.
-
ComfyUI's stable branch added native INT8 support, with claims that ConvRot quantization beats FP8 variants on speed/quality metrics while supporting wider GPU compatibility (2xxx-5xxx NVIDIA cards). This could democratize access to larger image generation models.
-
Community reaction to Dario Amodei's anti-open-source stance, with calls to download and archive models while they remain available. Reflects concern that open-source image models may face restrictions.
-
Side-by-side comparison of Krea 2 and Z-Image Turbo image generation models at 2MP resolution, providing practical insight into model quality differences for practitioners evaluating which to use.
AI Signal - June 23, 2026
- Krea 2 Turbo — Native ComfyUI Workflow + FP8 Weights (12GB, Drag & Drop) r/StableDiffusion Score: 373
Krea 2 now has native ComfyUI support built-in with FP8 quantized weights (24.76GB → 12.01GB). Careful quantization preserving critical layers while compressing weight matrices to float8_e4m3fn format. Makes high-quality image generation accessible on more modest hardware configurations.
- As promised Krea 2 Turbo + "Raw" Quantized in FP8, MXFP8, NVFP4, INT8 and Convrot INT8! r/StableDiffusion Score: 202
Community member released Krea 2 (Base & Turbo) quantized in multiple formats (FP8, MXFP8, NVFP4, INT8, ConvRot INT8) for different GPU tiers. Includes detailed comparison of Raw vs Turbo models and quantization tradeoffs. Demonstrates active open-source optimization ecosystem around new image models.
-
Demonstration of LTX-2.3 water simulation IC-LoRA applied to famous Joker stairs location. Wide shots work well, close-ups more challenging. Shows progress in specialized LoRA for physics simulation in video models, potentially useful for VFX and creative applications.
AI Signal - June 16, 2026
- How far away are we from feature-length AI films? I made this trailer in one week for under $100 r/ChatGPT Score: 832
Creator produced a 4K film trailer in one week for under $100 using Seedance 2.0, Runway, ElevenLabs, Adobe Premiere, and ChatGPT. Demonstrates the accessibility of AI filmmaking tools for independent creators with minimal budgets.
-
Demonstration of SCAIL-2 animation in ComfyUI using Z-Image Turbo character LoRA and TikTok dance clip as motion reference. Created helper node for longer clips to reduce identity drift. Workflow available, showcasing local animation capabilities.
-
Recreation of iconic 1980s horror posters using only Ideogram 4 prompts and bounding boxes—no image reference, controlnets, or LoRAs. Demonstrates impressive compositional control available through prompting alone in newer image generation models.
AI Signal - June 09, 2026
- Ideogram 4.0's Understanding of Characters and IP is Crazy for an Open Model r/StableDiffusion Score: 835
Ideogram 4.0 demonstrates exceptional character and IP knowledge without LoRAs, running locally in ComfyUI at 1.5 megapixels. Initial workflow issues and safety filters have been resolved, making it one of the most capable open image generation models. Generated at 1440x1024 using INT8 versions on consumer hardware.
-
Ideogram 4 running locally on RTX 3060 12GB with 64GB RAM producing high-quality results at ~80 seconds per 1MP image. Demonstrates that cutting-edge image generation is now viable on consumer hardware with careful optimization and cherry-picking.
-
Defense of Ideogram 4 as the closest open model to commercial quality (NB/GPT Image), surpassing recent releases like Ernie, MS Lens, and HiDream. Author emphasizes this is the first model since Z-Image to genuinely impress, suggesting it represents a quality tier shift for open image models.
- How to bypass Ideogram 4's "Image blocked by safety filter" for swimwear/beachwear (Understanding the filter mechanics) r/StableDiffusion Score: 176
Technical analysis of Ideogram 4's safety filter mechanics with methods to bypass for legitimate use cases like swimwear/beachwear photography. Demonstrates how subtle prompt and parameter adjustments can work around overly aggressive filtering while staying within acceptable use.
-
Experimenting with 17-megapixel Ideogram 4 generations taking 10-15 minutes per image. Demonstrates the model's capability at very high resolutions, though composition is hard to predict until deep into generation. Uses Qwen3.6-35B for prompt engineering.
- Ideogram 4: a solution for removing the annoying censorship has been found. r/StableDiffusion Score: 267
Two methods discovered to bypass Ideogram 4's safety filter: shifting first sigma step by +0.005 or +0.01, or using a custom preset with adjusted sigma values. Both methods work by slightly moving the starting point of the diffusion trajectory away from what triggers the filter.
-
Anima 2B model fine-tune (Photanima v2.1) generating quality images in ~2 seconds. Demonstrates exceptional speed and prompt adherence for a 2B model, showing the potential of small, specialized models for specific use cases.
- Lodestone is thinking about training ideogram! Prove him it's a good idea! r/StableDiffusion Score: 191
Community discussion encouraging Lodestone (creator of Chroma) to create a fine-tune or variant of Ideogram 4. Reflects community desire for specialized variants of the new base model to address specific use cases and aesthetic preferences.
AI Signal - June 02, 2026
-
Nvidia dropped a 64B parameter image-to-video model (Cosmos3-Super-Image2Video) on Hugging Face. The near-perfect 0.98 ratio and 132 comments indicate genuine excitement in the image generation community. At 64B parameters, this is a significant resource requirement for local inference but represents a meaningful step in open video generation capability.
- Does anyone else can't stand ComfyUI and prefers classic Automatic/Forge UI? r/StableDiffusion Score: 225
A user frustrated with ComfyUI's node-graph complexity asks for alternatives. The 265-comment thread surfaced SwarmUI (Automatic-style front end over ComfyUI) and Forge Neo as active, maintained alternatives. Represents an ongoing developer experience split in the image generation community: power users favor ComfyUI's programmability; others want the simpler form.