Partyzan
Member
Thanks for the feedback!Thank you so much! As a Stable 1.5 user who is kinda stuck in a rut, this is extremely helpful.
To answer your questions about the workflow: converting comic panels with this kind of raw or mature tone is basically impossible on mainstream commercial platforms (like Midjourney, DALL-E, or standard cloud APIs). Their safety filters flag the scene immediately. The only viable path is using uncensored open-source models running locally on your own hardware.
While many have moved on to newer architectures, Stable Diffusion 1.5 remains the workhorse for this task thanks to its mature, flexible ecosystem:
- Depth vs. Lineart and Canny in ControlNet:
- Depth (depth_zoe): The primary driver for realism. It captures volume, 3D perspective, and camera angles while completely ignoring 2D ink strokes.
- Canny: Great at edge detection, but it blindly grabs everything β including crosshatching, ink shading, and high-contrast lines. If cranked too high, it bakes hard black comic outlines right into the skin and clothes.
- Lineart (lineart_realistic): Slightly more forgiving than Canny because it softens outlines into structural lines. However, for a true 2D-to-photo conversion, both Lineart and Canny need to be kept at a low weight (around 0.3β0.4) and turned off early in the sampling process (ending step ~0.4) just to anchor the silhouette; otherwise, they kill the photorealism.
- ADetailer (face_yolov8s / hand_yolov8n):
In multi-character scenes, SD 1.5 will inevitably smudge distant faces and hands. ADetailer automates high-res inpainting on masked areas right after the main pass, adding real skin pores, sharp eyes, and proper anatomy without losing the original expressions. - Qwen-VL for Prompt Parsing:
A local vision-language model translates the panel into purely physical descriptions (directional lighting, fabric textures, distress expressions) while stripping out comic terminology, formatting it tightly for SD 1.5's CLIP encoder.
Itβs an extremely meticulous, frame-by-frame process with plenty of manual labor:
- Text & Speech Bubbles: Text and bubbles ruin the generation. SD 1.5 doesn't understand text context well enough to remove it cleanly, so you have to clean it up manually beforehand in Photoshop or through tedious inpainting passes (models like Flux handle text recognition and inpainting much better, but they lack this exact level of lightweight modular control).
- Character Consistency: My biggest unresolved bottleneck right now is keeping the exact same face and features across consecutive comic panels. General prompts always drift. Solving this across an entire comic sequence will likely require training dedicated character LoRAs or setting up a multi-reference pipeline (like IP-Adapter/ReActor), which is the next challenge on my list.














