You need 12GB minimum for modern Stable Diffusion, and 16GB is the real recommendation. SD 1.5 runs on 6GB. SDXL needs 8-12GB. Flux needs 12-16GB. Add ControlNet, and every number goes up by 2-4GB. Here is the exact breakdown for every workflow.
Check NVIDIA GeForce RTX 4070 Ti Super on Amazon→Buy on Shopee SG→Who this is for
This guide is for anyone planning a GPU purchase for Stable Diffusion image generation. Whether you run SD 1.5, SDXL, or Flux — and whether you use basic txt2img or complex multi-ControlNet pipelines — the VRAM requirement varies dramatically based on your workflow.
VRAM usage by model
| Model | Base VRAM (FP16) | With 1 ControlNet | With 2 ControlNets | With Upscaler |
|---|---|---|---|---|
| SD 1.5 (512px) | ~4GB | ~6GB | ~7.5GB | ~6GB |
| SD 1.5 (768px) | ~5.5GB | ~7GB | ~9GB | ~7.5GB |
| SDXL (1024px) | ~6.5GB | ~9GB | ~11.5GB | ~9GB |
| SDXL (1536px) | ~9GB | ~11.5GB | ~14GB | ~12GB |
| Flux dev (1024px, FP8) | ~12GB | ~15GB | ~18GB | ~15GB |
| Flux dev (1536px, FP8) | ~16GB | ~19GB | ~22GB | ~19GB |
Each ControlNet model adds 1.5-3GB depending on the architecture. Upscalers (ESRGAN, Tile ControlNet) add 1.5-3GB on top. The SD rows assume FP16 precision — running at FP32 doubles everything.
The Flux rows are marked FP8 for a reason. FLUX.1-dev is a 12-billion-parameter model published at BF16, so its transformer weights alone are roughly 24GB at full precision — which is what this page’s own summary means by “24GB FP16”. FP8 halves that to about 12GB, and that is the path almost everyone actually runs on a consumer card.
The VRAM calculator takes the same numbers and lets you vary them: pick SD 1.5, SDXL, SD 3.5 or Flux, set the precision and resolution, and see what is left for a ControlNet.
VRAM tiers explained
6-8GB — SD 1.5 only
Cards: RTX 4060 (8GB), RTX 3060 Ti (8GB)
You can run SD 1.5 at 512x512 with one ControlNet. SDXL technically loads but leaves no room for ControlNet or upscaling. Flux will not run. This tier is obsolete for modern workflows.
10-12GB — SDXL baseline
Cards: RTX 3060 12GB, RTX 5070 (12GB)
SDXL works at 1024x1024 with one ControlNet. Flux base generation is tight but possible. Multi-ControlNet workflows and high-res generation will OOM. Workable for casual use but limiting for creative professionals.
16GB — the sweet spot
Cards: RTX 4060 Ti 16GB, RTX 5070 Ti, RTX 4070 Ti Super, RTX 5080
Handles SDXL with multiple ControlNets, Flux at 1024px with ControlNet, LoRA training for SDXL, and high-res generation with upscalers. This is the tier where you stop fighting memory limits and start focusing on creative work.
24GB+ — no limits
Cards: RTX 4090 (24GB), RTX 5090 (32GB), RTX 3090 (24GB used)
Everything runs with headroom. Flux at 1536px with ControlNet, Dreambooth training, multi-ControlNet pipelines at high resolution. Overkill for basic generation but necessary for professional workflows and training.
Real-world VRAM snapshots
| Workflow | Actual VRAM Used | Minimum Card |
|---|---|---|
| SDXL txt2img (1024px) | 6.8GB | 8GB (tight) |
| SDXL + ControlNet Canny | 9.2GB | 12GB |
| SDXL + ControlNet + upscaler | 11.5GB | 12GB (tight) |
| SDXL + 2 ControlNets | 12.1GB | 16GB |
| Flux txt2img (1024px) | 10.4GB | 12GB |
| Flux + ControlNet | 13.8GB | 16GB |
| Flux + ControlNet + upscaler | 16.2GB | 24GB |
| SDXL LoRA training | 10.5GB | 12GB (tight) |
| SDXL Dreambooth training | 18GB | 24GB |
How to reduce VRAM usage
If you are on the edge of fitting a workflow:
- Use FP16 or BF16 precision — halves VRAM vs FP32 with no visible quality loss
- Enable VAE tiling — critical for high-res images, reduces decode VRAM dramatically
- Use Forge or ComfyUI — better memory management than original Automatic1111
- Enable —medvram or —lowvram flags — trades speed for lower peak VRAM
- Offload ControlNet to CPU — slower but frees 1.5-3GB per ControlNet
- Use FP8 quantized Flux — halves the weights against the released BF16, 23.8GB down to about 12GB. Not a trim you can skip: no consumer card holds the full-precision checkpoint
Which GPU should you buy?
SD 1.5 only, basic generation: Any 8GB card works. The RTX 4060 at $479 handles SD 1.5 comfortably.
SDXL with ControlNet (the current mainstream): Buy a 16GB card. The RTX 4070 Ti Super at $800 or RTX 4060 Ti 16GB at $425 covers all standard SDXL workflows. For a focused SDXL-only buying guide, see our best GPU for SDXL article — it covers SDXL-specific optimizations and budget recommendations.
Flux with ControlNet: 16GB is tight, 24GB is comfortable. The RTX 4090 at $2,200 gives you full Flux capability without compromise.
Training (LoRA, Dreambooth): 16GB for SDXL LoRA, 24GB for Dreambooth. Match your training target to the VRAM table above. For deeper picks, see our best GPU for LoRA training and best GPU for Dreambooth guides.
Chroma (Flux-based): Chroma’s Flux 1.dev fork has slightly different VRAM behavior. See our best GPU for Chroma AI guide for the specifics.
Common mistakes to avoid
- Buying 8GB for SDXL. It technically loads, but one ControlNet pushes you over. You will spend more time debugging OOM errors than generating images.
- Running at FP32 precision. This is the most common VRAM waste. FP16 is visually identical for generation and halves your memory footprint.
- Assuming VRAM numbers from 2024 guides are still accurate. Flux and SDXL with ControlNet use significantly more VRAM than SD 1.5. Check the tables above for current requirements.
- Forgetting that VRAM is shared with your display. Running a 4K desktop costs 200-500MB of VRAM. On an 8GB card, that matters.
Final verdict
| Workflow Target | VRAM Needed | Recommended GPU |
|---|---|---|
| SD 1.5 basic | 6-8GB | RTX 4060 ($479) |
| SDXL + ControlNet | 12-16GB | RTX 4070 Ti Super ($800) |
| Flux + ControlNet | 16-24GB | RTX 4090 ($2,200) |
| Training | 16-24GB | RTX 4090 ($2,200) |
For the majority of Stable Diffusion users in 2026, 16GB is the right target. It covers SDXL with ControlNet, basic Flux, and LoRA training — the three workflows that define modern image generation. See our full guides on Stable Diffusion GPUs and Flux GPUs for card-specific recommendations.
Buy 16GB and stop worrying about VRAM. The creative bottleneck should be your imagination, not your memory bus.