The right GPU for Kohya_ss depends on what you are training. LoRA fine-tuning for Stable Diffusion XL or Flux.1, DreamBooth for character consistency, and full model fine-tuning all have different hardware demands. This guide breaks down what you actually need for each scenario.
Quick answer: For LoRA training, 16GB VRAM hits the sweet spot. The RTX 4090 is the fastest consumer training GPU. The RTX 4060 Ti 16GB is the best budget pick for VRAM headroom. The RTX 3060 12GB works for light LoRA but gets tight fast.
NVIDIA GeForce RTX 4090
24GB GDDR6X24GB VRAM and top-tier compute. Trains SD XL LoRAs in minutes rather than hours. The professional's choice for Kohya_ss.
Affiliate links — we may earn a commission at no extra cost to you. Amazon ships globally; Shopee SG covers Singapore & ASEAN.
VRAM needs by training task
| Training task | Minimum VRAM | Recommended VRAM | Notes |
|---|---|---|---|
| SD 1.5 LoRA | 6GB | 8GB | Any modern GPU works |
| SD XL LoRA | 10GB | 12GB+ | 8GB requires aggressive gradient checkpointing |
| Flux.1 LoRA | 16GB | 24GB | Flux is memory-hungry |
| DreamBooth SD XL | 16GB | 24GB | Higher batch sizes need more VRAM |
| DreamBooth Flux.1 | 24GB | 32GB | Very demanding |
| Full model fine-tune | 24GB+ | 40GB+ | Rarely done on consumer hardware |
GPU recommendations by scenario
Scenario 1: Flux.1 LoRA training (most demanding popular task)
Flux.1 LoRA training in Kohya_ss is the current standard for high-quality character and style training. It needs 16GB minimum and runs best with 24GB. Most hobbyists who don’t want to fight the Kohya CLI use FluxGym — a Gradio wrapper that runs the same Kohya sd-scripts backend with a friendlier UX and identical hardware requirements.
Top pick: RTX 4090 — 24GB GDDR6X trains Flux.1 LoRAs comfortably. A typical 1500-step run completes in 20-30 minutes. Fast enough to iterate quickly.
Value pick: RTX 4070 Ti Super — 16GB is tight for Flux.1 but works with gradient checkpointing enabled. Training takes 50-80% longer than the 4090.
Check RTX 4090 on Amazon→Buy on Shopee SG→ Check RTX 4070 Ti Super on Amazon→Buy on Shopee SG→Scenario 2: SD XL LoRA training (most common task)
SD XL LoRA is forgiving. 12GB VRAM handles it, and 16GB gives comfortable headroom for higher resolution or larger batch sizes.
Top pick: RTX 4090 — Fastest training times, excellent VRAM headroom.
Value pick: RTX 4060 Ti 16GB — 16GB GDDR6 is exactly right for SD XL LoRA. Significantly cheaper than the 4090. Training is slower but perfectly usable for personal projects.
Check RTX 4060 Ti 16GB on Amazon→Buy on Shopee SG→Scenario 3: Budget SD 1.5 / SD 2.1 training
For older model training, the requirements drop significantly. A 12GB GPU handles everything comfortably.
Budget pick: RTX 3060 12GB — 12GB at the lowest price point. Trains SD 1.5 LoRAs without issues. Struggles with Flux.1 and higher-VRAM tasks, but for basic character LoRA work it gets the job done.
Check RTX 3060 12GB on Amazon→Buy on Shopee SG→Training time comparison (SD XL LoRA, 1500 steps, batch 1)
| GPU | VRAM | Approx. training time | Relative speed |
|---|---|---|---|
| RTX 4090 | 24GB | ~12 min | 1x (baseline) |
| RTX 4070 Ti Super | 16GB | ~20 min | 0.6x |
| RTX 4060 Ti 16GB | 16GB | ~30 min | 0.4x |
| RTX 3060 12GB | 12GB | ~55 min | 0.22x |
Estimates at 1024x1024 resolution with network cache enabled. Flux.1 LoRA times are 2-3x longer across all GPUs.
What about the RTX 5090?
The RTX 5090 trains faster than the 4090 — roughly 1.5-2x depending on the task. For pure training throughput, it is the fastest consumer option. But the 4090 is already fast enough that the extra speed rarely justifies $400+ more cost for personal training work.
See also: Best GPU for LoRA training, Best GPU for fine-tuning, and Best GPU for DreamBooth.
Which GPU should YOU buy?
- Training Flux.1 LoRAs regularly? RTX 4090 (24GB) is the minimum comfortable option. The 4060 Ti 16GB works but is slow.
- Mostly SD XL LoRA training? RTX 4060 Ti 16GB is the best value — 16GB VRAM, much cheaper than the 4090.
- On a tight budget running SD 1.5? RTX 3060 12GB works. Expect slower training times.
- Professional training pipeline with iteration speed critical? RTX 4090 or RTX 5090 (if budget allows).
- Training infrequently? Cloud GPUs are worth considering — RunPod hourly rates often beat buying hardware for light use.
Common mistakes to avoid
- Buying an 8GB GPU to save money then hitting constant out-of-memory errors in Kohya_ss — 16GB is the proper minimum for modern tasks
- Running Flux.1 LoRA training on 12GB and expecting a smooth experience — it technically runs but barely
- Ignoring gradient checkpointing settings — enabling them on a 16GB GPU can mean the difference between a training run working or failing
- Using a slow HDD for dataset storage — Kohya_ss reads training images repeatedly, and slow storage adds meaningful time across thousands of steps
- Skipping the network cache step — caching latents before training dramatically speeds up LoRA runs and is often overlooked by beginners
Final verdict
| Use case | Best pick | Budget pick |
|---|---|---|
| Flux.1 LoRA | RTX 4090 | RTX 4070 Ti Super |
| SD XL LoRA | RTX 4090 | RTX 4060 Ti 16GB |
| SD 1.5 LoRA | RTX 4060 Ti 16GB | RTX 3060 12GB |
| DreamBooth | RTX 4090 | RTX 4060 Ti 16GB |
For most Kohya_ss users, the RTX 4060 Ti 16GB offers the best balance of VRAM capacity and price. If you train Flux.1 seriously, save up for the RTX 4090.
NVIDIA GeForce RTX 4060 Ti 16GB
16GB GDDR616GB GDDR6 handles SD XL and most Flux.1 LoRA training at a fraction of the 4090's price. The smart pick for serious hobbyists.
Affiliate links — we may earn a commission at no extra cost to you. Amazon ships globally; Shopee SG covers Singapore & ASEAN.
In Kohya_ss, VRAM capacity determines what you can run. Training speed determines how fast you can iterate. Both matter.
Common questions about Kohya_ss GPUs
How much VRAM does Kohya_ss need for SDXL LoRA training?
SDXL LoRA training needs 10GB minimum, with 12GB or more recommended. On 8GB cards it only works with aggressive gradient checkpointing. 16GB gives comfortable headroom for higher resolutions or larger batch sizes, which is why the RTX 4060 Ti 16GB is the value pick for this workload — it has exactly the right capacity at a fraction of the RTX 4090’s price.
Can Kohya_ss run on an 8GB or 12GB GPU?
8GB is enough only for SD 1.5 LoRA training, and buying an 8GB card for modern Kohya_ss work usually leads to constant out-of-memory errors — 16GB is the proper minimum. A 12GB card like the RTX 3060 trains SD 1.5 LoRAs without issues and can handle SDXL LoRA, but Flux.1 training on 12GB technically runs and barely at that.
How long does LoRA training take in Kohya_ss?
For an SDXL LoRA at 1500 steps and batch size 1, expect roughly 10-15 minutes on an RTX 4090, roughly half an hour on an RTX 4060 Ti 16GB, and closer to an hour on an RTX 3060 12GB. Flux.1 LoRA runs take roughly 2-3x longer across all cards. Caching latents before training speeds up runs considerably and is often overlooked.
What GPU do you need for Flux.1 LoRA training in Kohya_ss?
Flux.1 LoRA training needs 16GB VRAM minimum and runs best with 24GB. The RTX 4090 completes a typical 1500-step Flux run in roughly 20-30 minutes. A 16GB card like the RTX 4070 Ti Super works with gradient checkpointing enabled, but training takes 50-80% longer. DreamBooth on Flux.1 is even more demanding, needing 24GB or more.