How to pick the right models and dramatically speed up image and video generation speeds & What I wish I knew when starting on AMD hardware with ComfyUI
For the below I’ll be referring mostly to ComfyUI workflow and model efficiency (using examples for Radeon AI PRO R9700 (32GB) which has a bandwidth of 680GB/sec.
Proof1 community thread
Key points
- I’ve seen 10second minimax H3 image to video workflows go from 28 mins to 3 mins by taking all this info into account and then upscaling after.
- Three terms people mix up: weights are the fixed model file, a LoRA is a small patch to those weights, and the latent is the thing being refined into what your prompt intended before the VAE encodes it into the pixels of your output image or video.
- What uses VRAM: model weights, the text encoder if it shares the card, working memory that grows with resolution × frames, and a spike at VAE decode.
- Rule of thumb: keep weights to about 60-70% of VRAM, roughly 19-22GB on a 32GBVRAM card.
Sources (1)
- [1]How to pick the right models and dramatically speed up image and video generation speeds & What I wish I knew when starting on AMD hardware with ComfyUIr/StableDiffusion (top, day) · Oct 11, 07:47 AM
For the below I’ll be referring mostly to ComfyUI workflow and model efficiency (using examples for Radeon AI PRO R9700 (32GB) which has a bandwidth of 680GB/sec.
I’ve seen 10second minimax H3 image to video workflows go from 28 mins to 3 mins by taking all this info into account and then upscaling after.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 10, 2026Krea2 Turbo Distill 2 step LoRA - FINAL checkpoint released (chk51195)
- Oct 9, 2026Qwen/Qwen-Image-2.1-Turbo
- Oct 8, 2026unslothai/unsloth v0.1.905-beta: Sandboxing is here!
- Oct 8, 2026Parametric Trajectory Distillation for Few-Step Video Generation
- Oct 8, 2026WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
- Oct 7, 2026unslothai/unsloth v0.1.904-beta: Train your own Decision model
