Hardware

RTX 3090 MiniMax H3 Speed Comparison: FP8 Scaled vs INT8 ConvRot (W8A8)

Setup: GPU: RTX 3090 24GB RAM: 32GB ComfyUI 0.30.0 PyTorch 2.13.0+cu130 CUDA 13.0 SageAttention enabled (sageattention-2.2.0+cu130torch2.10.0andhigher.post6-cp310-abi3-win_amd64) Spectrum node + Euler

DGX agentreddit
hardwarer-stablediffusion

Setup: GPU: RTX 3090 24GB RAM: 32GB ComfyUI 0.30.0 PyTorch 2.13.0+cu130 CUDA 13.0 SageAttention enabled (sageattention-2.2.0+cu130torch2.10.0andhigher.post6-cp310-abi3-win_amd64) Spectrum node + Euler, 17 steps Resolution: 0.3 MP Duration: 2 seconds Results minimax_h3_ref2va_pruned_fp8_scaled (native loader) Generation Time 1st 484s 2nd 283s 3rd 270s Some iterations: 2/17 [00:47<05:53, 23.54s/it] 3/17 [01:10<05:31, 23.67s/it] 6/17 [01:58<00:02, 3.79it/s] 8/17 [02:21<00:02, 3.32it/s] 10/17 [02:45<00:02, 2.98it/s] 16/17 [03:57<00:00, 2.50it/s] ---------- minimax_h3_ref2va_pruned_int8_convrot + BobJohnson’s W8A8 node Generation Time 1st 229s 2nd 201s (I didn't do a third generation since it was already obvious who won here.) Some iterations: 2/17 [00:30<03:47, 15.16s/it] 3/17 [00:45<03:28, 14.88s/it] 10/17 [01:45<00:02, 2.55it/s] 14/17 [02:16<00:01, 2.72it/s] Well, I can't post the videos because they're not appropriate, haha, but I basically see no differences. It also has to do with the movements being slow and subtle. This was done in the reference workflow, using an image and video input for character replacement. --- Update: in a comment below I’m showing generation times with the FL2V model for I2V, and it’s a lot faster. Update 2: Since I'm using low steps and two boosters, I can't really spot any quality difference, but this gives a decent reference for render times. I need to run longer and higher-resolution videos for a real quality comparison, though GPUs and version differences might alter results anyway." submitted by /u/Nevaditew [link] [comments]

Related

Source: r/StableDiffusion | 2026-08-05

Loading related sources…