Repository navigation
ImageUpscaleWithModel fails on low-VRAM GPUs (4GB): "Input type (torch.cuda.FloatTensor) and weight type (torch.FloatTensor) should be the same" (v0.29+ regression, broken in 0.30/0.31) #15433
Description
Activity
🔗 Related PRs
#15448 - Fix device mismatch for Upscale Models (Spandrel) with V3 Execution (ROCm) [open]
#15456 - Free other models' memory before retrying handles_tiling VAE decode [open]
🧪 Issue enrichment is currently in open beta.
You can configure auto-planning by selecting labels in the issue_enrichment configuration.
To disable automatic issue enrichment, add the following to your .coderabbit.yaml:
issue_enrichment:
auto_enrich:
enabled: false💬 Have feedback or questions? Drop into our discord!
All of the above content comes from my Hermes. I do not understand technical issues. If there are any problems, please point them out.
@YiGeSama I don't know how to offer a fix, but why don't you try something like this: https://github.com/yuvraj108c/ComfyUI-Upscaler-Tensorrt ?
It uses tensorRT technology and it is super fast.
Thank you, but my Hermes has already helped me solve this problem. It seems pretty easy? Also, my graphics card is too weak — I don't think I can run the project you recommended.
@YiGeSama You have an RTX card, though with limited VRAM. I don't know how RT tensors work, but if you are able to build the engines, which is the only thing that needs VRAM and is done automatically, then you should be fine, I guess. You can also try the Nvidia RTX video upscaler, which is also used for images. It is way faster than any upscaler. It is using the RT tensors without the need of building any engine, thus extremely fast in upscaling. But I don't know if it matches your work or your GPU. Just trying to help. Cheers!
This looks fixed already, by 9eaba63 ("Fix upscale models breaking on non dynamic vram low vram", #15437), which landed on 2026-08-08 shortly after this was filed. I bisected it on comparable hardware to confirm.
Setup: Windows, RTX 3050 Laptop 4 GB (3.46 GB free), torch 2.10.0+cu130, RealESRGAN_x4plus_anime_6B.pth, driving ImageUpscaleWithModel directly with a 512x512 float32 image.
00d02f2 (9eaba63^, before the fix)
RESULT: FAILED -> RuntimeError: Input type (torch.cuda.FloatTensor) and
weight type (torch.FloatTensor) should be the same
9eaba63 (the fix)
RESULT: upscale SUCCEEDED, output torch.Size([1, 2048, 2048, 3])
b323a34 (current master)
RESULT: upscale SUCCEEDED
So the exact error in this report reproduces on the parent commit and is gone from the fix onward. The force_full_load=True that #15437 added to load_models_gpu is what does it: the weights now get materialised on the compute device instead of being left on CPU while in_img moves to CUDA.
One observation that survives the fix, in case it is worth its own issue. The estimate is still wildly out of proportion to the model:
actual model weights: 17.9 MB
memory_required estimate: 4.83 GB
That is (512*512*3) * 4 bytes * scale 4 * 384.0, where the 384.0 fudge factor dominates and the real weights are three orders of magnitude smaller. force_full_load=True means the estimate no longer gates the load, so nothing breaks today, but the number is still what made a 17.9 MB model look unloadable on a 4 GB card in the first place. The existing TODO: make it more accurate comment covers it. Happy to open a separate issue if that is useful rather than leaving it buried here.
@YiGeSama if you update to 0.31.1 or later this should work without the local patch.
Thanks a lot for taking the time to bisect this on comparable hardware — that's really thorough, and your explanation of what #15437 does makes it crystal clear.
I've since updated to v0.33.0 and can confirm the fix on my side too: the exact error is gone on my RTX 3050 Ti Laptop (4GB). The core loader path now upscales 512x512 -> 2048x2048 without any issue.
One thing I found along the way that might be worth a look on the ComfyUI side: ImageUpscaleWithModel.execute still hard-requires upscale_model.patcher on its first line. Third-party loaders like WAS Node Suite's "Upscale Model Loader" return a bare spandrel descriptor with no .patcher attached, so that path still throws AttributeError: 'ImageModelDescriptor' object has no attribute 'patcher'. I worked around it locally by attaching a CoreModelPatcher on the fly when it's missing — a small compatibility guard in the node itself would help everyone using third-party loaders, if that's something you think is worth fixing.
And yes, please do open the separate issue about the memory estimate if you think it's useful — it's exactly what made this whole situation confusing in the first place (a 17.9 MB model being "estimated" at 4.83 GB). Happy to confirm details there too.
Thanks again!
Bug Description
ImageUpscaleWithModel(core Upscale Image (using Model) node) crashes on low-VRAM GPUs in v0.31.0 with:This is a regression: the same workflow ran fine through v0.28 (older code called
upscale_model.to(device)unconditionally, so the weights were always moved to GPU). The breakage was introduced in v0.29.0 (commit f8a3fd9, PR #15063) when the node was switched to a ModelPatcher +load_models_gpu(memory_required=...); v0.30.0 and v0.31.0 have identical code.Environment
dd79c643a; not fixed in v0.31.1)RealESRGAN_x4plus_anime_6B.pth(~67 MB)Root Cause
v0.31 rewrote the node to load the upscale model through a
ModelPatcherwith an upfront VRAM estimate (comfy_extras/nodes_upscale_model.py):With a float32 image and scale 4 the estimate is ~4.5 GB (the
384.0fudge factor dominates). On a 4 GB card with ~3.4 GB free,load_models_gpuconcludes the model cannot fit and leaves the weights on CPU — see log:But the node still moves the input tensor to CUDA:
→
F.conv2dreceives CUDA input × CPU weights → the error above. The actual model is only ~67 MB; it's the estimate that doesn't fit, so the load is skipped entirely.Reproduction (verified locally)
Suggested Fix
Force the (tiny) weights onto the load device after the load call — restores pre-0.31 behavior:
Verified locally: weights land on
cuda:0and the node completes a 4× upscale of a 1344×768 image (5376×3072 output) with no error.The
384.0factor in thememory_requiredformula is also worth revisiting — for typical ESRGAN-style models it overestimates the footprint by orders of magnitude, and that inflated estimate is exactly what makesload_models_gpurefuse to load on small GPUs.