Skip to content

ImageUpscaleWithModel fails on low-VRAM GPUs (4GB): "Input type (torch.cuda.FloatTensor) and weight type (torch.FloatTensor) should be the same" (v0.29+ regression, broken in 0.30/0.31) #15433

Description

@YiGeSama

Bug Description

ImageUpscaleWithModel (core Upscale Image (using Model) node) crashes on low-VRAM GPUs in v0.31.0 with:

RuntimeError: Input type (torch.cuda.FloatTensor) and weight type (torch.FloatTensor) should be the same

This is a regression: the same workflow ran fine through v0.28 (older code called upscale_model.to(device) unconditionally, so the weights were always moved to GPU). The breakage was introduced in v0.29.0 (commit f8a3fd9, PR #15063) when the node was switched to a ModelPatcher + load_models_gpu(memory_required=...); v0.30.0 and v0.31.0 have identical code.

Environment

  • ComfyUI v0.31.0 (also present on current master dd79c643a; not fixed in v0.31.1)
  • Windows 10, NVIDIA GeForce RTX 3050 Ti Laptop, 4 GB VRAM (~3.2–3.4 GB free)
  • PyTorch 2.7.0+cu128, Python 3.11.6
  • Upscale model: RealESRGAN_x4plus_anime_6B.pth (~67 MB)

Root Cause

v0.31 rewrote the node to load the upscale model through a ModelPatcher with an upfront VRAM estimate (comfy_extras/nodes_upscale_model.py):

memory_required = (512 * 512 * 3) * image.element_size() * max(upscale_model.scale, 1.0) * 384.0 #The 384.0 is an estimate ...
memory_required += image.nelement() * image.element_size()
model_management.load_models_gpu([upscale_model.patcher], memory_required=memory_required)

With a float32 image and scale 4 the estimate is ~4.5 GB (the 384.0 fudge factor dominates). On a 4 GB card with ~3.4 GB free, load_models_gpu concludes the model cannot fit and leaves the weights on CPU — see log:

Requested to load RRDBNet
0 models unloaded.
loaded completely; 0.00 MB usable, 0.00 MB loaded, full load: False

But the node still moves the input tensor to CUDA:

in_img = image.movedim(-1,-3).to(device)

→ F.conv2d receives CUDA input × CPU weights → the error above. The actual model is only ~67 MB; it's the estimate that doesn't fit, so the load is skipped entirely.

Reproduction (verified locally)

# ups = UpscaleModelLoader.execute(...) via the real node code path
mm.load_models_gpu([ups.patcher], memory_required=memory_required)   # estimate ~4.5 GB, free ~3.2 GB
sorted({str(p.device) for p in ups.model.parameters()})              # -> ['cpu']  (nothing loaded)
ups(torch.rand(1, 768, 1344, 3).movedim(-1, -3).to('cuda').float())  # RuntimeError: Input type (torch.cuda.FloatTensor)
                                                                     # and weight type (torch.FloatTensor) should be the same

Suggested Fix

Force the (tiny) weights onto the load device after the load call — restores pre-0.31 behavior:

device = upscale_model.patcher.load_device
...
model_management.load_models_gpu([upscale_model.patcher], memory_required=memory_required)
upscale_model.patcher.model.to(device)

Verified locally: weights land on cuda:0 and the node completes a 4× upscale of a 1344×768 image (5376×3072 output) with no error.

The 384.0 factor in the memory_required formula is also worth revisiting — for typical ESRGAN-style models it overestimates the footprint by orders of magnitude, and that inflated estimate is exactly what makes load_models_gpu refuse to load on small GPUs.

Activity

coderabbitai commented on Aug 8, 2026

@coderabbitai
🔗 Related PRs

#15448 - Fix device mismatch for Upscale Models (Spandrel) with V3 Execution (ROCm) [open]
#15456 - Free other models' memory before retrying handles_tiling VAE decode [open]


🧪 Issue enrichment is currently in open beta.

You can configure auto-planning by selecting labels in the issue_enrichment configuration.

To disable automatic issue enrichment, add the following to your .coderabbit.yaml:

issue_enrichment:
  auto_enrich:
    enabled: false

💬 Have feedback or questions? Drop into our discord!

YiGeSama commented on Aug 8, 2026

@YiGeSama
Author

All of the above content comes from my Hermes. I do not understand technical issues. If there are any problems, please point them out.

aferventu commented on Aug 9, 2026

@aferventu

@YiGeSama I don't know how to offer a fix, but why don't you try something like this: https://github.com/yuvraj108c/ComfyUI-Upscaler-Tensorrt ?
It uses tensorRT technology and it is super fast.

YiGeSama commented on Aug 10, 2026

@YiGeSama
Author

Thank you, but my Hermes has already helped me solve this problem. It seems pretty easy? Also, my graphics card is too weak — I don't think I can run the project you recommended.

changed the title [-]ImageUpscaleWithModel fails on low-VRAM GPUs (4GB): "Input type (torch.cuda.FloatTensor) and weight type (torch.FloatTensor) should be the same" (v0.31 regression)[/-] [+]ImageUpscaleWithModel fails on low-VRAM GPUs (4GB): "Input type (torch.cuda.FloatTensor) and weight type (torch.FloatTensor) should be the same" (v0.29+ regression, broken in 0.30/0.31)[/+] on Aug 10, 2026

aferventu commented on Aug 11, 2026

@aferventu

@YiGeSama You have an RTX card, though with limited VRAM. I don't know how RT tensors work, but if you are able to build the engines, which is the only thing that needs VRAM and is done automatically, then you should be fine, I guess. You can also try the Nvidia RTX video upscaler, which is also used for images. It is way faster than any upscaler. It is using the RT tensors without the need of building any engine, thus extremely fast in upscaling. But I don't know if it matches your work or your GPU. Just trying to help. Cheers!

Dhevenddra commented on Aug 13, 2026

@Dhevenddra

This looks fixed already, by 9eaba63 ("Fix upscale models breaking on non dynamic vram low vram", #15437), which landed on 2026-08-08 shortly after this was filed. I bisected it on comparable hardware to confirm.

Setup: Windows, RTX 3050 Laptop 4 GB (3.46 GB free), torch 2.10.0+cu130, RealESRGAN_x4plus_anime_6B.pth, driving ImageUpscaleWithModel directly with a 512x512 float32 image.

00d02f2  (9eaba63^, before the fix)
  RESULT: FAILED -> RuntimeError: Input type (torch.cuda.FloatTensor) and
                    weight type (torch.FloatTensor) should be the same

9eaba63  (the fix)
  RESULT: upscale SUCCEEDED, output torch.Size([1, 2048, 2048, 3])

b323a34  (current master)
  RESULT: upscale SUCCEEDED

So the exact error in this report reproduces on the parent commit and is gone from the fix onward. The force_full_load=True that #15437 added to load_models_gpu is what does it: the weights now get materialised on the compute device instead of being left on CPU while in_img moves to CUDA.

One observation that survives the fix, in case it is worth its own issue. The estimate is still wildly out of proportion to the model:

actual model weights:     17.9 MB
memory_required estimate:  4.83 GB

That is (512*512*3) * 4 bytes * scale 4 * 384.0, where the 384.0 fudge factor dominates and the real weights are three orders of magnitude smaller. force_full_load=True means the estimate no longer gates the load, so nothing breaks today, but the number is still what made a 17.9 MB model look unloadable on a 4 GB card in the first place. The existing TODO: make it more accurate comment covers it. Happy to open a separate issue if that is useful rather than leaving it buried here.

@YiGeSama if you update to 0.31.1 or later this should work without the local patch.

YiGeSama commented on Aug 14, 2026

@YiGeSama
Author

Thanks a lot for taking the time to bisect this on comparable hardware — that's really thorough, and your explanation of what #15437 does makes it crystal clear.

I've since updated to v0.33.0 and can confirm the fix on my side too: the exact error is gone on my RTX 3050 Ti Laptop (4GB). The core loader path now upscales 512x512 -> 2048x2048 without any issue.

One thing I found along the way that might be worth a look on the ComfyUI side: ImageUpscaleWithModel.execute still hard-requires upscale_model.patcher on its first line. Third-party loaders like WAS Node Suite's "Upscale Model Loader" return a bare spandrel descriptor with no .patcher attached, so that path still throws AttributeError: 'ImageModelDescriptor' object has no attribute 'patcher'. I worked around it locally by attaching a CoreModelPatcher on the fly when it's missing — a small compatibility guard in the node itself would help everyone using third-party loaders, if that's something you think is worth fixing.

And yes, please do open the separate issue about the memory estimate if you think it's useful — it's exactly what made this whole situation confusing in the first place (a 17.9 MB model being "estimated" at 4.83 GB). Happy to confirm details there too.

Thanks again!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions