Repository navigation
[Bug] MiniMax H3 image-to-video crashes (SIGSEGV) in Qwen3-VL vision encoder conv on ROCm/AMD #15895
Description
Activity
The traceback here is the Conv3d in the Qwen3-VL vision patch embedding, not a Conv2d. On v0.34.0 the
comfy/ops.pyframes in the report (620_conv_forward, 624forward_comfy_cast_weights, 629forward) belong toops.Conv3d, and the two frames above them areqwen35.py:649(x = self.patch_embed(x)inQwen35VisionModel.forward) andqwen35.py:450(return self.proj(x).view(...)inQwen35VisionPatchEmbed.forward). The call isconv3dwith input(N, 3, 2, 16, 16), weight(1152, 3, 2, 16, 16)andstride == kernel_size. For the 460x726 input that is a 28x46 grid, soN = 1288.That also accounts for T2V working: with no image there is no
preprocess_embed, so the vision tower and this conv are never reached. The quant format, the VAE and the diffusion model are not involved, which matches the report that swapping any of them changes nothing.This is the same crash as #14215, filed in June against a different model (HiDream o1) on gfx1150 with
torch 2.12.0+rocm7.2. Same frames:qwen35.py:650->qwen35.py:455->ops.pyConv3d ->torch/nn/modules/conv.py:730. That PR replaces the conv withF.linearand is still open.Upstream has it too. pytorch/pytorch#165141 is this conv shape (
in_channels=3,out_channels=1152,kernel=stride=(2,14,14), large batch) on an RX 7900 XTX, crashing or degrading by three orders of magnitude, still open and labeledmodule: rocm. pytorch/pytorch#169857 is the same patch-embed conv failing to launch on MI325X withinvalid configuration argument.Worth noting which conv path gfx1100 actually takes:
comfy/model_management.py:477-485setstorch.backends.cudnn.enabled = Falseon AMD unless the arch is inAMD_RDNA2_AND_OLDER_ARCH, and gfx1100 is not in that list, so MIOpen is already off and the fault is in torch's native 5D conv fallback rather than in MIOpen.One catch if #14215 is taken as the fix: it gates the linear path on
x.dtype in (torch.float16, torch.bfloat16), and this path is fp32.qwen3vl.py:66callsself.visual(image.to(device, dtype=torch.float32), grid), andcast_bias_weighttakes its dtype from the input, so the conv runs in fp32 and the gate never fires. The dtype condition needs to be dropped, or extended to fp32, to cover MiniMax H3.Two things that would confirm this: applying the #14215 change with the dtype condition removed, and running once with
COMFYUI_ENABLE_MIOPEN=1to see whether the MIOpen path fails differently. Per pytorch/pytorch#165141 it may still fail, but it separates a MIOpen fault from a native-fallback fault. Unrelated, but the report lists system ROCm 6.2.0 under a+rocm7.2wheel, so it is also worth confirmingLD_LIBRARY_PATHis not putting/opt/rocm6.2 libraries ahead of the bundled ones intorch/lib.One catch if #14215 is taken as the fix: it gates the linear path on x.dtype in (torch.float16, torch.bfloat16), and this path is fp32. qwen3vl.py:66 calls self.visual(image.to(device, dtype=torch.float32), grid), and cast_bias_weight takes its dtype from the input, so the conv runs in fp32 and the gate never fires. The dtype condition needs to be dropped, or extended to fp32, to cover MiniMax H3.
I noticed the same issue as well, I'm actually adjusting it now
Confirmed on Strix Halo / Radeon 8060S (gfx1151).
Environment:
- ComfyUI 0.34.0
- PyTorch 2.12.1+rocm7.2
- Python 3.12.3
T2VA worked normally, including multiple successful generations, but I2VA consistently aborted in the Qwen3-VL vision patch Conv3d path with:
terminate called after throwing an instance of 'std::bad_variant_access' what(): std::get: wrong index for variant Fatal Python error: AbortedThe stack reached:
torch/nn/modules/conv.py:_conv_forwardcomfy/ops.pycomfy/text_encoders/qwen35.pycomfy/text_encoders/qwen3vl.py:preprocess_embedcomfy/text_encoders/minimax.py:preprocess_embed
Applied the current #14215 patch and the same I2VA workflow now works successfully.
So the
F.linearworkaround also fixes this on gfx1151 / Radeon 8060S.Reacted by Mircea Dogaru
[Bug] MiniMax H3 image-to-video (I2V/R2V) crashes with SIGSEGV in Qwen3-VL vision encoder conv on ROCm (AMD)
Environment
12d52794)minimax_h3_fl2va_pruned_int8_convrot.safetensors,qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors(also tried the int8 TE),minimax_h3_video_vae_fp16.safetensors,minimax_h3_audio_vae_fp32.safetensors, turbo 8-step LoRABug
MiniMax H3 text-to-video works fine, but image-to-video (and reference mode) crashes ComfyUI with a hard SIGSEGV (sometimes SIGABRT) during prompt/image conditioning. The crash happens inside the Qwen3-VL vision encoder's Conv2d when encoding the input image — before the diffusion model is even loaded for sampling.
Full Python crash stack:
A second capture (with the image pre-scaled to a multiple of 32) shows the crash deeper in the Conv2d path:
Native crash symbols reference
libtorch_hip.so(at::native::copy_device_to_device/at::native::linspace_cuda_out) andlibamdhip64.so.Reproduction
Fresh ComfyUI v0.34.0 with the models above (official layout:
models/diffusion_models,models/text_encoders,models/vae,models/loras).Use either the official workflow template
video_minimax_h3_i2v.json(replace input image) or a minimal core-node workflow:Queue the prompt with an image connected to
first_frame.ComfyUI segfaults ~4-8 s in, during
MiniMaxH3ImageToVideo's clip conditioning (before any sampling progress is shown).Text-to-video with the exact same models (same workflow, no image input) completes successfully.
What I tried (all crash identically)
ComfyUI-MiniMaxH3-Easy) vs. core nodes (MiniMaxH3ImageToVideo)comfy/ops.pyCastBiasWeightContext(offloadable=False)patch for Conv2dqwen3vl.pypreprocess_embedinput dtypefloat32->bfloat16None avoid the crash. It is 100% reproducible with any image input to the H3 CLIP.
Expected behavior
I2V/R2V should encode the reference image and generate the video, like T2V does on the same stack.
If useful, I can attach the full core-dump symbol list or run additional diagnostics. Happy to test a fix branch on this AMD setup.