Skip to content

[Bug] MiniMax H3 image-to-video crashes (SIGSEGV) in Qwen3-VL vision encoder conv on ROCm/AMD #15895

Description

@niuadou

[Bug] MiniMax H3 image-to-video (I2V/R2V) crashes with SIGSEGV in Qwen3-VL vision encoder conv on ROCm (AMD)

Environment

  • OS: Ubuntu (kernel 7.0.0-28-generic)
  • GPU: AMD Radeon RX 7900 XTX (24 GB VRAM)
  • ROCm: 6.2.0 (system) + bundled ROCm in PyTorch wheels
  • PyTorch: 2.13.0+rocm7.2 (Python 3.14)
  • ComfyUI: v0.34.0 (commit 12d52794)
  • Models: minimax_h3_fl2va_pruned_int8_convrot.safetensors, qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (also tried the int8 TE), minimax_h3_video_vae_fp16.safetensors, minimax_h3_audio_vae_fp32.safetensors, turbo 8-step LoRA

Bug

MiniMax H3 text-to-video works fine, but image-to-video (and reference mode) crashes ComfyUI with a hard SIGSEGV (sometimes SIGABRT) during prompt/image conditioning. The crash happens inside the Qwen3-VL vision encoder's Conv2d when encoding the input image — before the diffusion model is even loaded for sampling.

Full Python crash stack:

Fatal Python error: Segmentation fault
  File "comfy/text_encoders/qwen35.py", line 649, in forward
  File "torch/nn/modules/module.py", line 1789, in _call_impl
  File "torch/nn/modules/module.py", line 1778, in _wrapped_call_impl
  File "comfy/text_encoders/qwen3vl.py", line 66, in preprocess_embed
  File "comfy/text_encoders/minimax.py", line 93, in preprocess_embed
  File "comfy/sd1_clip.py", line 228, in process_tokens
  File "comfy/sd1_clip.py", line 266, in forward
  File "comfy/sd1_clip.py", line 306, in encode
  File "comfy/sd1_clip.py", line 45, in encode_token_weights
  File "comfy/text_encoders/minimax.py", line 114, in encode_token_weights
  File "comfy/sd1_clip.py", line 743, in encode_token_weights
  File "comfy/sd.py", line 410, in encode_from_tokens
  File "comfy/sd.py", line 341, in encode_from_tokens_scheduled

A second capture (with the image pre-scaled to a multiple of 32) shows the crash deeper in the Conv2d path:

  File "torch/nn/modules/conv.py", line 730, in _conv_forward
  File "comfy/ops.py", line 620, in _conv_forward
  File "comfy/ops.py", line 624, in forward_comfy_cast_weights
  File "comfy/ops.py", line 629, in forward
  File "comfy/text_encoders/qwen35.py", line 450, in forward

Native crash symbols reference libtorch_hip.so (at::native::copy_device_to_device / at::native::linspace_cuda_out) and libamdhip64.so.

Reproduction

  1. Fresh ComfyUI v0.34.0 with the models above (official layout: models/diffusion_models, models/text_encoders, models/vae, models/loras).

  2. Use either the official workflow template video_minimax_h3_i2v.json (replace input image) or a minimal core-node workflow:

    LoadImage -> MiniMaxH3ImageToVideo(first_frame) -> [sigma shift -> lora -> BasicGuider]
    + CLIPLoader(type=minimax) + VAELoader(video_vae) + BasicScheduler + KSamplerSelect + SamplerCustomAdvanced
    
  3. Queue the prompt with an image connected to first_frame.

  4. ComfyUI segfaults ~4-8 s in, during MiniMaxH3ImageToVideo's clip conditioning (before any sampling progress is shown).

Text-to-video with the exact same models (same workflow, no image input) completes successfully.

What I tried (all crash identically)

  • Easy custom node (ComfyUI-MiniMaxH3-Easy) vs. core nodes (MiniMaxH3ImageToVideo)
  • nvfp4 TE vs. int8 TE
  • fp16 video VAE vs. int8 video VAE
  • Raw 460x726 image vs. pre-scaled image (ImageScaleToTotalPixels, multiples of 32)
  • comfy/ops.py CastBiasWeightContext(offloadable=False) patch for Conv2d
  • qwen3vl.py preprocess_embed input dtype float32 -> bfloat16

None avoid the crash. It is 100% reproducible with any image input to the H3 CLIP.

Expected behavior

I2V/R2V should encode the reference image and generate the video, like T2V does on the same stack.


If useful, I can attach the full core-dump symbol list or run additional diagnostics. Happy to test a fix branch on this AMD setup.

Activity

  1. 0xDELUXA commented on Aug 28, 2026

    @0xDELUXA
    Contributor

    The traceback here is the Conv3d in the Qwen3-VL vision patch embedding, not a Conv2d. On v0.34.0 the comfy/ops.py frames in the report (620 _conv_forward, 624 forward_comfy_cast_weights, 629 forward) belong to ops.Conv3d, and the two frames above them are qwen35.py:649 (x = self.patch_embed(x) in Qwen35VisionModel.forward) and qwen35.py:450 (return self.proj(x).view(...) in Qwen35VisionPatchEmbed.forward). The call is conv3d with input (N, 3, 2, 16, 16), weight (1152, 3, 2, 16, 16) and stride == kernel_size. For the 460x726 input that is a 28x46 grid, so N = 1288.

    That also accounts for T2V working: with no image there is no preprocess_embed, so the vision tower and this conv are never reached. The quant format, the VAE and the diffusion model are not involved, which matches the report that swapping any of them changes nothing.

    This is the same crash as #14215, filed in June against a different model (HiDream o1) on gfx1150 with torch 2.12.0+rocm7.2. Same frames: qwen35.py:650 -> qwen35.py:455 -> ops.py Conv3d -> torch/nn/modules/conv.py:730. That PR replaces the conv with F.linear and is still open.

    Upstream has it too. pytorch/pytorch#165141 is this conv shape (in_channels=3, out_channels=1152, kernel=stride=(2,14,14), large batch) on an RX 7900 XTX, crashing or degrading by three orders of magnitude, still open and labeled module: rocm. pytorch/pytorch#169857 is the same patch-embed conv failing to launch on MI325X with invalid configuration argument.

    Worth noting which conv path gfx1100 actually takes: comfy/model_management.py:477-485 sets torch.backends.cudnn.enabled = False on AMD unless the arch is in AMD_RDNA2_AND_OLDER_ARCH, and gfx1100 is not in that list, so MIOpen is already off and the fault is in torch's native 5D conv fallback rather than in MIOpen.

    One catch if #14215 is taken as the fix: it gates the linear path on x.dtype in (torch.float16, torch.bfloat16), and this path is fp32. qwen3vl.py:66 calls self.visual(image.to(device, dtype=torch.float32), grid), and cast_bias_weight takes its dtype from the input, so the conv runs in fp32 and the gate never fires. The dtype condition needs to be dropped, or extended to fp32, to cover MiniMax H3.

    Two things that would confirm this: applying the #14215 change with the dtype condition removed, and running once with COMFYUI_ENABLE_MIOPEN=1 to see whether the MIOpen path fails differently. Per pytorch/pytorch#165141 it may still fail, but it separates a MIOpen fault from a native-fallback fault. Unrelated, but the report lists system ROCm 6.2.0 under a +rocm7.2 wheel, so it is also worth confirming LD_LIBRARY_PATH is not putting /opt/rocm 6.2 libraries ahead of the bundled ones in torch/lib.

  2. peterwilli commented on Aug 29, 2026

    @peterwilli

    One catch if #14215 is taken as the fix: it gates the linear path on x.dtype in (torch.float16, torch.bfloat16), and this path is fp32. qwen3vl.py:66 calls self.visual(image.to(device, dtype=torch.float32), grid), and cast_bias_weight takes its dtype from the input, so the conv runs in fp32 and the gate never fires. The dtype condition needs to be dropped, or extended to fp32, to cover MiniMax H3.

    I noticed the same issue as well, I'm actually adjusting it now

  3. Aelvryx commented on Sep 4, 2026

    @Aelvryx

    Confirmed on Strix Halo / Radeon 8060S (gfx1151).

    Environment:

    • ComfyUI 0.34.0
    • PyTorch 2.12.1+rocm7.2
    • Python 3.12.3

    T2VA worked normally, including multiple successful generations, but I2VA consistently aborted in the Qwen3-VL vision patch Conv3d path with:

    terminate called after throwing an instance of 'std::bad_variant_access'
    what(): std::get: wrong index for variant
    Fatal Python error: Aborted
    

    The stack reached:

    • torch/nn/modules/conv.py:_conv_forward
    • comfy/ops.py
    • comfy/text_encoders/qwen35.py
    • comfy/text_encoders/qwen3vl.py:preprocess_embed
    • comfy/text_encoders/minimax.py:preprocess_embed

    Applied the current #14215 patch and the same I2VA workflow now works successfully.

    So the F.linear workaround also fixes this on gfx1151 / Radeon 8060S.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions