Custom Node Testing
Expected Behavior
When a multi-frame IMAGE batch (video frames) is connected to MiniMaxH3AddGuide (frame_idx = 0) with the MiniMax H3 video VAE, the model should use that sequence as an aligned guide latent, the same way it does on CUDA.
The same expectation applies to workflows that inject a video batch via nodes like H3 Inject Video Latent (img2img) for face refine / v2v-style passes.
A single still image on Add Guide should continue to work (this already works on MPS).
Actual Behavior
On Apple Silicon (MPS):
Multi-frame guide input is effectively ignored.
Generation behaves as if only text conditioning (and any still image references) were present.
Example: head-swap workflows that pass a body video into Add Guide only follow the face still; body motion/identity from the video guide is missing.
Sampling completes without error (silent failure).
On CUDA (same models, same workflow, e.g. RunPod):
The same multi-frame guide is respected and head-swap / LMS-style guide passes work as expected.
A single-frame image on MiniMaxH3AddGuide does work on MPS.
Verified that the IMAGE tensor reaching Add Guide is a real multi-frame batch (not collapsed to 1 frame), e.g. 56 frames @ 672×1024 (Get Image Size & Count: width 672, height 1024, count 56). Resolution is multiple of 32; frame count is on the H3 17n+5 grid.
Issue persists after disabling comfyui-applesilicon-fp8. ComfyUI is started with --use-pytorch-cross-attention. Correct MiniMax H3 video VAE is used (not another model’s VAE).
Steps to Reproduce
Apple Silicon Mac, ComfyUI with MPS backend, launch with --use-pytorch-cross-attention.
Load MiniMax H3 ref2va checkpoint + MiniMax H3 video VAE + text encoder.
Use a minimal graph (or Alissonerdx head-swap / LMS workflow):
EmptyMiniMaxH3LatentAV (or equivalent) matching source size/length
CLIPTextEncode with a fixed prompt
Source video → IMAGE batch (≥5 frames, ideally 22/56/107/124…) → MiniMaxH3AddGuide (frame_idx=0, video VAE connected)
Guider + sampler, denoise 1.0
Run on MPS → output ignores the video guide (only still/text behavior).
Run the same workflow/models on CUDA → video guide is applied correctly.
Optional: replace the video batch with a single image on Add Guide on MPS → that still image is respected.
Debug Logs
MPS (fails – guide ignored, run completes):
[INFO] got prompt
[INFO] Model storage policy: fast_disk=False paths=['/Volumes/Civi/IA/ComfyUI/models/vae/minimax_h3_video_vae_fp16.safetensors']
[INFO] VAE load device: mps, offload device: cpu, dtype: torch.float16
[INFO] Model storage policy: fast_disk=False paths=['/Volumes/Civi/IA/ComfyUI/models/vae/minimax_h3_audio_vae_fp32.safetensors']
[INFO] VAE load device: mps, offload device: cpu, dtype: torch.float32
[INFO] gguf qtypes: Q4_K (390), F32 (433), Q6_K (50), Q5_K (27), F16 (2)
[INFO] Model storage policy: fast_disk=False paths=[]
[INFO] CLIP/text encoder model load device: mps, offload device: cpu, current: cpu, dtype: torch.float16
[INFO] Requested to load MiniMaxH3VideoVAE
[INFO] loaded completely; 4966.19 MB loaded, full load: True
[INFO] Requested to load MiniMaxH3TEModel_
[INFO] loaded completely; 15393.81 MB loaded, full load: True
[INFO] Requested to load MiniMaxH3AudioVAE
[INFO] loaded completely; 577.08 MB loaded, full load: True
[INFO] gguf qtypes: F32 (319), Q8_0 (213)
[INFO] model weight dtype torch.bfloat16, manual cast: None
[INFO] model_type FLOW_AV
[INFO] Model storage policy: fast_disk=False paths=[]
[INFO] Requested to load MiniMaxH3
[INFO] 0 models unloaded.
[INFO] loaded completely; 20795.32 MB loaded, full load: True
100%|███████████████████████████████████████████████████████████████████████████████████| 4/4 [08:17<00:00, 124.48s/it]
[INFO] Requested to load MiniMaxH3VideoVAE
[INFO] loaded completely; 4966.19 MB loaded, full load: True
[INFO] Prompt executed in 00:12:13
No exception; wrong semantic result (video guide not used).
CUDA (works):
[INFO] got prompt
[INFO] Model storage policy: fast_disk=False paths=['/workspace/models_local/vae/minimax_h3_video_vae_fp16.safetensors']
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float16
[INFO] Model storage policy: fast_disk=False paths=['/workspace/models_local/vae/minimax_h3_audio_vae_fp32.safetensors']
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.float32
[INFO] gguf qtypes: Q4_K (390), F32 (433), Q6_K (50), Q5_K (27), F16 (2)
[INFO] Model storage policy: fast_disk=False paths=[]
[INFO] CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16
[INFO] Requested to load MiniMaxH3VideoVAE
[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 4965MB Staged. 0 patches attached. Force pre-loaded 128 weights: 352 KB.
[INFO] Requested to load MiniMaxH3TEModel_
[INFO] loaded completely; 43425.97 MB usable, 15385.37 MB loaded, full load: True
[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 4965MB Staged. 0 patches attached. Force pre-loaded 128 weights: 352 KB.
[INFO] Requested to load MiniMaxH3AudioVAE
[INFO] Model MiniMaxH3AudioVAE prepared for dynamic VRAM loading. 576MB Staged. 0 patches attached. Force pre-loaded 401 weights: 539 KB.
[INFO] gguf qtypes: F32 (319), Q8_0 (213)
[INFO] model weight dtype torch.bfloat16, manual cast: None
[INFO] model_type FLOW_AV
[INFO] Model storage policy: fast_disk=False paths=[]
[INFO] Requested to load MiniMaxH3
[INFO] loaded completely; 27355.16 MB usable, 20796.43 MB loaded, full load: True
0%| | 0/4 [00:00<?, ?it/s, Model Initializing ... ]/usr/local/lib/python3.12/dist-packages/torch/nn/functional.py:2954: UserWarning: Mismatch dtype between input and weight: input dtype = c10::BFloat16, weight dtype = float, Cannot dispatch to fused implementation. (Triggered internally at /pytorch/aten/src/ATen/native/layer_norm.cpp:344.)
return torch.rms_norm(input, normalized_shape, weight, eps)
100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [01:27<00:00, 21.94s/it]
[INFO] Comfy model compiler graph breaks: 2, rogues: 0
[INFO] Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading. 4965MB Staged. 0 patches attached. Force pre-loaded 128 weights: 352 KB.
[INFO] Prompt executed in 207.32 seconds
Other
Hardware: Mac M5 Max, 48 GB unified memory (Total VRAM 49152 MB / Total RAM 49152 MB, SHARED)
OS: macOS Tahoe 26.5.2 (Mac Version (26, 5, 2))
Backend: MPS
PyTorch: 2.13.0
Attention: pytorch attention (--use-pytorch-cross-attention / “Using pytorch attention”)
ComfyUI: 0.37.4
Frontend: ComfyUI_frontend v1.52.7
Templates: v0.11.69
Custom nodes present (issue also with AppleSilicon FP8 disabled): LoRA Manager v1.2.1-stable, ComfyUI-AusBoss v2.4.0, EasyUse v1.3.6, ComfyUI-Manager V3.41, rgthree-comfy v1.0.2608210019, KJNodes (used to verify batch count)
Models: MiniMax H3 ref2va + official MiniMax H3 video VAE
Related workflows: Alissonerdx MiniMax H3 LMS / head-swap (MiniMaxH3AddGuide), face-refine via video latent inject
Not a wiring/shape issue: IMAGE batch size and spatial size checked immediately before Add Guide
Custom Node Testing
Expected Behavior
When a multi-frame IMAGE batch (video frames) is connected to MiniMaxH3AddGuide (frame_idx = 0) with the MiniMax H3 video VAE, the model should use that sequence as an aligned guide latent, the same way it does on CUDA.
The same expectation applies to workflows that inject a video batch via nodes like H3 Inject Video Latent (img2img) for face refine / v2v-style passes.
A single still image on Add Guide should continue to work (this already works on MPS).
Actual Behavior
On Apple Silicon (MPS):
Multi-frame guide input is effectively ignored.
Generation behaves as if only text conditioning (and any still image references) were present.
Example: head-swap workflows that pass a body video into Add Guide only follow the face still; body motion/identity from the video guide is missing.
Sampling completes without error (silent failure).
On CUDA (same models, same workflow, e.g. RunPod):
The same multi-frame guide is respected and head-swap / LMS-style guide passes work as expected.
A single-frame image on MiniMaxH3AddGuide does work on MPS.
Verified that the IMAGE tensor reaching Add Guide is a real multi-frame batch (not collapsed to 1 frame), e.g. 56 frames @ 672×1024 (Get Image Size & Count: width 672, height 1024, count 56). Resolution is multiple of 32; frame count is on the H3 17n+5 grid.
Issue persists after disabling comfyui-applesilicon-fp8. ComfyUI is started with --use-pytorch-cross-attention. Correct MiniMax H3 video VAE is used (not another model’s VAE).
Steps to Reproduce
Apple Silicon Mac, ComfyUI with MPS backend, launch with --use-pytorch-cross-attention.
Load MiniMax H3 ref2va checkpoint + MiniMax H3 video VAE + text encoder.
Use a minimal graph (or Alissonerdx head-swap / LMS workflow):
EmptyMiniMaxH3LatentAV (or equivalent) matching source size/length
CLIPTextEncode with a fixed prompt
Source video → IMAGE batch (≥5 frames, ideally 22/56/107/124…) → MiniMaxH3AddGuide (frame_idx=0, video VAE connected)
Guider + sampler, denoise 1.0
Run on MPS → output ignores the video guide (only still/text behavior).
Run the same workflow/models on CUDA → video guide is applied correctly.
Optional: replace the video batch with a single image on Add Guide on MPS → that still image is respected.
Debug Logs
Other
Hardware: Mac M5 Max, 48 GB unified memory (Total VRAM 49152 MB / Total RAM 49152 MB, SHARED)
OS: macOS Tahoe 26.5.2 (Mac Version (26, 5, 2))
Backend: MPS
PyTorch: 2.13.0
Attention: pytorch attention (--use-pytorch-cross-attention / “Using pytorch attention”)
ComfyUI: 0.37.4
Frontend: ComfyUI_frontend v1.52.7
Templates: v0.11.69
Custom nodes present (issue also with AppleSilicon FP8 disabled): LoRA Manager v1.2.1-stable, ComfyUI-AusBoss v2.4.0, EasyUse v1.3.6, ComfyUI-Manager V3.41, rgthree-comfy v1.0.2608210019, KJNodes (used to verify batch count)
Models: MiniMax H3 ref2va + official MiniMax H3 video VAE
Related workflows: Alissonerdx MiniMax H3 LMS / head-swap (MiniMaxH3AddGuide), face-refine via video latent inject
Not a wiring/shape issue: IMAGE batch size and spatial size checked immediately before Add Guide