Skip to content

SAM3: moving the text encoder with "Select CLIP Device" fails with 'NoneType' object has no attribute 'keys' #16675

Description

@DaWasteh

Expected Behavior

SelectCLIPDevice can move the CLIP of a SAM3 / SAM 3.1 checkpoint (CheckpointLoaderSimple) to a second GPU, like any other CLIP.

Actual Behavior

When the selected device differs from the load device, the node fails:

File "comfy_extras/nodes_multigpu.py", line 276, in execute
    clip.patcher = _apply_patcher_device(clip.patcher, resolved)
File "comfy/model_patcher.py", line 521, in deepclone_multigpu
    temp_model_patcher = self.cached_patcher_init[0](*self.cached_patcher_init[1])
File "comfy/sd.py", line 2179, in load_checkpoint_clip_patcher
    _, clip, _, _ = load_checkpoint_guess_config(...)
File "comfy/sd.py", line 2309, in load_state_dict_guess_config
    clip_sd = model_config.process_clip_state_dict(sd)
File "comfy/supported_models.py", line 2464, in process_clip_state_dict
    clip_keys = utils.state_dict_prefix_replace(clip_keys, {...}, filter_keys=True)
File "comfy/utils.py", line 245, in state_dict_prefix_replace
AttributeError: 'NoneType' object has no attribute 'keys'

Cause

SAM3.process_unet_state_dict moves the language_backbone keys into self._clip_stash, and SAM3.process_clip_state_dict reads them back with getattr(self, "_clip_stash", {}). The CLIP-only reload (load_checkpoint_clip_patcher, output_model=False) never calls process_unet_state_dict, so the attribute is missing. The default {} is never used, because supported_models_base.BASE.__getattr__ returns None for every unknown attribute. (With {} the reload would silently build an empty text encoder instead.)

Minimal reproduction (core only)

import comfy.sd, folder_paths
ckpt = folder_paths.get_full_path_or_raise("checkpoints", "sam3.1_multiplex_fp16.safetensors")
comfy.sd.load_checkpoint_clip_patcher(ckpt)   # AttributeError: 'NoneType' object has no attribute 'keys'

Suggested fix

Collect the stash from the checkpoint when it is missing, as process_unet_state_dict does:

def process_clip_state_dict(self, state_dict):
    if self.__dict__.get("_clip_stash") is None:  # CLIP-only reload
        self._clip_stash = {k: state_dict.pop(k) for k in list(state_dict.keys())
                            if "language_backbone" in k and "resizer" not in k}
    clip_keys = self._clip_stash
    ...

With this change the reload above returns the SAM3 text encoder (SAM3ClipModelWrapper, 354M parameters), and the normal load path is unchanged.

Environment

ComfyUI 0.37.0 (a716932; code unchanged on master 8cfe5e1), frontend 1.53.6, Windows 11, PyTorch 2.13 + ROCm, two AMD GPUs (Radeon AI PRO R9700 + RX 9070 XT). Custom nodes are not involved: the reproduction above uses core code only.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions