Skip to content

[Bug] Qwen3-VL-4B-Instruct-Lora 可视化微调案例-李秀奇-'mm_token_type_ids' is missing. #506

Description

@ACED26

出bug的具体模型

Qwen3-VL-4B-Instruct

出bug的具体模型教程

Qwen/Qwen3-VL-4B-Instruct Lora 可视化微调案例 - LaTexOCR

教程负责人

李秀奇

Bug描述

您好,首先感谢您的教程。
我尝试使用教程里的Qwen2.5-VL-3B-Instruct,Qwen3-VL-4B-Instruct,此外也尝试了Qwen3-VL-2B-Instruct。
我发现对于Qwen2.5-VL-3B-Instruct,代码运行成功,
但是对于Qwen3-VL-2B-Instruct和Qwen3-VL-4B-Instruct,我发现报错:
ValueError: Multimodal data was passed (via image_grid_thw or video_grid_thw) but mm_token_type_ids is missing. Please pass mm_token_type_ids to the model so that multimodal RoPE (M-RoPE) can be computed correctly. mm_token_type_ids is returned by the processor alongside input_ids.
这似乎是因为Qwen3-VL系列的新MRoPE导致的。
为了教程的可复现性和扩展性,请检查代码的数据处理部分,感谢!

复现步骤

执行教程里"微调的完整代码"的部分即出现报错。

期望行为

请检查代码的数据处理部分。

环境信息

ubuntu 22.04
Python 3.12
PyTorch 2.7.0
CUDA 12.8

accelerate==1.13.0
peft==0.19.1
qwen-vl-utils==0.0.14
tokenizers==0.22.1
transformers==5.7.0

其他信息

Image Image

确认事项 / Verification

  • 此问题未在过往Issue中被报告过 / This issue hasn't been reported before

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions