i think it's miscalculating the hidden reasoning tokens, I don't think it's different tokenizers, even I thought it was a two model router at first
it matches llama3 tokenizer (which seems to be on purpose, i think it's a joke) but image inputs are 17 tokens above glm 5.3 flash