Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation
In this post, we’ll take a closer look at the internal structure of multimodal Mixture-of-Experts (MoE) models. We’ll start by asking whether their experts naturally organize themselves around different modalities, semantic domains, and even finer-grained visual concepts. We’ll then look at how that structure can be uncovered from the model itself, without needing to probe it with domain-specific data. Finally, we’ll show how this emergent modularity can be put to practical use, using it to adapt multimodal MoEs more efficiently by updating the experts that matter most for a new domain.
Emergent Structure in Mixture-of-Experts VLMs
Let's start by asking a simple question: Do pre-trained multimodal MoEs develop meaningful internal structure? To explore this, we look at expert specialization at three increasingly fine-grained levels: first across modalities, then across semantic domains, and finally across individual visual concepts. We find that this organization appears surprisingly consistently, suggesting that experts are far from interchangeable. We close this section by introducing ExpertLens, a way to uncover this structure directly from the model’s pretrained weights, without needing any domain-specific data.
Our findings on multimodal MoE models' semantic organization add new context to previous works in MoE interpretability. In language-only models, Jiang et al., Antoine et al., and Li et al. have shown that experts can develop preferences for different token types, parts of speech, and other linguistic patterns, without as much specialization for semantic domains. Wang et al. explicitly train text-only MoEs to encourage stronger domain-level modularity, treating modularity as something to build into the training objective rather than an emergent property. By comparison, much less work has looked at multimodal MoEs. Recent studies from Xu et al. and Zhang et al. have started exploring how routing changes across visual, textual, and reasoning stages, but the semantic organization of multimodal experts, which is the focus of our work, is still relatively unexplored.
Specialization by modality
We'll start with the most basic distinction in a multimodal model: vision versus language. Since the same MoE layers process both image patches and text tokens, we asked whether individual experts treat these two modalities differently, or whether they are used more or less interchangeably.
To test this, we can look at how often each expert was activated by image patches versus text tokens when an equal number of both is fed to the model. For Qwen3-VL-30B-A3B, many experts show a clear preference for one modality or the other. In the figure below, each vertical bar represents an expert, with orange showing image-patch activations and purple showing text-token activations. Across layers, experts are rarely perfectly balanced. Many lean strongly toward either vision or language.
This structure is particularly visible when we consider all experts at once (B). Instead of a single unimodal peak for modality preference, the distribution is visibly bimodal: one group of experts is strongly image-oriented, while another peak appears to be more text-oriented. This bimodality is confirmed with a Hartigan dip of 0.0136 (p < 0.001). In other words, even though the model was not explicitly designed with separate "vision experts" and "language experts", that division begins to emerge naturally during pre-training.
(The above figures show Qwen3-VL-30B-A3B. We find Gemma-4-26B-A4B is similarly bimodal, while Kimi-VL-A3B is more nearly unimodal, likely due to its smaller expert pool.)
Specialization by domain
So far we have seen that experts are not used interchangeably across vision and language. We next ask whether this specialization extends within each modality: do multimodal MoEs also organize their representations and experts around different semantic domains?
We study this using datasets from three very different settings: medical imaging, mathematical reasoning, and remote sensing. We start by looking at the token representations that drive routing. When we reduce the dimensionality of image-patch and text-token activations from these datasets through a t-SNE projection, we see some clear, interesting patterns. First, image representations create much tighter domain-specific clusters, while text representations are generally more diffuse. Further, we can see that datasets from the same domain often neighbor each other in the t-SNE projection. This suggests the visible organization reflects broader semantic structure instead of individual dataset identity. Below we include an interactive projection from Layer 18 of Qwen3-VL-30B-A3B, but we can find similar latent structure in other layers of Qwen3-VL-30B-A3B and other multimodal MoEs.
To characterize this organization more systematically, we can measure the local structure of the representation space using nearest-neighbor purity. Basically, we want to see whether grouping by semantic category, versus low-level image or text statistics, results in a better explanation of the organization of intermediate image-patch and text-token activations.
For image patches, representations are organized much more strongly by dataset than by simple visual properties such as average color or spatial patch position. On the other hand, we found that text tokens' representations align more strongly with linguistic structure, such as part of speech, than with dataset identity.
Together, these results point to a particularly strong domain-level organization in the visual representations of multimodal MoEs.
Specialization by object semantics
The domain-level results suggest that multimodal MoEs are organized around high-level visual structure. Now we can ask whether this specialization goes one step further: within a domain, are experts also sensitive to object-level semantics? In other words, do semantically related objects tend to recruit similar experts?
To isolate semantics from dataset effects, we ran the analysis entirely within TinyImageNet. For each object class, we tracked which experts were consistently activated by its image patches. We then measured how much those expert sets overlapped between pairs of classes. We represent the degree of "expert overlap" betwee object classes as Z(a,b).
Across all class pairs, we can see a clear pattern: the farther apart two classes are semantically, the less their expert sets overlap. (A) shows this relationship over all TinyImageNet class pairs, with "semantic distance" represented by WordNet Wu-Palmer distance. The larger the Wu-Palmer distance between two classes, the fewer the number of experts they both recruit.
(B) gives a more concrete view of the same pattern. We consider ten TinyImageNet classes grouped into five semantically related pairs, shown on the right. Each colored solid line corresponds to one related pair, such as a golden retriever and tabby or train and freight car, and shows how strongly their expert sets overlap across layers. The gray dashed line is the baseline overlap for unrelated class pairs, so whenever a colored line sits above it, that semantic pair is recruiting more similar experts than we would expect for unrelated objects. Most pairs remain well above this baseline.
What is particularly surprising is that this is not limited to visually similar pairings. Trains and freight cars share obvious visual structure, but golden retrievers and tabbies, or sea slugs and jellyfish, look very different and still recruit overlapping experts.
Zooming back out, these results suggest that multimodal MoEs develop strong internal structure that spans modality preferences, distinct domain specialization, and sensitivity to finer-grained object semantics.
ExpertLens
So far, we’ve found specialization by looking at how experts respond to data. But if we want to actually use this structure for adaptation, it would be much more useful to identify the relevant experts from the pretrained model itself, without first needing a representative dataset from all the domains we might want to adapt to.
This is the idea behind our method, ExpertLens. ExpertLens recovers expert semantics directly from pretrained weights by decoding router weights into semantically meaningful vocabulary tokens, a la logit lens. We illustrate our method in the animation below.
For a given expert, we take its router weight and pass it through the model’s language-model head. This gives us a ranked list of vocabulary tokens associated with that expert. A math expert, for example, might decode to words like algebra, equation, or Euler; a medical expert might decode to terms related to DNA or patient. We then use those decoded tokens to decide whether an expert is associated with a target domain. The nice part is that this does not require any domain examples. ExpertLens reads the specialization directly from the pretrained model weights.
Now we've identified a set of experts that look domain-specialized from their weights, but do they actually behave that way on real data? To find out, we group the experts that ExpertLens labels as medical, math, or remote sensing specialists and measure how often each group activates on datasets from those domains.
We can visualize these activations, as well as the logits that each expert decodes to. Below we show three experts from layer 18 of Qwen3-VL-30B-A3B that ExpertLens labels as medical (E121), math (E76), and remote sensing (E93), along with the vocabulary tokens decoded from their router weights and the six images whose patches most often routed to each expert.
We find medical experts are preferentially activated by medical data, math experts by math data, and remote sensing experts by remote sensing data, with the clearest alignment on image patches. In other words, the semantic preferences recovered from the weights line up closely with the experts’ actual routing behavior.
Domain-Specialized Experts for Efficient Adaptation
Up to this point, we've mostly treated the internal structure of MoEs as something to understand. Now, given our understanding that experts are emergently specialized by domain, we wonder if we can use that specialization to adapt the model more efficiently?
We test this by adapting a single Qwen3-VL-30B-A3B model sequentially across three domains: math, medical, and remote sensing (same as before). Rather than resetting back to the pretrained model each time, we keep adapting the same model from one domain to the next. This gives us a setting where the model has to repeatedly pick up new domain-specific behavior, while letting us compare different ways of deciding which parameters to update. Our goal is to understand: if we update only the experts that already look relevant to a domain, can we match full fine-tuning with substantially less training time?
For ExpertLens, we use the domain specialists identified from the pretrained weights and update only those expert MLPs, while freezing the rest. We still train the shared decoder components that help route and shape information through the model, including the router, attention projections, and LM head. We compare this against full fine-tuning, where all expert MLPs are updated, and LoRA, a standard parameter-efficient baseline. All methods use the same data, training objective, and sequential adaptation setup.
ExpertLens for efficient adaptation
Results of comparing ExpertLens to full fine-tuning and LoRA are promising! The first thing to notice is that selectively updating the experts identified by ExpertLens does not come at the cost of adaptation quality. Across math, medical, and remote sensing, ExpertLens matches or slightly surpasses full fine-tuning, while LoRA consistently trails behind.
The animation below plays the three methods at their measured training wall-clock speeds. The top row tracks performance over training steps, while the bottom row shows the corresponding in-domain accuracy against actual training time. Depending on the domain, ExpertLens updates only 21.7–47.0% of the model parameters, which substantially reduces training time.
We see that ExpertLens delivers a substantial speedup in training time. Compared with full fine-tuning, ExpertLens is 3.2× faster on math, 4.0× faster on medical, and 4.8× faster on remote sensing, for a 4.0× average speedup across the three domains. It is also 2.9× faster than LoRA, while reaching higher accuracy.
This means that, by leveraging the emergent structure of multimodal MoEs, we can match the performance of full fine-tuning at substantially reduced cost. This is exciting!
Importance of expert selection
In the previous experiment we saw that updating fewer experts can work well, but is it possible that we would have seen the same speedup and performance without going through the ExpertLens selection process? To explore the importance of selecting the right experts, we compare ExpertLens against random expert subsets with exactly the same expert budget.
We use two random baselines. Random (full) samples the same number of experts from the entire expert pool, while Random (disjoint) samples only from experts that ExpertLens did not select. We resample the random subsets for each seed, so the comparison is about expert identity rather than simply how many parameters are being updated. The animation below shows how the three methods behave throughout the same sequential adaptation setup.
With Random (full), we can immediately see that even though it updates the same number of experts, it consistently falls behind ExpertLens, especially on math and medical. This tells us that the benefit of ExpertLens training is not just coming from sparse fine-tuning. Instead, the experts we update seem to matter. Random (disjoint) is a bit more interesting. On math, it can keep up with ExpertLens, which tells us there is some redundancy in the model and more than one subset of experts can work. But on medical and remote sensing it still falls behind the ExpertLens selection. Overall, the picture is that sparse adaptation helps, but picking experts that are already aligned with the target domain gets us the biggest bang for our buck.
Takeaways
The main takeaway from this post is that multimodal MoEs are much more organized internally than they might first appear. Even though their training objective was to route tokens efficiently, their experts develop meaningful specialization across modalities, domains, and even individual visual concepts.
Once we understand this structure, we can take advantage of it. With ExpertLens, we can recover domain-specialized experts directly from the pretrained weights, without needing examples from the target domain. And once we know where that specialization lives, we can use it for adaptation: updating only the relevant experts is enough to match full fine-tuning while cutting training time substantially.
More broadly, we're excited at the prospect of designing algorithms that build upon properties of model architectures. In multimodal MoEs, sparsity induces a natural modularity, and ExpertLens is one example of what becomes possible when the adaptation method is designed around that architectural property.
If you found this post interesting, please consider citing our paper!
@article{marsili2026expertlens,
title = {Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation},
author = {Marsili, Damiano and Kang, Raphi and Mehta, Aditya and Perona, Pietro and Gkioxari, Georgia},
journal = {arXiv preprint arXiv:COMINGSOON},
year = {2026}
}
The paper is not on arxiv yet, but will be soon. Please check back soon!