Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over many tokens, enabling more reliable cross-lingual comparisons at the sequence-level. Controlled continual pre-training experiments on four open-source MoEs show that incorporating this routing loss also aligns the underlying hidden representations across languages. Most importantly, this loss improves multilingual performance on our diverse evaluation suite, demonstrating the potential of cross-lingual MoE router alignment.
Alex: When a multilingual mixture-of-experts model is pushed to send parallel sentences in different languages through the same experts, cross-lingual transfer improves. That is the work from Bandarkar and colleagues. The supervision goes on the router rather than on the hidden states.
Sam: Why the router? Aligning hidden states directly seems like the obvious route.
Alex: Hidden states are high-dimensional, and pooling them is lossy. When you average across tokens, features tend to cancel. Routing distributions behave more like a sequence-level summary of what the input is about. You can compare the summary for a sentence and its translation without needing token-level correspondence, which is hard to get in decoder-only models.
Sam: So what does the training signal look like?
Alex: A cross-lingual routing loss. It measures the divergence between the routing distributions of the two sides of a parallel pair and pushes them together. The cost is kept low with a shared packed forward pass, where source and target sequences go through the model together. That avoids redundant computation, so the overhead stays minimal.
Sam: My worry is the obvious one. If you force English and Sinhala through the same experts, you may strip out the pathways that handle language-specific structure.
Alex: That is the central trade-off. The main evidence is that the routing loss beats the baseline in nearly every case, across tasks from physical reasoning to knowledge transfer. The average gain is about one point. That is consistent rather than large, and the consistency is what the claim rests on.
Sam: And on the specificity worry?
Alex: The authors' results suggest the shared representational space helps transfer more than the loss of language-specific structure hurts. But that rests on aggregate gains. They did not convincingly establish that router alignment helps reasoning or knowledge tasks more than it helps simple translation. So the "why it helps" is less settled than the "that it helps."
Sam: Does the alignment survive training? Or is it a surface effect on the gating, with the representations unchanged?
Alex: This is the strongest supporting evidence. The loss touches only the router outputs, yet the hidden representations end up aligned too. So the effect is not confined to the gate.
Sam: Here is my confound question, though. How do we know the auxiliary loss is not simply helping optimization, independent of the semantics?
Alex: The authors raise that. They argue the routing divergence loss gives more direct supervision to early-layer parameters, which stabilizes the training signal. So there are plausibly two contributions. One is semantic alignment, and the other is a better gradient path.
Sam: Which a careful referee would want separated.
Alex: Yes. A control with an auxiliary loss that carries no cross-lingual information would separate them. I would not read the one-point gain as purely semantic until that is shown.
Sam: What about the load-balancing loss? The router is already being used for balancing.
Alex: The authors did not test removing it. They rely on load balancing to prevent representation collapse. So it is an open question whether the alignment loss could work independently. My guess is that without balancing, experts could specialize by language, which would undercut the purpose. That is a guess, not a result.
Sam: And the limitation that most constrains the work?
Alex: Parallel data. High-quality parallel corpora are scarce and often cover a narrow distribution of text. The authors acknowledge the method cannot be the sole driver of multilingual performance. They treat it as one component of a broader training curriculum.
Sam: So how far does it generalize to non-parallel, real-world data? The alignment is learned entirely from sentence pairs.
Alex: That is not answered by the headline numbers. The gains are measured on benchmarks, and the method's reach beyond the pairs it was trained on is the part a referee would press hardest.
Sam: So a consistent but modest gain, evidence that the alignment propagates into the hidden states, an unseparated optimization confound, and a hard dependency on parallel data.
Alex: That is fair. The router works as an alignment interface, and it offers a low-overhead way to add sequence-level supervision in decoder-only models without token-level alignment. How much of the benefit is semantic is still to be shown.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.