Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal, Nanyun Peng
4 min
Abstract
Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over many tokens, enabling more reliable cross-lingual comparisons at the sequence-level. Controlled continual pre-training experiments on four open-source MoEs show that incorporating this routing loss also aligns the underlying hidden representations across languages. Most importantly, this loss improves multilingual performance on our diverse evaluation suite, demonstrating the potential of cross-lingual MoE router alignment.
Sam: Which a careful referee would want separated.
Alex: Yes. A control with an auxiliary loss that carries no cross-lingual information would separate them. I would not read the one-point gain as purely semantic until that is shown.
Sam: What about the load-balancing loss? The router is already being used for balancing.
Alex: The authors did not test removing it. They rely on load balancing to prevent representation collapse. So it is an open question whether the alignment loss could work independently. My guess is that without balancing, experts could specialize by language, which would undercut the purpose. That is a guess, not a result.
Sam: And the limitation that most constrains the work?
Alex: Parallel data. High-quality parallel corpora are scarce and often cover a narrow distribution of text. The authors acknowledge the method cannot be the sole driver of multilingual performance. They treat it as one component of a broader training curriculum.
Sam: So how far does it generalize to non-parallel, real-world data? The alignment is learned entirely from sentence pairs.
Alex: That is not answered by the headline numbers. The gains are measured on benchmarks, and the method's reach beyond the pairs it was trained on is the part a referee would press hardest.
Sam: So a consistent but modest gain, evidence that the alignment propagates into the hidden states, an unseparated optimization confound, and a hard dependency on parallel data.
Alex: That is fair. The router works as an alignment interface, and it offers a low-overhead way to add sequence-level supervision in decoder-only models without token-level alignment. How much of the benefit is semantic is still to be shown.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.