Oriol Vinyals, Samy Bengio, Manjunath Kudlur
6 min
The "OrderMatters" paper investigates the limitations of standard sequence-to-sequence (seq2seq) models when applied to data that does not naturally exist as a sequence. While seq2seq frameworks are highly effective for tasks like machine translation, they rely on the chain rule to process data in a specific order. The authors argue that for many problems—such as sorting a set of numbers or modeling the joint probability of random variables—the order in which data is presented to the model significantly impacts training performance and final accuracy.
The researchers propose a new architecture, the Read-Process-Write model, designed to handle input sets in a permutation-invariant way. This model uses an attention-based mechanism to "read" input elements into a memory, a "process" block that performs computation over these memories without being constrained by input order, and a "write" block that generates the output. By decoupling the internal representation from the input sequence order, the model can effectively process unordered sets.
Furthermore, the authors address the challenge of outputting sets where no natural order exists. They demonstrate that even when the chain rule theoretically allows for any ordering, the choice of output sequence (e.g., depth-first vs. breadth-first traversal for parsing trees) drastically affects the model's ability to learn. They conclude that because deep learning models are sensitive to the optimization landscape, choosing an appropriate, structured ordering for both inputs and outputs is a critical component of successful model design.
Sequences have become first class citizens in supervised learning thanks to the resurgence of recurrent neural networks. Many complex tasks that require mapping from or to a sequence of observations can now be formulated with the sequence-to-sequence (seq2seq) framework which employs the chain rule to efficiently represent the joint probability of sequences. In many cases, however, variable sized inputs and/or outputs might not be naturally expressed as sequences. For instance, it is not clear how to input a set of numbers into a model where the task is to sort them; similarly, we do not know how to organize outputs when they correspond to random variables and the task is to model their unknown joint probability. In this paper, we first show using various examples that the order in which we organize input and/or output data matters significantly when learning an underlying model. We then discuss an extension of the seq2seq framework that goes beyond sequences and handles input sets in a principled way. In addition, we propose a loss which, by searching over possible orders during training, deals with the lack of structure of output sets. We show empirical evidence of our claims regarding ordering, and on the modifications to the seq2seq framework on benchmark language modeling and parsing tasks, as well as two artificial tasks -- sorting numbers and estimating the joint probability of unknown graphical models.
Sam: That's the core tension, and it's worth unpacking carefully. In theory, yes—the math says the model should be order-independent. But learning isn't just about the math. It's about the process of getting there. Think of it like hiking to a mountain summit. There might be dozens of routes, and they all end at the same peak. But some paths are smooth and well-marked, while others are full of dead ends and cliffs. If you start on a bad path, you might get stuck long before you reach the top.
Alex: So the model could theoretically reach the right answer from any starting order—but a bad starting order makes the journey much harder?
Sam: Precisely. The researchers found that the order in which you present data during training has a real effect on how well the model learns. Present it poorly, and the model settles into a suboptimal state—it finds a local solution that looks good enough, but isn't the best it could do.
Alex: How did they address that?
Sam: They introduced something called a "Dynamic Ordering Loss." Here's how it works: during training, instead of committing to one fixed output order, the model tries several different orderings and asks, "which of these makes the correct answer most likely?" It then reinforces that ordering—essentially rewarding itself for finding a more efficient path through the data. Over many training rounds, it learns to naturally gravitate toward orderings that make the problem easier to solve.
Alex: So they're not just building a smarter model. They're teaching it to find better ways of thinking about the problem itself.
Sam: That's a good way to put it. And it points to something the paper makes explicit: even with a theoretically order-independent architecture, the practical process of learning is still sensitive to how data is presented. The researchers are careful to note that this is a real risk—the optimization process can get stuck, and no architecture fully eliminates that. What the Dynamic Ordering Loss does is reduce that risk by making the training process more adaptive.
Alex: So the title is almost a warning as much as a finding—"order matters" even when you've designed a system that's supposed to not care about order.
Sam: Exactly. The contribution isn't just a new architecture. It's a clearer understanding of where the bottleneck actually lives. It's not in the data. It's in our insistence on imposing a fixed sequence during training—and in the model's tendency to get comfortable with whatever path it finds first, even if better paths exist.
Alex: That's a meaningful shift in how you'd think about designing these systems.
Sam: It is. And it has practical implications. The researchers tested this on tasks like sorting sets of numbers—problems where there's no obvious "right" order to present the data. A standard model trained with a fixed output sequence struggled as the sets grew larger. The model using Dynamic Ordering Loss remained more reliable, because it had learned to find a sequence that made the problem tractable, rather than being handed one arbitrarily.
Alex: It's almost like the difference between a student who memorizes answers in a fixed order versus one who understands the material well enough to approach questions from any angle.
Sam: That's a fair analogy. And it's why the researchers frame this not just as a technical fix, but as a rethinking of a basic assumption—that the sequence we impose on training data is neutral. It isn't. The sequence is a design choice, and like any design choice, it can be made well or poorly.
Alex: So the broader takeaway is that when we build AI systems, we should be asking not just "what data are we feeding in," but "in what order, and does that order reflect something real—or is it just a habit?"
Sam: That's precisely what the paper argues. Order is a variable, not a given. And treating it that way—letting the model participate in finding the best sequence rather than having one imposed on it—leads to more capable, more robust systems.
Alex: Thanks for walking through that. And thanks to everyone listening to ResearchPod.