Andrey Bochkov
5 min
Abstract
A trainable input embedding table assigns each vocabulary item an independently adjustable vector. We investigate whether this token-specific parameterization is required for substantial language-modeling capability, or whether a shared Transformer can learn from fixed token identities. We compare three decoder-only language models trained from scratch with the same tokenizer, contextual backbone, untied output-head architecture, and training recipe, with a target budget of 100 billion prediction tokens per model. Their input interfaces are a learned table, canonical 16-bit token-ID codes, and one fixed invertible recoding over GF(2). The fixed codes are repeated to model width without an additional trainable input projection. Both fixed-code models acquire substantial capabilities: canonical codes achieve 52.40\% HellaSwag normalized accuracy, 70.51\% PIQA accuracy, and 42.75\% LAMBADA accuracy. The learned-input control performs better on several evaluations, including HellaSwag and LAMBADA, so these results establish viability rather than performance parity. The fixed interfaces remove 100.7 million trainable parameters, yielding 1.711B-parameter models, but parameter reduction is not the central result. These single-run experiments distinguish architectural necessity from empirical utility: independently trainable token-specific input vectors are not required for the observed capabilities. A fixed identity interface also provides a controlled setting for studying representation learning downstream of an immutable input, without establishing where particular capabilities are localized.
Alex: They point to the tokenizer. It carries incidental structure, such as irregularities in how pieces of text get split. A learned table can adjust coordinates to smooth those out. With frozen coordinates, the backbone can't, so it has to absorb that irregularity in its contextual processing.
Sam: So the attention layers carry more of the disambiguation. Is the fixed interface effectively a stress test for context use?
Alex: That's a plausible mechanism. The model can't nudge input coordinates to separate similar tokens, so it has to do that work contextually. But I'd call it a hypothesis the setup invites, not something the paper isolates.
Sam: Still, if performance on tasks like PIQA stays close to the controls, that fits the picture that internal representations aren't just reflections of the input embeddings.
Alex: With one caveat: the output head remains fully trainable.
Sam: That's the point I'd flag. Learned lexical storage hasn't been removed, only moved to the output side.
Alex: Yes, the intervention is specific to the input interface. The output projection still carries significant capacity, so "the embedding table is unnecessary" is narrower than it sounds.
Sam: And the training protocol? The study relies on single runs. Could some of these gaps just be initialization noise?
Alex: The authors acknowledge that. With no multiple runs, the numbers are specific checkpoints, not an estimated distribution of outcomes.
Sam: So we can't separate the fixed codes, the missing parameters, and the geometry of the binary interface as sources of the gap.
Alex: No, and that's the limitation that constrains the work most. The evidence supports viability. It doesn't support any ranking of interface designs.
Sam: Then the value is mostly as an instrument. With input identity held constant, you can run cleaner ablations on how the backbone builds representations, without a shifting embedding table as a confound.
Alex: I'd agree. It opens a route to studying how lexical and contextual distinctions emerge across depth, provided follow-up work uses causal interventions rather than correlational probing.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.