A trainable input embedding table assigns each vocabulary item an independently adjustable vector. We investigate whether this token-specific parameterization is required for substantial language-modeling capability, or whether a shared Transformer can learn from fixed token identities. We compare three decoder-only language models trained from scratch with the same tokenizer, contextual backbone, untied output-head architecture, and training recipe, with a target budget of 100 billion prediction tokens per model. Their input interfaces are a learned table, canonical 16-bit token-ID codes, and one fixed invertible recoding over GF(2). The fixed codes are repeated to model width without an additional trainable input projection. Both fixed-code models acquire substantial capabilities: canonical codes achieve 52.40\% HellaSwag normalized accuracy, 70.51\% PIQA accuracy, and 42.75\% LAMBADA accuracy. The learned-input control performs better on several evaluations, including HellaSwag and LAMBADA, so these results establish viability rather than performance parity. The fixed interfaces remove 100.7 million trainable parameters, yielding 1.711B-parameter models, but parameter reduction is not the central result. These single-run experiments distinguish architectural necessity from empirical utility: independently trainable token-specific input vectors are not required for the observed capabilities. A fixed identity interface also provides a controlled setting for studying representation learning downstream of an immutable input, without establishing where particular capabilities are localized.
Alex: A Transformer can reach substantial language modeling performance without a trainable input embedding table. That's what Andrey Bochkov reports in a recent study of fixed-identity interfaces.
Sam: The embedding table is normally where tokens get mapped into a continuous space. If you remove it, how does the model tell tokens apart?
Alex: The learned table is replaced by a fixed, deterministic code. Each token gets a unique sixteen-bit binary identity, repeated to fill the model's hidden dimension.
Sam: So every token arrives as a static barcode. The backbone has to learn to read those codes from scratch, with no pre-optimized lookup to lean on?
Alex: Yes. The backbone stays fully trainable, but the lexical entry point is frozen. Any semantic structure has to be built downstream of those fixed inputs.
Sam: I'd expect a cost for losing flexibility at the vocabulary boundary.
Alex: There is one. These models are viable, but they trail learned-input controls on benchmarks like HellaSwag. The study's position is that a separately parameterized input table is useful but not strictly necessary.
Sam: Then my first worry is the tokenizer. If the identity is fixed, is the model sensitive to how the text gets partitioned?
Alex: The study holds the tokenizer constant to isolate the interface, so that question isn't tested. What they did test is an invertible linear recoding of the codes over a Galois field. Performance doesn't appear tied to the particular binary assignment.
Sam: That would suggest the backbone is picking up structure in the identity, not memorizing bit patterns.
Alex: That's a reasonable reading, though I'd hold it loosely. The recoding check rules out dependence on one specific code. It doesn't tell you what the backbone is actually extracting.
Sam: The parameter accounting bothers me more. Dropping the table removes over a hundred million parameters. Does the model compensate somewhere, or is it just smaller?
Alex: It's just smaller. The authors describe this as a backbone-matched intervention. They didn't reallocate capacity or adjust the training recipe to make up for the loss.
Sam: Then the gap to the controls could be a parameter-count effect rather than a limit of the fixed interface. That's a serious confound.
Alex: It is, and the study doesn't claim fixed codes are better or even equal. It claims viability: the model can still converge on meaningful linguistic structure.
Sam: What do the authors offer as the reason for the remaining gap?
Alex: They point to the tokenizer. It carries incidental structure, such as irregularities in how pieces of text get split. A learned table can adjust coordinates to smooth those out. With frozen coordinates, the backbone can't, so it has to absorb that irregularity in its contextual processing.
Sam: So the attention layers carry more of the disambiguation. Is the fixed interface effectively a stress test for context use?
Alex: That's a plausible mechanism. The model can't nudge input coordinates to separate similar tokens, so it has to do that work contextually. But I'd call it a hypothesis the setup invites, not something the paper isolates.
Sam: Still, if performance on tasks like PIQA stays close to the controls, that fits the picture that internal representations aren't just reflections of the input embeddings.
Alex: With one caveat: the output head remains fully trainable.
Sam: That's the point I'd flag. Learned lexical storage hasn't been removed, only moved to the output side.
Alex: Yes, the intervention is specific to the input interface. The output projection still carries significant capacity, so "the embedding table is unnecessary" is narrower than it sounds.
Sam: And the training protocol? The study relies on single runs. Could some of these gaps just be initialization noise?
Alex: The authors acknowledge that. With no multiple runs, the numbers are specific checkpoints, not an estimated distribution of outcomes.
Sam: So we can't separate the fixed codes, the missing parameters, and the geometry of the binary interface as sources of the gap.
Alex: No, and that's the limitation that constrains the work most. The evidence supports viability. It doesn't support any ranking of interface designs.
Sam: Then the value is mostly as an instrument. With input identity held constant, you can run cleaner ablations on how the backbone builds representations, without a shifting embedding table as a confound.
Alex: I'd agree. It opens a route to studying how lexical and contextual distinctions emerge across depth, provided follow-up work uses causal interventions rather than correlational probing.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.