Tianlai Chen, Zachary Quinn, Madeleine Dumas, Christina Yi Peng, Lauren Hong, Moises Lopez-Gonzalez, Alexander Mestre, Rio Watson, Sophia Vincoff, Lin Zhao, Jianli Wu, Audrey Stavrand, Mayumi Schaepers-Cheu, Tian Wang, Divya Srijay, Connor Monticello, Pranay Vure, Rishab Pulugurta, Sarah Pertsemlidis, Kseniia Kholina, Shrey Goel, Matthew P. DeLisa, Jen‐Tsan Chi, Ray Truant, Hector C. Aguilar, Pranam Chatterjee
5 min
Computational protein binder design has traditionally relied on structural information, which limits the ability to target conformationally disordered proteins or those lacking known binding pockets. PepMLM addresses this by leveraging the ESM-2 protein language model to design peptide binders based solely on the target protein's amino acid sequence. By masking the peptide region at the C terminus of the target protein during training, the model learns to reconstruct the binder, effectively capturing the biochemical and evolutionary context required for specific binding.
The researchers finetuned the ESM-2-650M model using a curated dataset of known peptide-protein interactions. During the generation phase, the model accepts a target protein sequence and uses a masking strategy to generate candidate peptides. The researchers implemented both greedy decoding and top-k sampling to balance sequence diversity with model confidence. They validated the model's performance through extensive in silico benchmarking against RFdiffusion, using AlphaFold-Multimer to assess binding affinity via ipTM and pLDDT scores.
PepMLM demonstrated superior performance in generating binders compared to RFdiffusion, achieving higher hit rates in both in silico and experimental assays. The model successfully designed peptides that bound to disease-relevant targets such as NCAM1 and AMHR2. Furthermore, when fused to E3 ubiquitin ligase catalytic domains to form ubiquibodies (uAbs), these peptides effectively induced the degradation of pathogenic proteins, including mutant huntingtin (mHTT) and viral phosphoproteins from Nipah, Hendra, and human metapneumovirus. The results suggest that PepMLM provides a robust, sequence-based platform for designing therapeutic binders against diverse and difficult-to-drug targets.
The computational design of protein-based binders presents unique opportunities to access 'undruggable' targets, but effective binder design often relies on stable three-dimensional structures or structure-influenced latent spaces. Here we introduce PepMLM, a target sequence-conditioned designer of de novo linear peptide binders. Using a masking strategy that positions cognate peptide sequences at the C terminus of target protein sequences, PepMLM finetunes the ESM-2 protein language model to fully reconstruct the binder region, achieving low perplexities matching or improving upon validated peptide-protein sequence pairs. After successful in silico benchmarking with AlphaFold-based docking, we experimentally validate the efficacy of PepMLM through both binding and degradation assays. PepMLM-derived peptides demonstrate sequence-specific binding to cancer and reproductive targets, including NCAM1 and AMHR2, and enable targeted degradation of proteins across diverse disease contexts, from Huntington's disease to live viral infections. Altogether, PepMLM enables the design of candidate binders to any target protein, without requiring structural input, facilitating broad applications in therapeutic development.
Alex: [processing, connecting dots] And they saw something similar with viral phosphoproteins?
Sam: [steady, calm] They screened twenty designed constructs against phosphoproteins from Nipah, Hendra, and human metapneumovirus, and got roughly a sixty-three percent hit rate. That's worth noting because these are highly homologous viral proteins — the model captured enough of the underlying sequence grammar to generalize across related strains, without being retrained for each one. [[RP_SECTION:limitations-and-future-outlook|Limitations and Future Outlook]]
Alex: [thoughtful, skeptical] Sixty-three percent is solid, but not close to complete. Where do the failures actually come from?
Sam: [measured, acknowledging limitations] The paper doesn't map the failure modes explicitly, so this is inference. The model decodes greedily, so some sequences likely settle into local minima — plausible-looking peptides that lack the affinity to outcompete whatever the target already binds to inside the cell. And PepMLM is agnostic to the cellular environment: it has no notion of steric hindrance or the broader protein interaction network it's dropping into. That connects to the deeper tension a referee would raise — the model is anchored to the language model's latent space. If a real binder needs an induced-fit conformation to work, the AlphaFold-based validation can end up overestimating how stable that interaction actually is. Telling a genuine binding mode apart from a model artifact remains an open problem.
Alex: [nodding, summarizing] So it's a high-throughput design tool, but still fundamentally probabilistic — it's not reasoning about the full physical context of the cell, just the grammar of the interface.
Sam: [concluding, reflective] That's a fair way to put it. It's a meaningful step toward programmable protein modulation, with real cellular validation behind it — but the gap between sequence-level confidence and physical binding is still the thing to watch.
Alex: If you want the benchmark details and the specific degradation assays we moved through quickly, you can generate a deep dive of this paper — the paper has the rest either way.
Sam: Thanks for listening.