MOTIVATION LETTER
The standard method for predicting drug resistance mutations requires a crystal structure of the target protein. This excludes roughly 82 percent of clinically relevant mutations because no structure exists. My venture removes that bottleneck entirely. Using ESM-2 protein language model delta-embeddings combined with ECFP4 drug fingerprints and a Random Forest classifier, we predict resistance from protein sequence alone. On the Platinum benchmark of 553 mutations with protein-grouped cross-validation, we achieve an AUROC of 0.804 plus or minus 0.025. The published SOTA, mCSM-lig, reports an AUROC of approximately 0.70. We beat it by over 10 points while covering 100 percent of mutations versus roughly 18 percent for structure-dependent tools.
I am a pharmacist turned machine learning engineer. I built this system as a sole author. I validated it on SKEMPI 2.0, where we scored 0.634, and I am now fine-tuning ESM-2 on the SKEMPI 3K mutation set to push that above 0.70. My roadmap leads directly to a pilot with Servier in Suresnes, followed by recurring revenue. I have already initiated meetings with SEMIA and Quest for Health, submitted applications to IncubAlliance and AI House, and secured a slot in WILCO One BioTech for October 2026. Named partners include Servier, Paris-Saclay I2BC, Institut Pasteur, and Sanofi in Gentilly.
Y Combinator S26 selects companies where removing AI would break the product. That describes my venture exactly. Without the protein language model embeddings, there is no prediction. The core value proposition is the AI itself. I am applying for the standard 500,000 dollar investment to incorporate in the United States, build a small team, and execute the Servier pilot while expanding the training dataset to cover all major drug classes in oncology, antivirals, and antimicrobial resistance.
I am a solo founder. Y Combinator scrutinizes solo founders more heavily. I accept that. My background spans both the biology and the engineering: a pharmacy degree, clinical experience, and a full transition into machine learning engineering with published work in computational pharmacology. I know the drug development pipeline from the bench to the bedside. I know the data formats, the regulatory constraints, and the commercial incentives. I do not need a co-founder to explain the biology or the code. I need capital, compute credits, and network access to pharma partners. Y Combinator provides all three.
The market is large and growing. Antimicrobial resistance alone causes over 1.2 million deaths per year globally. Oncology resistance costs the healthcare system billions in wasted therapy. Every pharmaceutical company developing small molecules needs resistance prediction. My venture delivers it without requiring crystallography labs, without months of structure determination, and without the 82 percent blind spot.
SHORT ESSAY: WHY THIS VENTURE NOW
Three converging trends make this venture viable today that would have been impossible five years ago. First, protein language models like ESM-2 have reached sufficient accuracy to capture mutational effects from sequence alone. Second, the cost of compute has dropped enough that fine-tuning a 650-million-parameter model on a few thousand mutations is feasible for a pre-seed startup. Third, the pharmaceutical industry has accepted AI-driven discovery as a standard tool, opening the door for partnerships with companies like Servier and Sanofi.
The timing is also driven by a specific gap. Existing tools such as mCSM-lig, FoldX, and Rosetta require a protein crystal structure. For most resistance mutations, especially those in emerging viral variants or understudied bacterial targets, no structure exists. My venture fills that gap with a method that works on every protein for which a sequence is known. That is essentially every protein.
I have already validated the approach on a public benchmark. The next step is fine-tuning on a larger, more diverse mutation set to reach an AUROC of 0.70 or higher on SKEMPI 2.0. That milestone unlocks the Servier pilot. The pilot generates the first revenue and the first reference customer. From there, the venture expands to additional pharma partners and a subscription-based software-as-a-service model.
SHORT ESSAY: MARKET AND COMPETITION
The total addressable market includes every pharmaceutical and biotechnology company developing small-molecule drugs. The global drug discovery market is valued at over 70 billion dollars annually. Within that, computational drug discovery is the fastest-growing segment, driven by AI adoption. My venture targets a specific subsegment: resistance prediction for preclinical and clinical-stage compounds.
Direct competitors include mCSM-lig, Rosetta, FoldX, and newer AI tools like AlphaFold-based mutation predictors. All of them require a crystal structure or a high-quality homology model. My venture does not. That is the fundamental differentiator. On the Platinum benchmark, my venture achieves an AUROC of 0.804 versus approximately 0.70 for mCSM-lig. On coverage, my venture scores 100 percent versus approximately 18 percent for structure-limited tools.
Indirect competitors include wet-lab resistance testing, which is slow and expensive, and clinical surveillance, which is reactive. My venture is predictive, fast, and cheap. A single prediction costs pennies in compute and returns results in seconds.
The defensibility comes from the training data and the fine-tuned model. As I add more mutations from pharma partnerships, the model improves and the barrier to entry rises. No competitor can replicate the dataset without access to the same proprietary mutation libraries.
RESEARCH STATEMENT
My venture predicts drug resistance mutations from protein sequence alone using a machine learning pipeline built on ESM-2 protein language model delta-embeddings and ECFP4 drug fingerprints classified by a Random Forest. The input is a protein sequence and a drug structure. The output is a binary classification: resistant or not resistant.
The technical approach has three stages. First, the wild-type protein sequence and each mutant sequence are passed through ESM-2 to generate per-residue embeddings. The delta between the wild-type and mutant embeddings captures the structural and functional effect of the mutation. Second, the drug is encoded as an ECFP4 fingerprint, a standard molecular representation. Third, the delta embeddings and the fingerprint are concatenated and fed into a Random Forest classifier trained on known resistance mutations.
The validation dataset is the Platinum benchmark, which contains 553 mutations across multiple protein targets. I used protein-grouped cross-validation to avoid data leakage between training and test sets. The result is an AUROC of 0.804 with a standard deviation of 0.025. For comparison, the published SOTA mCSM-lig reports an AUROC of approximately 0.70 on the same benchmark. On SKEMPI 2.0, a more challenging dataset of protein-protein interface mutations, my venture scores 0.634.
The current limitation is the SKEMPI 2.0 score. The roadmap addresses this by fine-tuning ESM-2 on the SKEMPI 3K mutation set, which is three times larger and more diverse. I expect this fine-tuning to raise the SKEMPI 2.0 AUROC to 0.70 or higher, matching the performance of structure-dependent tools on their own benchmarks. At that point, the venture has a clear commercial advantage: equivalent accuracy with 100 percent coverage.
The long-term research goal is to extend the method to predict resistance mechanisms, not just binary resistance. This requires interpretability techniques to identify which residues drive the prediction and what structural changes they cause. I plan to integrate attention-based attribution methods from the ESM-2 model to generate mechanistic hypotheses that experimental biologists can test.
CHECKLIST
- [ ] Y Combinator S26 application form completed on the YC website
- [ ] One-minute founder video recorded and uploaded
- [ ] Motivation letter as written above
- [ ] Short essay on why this venture now as written above
- [ ] Short essay on market and competition as written above
- [ ] Research statement as written above
- [ ] Proof-of-concept results summary (AUROC 0.804 on Platinum, 0.634 on SKEMPI 2.0)
- [ ] Founder CV highlighting pharmacy degree, ML engineering experience, and published work
- [ ] List of named partners and current status of each relationship (Servier, Sanofi, Paris-Saclay, Institut Pasteur)
- [ ] Roadmap document with milestones and timeline to Servier pilot and ARR
- [ ] Confirmation of US incorporation plan or existing entity
- [ ] Compute budget and infrastructure plan for fine-tuning ESM-2 on SKEMPI 3K
EDITOR NOTES
- Eligibility risk: Y Combinator typically requires a US-based company. Eniola plans to target EU geography. The application should clarify whether she intends to incorporate in Delaware or another US state and relocate temporarily, or whether she has a US co-founder or employee. This must be addressed explicitly in the application or video.
- Fact verification: The claim that mCSM-lig achieves AUROC approximately 0.70 on Platinum should be double-checked against the most recent published benchmark. If the number has changed, update the application accordingly.
- Gap to fill: The application does not mention any team members or advisors. Y Combinator may ask about plans to hire or recruit a co-founder. Eniola should prepare a specific answer about her hiring timeline and the roles she would fill first (e.g., a second ML engineer, a computational biologist, or a business development lead).
- Personal detail needed: The one-minute video must include Eniola's face, voice, and a demonstration of her domain expertise. The written materials are strong, but the video is a separate deliverable that requires authenticity and clarity. She should practice explaining the technology in 30 seconds to a non-expert audience.