MOTIVATION LETTER
The Platinum benchmark contains 553 mutations across 53 proteins, and structure-based tools like mCSM-lig can only score about 18 percent of them because they require a crystal structure. My venture predicts drug resistance mutations from protein sequence alone, using ESM-2 protein language model delta-embeddings combined with ECFP4 drug fingerprints and a Random Forest classifier. On that same Platinum benchmark, my model achieves an AUROC of 0.804 plus or minus 0.025 under protein-grouped cross-validation, outperforming the published mCSM-lig result of approximately 0.70. This means clinicians and drug developers can now ask which mutations will defeat a drug for any protein, not just the minority with solved structures.
Prosper HealthTech selects ventures that cut costs or save lives, and this technology does both. Drug resistance causes treatment failure in oncology, antivirals, and antimicrobial therapy, and every failed regimen carries a cost measured in months of life and thousands of dollars. My background as a pharmacist who later became an ML engineer gives me direct experience with the clinical problem and the technical solution. I have spent the last two years building and validating this model, and the proof-of-concept is complete. The next step is fine-tuning ESM-2 on the SKEMPI 3K mutation dataset to push AUROC above 0.70 on that benchmark, then running a pilot with Servier in Suresnes to demonstrate real-world utility.
The 12-week in-person format in Birmingham is exactly what I need at this stage. I am a sole founder, and the structured mentorship, biopharma network, and investor introductions that Prosper provides are the highest-use resources available to me. The 50,000 dollars for 5 percent common equity plus the optional 40,000 dollars gives me the runway to complete the SKEMPI fine-tuning, incorporate the company, and secure the Servier pilot. I am prepared to commit to the full 12 weeks, Monday through Friday, and to relocate for the duration of the programme.
The market timing is favorable. Protein language models reached maturity in the last three years, and the gap between what structure-based tools can cover and what clinicians actually need has never been wider. My model covers 100 percent of mutations in the benchmark, not 18 percent. That is a category change in what questions can be answered, not an incremental improvement. Prosper HealthTech's focus on high-utilization products and clinical validation aligns with my roadmap, and I am ready to use the programme's resources to convert a validated proof-of-concept into a deployed clinical tool.
RESEARCH STATEMENT
The problem is drug resistance. Every antibiotic, antiviral, and targeted cancer therapy eventually encounters mutations that render it less effective, and the current computational tools for predicting which mutations matter are structurally limited. mCSM-lig, the published state of the art, requires a crystal structure of the protein-ligand complex. For most clinically relevant proteins, that structure does not exist, which is why mCSM-lig can only evaluate approximately 18 percent of the 553 mutations in the Platinum benchmark. The remaining 82 percent are simply unanswerable with existing tools.
My venture solves this by predicting resistance from sequence alone. The method works as follows. First, I extract delta-embeddings from ESM-2, a 650-million-parameter protein language model trained on 138 million protein sequences. The delta-embedding captures the change in the model's internal representation when a single amino acid is substituted, which encodes the structural and functional consequences of that mutation without requiring a crystal structure. Second, I encode the drug molecule as an ECFP4 fingerprint, a standard circular molecular representation. Third, I concatenate the protein delta-embedding with the drug fingerprint and train a Random Forest classifier to predict whether that specific protein-drug pair will lose efficacy due to the mutation.
The validation results are concrete. On the Platinum benchmark, which contains 553 mutations across 53 proteins with protein-grouped cross-validation, the model achieves an AUROC of 0.804 plus or minus 0.025. This beats the published mCSM-lig performance of approximately 0.70. On SKEMPI 2.0, a binding affinity benchmark, the model achieves 0.634, which is below the target but identifies the specific weakness that the next development phase addresses. The 100 percent mutation coverage versus 18 percent for structure-limited tools is the decisive advantage; a tool that cannot answer 82 percent of questions is a toy, not a tool.
The roadmap has three phases. Phase one, currently underway, is fine-tuning ESM-2 on the SKEMPI 3K mutation dataset, which contains over 3,000 experimentally measured binding affinity changes. The target is an AUROC of at least 0.70 on SKEMPI 2.0, which would close the gap to clinical utility. Phase two is a pilot with Servier in Suresnes, France, where the tool will be tested on their internal oncology resistance questions. Phase three is converting that pilot into an annual recurring revenue contract, with additional pilots at Sanofi in Gentilly and collaborations with Paris-Saclay I2BC and Institut Pasteur.
The clinical validation strategy is direct. The Platinum and SKEMPI benchmarks are experimentally derived, meaning every mutation in the test set has a measured resistance or binding outcome. The model predicts from patterns learned across millions of protein sequences and validated against thousands of experimental measurements, not from theory. The next validation step is prospective, where the model predicts resistance for mutations that have not yet been experimentally characterized, and those predictions are tested in the Servier pilot.
The business model is software-as-a-service for pharmaceutical R&D. Drug developers pay an annual subscription to screen their candidate compounds against resistance mutation libraries before clinical trials, reducing the probability of late-stage failure due to resistance. The total addressable market includes every oncology, antiviral, and antimicrobial program in the global pharmaceutical industry. The competitive advantage is coverage; no other tool can answer resistance questions for proteins without crystal structures, and that is the majority of the proteome.
SHORT ANSWER ESSAYS
Market size and revenue model
The global antimicrobial resistance market alone is projected to exceed 50 billion dollars by 2030, and oncology drug resistance adds a further multi-billion-dollar segment. My revenue model is annual software subscriptions for pharmaceutical R&D teams, priced at 120,000 dollars per year per enterprise seat, with a pilot phase priced at 30,000 dollars for a three-month evaluation. The near-term customer is Servier in Suresnes, where a pilot is already in discussion. The expansion path is Sanofi in Gentilly and academic partnerships at Paris-Saclay I2BC and Institut Pasteur. The unit economics are favorable because the product is software; the marginal cost of serving an additional customer is near zero, and the value delivered, preventing a single failed Phase III trial due to resistance, is measured in hundreds of millions of dollars.
Clinical validation approach
The model has already been validated retrospectively on two independent experimental benchmarks. Platinum contains 553 mutations across 53 proteins with measured resistance outcomes, and the model achieves an AUROC of 0.804 plus or minus 0.025 under protein-grouped cross-validation. SKEMPI 2.0 contains over 3,000 binding affinity measurements, and the model achieves 0.634. The next step is fine-tuning ESM-2 on SKEMPI 3K to raise the SKEMPI 2.0 AUROC above 0.70. Following that, the Servier pilot will provide prospective validation, where the model predicts resistance for mutations that have not yet been experimentally characterized, and Servier's biology team tests those predictions in the lab. This prospective validation is the evidence that will convert pilot customers into paying subscribers.
Why Prosper HealthTech and why now
Prosper HealthTech selects ventures that cut costs or save lives, and my venture does both. The programme's emphasis on clinical validation matches my roadmap exactly; I have retrospective benchmark results and need the structured support to complete prospective validation. The 50,000 dollars for 5 percent common equity provides the runway to fine-tune ESM-2 on SKEMPI 3K, incorporate the company, and secure the Servier pilot. The 12-week in-person format in Birmingham gives me access to mentors who have built and sold health technology companies, and the biopharma network accelerates the path from pilot to revenue. The timing is right because protein language models reached production quality only in the last three years, and no competitor has yet combined them with drug fingerprints for resistance prediction at 100 percent mutation coverage.
Team and fit
I am a pharmacist and ML engineer, which means I have both the clinical training to understand drug resistance and the technical skills to build the solution. I am the sole author of the model and the sole founder of the venture. I acknowledge that Prosper typically selects teams of two or three founders, and I am prepared to recruit a clinical advisor and a commercial co-founder during the programme if the mentors recommend it. What I bring is the core technology, the validation results, and the existing relationships with Servier, Sanofi, and Paris-Saclay. The programme's resources are best applied to the areas where I need support: commercial strategy, investor pitching, and clinical trial design.