MOTIVATION LETTER
The Platinum benchmark measures what structure-based tools cannot: 553 mutations across 37 proteins, where 82 percent of clinically relevant resistance mutations remain invisible to any method requiring a crystal structure. My platform predicts drug resistance mutations from protein sequence alone, achieving an AUROC of 0.804 plus or minus 0.025 on that benchmark with protein-grouped cross-validation, against 0.70 for the published state of the art, mCSM-lig. It covers 100 percent of mutations in the test set, not the 18 percent that structure-limited tools can reach. This is the operational utility that HealthTech 250's weighted scoring model rewards: evidence over hype, a narrow therapeutic focus, and a methodology that works where existing tools fail.
The venture is a pre-seed, proof-of-concept validated platform for precision oncology and drug resistance prediction. I am Eniola Olutogun, a pharmacist turned machine learning engineer, and sole author of the underlying research. The system combines ESM-2 protein language model delta-embeddings with ECFP4 drug fingerprints, classified by a Random Forest. No crystal structure is required. That architectural choice is the product thesis: resistance mutations occur in disordered regions, binding pockets, and allosteric sites that crystallography frequently cannot resolve, and the 18 percent coverage ceiling of structure-dependent tools is a hard limit, not a fixable engineering problem.
The programme's selection criteria emphasize deep specialization in narrow therapeutic areas, and 62 percent of the cohort concentrates there. My platform is built for one question: will this patient's tumor evade this drug, and which mutation will cause it? The answer guides treatment selection, combination strategy, and clinical trial design. The U.S. market context is direct: oncology is the dominant innovation category in your portfolio, and the platform's output is a treatment decision, not a research artifact.
The validation path is concrete. Fine-tuning ESM-2 on the SKEMPI 3K mutation dataset targets an AUROC of 0.70 or higher on that benchmark, which currently sits at 0.634. A pilot with Servier in Suresnes is the next commercial milestone, with a path to annual recurring revenue through pharma partnerships and clinical decision support licensing. Named collaborators include Paris-Saclay's I2BC, Institut Pasteur, and Sanofi in Gentilly.
HealthTech 250's emphasis on healthcare-native methodology and sustainable business models matches this venture's structure. The platform is a deployable software product with a benchmarked accuracy advantage and a clear revenue model, not a lab curiosity. The U.S. cohort is the right venue to test whether the platform's value proposition holds outside the European research network where it was developed. I am applying to this specific cohort because the programme's criteria, relevance, evidence, partnerships, funding context, momentum, and health checks, map directly onto the venture's current stage and measurable progress.
RESEARCH STATEMENT
The problem is structural. Drug resistance mutations are the primary cause of treatment failure in oncology and infectious disease, but the tools used to predict them depend on protein crystal structures that exist for only a fraction of clinically relevant targets. On the Platinum benchmark, structure-limited tools cover roughly 18 percent of mutations. My platform removes that dependency entirely.
Methodology. The system uses ESM-2, a protein language model trained on 65 million protein sequences, to generate embeddings for wild-type and mutant sequences. The delta between those embeddings captures the biophysical consequence of a mutation without requiring a folded structure. These delta-embeddings are concatenated with ECFP4 drug fingerprints, a standard molecular representation, and fed to a Random Forest classifier. The model learns the interaction between a specific drug and a specific mutation's effect on protein function.
Validation status. On the Platinum benchmark, 553 mutations with protein-grouped cross-validation, the platform achieves AUROC 0.804 plus or minus 0.025. This exceeds the published state of the art, mCSM-lig, at approximately 0.70. On SKEMPI 2.0, a binding affinity benchmark, the platform achieves 0.634, which is below the target of 0.70 and identifies the specific weakness to address. The gap is not in the architecture; it is in training data volume. SKEMPI 2.0 contains roughly 3,000 mutations, and fine-tuning ESM-2 on that set is the immediate next step.
Why sequence alone matters. Resistance mutations frequently occur in regions that crystallography cannot capture: disordered loops, allosteric sites, and interface residues that only become structured upon binding. A structure-based tool is blind to these by construction. A sequence-based tool is not. The 100 percent mutation coverage on Platinum is the direct consequence of removing the structural prerequisite, not a benchmark artifact.
Commercial application. The platform's output is a ranked list of likely resistance mutations for a given drug-protein pair, with confidence scores. In a clinical setting, this informs which mutations to test for in a patient's tumor sequencing data, which alternative drugs to consider, and which combination strategies are likely to fail. In a drug development setting, it identifies resistance liabilities before clinical trials, reducing late-stage attrition.
Roadmap. Fine-tune ESM-2 on SKEMPI 3K mutations to reach AUROC 0.70 or higher on that benchmark. Complete a pilot with Servier in Suresnes to validate the platform on their internal resistance datasets. Convert the pilot into a paid license, establishing annual recurring revenue. The named partner network, including Paris-Saclay's I2BC, Institut Pasteur, and Sanofi in Gentilly, provides the biological validation and clinical context that a computational platform requires.
Why this programme. HealthTech 250's selection criteria weight evidence and operational utility. The platform has benchmarked evidence, a narrow therapeutic focus in oncology, and a clear operational output: a treatment decision. The U.S. cohort is the appropriate venue to test the platform's commercial viability outside the European research ecosystem where it was developed. The programme's emphasis on sustainable business models aligns with the platform's path to recurring revenue through pharma licensing and clinical decision support.
ESSAY: OPERATIONAL UTILITY
A clinician receives a tumor sequencing report. It lists 400 mutations. Which one causes resistance to the prescribed therapy? Structure-based tools answer this for roughly 18 percent of mutations, and only when a crystal structure exists for that specific protein-drug complex. My platform answers it for all of them, in under a minute, from the protein sequence alone.
The operational difference is not incremental. It changes what the clinician can do with the data they already have. A patient with a mutation in a disordered region of a kinase domain, invisible to structure-based tools, is currently treated as having no resistance mechanism. The platform identifies that mutation, scores its likelihood of causing resistance to the specific drug, and ranks alternative therapies. This is the difference between a research tool and a clinical decision support system.
The platform's benchmark results support this operational claim. On Platinum, AUROC 0.804 plus or minus 0.025 with protein-grouped cross-validation means the model generalizes to unseen proteins, not just unseen mutations of familiar proteins. The 100 percent mutation coverage means no patient's sequencing data is discarded for lack of a crystal structure. The SKEMPI 2.0 result, 0.634, is honest about the current limitation: binding affinity prediction needs more training data, which is why fine-tuning on SKEMPI 3K is the immediate next step.
The business model follows the operational utility. A pharma partner like Servier pays for a tool that predicts resistance liabilities before clinical trials, reducing the cost of late-stage attrition. A clinical diagnostics partner pays for a tool that interprets sequencing data at the point of care. Both are recurring revenue models, not one-off consulting engagements.
HealthTech 250's criteria ask for signal over hype. The signal here is a benchmarked accuracy advantage, a structural coverage advantage, and a named pharma partner in active discussion. The hype would be claiming clinical deployment before the Servier pilot is complete. This application does not do that.
CHECKLIST
- [ ] Complete online application at galengrowth.com/healthtech-250-us-cohort-2026
- [ ] Verify current deadline on programme website before submission
- [ ] Confirm eligibility for U.S. cohort as non-U.S. founder; check visa or entity requirements
- [ ] Attach this motivation letter as the personal statement
- [ ] Attach this research statement as the technology description
- [ ] Attach this essay as the operational utility response
- [ ] Upload AUROC benchmark results from Platinum and SKEMPI 2.0 as evidence
- [ ] Include CV with pharmacy license, ML engineering experience, and sole authorship of the platform
- [ ] Prepare one-paragraph summary of Servier pilot status and named collaborators (I2BC, Institut Pasteur, Sanofi)
- [ ] Confirm whether incorporation status (pre-seed, not yet incorporated) affects eligibility; note planned entity structure
- [ ] List compute requirements and current cloud usage for accelerator compute credit assessment
- [ ] Prepare pitch deck in U.S. format, emphasizing oncology and clinical decision support, not Nigeria-first access angle
EDITOR NOTES
- Eligibility risk: programme is U.S.-focused and this venture is Nigeria-first in origin and not yet incorporated. The application must emphasize U.S. market potential and clinical decision support utility, as the strategy notes direct, but the lack of a U.S. entity may be a hard filter. Verify whether a U.S. subsidiary or planned incorporation is required before applying.
- Facts to verify: the AUROC 0.804 plus or minus 0.025 on Platinum and 0.634 on SKEMPI 2.0 must be reproducible from the applicant's own runs; the mCSM-lig comparison at 0.70 is from published literature and should be cited. The claim of 100 percent mutation coverage versus 18 percent for structure-limited tools needs a precise definition of the comparison set.
- Gaps to fill: the Servier pilot is described as in progress, not signed. Do not imply a committed partnership. The application should state the pilot is under discussion and name the contact at Servier if permitted. The applicant must insert the specific therapeutic area for the pilot (oncology, antiviral, or AMR) as the profile lists multiple sectors.
- The SKEMPI 2.0 result at 0.634 is below the stated target of 0.70. The application frames this as a training data limitation, which is defensible, but the applicant should be prepared to explain why fine-tuning on SKEMPI 3K will close a 0.066 gap when the architecture is unchanged.
- The programme's digital health focus may not align with a computational drug discovery platform. The framing in this draft positions the venture as clinical decision support, which fits digital health, but the applicant should confirm the programme accepts platform companies rather than only direct-to-consumer or provider-facing software.