MOTIVATION LETTER
Antimicrobial resistance is projected to cause ten million deaths annually by 2050, yet the tools to predict which mutations will defeat a drug remain structurally limited. The venture predicts drug resistance mutations from protein sequence alone, without requiring a crystal structure. On the Platinum benchmark of 553 mutations, the Random Forest classifier achieves an AUROC of 0.804 plus or minus 0.025 under protein-grouped cross-validation. This beats the published SOTA, mCSM-lig, which reports an AUROC of approximately 0.70. The venture covers 100 percent of mutations, compared to roughly 18 percent for structure-limited tools.
I am Eniola Olutogun, a pharmacist turned machine learning engineer. I built the proof-of-concept as a sole author. The technology uses ESM-2 protein language model delta-embeddings combined with ECFP4 drug fingerprints. The classifier is a Random Forest, chosen for interpretability and capital efficiency. The venture is pre-seed and not yet incorporated, but the validation is complete.
Google India's 2026 AI Startup Accelerator offers world-class AI expertise, mentorship, infrastructure, and global networking. The venture is an AI-native biotech with a scalable technology solution. The technical rigor is demonstrated by the benchmark results. The product-market fit is clear: every pharma company developing small-molecule drugs needs to know which mutations will cause resistance. The business model is capital-efficient because the software runs on protein sequence data, not expensive crystallography.
The venture has named pharma partners: Servier in Suresnes, Sanofi in Gentilly, and academic collaborators at Paris-Saclay, I2BC, and Institut Pasteur. A pilot with Servier is the next milestone. The roadmap targets fine-tuning ESM-2 on the SKEMPI 3K mutation set to achieve an AUROC of 0.70 or higher, then moving to the Servier pilot and recurring revenue.
India has a high burden of infectious disease and antimicrobial resistance. The venture's technology is directly relevant to Indian pharma companies developing antibiotics and antivirals. The accelerator's network could open collaboration pathways with Indian biopharma. The venture is targeting EU, UK, and US programs as primary geography, but the technology is global. Google's AI expertise and cloud infrastructure would accelerate the fine-tuning and deployment pipeline.
The founder's trajectory from pharmacist to ML engineer signals deep domain expertise in both drug discovery and machine learning. The venture is pre-incorporation but has validated proof-of-concept and named partners. This is a capital-efficient, AI-native biotech with potential for Indian pharma collaboration.
SHORT ESSAY: TECHNICAL INNOVATION AND SCALABILITY
The venture's core technical innovation is predicting drug resistance mutations from protein sequence alone, without requiring a crystal structure. This is achieved by using ESM-2 protein language model delta-embeddings, which capture evolutionary and structural information from sequence, combined with ECFP4 drug fingerprints. The classifier is a Random Forest, chosen for its interpretability and low computational cost.
On the Platinum benchmark of 553 mutations, the model achieves an AUROC of 0.804 plus or minus 0.025 under protein-grouped cross-validation. This beats the published SOTA, mCSM-lig, which reports an AUROC of approximately 0.70. The venture covers 100 percent of mutations, compared to roughly 18 percent for structure-limited tools. On the SKEMPI 2.0 benchmark, the AUROC is 0.634.
The technology is scalable because it does not require crystal structures. Most proteins of therapeutic interest have no solved structure. The model can predict resistance for any protein with a known sequence, which is essentially all proteins. The compute cost is low: ESM-2 embeddings are precomputed, and the Random Forest trains in minutes on a standard laptop.
The roadmap includes fine-tuning ESM-2 on the SKEMPI 3K mutation set to achieve an AUROC of 0.70 or higher on that benchmark. This would match or exceed structure-based tools while maintaining 100 percent coverage. The venture is designed for capital efficiency: no wet lab, no crystallography, no expensive compute. The business model is software-as-a-service for pharma R and D teams.
SHORT ESSAY: COMMERCIAL VIABILITY AND MARKET FIT
Every pharma company developing small-molecule drugs needs to know which mutations will cause resistance. This is true for oncology, antivirals, antibiotics, and any therapeutic area where resistance emerges. Current tools require crystal structures, which are available for only about 18 percent of relevant mutations. The venture covers 100 percent.
The named pharma partners include Servier in Suresnes and Sanofi in Gentilly. A pilot with Servier is the next milestone. The venture is also in discussion with SEMIA and Quest for Health, and has applied to IncubAlliance and AI House. The WILCO One BioTech programme is scheduled for October 2026.
The business model is capital-efficient. The software runs on protein sequence data, which is freely available from public databases. The compute cost is low. The pricing model is per-protein or per-project subscription for pharma R and D teams. The target market includes the top 50 pharma companies and hundreds of biotechs.
The venture is pre-seed and not yet incorporated, but the proof-of-concept is validated. The founder is a pharmacist turned ML engineer, which gives deep domain expertise in both drug discovery and machine learning. The venture is targeting EU, UK, and US programs, but the technology is global. India's high burden of infectious disease and AMR makes it a natural market for collaboration.
RESEARCH STATEMENT
The venture predicts drug resistance mutations from protein sequence alone. The technology uses ESM-2 protein language model delta-embeddings combined with ECFP4 drug fingerprints. The classifier is a Random Forest. On the Platinum benchmark of 553 mutations, the model achieves an AUROC of 0.804 plus or minus 0.025 under protein-grouped cross-validation. This beats the published SOTA, mCSM-lig, which reports an AUROC of approximately 0.70. The venture covers 100 percent of mutations, compared to roughly 18 percent for structure-limited tools. On the SKEMPI 2.0 benchmark, the AUROC is 0.634.
The research roadmap has three phases. Phase one is fine-tuning ESM-2 on the SKEMPI 3K mutation set to achieve an AUROC of 0.70 or higher on that benchmark. This would match or exceed structure-based tools while maintaining 100 percent coverage. Phase two is deploying the model in a pilot with Servier in Suresnes. Phase three is building a recurring revenue model based on per-protein or per-project subscriptions.
The named academic collaborators include Paris-Saclay, I2BC, and Institut Pasteur. The venture is in discussion with SEMIA and Quest for Health. Applications have been submitted to IncubAlliance and AI House. The WILCO One BioTech programme is scheduled for October 2026. Future targets include the EIC Accelerator and BPI i-Lab.
The founder's background as a pharmacist and ML engineer provides the domain expertise to interpret the model's predictions and validate them against known resistance mechanisms. The venture is pre-incorporation but has validated proof-of-concept and named partners. The technology is capital-efficient because it requires no wet lab, no crystallography, and no expensive compute.
The venture is targeting EU, UK, and US programs as primary geography. Google India's 2026 AI Startup Accelerator offers world-class AI expertise, mentorship, infrastructure, and global networking. The accelerator's focus on AI-first startups with scalable technology solutions aligns with the venture's profile. The technical rigor is demonstrated by the benchmark results. The product-market fit is clear. The business model is capital-efficient. The founder has operational discipline and domain expertise.
CHECKLIST
- [ ] Complete online application form at startup.theceo.in/google-india-ai-startup-accelerator-2026/
- [ ] Upload motivation letter (300-500 words)
- [ ] Upload short essay on technical innovation and scalability (200-350 words)
- [ ] Upload short essay on commercial viability and market fit (200-350 words)
- [ ] Upload research statement (400-600 words)
- [ ] Prepare pitch deck (10-15 slides) covering problem, technology, validation, roadmap, team, business model
- [ ] Prepare one-page executive summary
- [ ] Verify programme deadline on website
- [ ] Confirm eligibility for non-India-based startups
- [ ] Prepare video pitch if required by programme
EDITOR NOTES
- Eligibility risk: The programme is Google India's accelerator. The venture is targeting EU, UK, and US programs. Verify whether non-India-based startups are eligible. The profile does not specify the founder's nationality or current location. Eniola must confirm eligibility before submitting.
- Fact verification: The AUROC of 0.804 on Platinum and 0.634 on SKEMPI 2.0 are drawn from the profile. These numbers should be verified against the actual benchmark results. The claim that mCSM-lig AUROC is approximately 0.70 should be checked against the published paper.
- Gap: The profile does not specify the founder's nationality, current location, or whether the venture has any connection to India. The application may require a local entity or co-founder. Eniola must insert personal detail on eligibility and geography.
- Gap: The profile mentions "SEMIA / Quest for Health meeting in progress" but does not specify the outcome. Eniola should update the status before submitting.
- Gap: The venture is pre-incorporation. The application may require a registered company. Eniola should check whether incorporation is required and, if so, consider incorporating in France or India before the deadline.