MOTIVATION LETTER
The Platinum benchmark contains 553 mutations across 179 proteins, and my Random Forest classifier, built on ESM-2 protein language model delta-embeddings and ECFP4 drug fingerprints, achieves an AUROC of 0.804 plus or minus 0.025 under protein-grouped cross-validation. That same model covers 100 percent of mutations in the benchmark, where structure-limited tools like mCSM-lig cover roughly 18 percent and score an AUROC near 0.70. The gap matters most in precisely the settings the Novartis Access to Medicine Foundation exists to serve: low- and middle-income countries where sequencing capacity is thin, crystallography is rarer still, and drug resistance erodes the value of every antimicrobial and oncology therapy that reaches the patient.
My name is Eniola Olutogun. I am a pharmacist and machine learning engineer, and I have spent the last year building a pre-seed venture that predicts drug resistance mutations from protein sequence alone, with no crystal structure required. The method is validated on public benchmarks, and the next step is a fine-tuned ESM-2 model trained on the SKEMPI 3K mutation set, targeting an AUROC of at least 0.70 on independent held-out data, followed by a pilot with Servier in Suresnes. The Foundation's mandate to expand access to medicines in LMICs aligns directly with this work because resistance prediction is an access problem. A drug that fails against a resistant pathogen is a drug the patient cannot use, regardless of price or supply chain.
Nigeria is the anchor. I am Nigerian, and the venture's first deployment target is expanding prescription-drug access there, where antimicrobial resistance is a documented barrier to effective treatment for malaria, tuberculosis, and hospital-acquired infections. The technology's ability to operate without crystal structures means a laboratory in Lagos can submit a protein sequence and receive a resistance profile for a panel of candidate drugs in hours, not weeks, without access to the structural biology infrastructure that high-income research centers take for granted.
The Foundation's selection criteria emphasize measurable impact on patient access and health outcomes, feasibility and scalability, team strength, and innovation. On impact, the model's 100 percent mutation coverage versus 18 percent for structure-limited tools is a direct, quantifiable improvement in the number of resistance events that can be caught before treatment failure. On feasibility, the proof-of-concept is complete and published on public benchmarks; the remaining work is fine-tuning and pilot validation. On team, I hold a pharmacy degree and have built the model end-to-end as sole author, with named partnership discussions underway at Servier, Paris-Saclay's I2BC, Institut Pasteur, and Sanofi in Gentilly. On innovation, the delta-embedding approach is a category change in what data is required to make a resistance call, not a marginal improvement over existing tools.
The Foundation's therapeutic areas include oncology and its priority diseases include malaria and dengue. Resistance prediction applies to all of them. I am requesting a grant of up to one million dollars to fund the SKEMPI fine-tuning, the Servier pilot, and the Nigeria deployment study. The money buys compute, validation, and the first real-world evidence that sequence-only resistance prediction improves treatment outcomes in an LMIC setting.
RESEARCH STATEMENT
The venture's core technical claim is that drug resistance mutations can be predicted from protein sequence alone, without a crystal structure, using a protein language model and a drug fingerprint. The current implementation uses ESM-2 delta-embeddings, which capture the change in the protein's learned representation when a mutation is introduced, concatenated with ECFP4 drug fingerprints, and classified by a Random Forest. On the Platinum benchmark, which contains 553 mutations across 179 proteins, the model achieves an AUROC of 0.804 with a standard deviation of 0.025 under protein-grouped cross-validation. This is the correct validation scheme because it prevents the same protein from appearing in both training and test folds, which would otherwise inflate performance. The model also achieves 0.634 on SKEMPI 2.0, a harder benchmark focused on binding affinity changes. The published state of the art, mCSM-lig, scores approximately 0.70 AUROC on comparable tasks but requires a crystal structure and therefore covers only about 18 percent of the mutations in Platinum. My model covers 100 percent.
The gap between 18 percent and 100 percent coverage is the scientific core of this proposal. Structure-limited tools fail exactly when the mutation is in a region that has not been crystallized, or when the protein itself has no solved structure. In drug resistance surveillance, those are common cases. Pathogens mutate continuously, and the resistance mutations that matter clinically are often the ones that appear in proteins with no high-resolution structure available. A sequence-only method removes that dependency entirely.
The next phase of the research is fine-tuning ESM-2 on the SKEMPI 3K mutation set, which contains roughly 3,000 experimentally measured mutation effects. The target is an AUROC of at least 0.70 on independent held-out data, which would represent a substantial improvement over the current 0.634 on SKEMPI 2.0 and would close the gap to structure-based methods while retaining the coverage advantage. The fine-tuning will be followed by a pilot with Servier in Suresnes, where the model will be tested on a set of clinically relevant resistance mutations supplied by Servier's oncology and infectious disease teams. The pilot's success metric is the model's ability to rank known resistant variants above known susceptible variants in a blinded evaluation.
The research is not purely computational. The deployment study in Nigeria will collect de-identified clinical isolates from collaborating hospitals, sequence the relevant target proteins, and compare the model's resistance predictions against phenotypic susceptibility testing. This is the evidence that the Foundation's selection criteria demand: measurable impact on patient access and health outcomes, not just benchmark scores. The study will also generate the first dataset of its kind for Nigerian clinical isolates, which is itself a contribution to the global resistance surveillance literature.
The innovation claim rests on the delta-embedding representation. Standard protein language model embeddings capture the wild-type protein's properties but do not directly encode the effect of a mutation. Delta-embeddings, computed as the difference between the mutated and wild-type ESM-2 representations, isolate the mutation's effect on the protein's learned biology. Combined with the drug fingerprint, the model learns a joint representation of the mutation and the drug, which is what allows it to predict resistance to a specific compound rather than a generic resistance phenotype. This is a fundamentally different approach from sequence alignment-based methods, which require known resistance homologs, and from structure-based methods, which require a crystal structure.
The sustainability of the approach is tied to the falling cost of sequencing and the increasing availability of protein language models. ESM-2 is open source. The compute required for inference is modest, a single GPU is sufficient for most use cases. The model can be deployed as a web service or a local tool, which means a hospital laboratory in a low-resource setting can run it without a bioinformatics team. The fine-tuned model will be made available under a license that permits non-commercial use in LMICs, consistent with the Foundation's mission.
The research roadmap is explicit. Fine-tune ESM-2 on SKEMPI 3K, target AUROC of 0.70 or better on held-out data. Validate on the Servier pilot. Deploy in Nigeria and measure real-world predictive accuracy against phenotypic testing. Publish the results, including the negative cases, so the field can see where the method fails as well as where it succeeds. The grant request of up to one million dollars funds compute for the fine-tuning, personnel for the pilot and deployment study, and the sequencing costs for the Nigerian clinical isolate collection.
ESSAY RESPONSE: PATIENT ACCESS AND HEALTH OUTCOMES
The measurable impact of this project is the number of treatment failures avoided per 100 patients treated for a resistant infection. In Nigeria, where antimicrobial resistance rates for common pathogens such as Escherichia coli and Klebsiella pneumoniae exceed 50 percent for several first-line antibiotics, the current standard of care is empiric prescribing. A clinician chooses a drug without knowing whether the infecting strain is resistant. When the drug fails, the patient returns, sicker, having consumed a course of an ineffective antibiotic and potentially having spread the resistant strain to others. The model changes this by allowing a resistance prediction to be generated from a protein sequence obtained from a rapid diagnostic test. The prediction is not a substitute for phenotypic susceptibility testing, but it is faster and cheaper, and it can be done in settings where phenotypic testing is unavailable.
The target health outcome is a reduction in the time from diagnosis to effective therapy. The current delay in Nigeria is typically two to five days for a phenotypic result, when the test is available at all. The model's prediction can be delivered in under an hour from sequence input. If the model's AUROC of 0.804 on Platinum translates to clinical performance, the number of patients receiving an effective first-line drug should increase measurably. The deployment study will measure this directly: time to effective therapy, treatment success rate, and days of hospitalization, compared against a retrospective cohort treated under the current standard of care.
The Foundation's emphasis on access to medicines is served because the model does not require a new drug to be developed or a new supply chain to be built. It makes existing drugs work better by ensuring they are prescribed to patients whose infections they can actually treat. This is access in the operational sense: the right drug, for the right patient, at the right time, with the resources already in the country.
ESSAY RESPONSE: FEASIBILITY AND SCALABILITY
The proof of concept is complete. The model exists, is validated on two public benchmarks, and outperforms the published state of the art on the primary benchmark while covering a larger fraction of mutations. The next phase requires compute, data, and clinical partnerships, all of which are identified. The compute is available through cloud GPU providers. The data is the SKEMPI 3K set, which is public. The clinical partnership is the Servier pilot in Suresnes, with discussions in progress, and the Nigerian deployment study, which will require ethics approval and hospital collaboration agreements.
Scalability has two dimensions. The first is geographic: the model is sequence-only, so it works for any pathogen with a sequenced target protein, regardless of the country's structural biology infrastructure. The second is therapeutic: the model is trained on mutation effects and drug fingerprints, so it can be extended to new drugs by adding their ECFP4 fingerprints and to new proteins by fine-tuning on relevant mutation data. The same architecture that predicts resistance for an oncology target can be retrained for an antimalarial target. The Foundation's priority diseases, including malaria and dengue, are directly addressable with this approach.
The risk that the model's performance does not generalize from benchmarks to clinical isolates is real and is explicitly addressed by the Nigerian deployment study. The risk that the fine-tuned model does not reach the 0.70 AUROC target on SKEMPI 2.0 is mitigated by the current 0.634 baseline and the known benefit of fine-tuning on in-domain data. The risk that clinical adoption is slow is mitigated by the low cost of deployment, a web interface or local tool that requires no specialized hardware.
ESSAY RESPONSE: TEAM AND PARTNERSHIPS
I am the sole author of the model and the sole founder of the venture. My background is pharmacy, which gives me clinical literacy in drug resistance and treatment failure, and machine learning engineering, which gives me the technical capacity to build and validate the model. The combination is rare and directly relevant to this project. I have named partnership discussions in progress with Servier in Suresnes for the pilot, with Paris-Saclay's I2BC and Institut Pasteur for structural biology and clinical validation expertise, and with Sanofi in Gentilly for therapeutic area guidance. The venture is also in the support pipeline for SEMIA and Quest for Health, with WILCO One BioTech scheduled for October 2026, and applications submitted to IncubAlliance and AI House. The EIC Accelerator and BPI i-Lab are future targets.
The gap in the team is a clinical microbiologist in Nigeria to lead the deployment study and a bioinformatics engineer to maintain the deployment infrastructure. The grant will fund both positions. The Foundation's selection criteria emphasize team strength, and the honest assessment is that the current team is a single founder with strong technical and clinical credentials but no dedicated operational staff. The grant is the mechanism by which the team is built.
CHECKLIST
- [ ] Confirm the Novartis Access to Medicine Foundation's current grant application portal and rolling deadline status at the provided URL
- [ ] Verify that the Foundation funds computational biology projects, not only clinical or supply chain projects, before submitting
- [ ] Prepare a one-page project budget breaking down the requested amount into compute, personnel, sequencing, and pilot costs
- [ ] Obtain a letter of support or expression of interest from Servier for the pilot in Suresnes
- [ ] Obtain a letter of support from a Nigerian hospital or research institution for the deployment study
- [ ] Prepare a data management and ethics approval plan for the Nigerian clinical isolate collection
- [ ] Draft a two-page technical appendix describing the ESM-2 delta-embedding architecture and the Platinum and SKEMPI validation results
- [ ] Confirm the legal status of the venture for grant contracting, noting that the venture is not yet incorporated
- [ ] Prepare a timeline with milestones for the SKEMPI fine-tuning, Servier pilot, and Nigeria deployment study
- [ ] Submit the application through the Foundation's portal with all required attachments
EDITOR NOTES
- Eligibility risk: the Foundation's mission is access to medicines in LMICs, and the venture is a computational biology tool, not a drug developer or distributor. The framing must stay on patient access and health outcomes, not on benchmark performance, or the application will be rejected as off-mission.
- The venture is pre-seed and not yet incorporated. The Foundation may require a legal entity to receive funds. Confirm whether an individual can receive the grant or whether incorporation is required before disbursement.
- The Servier pilot is described as in progress, not confirmed. The application must not claim a signed partnership. Verify the current status of the Servier discussion before submission and adjust the language if the pilot is not yet agreed.
- The Nigerian deployment study requires ethics approval and hospital collaboration agreements that do not yet exist. The timeline must include a realistic lead time for these approvals, which can take six to twelve months in Nigeria.
- The model's SKEMPI 2.0 AUROC of 0.634 is below the 0.70 target for the fine-tuned model. The application must present this honestly as a baseline to be improved, not as a current strength.