MOTIVATION LETTER
The Grand Challenges programme funds bold ideas that can scale in low-resource settings. My research addresses a specific failure point in that mission: antimicrobial resistance (AMR) prediction tools require protein structures that do not exist for most clinically relevant mutations, and African healthcare systems rarely have the computational infrastructure to generate them. My TOPOLOGIX project solves this by predicting drug-resistance mutations from amino acid sequence alone, using protein language model embeddings and chemical fingerprints. The system achieves an AUROC of 0.804 plus or minus 0.025 on the Platinum benchmark of 553 mutations, covering 100 percent of mutations tested. Structure-based tools like mCSM-lig reach roughly 0.70 AUROC but cover only about 18 percent of mutations because they require crystallographic structures. In African settings, where sequencing is increasingly available but structural biology is not, sequence-only prediction is the difference between a tool that works and a tool that cannot run.
The project began with a falsified hypothesis. I tested whether bipartite persistent homology of protein-ligand interfaces could predict hERG cardiotoxicity. A pre-registered, powered replication showed topological features do not beat a plain descriptor baseline: AUROC 0.8426 versus 0.8782. I then applied the same topological constructs to drug-resistance prediction and found they carried almost no signal, AUROC 0.425 and 0.485 on the Platinum benchmark. Those negative results ruled out interface geometry as the driver and motivated the sequence-representation approach now used in TOPOLOGIX. The current system is a Random Forest classifier trained on ESM-2 delta-embeddings and Morgan/ECFP fingerprints. It is computationally light enough to run on a standard laptop, which matters for deployment in settings without HPC access.
The Grand Challenges criteria emphasize cost reduction, scalability, and measurable health outcomes. TOPOLOGIX addresses all three. The tool requires no structural data, no specialized hardware, and no proprietary software. It can be retrained on local genomic surveillance data as it accumulates. The methodology is fully open and reproducible, with the codebase on GitHub and the validation pipeline pre-registered. My background supports the translation path: I am a licensed pharmacist with a B.Pharm from the University of Ibadan, I have built AMR surveillance pipelines with the Genomic Health Research Unit at GHRU-GSAR, and I am enrolled in the M.Sc. Digital Health programme at the Hasso Plattner Institute in Potsdam starting Winter Semester 2026/27. I am currently a National Product Manager at Synthcare, where I manage pharmaceutical product strategy across Nigerian markets.
The next phase is a prospective validation study using Nigerian clinical isolates, where genomic data exists but structural data does not. This requires sequencing costs, sample collection logistics, and a modest compute budget. A Grand Challenges grant would fund exactly that validation, producing the evidence base needed for adoption by national AMR surveillance programs. The work fits the programme's mandate for diagnostic and screening innovation in low-resource settings, and it is ready to move from benchmark to field.
RESEARCH STATEMENT
TOPOLOGIX is a sequence-based machine learning system for predicting drug-resistance mutations. It was developed in response to a specific gap: existing resistance prediction tools depend on protein structures that are unavailable for most mutations of clinical interest. The Platinum benchmark, a standard dataset of 553 resistance-associated mutations, shows that structure-based tools like mCSM-lig achieve roughly 0.70 AUROC but can only score about 18 percent of mutations because the remaining 82 percent lack crystallographic structures. TOPOLOGIX uses ESM-2 protein language model delta-embeddings combined with Morgan/ECFP drug fingerprints and a Random Forest classifier to predict resistance from sequence alone. It achieves 0.804 AUROC plus or minus 0.025 on Platinum and 0.634 on SKEMPI 2.0, covering 100 percent of mutations in both benchmarks.
The design rationale comes from two negative results I published as preprints. First, a pre-registered replication study of bipartite persistent homology for hERG cardiotoxicity prediction found that topological features of protein-ligand interfaces do not outperform a plain descriptor baseline: AUROC 0.8426 versus 0.8782. Second, applying the same interface-topology constructs to drug-resistance prediction on the Platinum benchmark produced AUROC values of 0.425 and 0.485, indicating no signal. These results ruled out interface geometry as the primary driver of resistance and motivated a shift to sequence representations, which capture evolutionary and biophysical information that structure-based methods discard when structures are missing.
The methodology is deliberately simple. ESM-2 embeddings are computed for wild-type and mutant sequences, and the delta between them serves as the mutation representation. Morgan fingerprints encode the drug. A Random Forest classifier learns the mapping from these features to resistance labels. The system requires no structural biology, no molecular dynamics, and no specialized hardware. Training and inference run on a standard laptop. This is a deliberate design choice: the tool must work in settings where structural biology infrastructure does not exist.
Validation followed a pre-registered protocol. The Platinum benchmark was used for primary evaluation with cross-validation. SKEMPI 2.0 served as an external generalization check. Results were compared against published structure-based baselines on the same benchmarks. The system outperforms those baselines on coverage while matching or exceeding them on discrimination. All code, data splits, and pre-registration documents are available on GitHub and OSF.
The next phase is prospective validation on Nigerian clinical isolates. Nigeria's AMR surveillance programs generate genomic sequencing data, but structural data for resistance-associated proteins is almost entirely absent. TOPOLOGIX can be applied directly to this data, and its predictions can be compared against phenotypic resistance testing. This validation would produce the first sequence-only resistance prediction benchmark for West African isolates and would test whether models trained on global data transfer to local populations.
The broader research program includes the CCT model for addiction neuroscience, a tripartite pharmacological framework for reward-memory encoding prevention, and neurocascade, a receptor-to-behavior brain-circuit simulation engine. These projects share a methodological commitment to pre-registration, Bayesian calibration, and honest reporting of negative results. The ergofluids project, which tests Koopman operator methods for drug transport in tumor tissue, passed synthetic-data validation gates but did not meet its primary pre-registered criterion on real data; that result was reported directly rather than reframed. This track record of transparent reporting is the foundation for the TOPOLOGIX proposal.
The specific aims for the Grand Challenges funding period are: first, retrain TOPOLOGIX on a combined dataset of Platinum, SKEMPI 2.0, and newly curated Nigerian isolate data; second, validate predictions against phenotypic resistance testing for a panel of clinically relevant antibiotics; third, package the system as a deployable tool with a simple command-line interface and documentation for non-specialist users; fourth, publish the validation results and release the trained model weights under an open license. The measurable outcome is a validated, sequence-only resistance prediction tool with demonstrated performance on African clinical isolates, ready for integration into national AMR surveillance pipelines.
ESSAY: SCALABILITY AND SUSTAINABILITY IN LOW-RESOURCE SETTINGS
TOPOLOGIX is designed for environments where structural biology infrastructure does not exist. The system requires only amino acid sequence data, which is increasingly available from genomic surveillance programs across Africa. Nigeria's AMR surveillance network, which I have worked with through GHRU-GSAR, generates sequencing data as part of routine monitoring. The computational requirements are minimal: ESM-2 inference and Random Forest classification run on a standard laptop with no GPU. This contrasts sharply with structure-based tools that require crystallographic data, molecular dynamics simulations, and HPC access.
Sustainability is built into the architecture. The model can be retrained on local data as it accumulates, allowing continuous improvement without external dependency. The codebase is open source, the validation pipeline is pre-registered, and the trained model weights will be released under an open license. This means national programs can adopt the tool, adapt it to local resistance patterns, and maintain it with local expertise. The training pipeline is documented so that a bioinformatician with standard Python skills can reproduce the full workflow.
Cost reduction is substantial. Structure-based resistance prediction requires either experimental structure determination or expensive homology modeling pipelines. TOPOLOGIX eliminates both. The marginal cost per prediction is effectively zero once the model is trained. For a country like Nigeria, where AMR surveillance budgets are constrained, this is the difference between a tool that can be deployed and a tool that remains a research artifact.
The scalability path is clear. The same sequence-only architecture applies to any pathogen with genomic surveillance data. The initial validation focuses on bacterial resistance, but the methodology transfers to viral resistance, including HIV and hepatitis, where sequence data is the primary data type available in African settings. The Grand Challenges emphasis on scalable, cost-effective health solutions matches this design exactly.
ESSAY: MEASURABLE HEALTH OUTCOMES AND CLEAR METRICS
The primary outcome is prediction accuracy for drug-resistance mutations, measured as AUROC on held-out benchmark data. The current system achieves 0.804 plus or minus 0.025 on the Platinum benchmark and 0.634 on SKEMPI 2.0. The target for the Grand Challenges funding period is a prospective validation on Nigerian clinical isolates with phenotypic resistance testing as the ground truth. The success criterion is an AUROC of at least 0.75 on this external validation set, which would demonstrate that the model transfers from global benchmarks to local populations.
Secondary outcomes include coverage, defined as the proportion of mutations that can be scored. TOPOLOGIX covers 100 percent of mutations in both benchmarks, compared to roughly 18 percent for structure-based tools. This coverage metric is directly relevant to clinical utility: a tool that cannot score a mutation is useless for that patient, regardless of its accuracy on the mutations it can handle.
The validation protocol is pre-registered. Sample size calculations are based on the expected effect size from the benchmark results. Phenotypic testing will follow Clinical and Laboratory Standards Institute guidelines. The analysis plan specifies the exact statistical tests, the handling of missing data, and the criteria for declaring the validation successful or unsuccessful. If the validation fails, that result will be reported directly, consistent with my track record on the hERG topology replication and the ergofluids real-data gate.
The downstream health outcome is improved antibiotic stewardship. When resistance is predicted accurately from sequence data, clinicians can choose effective antibiotics earlier, reducing treatment failure and slowing the spread of resistance. The measurable proxy for this outcome is the proportion of resistance predictions that match phenotypic testing results, which the validation study will quantify directly.
CHECKLIST
- [ ] Confirm eligibility for Grand Challenges 2026 as an independent researcher based in Nigeria
- [ ] Verify the specific call themes for 2026 and confirm TOPOLOGIX aligns with the diagnostic and screening innovation track
- [ ] Check whether the programme requires institutional affiliation or allows independent applicants
- [ ] Prepare a detailed budget for the prospective validation study, including sequencing costs, sample collection, and compute
- [ ] Obtain a letter of support from a Nigerian AMR surveillance program or collaborating institution
- [ ] Update the TOPOLOGIX GitHub repository with the latest benchmark results and pre-registration documents
- [ ] Prepare the application form responses based on the essays drafted above
- [ ] Submit the application before the programme deadline
- [ ] Confirm whether the programme requires a CV or publication list, and prepare both
- [ ] Verify the ORCID record is current and includes all preprints and publications
EDITOR NOTES
- The chosen research line is TOPOLOGIX, not the CCT addiction model or neurocascade, because the Grand Challenges programme explicitly funds health innovation, diagnostics, and AI for social impact in low-resource settings. TOPOLOGIX is the only active research line that directly matches those criteria. The CCT model is strong work but does not fit a global health funding call.
- The negative results from the hERG topology study and the Platinum resistance topology study are presented as motivation for TOPOLOGIX, not as current claims. This is accurate and consistent with the profile. Do not let anyone rewrite these as positive results.
- The prospective validation on Nigerian clinical isolates is a future plan, not an existing dataset. The profile does not indicate that any Nigerian isolate data has been collected or analyzed yet. The application must not imply that this validation has begun.
- The M.Sc. enrollment at Hasso Plattner Institute is for Winter Semester 2026/27, which may overlap with the funding period. Confirm that the programme allows concurrent enrollment and that the applicant can commit to the proposed timeline.
- The employment status as National Product Manager at Synthcare may raise questions about time commitment. The application should clarify how the research work is structured around this role, or whether the grant would fund a transition to full-time research.