MOTIVATION LETTER
The Challenge Programme of the Novo Nordisk Foundation funds research that confronts major societal challenges in human health. Antimicrobial resistance is one such challenge, responsible for an estimated 1.27 million direct deaths in 2019, a figure the WHO projects will rise to 10 million annually by 2050 without intervention. My research addresses a specific bottleneck in this crisis: the prediction of drug resistance mutations before they emerge clinically. My TOPOLOGIX system predicts resistance mutations from protein sequence alone, achieving an AUROC of 0.804 plus or minus 0.025 on the Platinum benchmark of 553 mutations, and 0.634 on SKEMPI 2.0. This performance exceeds structure-based tools such as mCSM-lig, which scores approximately 0.70, while covering 100 percent of mutations compared to roughly 18 percent for structure-limited methods.
The scientific novelty lies in the methodological approach. TOPOLOGIX uses ESM-2 protein language model delta-embeddings combined with Morgan/ECFP drug fingerprints fed into a Random Forest classifier. This data-centric strategy bypasses the need for resolved protein structures, which are unavailable for the majority of clinically relevant targets. My prior work established the necessity of this approach through negative results. A pre-registered, powered replication on hERG cardiotoxicity showed that bipartite persistent homology features did not beat a plain descriptor baseline, AUROC 0.8426 versus 0.8782. A subsequent study on drug resistance prediction using interface topology produced AUROC values of 0.425 and 0.485 on the Platinum benchmark, ruling out interface geometry as the driver of resistance signal. These findings, reported directly rather than reframed, motivated the sequence-representation approach that TOPOLOGIX now implements.
The Foundation's emphasis on scientific excellence and societal impact aligns with my track record. I am a licensed pharmacist with a B.Pharm from the University of Ibadan, CGPA 5.1 out of 7.0, German equivalent 1.9, and I am enrolled in the M.Sc. Digital Health programme at the Hasso Plattner Institute and University of Potsdam for Winter Semester 2026/27. My independent research spans addiction neuroscience, where my Conjunctive Consolidation Threshold model, a tripartite pharmacological framework with 14 free parameters calibrated via Bayesian MCMC against a literature screen of 1,847 records, confirmed all five pre-registered hypotheses with posterior super-additivity of 13 to 22 percentage points. Three sole-authored preprints are under review at peer-reviewed journals. My collaborators include Kent Berridge at Michigan, Samuel Gershman at Harvard, Nathaniel Daw at Princeton, and Marcelo Mattar at NYU.
The feasibility of the proposed research plan rests on demonstrated technical execution. I have built and validated the neurocascade brain-circuit simulation engine with 62 of 62 tests passing, and the ergofluids Koopman-operator framework with a pre-registered gated validation pipeline. My infrastructure includes four independent DuckDB-based ingest-to-analyte pipelines and self-hosted local LLM serving. The TOPOLOGIX system is operational and benchmarked. What the Challenge Programme would enable is the expansion of this system to larger clinical datasets, prospective validation against emerging resistance mutations, and the development of a deployable tool for clinical microbiology laboratories in low-resource settings, including Nigeria, where resistance surveillance is sparse. This is a foundational contribution to precision medicine and pandemic preparedness, and I request the Foundation's support to advance it.
RESEARCH STATEMENT
The global burden of antimicrobial resistance is not evenly distributed. Sub-Saharan Africa carries a disproportionate share, with resistance rates to common antibiotics exceeding 50 percent in several countries, yet genomic surveillance infrastructure remains minimal. My research programme addresses this disparity through computational methods that do not depend on infrastructure-intensive laboratory workflows. The central question is whether drug resistance mutations can be predicted from protein sequence alone, without resolved structures, and whether such predictions can be made actionable for clinical and pharmaceutical decision-making.
TOPOLOGIX is the current embodiment of this research line. The system combines ESM-2 protein language model delta-embeddings, which capture evolutionary and functional constraints in sequence space, with Morgan/ECFP drug fingerprints representing the chemical features of the therapeutic agent. A Random Forest classifier integrates these inputs to predict whether a given mutation confers resistance. On the Platinum benchmark of 553 mutations, the system achieves an AUROC of 0.804 with a standard deviation of 0.025. On SKEMPI 2.0, the AUROC is 0.634. The critical advantage is coverage: structure-based tools require a resolved protein-ligand complex, which exists for only a fraction of clinically relevant targets. TOPOLOGIX covers 100 percent of mutations in the benchmark, compared to approximately 18 percent for structure-limited tools.
The methodological trajectory that led to TOPOLOGIX is as important as the system itself. My first study in this area tested whether bipartite persistent homology, an opposition-distance metric computed with Ripser and GUDHI, could predict hERG cardiotoxicity from protein-ligand interface geometry. This was a pre-registered, powered replication of claims that had never been directly tested. The result was negative: topological features achieved an AUROC of 0.8426, while a plain descriptor baseline achieved 0.8782. The second study applied the same topological constructs to drug resistance prediction, producing AUROC values of 0.425 and 0.485 on the Platinum benchmark. These results ruled out interface geometry as the driver of resistance signal. Both negative results were reported directly, without reframing, because they are scientifically informative. They redirected the research towards sequence-based representations, which proved superior.
The proposed next phase has three components. First, expansion of the training and validation data beyond the Platinum and SKEMPI benchmarks to include clinical resistance datasets from public repositories such as CARD and ResFinder. Second, prospective validation: the system will be used to predict resistance mutations in newly sequenced clinical isolates, with predictions compared against phenotypic susceptibility testing. Third, deployment engineering: the model will be packaged as a lightweight, offline-capable tool that can run on standard laboratory hardware, with a web interface for non-specialist users. This deployment target is informed by my experience building production systems, including Linux VPS operations, systemd services, Caddy TLS, CI/CD pipelines, and automated backup and disaster-recovery procedures.
The broader research context includes my work in computational pharmacology and dynamical systems. The Conjunctive Consolidation Threshold model, a tripartite framework for reward-memory encoding prevention in addiction, couples dopaminergic reward prediction error, NMDAR-dependent long-term potentiation, and affective contrast in a system of ODEs solved with RK45. Bayesian calibration with PyMC DEMetropolisZ against literature-elicited priors from a screen of 1,847 records confirmed all five pre-registered hypotheses, with posterior super-additivity of 13 to 22 percentage points across model versions. The neurocascade engine couples pharmacokinetics to receptor binding to Wilson-Cowan circuit dynamics to behavioral readouts, with 62 of 62 tests passing. The ergofluids project extends Koopman operator methods with a Mori-Zwanzig memory kernel for macromolecular transport in tumor tissue, with a pre-registered gated validation pipeline that reported a failed real-data gate directly. These projects share a methodological commitment: quantitative models with explicit uncertainty, pre-registered hypotheses, and honest reporting of negative results.
The Novo Nordisk Foundation Challenge Programme supports research that addresses major societal challenges in human health. Antimicrobial resistance is a named priority in global health policy, and the Foundation's mission to support transformative research with clear societal impact is directly served by a tool that makes resistance prediction accessible to laboratories without structural biology capacity. The interdisciplinary nature of this work, combining protein language models, cheminformatics, and machine learning, aligns with the Foundation's emphasis on collaborative and cross-cutting research. My qualifications include a B.Pharm from the University of Ibadan, PCN licensure, enrollment in the M.Sc. Digital Health programme at Hasso Plattner Institute and University of Potsdam, and endorsements from Kent Berridge, Samuel Gershman, Nathaniel Daw, and Marcelo Mattar. The research plan is feasible within a 24-month timeframe, with clear milestones and deliverables at each stage.
SHORT ESSAY: SOCIETAL CHALLENGE
Antimicrobial resistance is a societal challenge because it erodes the foundation of modern medicine. Routine procedures, from cesarean sections to chemotherapy, depend on effective antibiotics. When resistance outpaces drug development, these procedures become life-threatening. The WHO estimates 10 million annual deaths from AMR by 2050, exceeding current cancer mortality. The burden falls hardest on low and middle-income countries, where surveillance infrastructure is weakest and resistance rates are highest. In Nigeria, where I am from, clinical microbiology laboratories capable of resistance testing are concentrated in a few urban centers, leaving most of the population without diagnostic access.
My research addresses this challenge by making resistance prediction computationally accessible. TOPOLOGIX requires only a protein sequence and a drug fingerprint, both of which can be obtained from widely available genomic sequencing and chemical databases. It does not require a resolved protein structure, which is the limiting factor for existing tools. This means that a laboratory in Lagos or Nairobi, equipped with a standard computer and an internet connection, can generate resistance predictions for clinical isolates without specialized structural biology expertise. The system's coverage of 100 percent of benchmark mutations, compared to 18 percent for structure-limited tools, directly addresses the equity gap in resistance surveillance.
The societal impact extends beyond diagnostics. Pharmaceutical companies developing new antibiotics need to understand resistance mechanisms early in the drug development pipeline. A tool that predicts resistance mutations from sequence alone can inform lead optimization, reducing the likelihood of developing drugs that will rapidly lose efficacy. This is a contribution to pandemic preparedness, as the same computational framework can be adapted to predict resistance for antiviral and antifungal agents. The Foundation's commitment to human health is served by research that strengthens both clinical practice and drug development against a global threat.
SHORT ESSAY: INTERDISCIPLINARY APPROACH
The TOPOLOGIX project is interdisciplinary by construction, integrating three distinct methodological traditions. Protein language models, specifically ESM-2, represent the application of deep learning to evolutionary biology. These models are trained on millions of protein sequences and learn representations that capture functional constraints without explicit structural input. Drug fingerprints, specifically Morgan/ECFP circular fingerprints, represent the cheminformatics tradition of encoding molecular structure as bit vectors for machine learning. The Random Forest classifier represents the statistical learning tradition, providing interpretable feature importance and strong performance on tabular data.
My training spans these domains. As a pharmacist, I understand the chemical and pharmacological context of drug resistance. As a computational modeler, I have built and calibrated ODE systems with Bayesian inference, including the CCT model with 14 free parameters and the neurocascade engine with 62 passing tests. As a software engineer, I have deployed production systems, including DuckDB-based data pipelines and self-hosted LLM serving. This combination allows me to make principled decisions about feature engineering, model selection, and validation that a specialist in any single domain might miss.
The negative results from my prior topological studies are an example of interdisciplinary rigor. The hERG cardiotoxicity study and the resistance prediction study both tested hypotheses that were plausible given the literature, and both produced null results. Reporting these results directly, rather than reframing them as positive findings, is a commitment to scientific integrity that the Foundation's emphasis on excellence should value. The interdisciplinary approach also extends to collaboration. My research has been endorsed by Kent Berridge in addiction neuroscience, Samuel Gershman in computational cognitive science, Nathaniel Daw in reinforcement learning, and Marcelo Mattar in memory and decision-making. These connections provide access to diverse expertise and validation of my methods.
SHORT ESSAY: FEASIBILITY AND METHODS
The research plan is feasible because the core system is already built and benchmarked. TOPOLOGIX achieves an AUROC of 0.804 plus or minus 0.025 on the Platinum benchmark and 0.634 on SKEMPI 2.0. The computational infrastructure is operational, including the ESM-2 embedding pipeline, the RDKit fingerprint generation, and the Random Forest training and evaluation code. The next phase requires data acquisition, model retraining, and deployment engineering, all of which are within my demonstrated skill set.
Data acquisition will target public repositories including CARD, ResFinder, and the NCBI BioProject database for clinical resistance datasets. These datasets include phenotypic susceptibility testing results paired with genomic sequences, enabling supervised training. The model will be retrained on this expanded corpus, with cross-validation to assess generalization. Prospective validation will involve collaboration with a clinical microbiology laboratory, which I will identify through my network at the University of Ibadan and the Ghanaian GHRU-GSAR programme where I previously worked on AMR genomics and surveillance pipelines.
Deployment engineering will package the model as a lightweight Python application with a web interface, deployable on a standard Linux server or a laptop. The system will include an offline mode for laboratories without reliable internet, using a pre-computed embedding database. The timeline is 24 months: months 1 to 6 for data acquisition and curation, months 7 to 12 for model retraining and validation, months 13 to 18 for prospective validation, and months 19 to 24 for deployment and documentation. Each phase has clear deliverables and success criteria, consistent with my practice of pre-registration and gated validation as demonstrated in the ergofluids project.
CHECKLIST
- [ ] Verify the Novo Nordisk Foundation Challenge Programme application portal and current deadline on the official Foundation website
- [ ] Confirm eligibility for independent researchers without a university faculty appointment
- [ ] Confirm whether the programme requires a host institution or allows direct application
- [ ] Prepare the full application form with personal details, ORCID 0009-0001-9272-6735, and affiliation as independent researcher
- [ ] Upload the motivation letter as prepared above
- [ ] Upload the research statement as prepared above
- [ ] Upload the short essay on societal challenge as prepared above
- [ ] Upload the short essay on interdisciplinary approach as prepared above
- [ ] Upload the short essay on feasibility and methods as prepared above
- [ ] Prepare a CV in the Foundation's required format, including B.Pharm, M.Sc. enrollment, publications, and technical skills
- [ ] Obtain and upload letters of support from Kent Berridge, Samuel Gershman, Nathaniel Daw, or Marcelo Mattar
- [ ] Prepare a budget justification for the requested amount, including personnel, computing resources, and data acquisition costs
- [ ] Verify the exact word and character limits for each essay on the application portal and adjust accordingly
- [ ] Confirm whether the programme requires a project timeline or Gantt chart and prepare if needed
- [ ] Check whether the programme requires a data management plan and prepare if needed
- [ ] Confirm whether the programme requires ethical approval documentation for clinical data use and prepare if needed
- [ ] Submit the application before the deadline and save the confirmation receipt
EDITOR NOTES
- Eligibility risk: The Novo Nordisk Foundation Challenge Programme may require affiliation with a Danish or Nordic research institution. The applicant is an independent researcher and enrolled in a German M.Sc. programme. Verify the eligibility rules on the official Foundation website before investing further time. The URL provided in the programme listing points to a FU Berlin page, not the Foundation itself, which is a red flag for accuracy.
- Amount and deadline are not specified in the programme listing. The applicant must research the current Challenge Programme call on the Novo Nordisk Foundation website and confirm the funding range and submission date. The strategy notes mention a pre-seed AI/biotech startup framing, but the applicant is an independent researcher, not a startup founder. Clarify which track is being applied for.
- The motivation letter references the Foundation's mission and selection criteria, but the applicant should verify the exact language of the current call for proposals and align terminology accordingly. The Foundation may have specific priority areas for the current cycle that should be referenced explicitly.
- The research statement mentions prospective validation with a clinical microbiology laboratory, but no specific partner is named. The applicant must identify and, ideally, secure a letter of intent from a collaborating laboratory before submission.
- The applicant's employment as National Product Manager at Synthcare from March 2026 is noted in the profile but not referenced in the application materials. If this role is relevant to the research or creates a conflict of interest, it should be addressed. The M.Sc. enrollment at HPI/Potsdam for Winter Semester 2026/27 should be verified for exact dates and status.