MOTIVATION LETTER
The Platinum benchmark contains 553 drug-resistance mutations. Structure-based tools like mCSM-lig can only score about 18 percent of them because most lack a resolved crystal structure. My TOPOLOGIX system, which uses ESM-2 protein language model delta-embeddings combined with Morgan fingerprints and a Random Forest classifier, covers all 553 mutations and reaches an AUROC of 0.804 plus or minus 0.025. That result is the reason I am applying to the AIMS South Africa DeepMind AI for Science Master's Scholarship.
I am a Nigerian pharmacist and independent computational researcher. My work sits at the intersection of protein machine learning, dynamical systems, and neuroscience. The AIMS programme's mission, to apply AI to fundamental scientific problems, matches the core of my research practice. I build models to answer specific biological questions, and I test them against pre-registered criteria. When a method fails, I report the failure directly. My cardiotoxicity topology study, for example, tested whether bipartite persistent homology could predict hERG cardiotoxicity from protein-ligand interface geometry. The pre-registered, powered replication found that topological features did not beat a plain descriptor baseline, AUROC 0.8426 versus 0.8782. That negative result settled a comparison the literature had never actually run.
The AIMS selection criteria ask for a degree in a discipline with a strong computational or mathematical component. My B.Pharm from the University of Ibadan, with a German equivalent grade of 1.9, included computational chemistry and ADMET modeling. My independent research since then has required Bayesian MCMC calibration with PyMC, ODE solving with RK45, persistent homology with Ripser and GUDHI, and protein language model embeddings. I have also built neurocascade, a receptor-to-behavior brain-circuit simulation engine with coupled pharmacokinetic, receptor-binding, and Wilson-Cowan circuit layers, with 62 of 62 tests passing. This is mathematical modeling applied to biological systems, which is precisely the profile AIMS seeks.
The DeepMind scholarship would allow me to formalize and deepen the mathematical foundations of my methods. My TOPOLOGIX project currently uses a Random Forest on top of ESM-2 embeddings. I want to understand why sequence-based representations generalize where structure-based ones fail, and I want to develop the theory of when topological descriptors carry signal and when they do not. AIMS, with its focus on mathematical sciences and its African research community, is the right environment for that work. I am already enrolled in an M.Sc. in Digital Health at the Hasso Plattner Institute in Potsdam, but the AIMS programme offers something that programme does not: a dedicated AI-for-Science curriculum and a network of African mathematical scientists.
I am a citizen of Nigeria and a resident of Africa. I have not held a previous scholarship to study at an AIMS centre. I meet the degree requirements. I am ready to commit to the full programme.
RESEARCH STATEMENT
My research program asks a single question: when do computational representations of biological molecules carry real predictive signal, and when do they fail? I pursue this question across protein machine learning, addiction neuroscience, and dynamical systems. The flagship project, TOPOLOGIX, is the one most directly aligned with the AIMS DeepMind mission.
TOPOLOGIX predicts drug-resistance mutations from sequence alone. The problem is concrete. Resistance mutations in proteins like kinases and HIV proteases render therapies ineffective, and predicting which mutations will confer resistance is essential for drug design. The standard approach relies on protein structures, but most mutations occur in proteins without resolved structures. Structure-based tools like mCSM-lig cover only about 18 percent of the Platinum benchmark's 553 mutations. TOPOLOGIX uses ESM-2 protein language model delta-embeddings, which capture the evolutionary and biophysical context of a mutation from sequence, combined with Morgan/ECFP drug fingerprints and a Random Forest classifier. It achieves an AUROC of 0.804 plus or minus 0.025 on the Platinum benchmark and 0.634 on SKEMPI 2.0, while covering 100 percent of mutations. This beats structure-based baselines by a wide margin on coverage and by a meaningful margin on discrimination.
The project emerged from a falsified hypothesis. My earlier work tested whether bipartite persistent homology, an opposition-distance metric computed with Ripser and GUDHI, could predict hERG cardiotoxicity from protein-ligand interface geometry. The pre-registered, powered replication found it could not, AUROC 0.8426 versus 0.8782 for a plain descriptor baseline. I then applied the same topological constructs to drug-resistance prediction and found they carried almost no signal, AUROC 0.425 and 0.485 on the Platinum benchmark. These negative results ruled out interface geometry as the driver of resistance and motivated the sequence-representation approach that became TOPOLOGIX. I report these failures because they are the evidence that matters. A method that cannot beat a baseline is a hypothesis that has been tested.
My other research lines deepen the same theme. The CCT model, a tripartite pharmacological framework for reward-memory encoding prevention in addiction, couples dopaminergic reward prediction error, NMDAR-dependent long-term potentiation, and affective contrast in a system of ODEs solved with RK45. I calibrated it with Bayesian MCMC using PyMC's DEMetropolisZ sampler, with 14 free parameters and literature-elicited priors from a screen of 1,847 records. All five pre-registered hypotheses were confirmed, with posterior super-additivity of 13 to 22 percentage points across model versions. neurocascade extends this to a receptor-to-behavior simulation engine, coupling pharmacokinetics to receptor binding to Wilson-Cowan circuit dynamics, with Bayesian calibration and 62 of 62 tests passing. ergofluids tests whether Koopman-operator methods with a Mori-Zwanzig memory kernel can model drug transport through dense tumor tissue. Its first real-data gate did not meet its pre-registered criterion, and I reported that directly.
What AIMS offers me is the mathematical depth to push these methods further. My current tools are effective but I want to understand the theoretical conditions under which protein language model embeddings capture biophysical signal, and when topological descriptors are informative versus redundant. The AIMS curriculum in mathematical sciences, combined with the DeepMind AI for Science focus, is the right training ground. I bring a track record of independent, rigorous, multi-domain research. AIMS brings the formal foundation and the African research community to take it to the next level.
ESSAY: MOTIVATION FOR AI FOR SCIENCE
The phrase AI for science can mean many things. For me it means one specific discipline: building models whose failures are as informative as their successes. My cardiotoxicity study is the clearest example. I tested whether topological data analysis could predict hERG cardiotoxicity from protein-ligand interface geometry. The hypothesis was reasonable, the methods were standard, and the pre-registration was rigorous. The result was negative. Topological features did not beat a plain descriptor baseline. That negative result is a contribution because it settles a question the literature had never actually tested. My subsequent resistance study found the same: interface topology carries almost no signal for drug-resistance prediction. These failures redirected my research toward sequence-based representations, which led to TOPOLOGIX and its AUROC of 0.804 on the Platinum benchmark.
This is the practice of science. Hypotheses are tested, most fail, and the ones that survive are the ones worth building on. AI for science, done honestly, is a discipline of pre-registration, baseline comparison, and direct reporting of negative results. I have built my entire independent research career on that discipline. The AIMS DeepMind programme is attractive to me because it trains researchers in the mathematical foundations that make this discipline possible. I want to understand the theory of when representations carry signal, not just which representation happens to work on a given benchmark. That understanding is what will let me build the next generation of models, for drug resistance, for addiction neuroscience, and for the other biological questions I work on.
I am also motivated by the African research context. AIMS is an African institution training African scientists. My work as a Nigerian researcher has been conducted independently, without institutional support, and I have built my own infrastructure: DuckDB-based pipelines, self-hosted LLM serving, and production systems operations. The AIMS community would connect me to a network of peers and mentors who share my context and my mathematical interests. That combination, rigorous training and African community, is unique. It is why I am applying.
CHECKLIST
- [ ] Complete online application at https://ai.aims.ac.za/apply
- [ ] Upload CV, including ORCID 0009-0001-9272-6735 and GitHub github.com/AmunRaPtah
- [ ] Upload academic transcripts from University of Ibadan (B.Pharm, 2014-2021)
- [ ] Upload proof of enrollment or acceptance for M.Sc. Digital Health at HPI/Potsdam (Winter Semester 2026/27)
- [ ] Submit motivation letter (approximately 500 words, as drafted above)
- [ ] Submit representative examples of work in mathematical sciences, including TOPOLOGIX preprint and CCT preprints (OSF/Zenodo links)
- [ ] Complete the coding problem as specified on the application portal
- [ ] Verify citizenship documentation for Nigeria
- [ ] Verify residency status in Africa at time of application
- [ ] Confirm no previous AIMS scholarship held
- [ ] Prepare for short online interview in May 2026
EDITOR NOTES
- Eligibility risk: the applicant is already enrolled in an M.Sc. at HPI/Potsdam. The AIMS programme does not prohibit concurrent enrollment, but the selection panel may question commitment. The motivation letter addresses this by framing AIMS as complementary, not redundant, but the applicant should be prepared to explain this in the interview.
- Verification needed: the applicant's enrollment at HPI/Potsdam for Winter Semester 2026/27 must be confirmed with official documentation. The application requires proof of degree completion by August 2025, which the B.Pharm satisfies, but the transcript must be official and include the German equivalent grade of 1.9.
- Gap to fill: the motivation letter does not mention specific AIMS faculty or research groups the applicant wishes to work with. The applicant should research AIMS South Africa's current AI for Science research themes and name at least one specific group or researcher in the final version.
- The coding problem is a required component. The applicant should treat it as a serious technical assessment, not a formality, and should allocate several hours to it. The selection panel evaluates problem-solving approach and computational depth, so the code should be clean, documented, and efficient.