MOTIVATION LETTER
The Merck Neuroscience / AI Grant 2026 targets a specific gap in CNS drug development: the absence of scalable, open-source computational tools that predict drug-resistance mutations in neurological targets before they emerge in clinical trials. My independent research over the past two years has produced exactly such a platform. TOPOLOGIX, a protein-language-model-based classifier using ESM-2 delta-embeddings and Morgan fingerprints, achieves AUROC 0.804 on the Platinum benchmark of 553 drug-resistance mutations, covering 100% of mutations compared to approximately 18% for structure-limited tools like mCSM-lig. This platform, combined with my CCT dynamical-systems model for addiction pharmacotherapy, forms the basis of a computational pipeline that can de-risk Merck's CNS pipeline at the process-development stage, without requiring laboratory access.
My background spans pharmacology, computational modeling, and software engineering. I hold a B.Pharm from the University of Ibadan, am enrolled in the M.Sc. Digital Health programme at Hasso Plattner Institute / University of Potsdam, and have published five preprints and one co-authored paper under review at Alcohol (Elsevier). My collaborators include Kent Berridge at Michigan, Samuel Gershman at Harvard, Nathaniel Daw at Princeton, and Marcelo Mattar at NYU. I am an independent researcher based in Nigeria, which positions me to address underserved patient populations in LMIC contexts where addiction and neurodegenerative diseases carry disproportionate burden.
The grant would fund the extension of TOPOLOGIX to CNS-specific targets: GPCRs and ion channels relevant to addiction and neurodegeneration. It would also fund the integration of the CCT model as a translational biomarker framework for predicting treatment response. The work plan is feasible within twelve months, uses only open-source tools and publicly available data, and produces a validated, deployable platform. The budget of EUR 150,000 covers compute costs, a part-time research assistant, and dissemination. Merck's mission to combine neuroscience with AI aligns directly with the architecture I have already built and tested.
RESEARCH STATEMENT
The central problem this project addresses is the late-stage failure of CNS drugs due to resistance mutations that are not detected during preclinical development. Current structure-based tools for predicting resistance mutations require high-resolution protein structures, which are unavailable for approximately 82% of clinically relevant mutations. This leaves a blind spot in the drug-development pipeline, particularly for GPCRs and ion channels that are primary targets in addiction and neurodegenerative disease.
TOPOLOGIX solves this problem by using ESM-2 protein-language-model embeddings to represent mutations from sequence alone, combined with Morgan/ECFP drug fingerprints and a Random Forest classifier. On the Platinum benchmark of 553 drug-resistance mutations, TOPOLOGIX achieves AUROC 0.804 plus or minus 0.025, outperforming structure-based baselines such as mCSM-lig (approximately 0.70) while covering all mutations rather than the 18% that have solved structures. On the SKEMPI 2.0 benchmark, the platform achieves AUROC 0.634, indicating generalizability beyond the training distribution.
The CCT model provides a complementary translational biomarker framework. It is a tripartite pharmacological framework with three coupled ODE axes representing dopaminergic reward-prediction error, NMDAR-dependent long-term potentiation, and affective contrast. Bayesian MCMC calibration using PyMC with 14 free parameters and literature-elicited priors from an 1,847-record screen confirmed all five pre-registered hypotheses, with posterior super-additivity of 13 to 22 percentage points across model versions. This model can predict individual treatment response in addiction pharmacotherapy, enabling stratified clinical trial design.
The proposed work plan has three phases. Phase one (months 1-4): curate a CNS-specific resistance mutation dataset from public databases and literature, targeting GPCRs and ion channels. Phase two (months 5-8): retrain and validate TOPOLOGIX on this dataset, benchmarking against structure-based tools. Phase three (months 9-12): integrate the CCT model as a downstream predictor of treatment response, producing a unified platform that takes a drug candidate and a patient genotype and outputs resistance probability and expected treatment effect. All code, data, and models will be released under an open-source license on GitHub.
The feasibility of this plan rests on existing infrastructure. I have four independent DuckDB-based ingest-to-analyze pipelines operational across life-sciences, tech/AI, and social-science domains. I self-host local LLM serving using llama.cpp with on-demand model swapping. I manage production systems on Linux VPS with systemd, Caddy TLS, CI/CD, and automated backup and disaster-recovery. The computational requirements for this project are modest: approximately EUR 20,000 for cloud compute over twelve months, with the remainder of the EUR 150,000 budget allocated to a part-time research assistant (EUR 60,000), publication fees and conference travel (EUR 30,000), and software maintenance and data acquisition (EUR 40,000).
The translational impact is direct. A pharmaceutical company using this platform can identify resistance mutations for a CNS candidate drug within days rather than months, and can design clinical trials that stratify patients by predicted treatment response. For Merck's neuroscience pipeline, this means earlier termination of non-viable candidates, reduced clinical trial costs, and faster progression of drugs that are likely to succeed. For LMIC populations, where addiction and neurodegenerative diseases are understudied and undertreated, the platform enables drug development that accounts for genetic diversity often excluded from structure-based tools.
BUDGET JUSTIFICATION
The total requested amount is EUR 150,000 over twelve months. Compute costs of EUR 20,000 cover GPU instances for ESM-2 embedding generation and Random Forest hyperparameter optimization, plus CPU instances for Bayesian MCMC calibration of the CCT model. A part-time research assistant at EUR 5,000 per month for twelve months (EUR 60,000) will handle data curation, literature screening, and validation experiments. Publication fees and conference travel (EUR 30,000) support open-access publication in a peer-reviewed journal and presentation at one computational neuroscience and one AI-in-drug-discovery conference. Software maintenance and data acquisition (EUR 40,000) cover API costs for database access, cloud storage, and software licensing for tools not available under open-source licenses. No equipment or laboratory costs are included, as the project is entirely computational. The budget is conservative relative to the EUR 500,000 maximum and reflects the lean, independent-researcher model I have operated for the past two years.
CHECKLIST
- [ ] Complete online application form at Merck Open Innovation portal
- [ ] Upload motivation letter (this document)
- [ ] Upload research statement (this document)
- [ ] Upload budget justification (this document)
- [ ] Upload CV with ORCID, GitHub, and publication list
- [ ] Upload two reference letters: one from Kent Berridge (University of Michigan), one from Samuel Gershman (Harvard University)
- [ ] Verify eligibility: confirm that independent researchers without institutional affiliation are accepted in this track
- [ ] Confirm deadline: August 31, 2026, 23:59 CET
- [ ] Prepare one-page visual summary of TOPOLOGIX architecture and CCT model results for optional supplementary materials
EDITOR NOTES
- Eligibility risk: The programme may require institutional affiliation for grant administration. Eniola should confirm whether an independent researcher can receive funds directly, or whether a host institution (e.g., Hasso Plattner Institute, where he is enrolled for the M.Sc.) must administer the grant. If the latter, a letter of support from HPI is needed.
- Fact verification: The AUROC value of 0.804 on the Platinum benchmark should be double-checked against the most recent preprint version. The profile states 0.804 plus or minus 0.025, but the SKEMPI 2.0 value of 0.634 may have changed with additional validation.
- Gap: The application does not specify which CNS targets (specific GPCRs or ion channels) will be used for the retraining phase. Eniola should select 3-5 concrete targets from Merck's published pipeline or from the literature on addiction and neurodegeneration, and name them explicitly in the research statement.
- Gap: No mention of ethical approval or data privacy for patient-genotype data. If the CCT model will be validated against clinical trial data, Eniola should clarify whether publicly available de-identified datasets (e.g., from ClinicalTrials.gov or OpenNeuro) will be used, and confirm that no IRB approval is needed.
- Tone check: The motivation letter opens with the research problem rather than self-introduction, which follows the standing preference. However, the first sentence of the research statement could be more specific about the failure rate of CNS drugs in clinical trials. Adding a concrete number (e.g., "approximately 90% of CNS drugs fail in Phase II or III trials") would strengthen the framing.