← Student Researcher Program — Google DeepMind MODERATE Neuropharm/CCT
AI Draft — Student Researcher Program — Google DeepMind
For Eniola, the strongest angle is to position his TOPOLOGIX project as a direct fit for DeepMind's AI-for-science focus, leveraging his ESM-2 protein language model expertise and ML methods (Random Forest, delta-embeddings) to solve drug-resistance prediction—a high-impact biomedical problem. This line most directly matches DeepMind's stated interest in applying AI to scientific challenges (e.g., AlphaFold), and his technical skills in protein ML and Python are highly relevant. He should emphasize his independent research record and the concrete AUROC results, framing himself as a self-driven researcher who can contribute immediately to a team working on protein ML or AI-driven drug discovery.
Full Research →
Model: deepseek/auto
Tokens: 0
Generated: 2026-08-04 20:42
Profile: researcher
MOTIVATION LETTER The Student Researcher Program at Google DeepMind sits at the exact intersection where my independent research has been operating for the past two years: applying machine learning to hard problems in biology and neuroscience. My current project, TOPOLOGIX, predicts drug-resistance mutations from protein sequence alone using ESM-2 protein language model delta-embeddings combined with Morgan/ECFP drug fingerprints and a Random Forest classifier. The system achieves an AUROC of 0.804 plus or minus 0.025 on the Platinum benchmark of 553 mutations, and 0.634 on SKEMPI 2.0. It beats structure-based baselines such as mCSM-lig at approximately 0.70 while covering 100 percent of mutations, compared to roughly 18 percent for structure-limited tools. This work is a direct continuation of the kind of AI-for-science research DeepMind pioneered with AlphaFold, and it is the line of my research that most closely matches the program's focus. I am currently enrolled in the M.Sc. Digital Health program at the Hasso Plattner Institute and the University of Potsdam, starting Winter Semester 2026/27. I hold a B.Pharm from the University of Ibadan with a German equivalent grade of 1.9, and I am a licensed pharmacist in Nigeria. My research record includes three sole-authored preprints on the Conjunctive Consolidation Threshold model of addiction, a tripartite pharmacological framework for reward-memory encoding prevention, each under review at a peer-reviewed journal. I also completed a pre-registered, powered replication study on hERG cardiotoxicity topology that settled a comparison the literature had never actually run: bipartite persistent homology features did not beat a plain descriptor baseline, with AUROC 0.8426 versus 0.8782. That negative result taught me to report what the data says, not what I hoped it would say. My technical stack includes Python with scipy, numpy, PyMC for Bayesian MCMC calibration, Ripser and GUDHI for topological data analysis, RDKit for cheminformatics, and experience with NEURON and Brian2 for neural simulation. I have built four independent DuckDB-based ingest-to-analyze pipelines across life sciences, tech, and social science domains, and I self-host local LLM serving with llama.cpp. I am comfortable with HPC environments using Nextflow and SLURM. The DeepMind program's emphasis on demonstrated research experience, strong ML skills, and alignment with AI-for-science matches my profile directly. TOPOLOGIX is a running system with published benchmark results, not a proposal. I am looking for a host team where I can contribute immediately to protein ML or AI-driven drug discovery projects, and where my independent research record can be an asset rather than a liability. The rolling application structure suits my timeline, as I am available to collaborate in person at a Google office while completing my M.Sc. coursework. RESEARCH STATEMENT TOPOLOGIX: Sequence-Based Prediction of Drug-Resistance Mutations Using Protein Language Model Embeddings The problem. Drug resistance is a primary cause of treatment failure in infectious disease and oncology. Predicting which mutations confer resistance, before they emerge in a clinical setting, would allow earlier intervention and better drug design. Existing tools such as mCSM-lig rely on protein structures, which are unavailable for approximately 82 percent of clinically relevant mutations. This structural bottleneck leaves most of the sequence space unexplored. The approach. TOPOLOGIX uses ESM-2 protein language model delta-embeddings to represent the change in protein sequence context induced by a mutation, paired with Morgan/ECFP drug fingerprints to represent the ligand. A Random Forest classifier learns the mapping from this combined representation to a resistance label. The system requires only sequence data, eliminating the structure dependency that limits current tools. Results to date. On the Platinum benchmark of 553 mutations, TOPOLOGIX achieves AUROC 0.804 with a standard deviation of 0.025. On SKEMPI 2.0, it achieves 0.634. It outperforms structure-based baselines including mCSM-lig at approximately 0.70 while covering 100 percent of mutations in the benchmark, compared to roughly 18 percent for structure-limited tools. These results are from a pre-registered protocol, and the code is available on my GitHub repository. The path that led here. TOPOLOGIX is the third iteration of a research line that began with topological data analysis of protein-ligand interfaces. My first study tested whether bipartite persistent homology with an opposition-distance metric could predict hERG cardiotoxicity from interface geometry. It could not beat a plain descriptor baseline, with AUROC 0.8426 versus 0.8782. My second study applied the same topological constructs to drug-resistance prediction and found they carried almost no signal, with AUROC 0.425 and 0.485 on the Platinum benchmark. These negative results ruled out interface geometry as the driver of resistance and motivated the sequence-representation approach that became TOPOLOGIX. The current system's performance is a direct consequence of that falsification-driven pivot. Why this fits DeepMind. DeepMind's AlphaFold demonstrated that protein language models and learned representations can solve problems that structure-based biophysics could not. TOPOLOGIX operates in the same paradigm: it treats protein sequence as a language and learns resistance-relevant features from embeddings rather than from hand-crafted structural descriptors. The program's stated interest in AI for science, and specifically in protein structure prediction and related problems, makes this a direct fit. What I would contribute. I bring a working system with benchmark results, a pre-registered methodology, and a demonstrated ability to run rigorous validation pipelines. I also bring domain expertise as a licensed pharmacist, which informs how these predictions would be used in clinical and drug-development contexts. I am prepared to extend TOPOLOGIX in several directions depending on host team priorities: scaling to larger mutation datasets, integrating attention-based architectures beyond Random Forest, or applying the delta-embedding framework to other protein properties such as binding affinity or stability. Validation status. TOPOLOGIX has passed its pre-registered benchmarks on Platinum and SKEMPI 2.0. The methodology is published as open code, and the results are reproducible from my repository. I have not claimed clinical deployment or product status; this is research software at the validation stage. EDITOR NOTES - Framing choice: TOPOLOGIX is the selected research line because it directly matches DeepMind's AI-for-science focus and protein ML expertise, per the strategy notes. The hERG and resistance topology studies are presented as falsified prior work that motivated TOPOLOGIX, not as current claims. - Eligibility risk: The program requires enrollment in a Bachelor's, Master's, or PhD program at time of application. Eniola's M.Sc. starts Winter Semester 2026/27. If the application is reviewed before enrollment begins, he may not meet the strict eligibility gate. Verify the program's exact enrollment timing requirement before submitting. - Verification needed: Confirm the AUROC figures for TOPOLOGIX (0.804 plus or minus 0.025 on Platinum, 0.634 on SKEMPI 2.0) and the mCSM-lig baseline of approximately 0.70 against the actual benchmark papers. Confirm the 18 percent coverage figure for structure-limited tools. - Missing detail: The letter does not specify which Google DeepMind office Eniola would prefer or is available to work from. The program requires in-person collaboration. Insert location preference and availability window before submission. - Program fit caveat: The Student Researcher Program is primarily aimed at current students. Eniola's independent researcher status is unusual for this program. The letter emphasizes his enrollment in the M.Sc. program to satisfy eligibility, but the gap between application and enrollment start may be a problem. Consider contacting the program directly to confirm eligibility timing. CHECKLIST - [ ] Confirm enrollment status in M.Sc. Digital Health at HPI/Potsdam and obtain proof of enrollment or admission letter - [ ] Verify TOPOLOGIX benchmark results against Platinum and SKEMPI 2.0 papers - [ ] Verify mCSM-lig baseline AUROC and coverage figures - [ ] Specify preferred Google DeepMind office location and availability window - [ ] Prepare GitHub repository with TOPOLOGIX code, pre-registration documents, and README - [ ] Update ORCID record with TOPOLOGIX and CCT preprints - [ ] Obtain letters of recommendation from Kent Berridge, Samuel Gershman, or Nathaniel Daw if required - [ ] Prepare academic transcripts from University of Ibadan and HPI/Potsdam - [ ] Confirm PCN pharmacist license is current and include registration number - [ ] Draft responses to any program-specific short-answer questions not covered in this document - [ ] Submit application through the DeepMind Student Researcher Program portal - [ ] Save copies of all submitted materials and confirmation emails
Draft History
v2 — 2026-08-04 20:05 · 0 tokens · researcher
v1 — 2026-07-31 00:25 · 0 tokens · researcher