← ARAI - Transfyr AI Fellowship MODERATE General
AI Draft — ARAI - Transfyr AI Fellowship
The strongest angle for Eniola is to leverage the TOPOLOGIX project as the core research proposal, because it directly matches the fellowship's focus on building ML systems that reason over scientific execution data (here, protein sequences and drug-resistance outcomes) and on real-world evaluations (benchmarking against Platinum and SKEMPI). Eniola should frame TOPOLOGIX as a foundation for a larger vision: using multimodal scientific-execution data (e.g., lab protocols, instrument traces, outcomes) to predict and understand drug resistance in real time, aligning with the fellowship's 'scientific execution ontology' and 'real-world evaluations' themes. The rigorous pre-registration and honest reporting of negative results (e.g., hERG topology study, ergofluids) demonstrate the 'exceptional research taste' the fellowship seeks, and the multi-domain expertise (pharmacology, ML, dynamical systems) is a unique asset for tackling the ambiguity of real scientific workflows.
Full Research →
Model: deepseek/auto
Tokens: 0
Generated: 2026-08-04 21:01
Profile: researcher
MOTIVATION LETTER The Transfyr AI Fellowship asks for researchers who can build ML systems that reason over how science actually happens, not just over clean benchmark datasets. My research trajectory has been a deliberate preparation for exactly that problem. I am a pharmacist and computational researcher who has spent the last two years building, pre-registering, and honestly evaluating ML systems across protein science, neuroscience, and dynamical systems. The project I propose for this fellowship, TOPOLOGIX, directly targets the fellowship's stated themes of scientific execution ontology and real-world evaluations: predicting drug-resistance mutations from protein sequence alone, using ESM-2 protein language model embeddings combined with Morgan fingerprints and a Random Forest classifier. TOPOLOGIX currently achieves an AUROC of 0.804 plus or minus 0.025 on the Platinum benchmark of 553 mutations, and 0.634 on SKEMPI 2.0. It beats structure-based baselines such as mCSM-lig at approximately 0.70 while covering 100 percent of mutations, versus roughly 18 percent for structure-limited tools. This is a working system that takes a protein sequence and a drug fingerprint and returns a resistance prediction with quantified uncertainty. The fellowship's emphasis on exceptional research taste is where my record is strongest. I have published three sole-authored preprints on the Conjunctive Consolidation Threshold model of addiction, a tripartite pharmacological framework with a 14-parameter Bayesian MCMC calibration from a screen of 1,847 records, with all five pre-registered hypotheses confirmed. I have also published negative results that settle open questions. My pre-registered replication of a cardiotoxicity topology study found that bipartite persistent homology does not beat a plain descriptor baseline for hERG cardiotoxicity prediction, AUROC 0.8426 versus 0.8782. My interface-topology-for-resistance study found that the same topological constructs carry almost no signal for drug-resistance prediction, AUROC 0.425 and 0.485 on the Platinum benchmark. These results are published honestly, not reframed. The ergofluids project, a Koopman-operator method with Mori-Zwanzig memory kernels for drug transport in tumor tissue, passed its synthetic-data gates but failed its first real-data gate against digitized published figures. I reported that failure directly rather than redefining the success criterion. The fellowship's theme of comfort with ambiguity and curiosity about how science happens in physical environments matches my background as a clinical pharmacist. I have dispensed medications, managed supply chains, and watched resistance emerge in real patients. I know that the gap between a laboratory result and a clinical outcome is where most ML systems fail. TOPOLOGIX is my attempt to close part of that gap. The full-time commitment is feasible. I am enrolled in the M.Sc. Digital Health program at Hasso Plattner Institute and University of Potsdam for Winter Semester 2026/27, and I will self-certify leave for the fellowship year. I am available to work in person in Boston or Cambridge. I am a Nigerian citizen and will require visa sponsorship, which the fellowship indicates is available. I am applying to the Transfyr AI Fellowship because it is the only program I have found that explicitly funds the kind of work I do: building ML systems that reason over scientific execution data, evaluating them against real-world benchmarks, and reporting failures with the same rigor as successes. RESEARCH STATEMENT The problem: drug resistance kills. The World Health Organization estimates antimicrobial resistance directly caused 1.27 million deaths in 2019. In oncology, resistance to targeted therapies is the primary reason most advanced cancers eventually become untreatable. Yet the dominant computational tools for predicting resistance mutations still require a protein structure, which exists for only a fraction of clinically relevant targets. When a structure is unavailable, clinicians and researchers are flying blind. The project: TOPOLOGIX is a sequence-based resistance prediction system. It takes an ESM-2 protein language model delta-embedding for a mutation, concatenates it with a Morgan/ECFP fingerprint of the drug, and feeds both to a Random Forest classifier. The architecture is deliberately simple. The contribution is a rigorous, pre-registered evaluation showing that sequence-only representations can outperform structure-based tools while covering the full mutation space. Current results on the Platinum benchmark, 553 mutations: AUROC 0.804 plus or minus 0.025. On SKEMPI 2.0: AUROC 0.634. Structure-based baseline mCSM-lig achieves approximately 0.70 but covers only 18 percent of mutations because it requires a resolved structure. TOPOLOGIX covers 100 percent. This is the first system I am aware of that demonstrates this coverage-performance tradeoff can be broken. The trajectory I propose for the fellowship year has three phases. Phase one, months one to three: expand the training data. Platinum and SKEMPI are small. I will assemble a larger corpus from published deep mutational scanning experiments, drug susceptibility databases, and clinical resistance surveillance data. I have already built four independent DuckDB-based ingest-to-analyze pipelines across life-sciences and other domains, so the data infrastructure is in place. The output of this phase is a new benchmark dataset with documented provenance, versioning, and a pre-registered evaluation protocol. Phase two, months four to eight: move from classification to reasoning. The current system predicts resistance or not. The fellowship's theme of multimodal reasoning and conflicting evidence points to the next step: predicting the mechanism of resistance, not just its presence. I will extend TOPOLOGIX to output structured predictions over resistance mechanisms, such as steric clash, charge disruption, or allosteric modulation, and to flag when different evidence streams, sequence features versus structural features when available, disagree. Disagreement becomes an explicit output, not an averaged-away nuisance. This builds directly on my psyche-twin architecture, where disagreement between evidence streams becomes an explicit graph edge. Phase three, months nine to twelve: real-world evaluation. The fellowship asks for real-world evaluations, not just benchmarks. I will partner with a clinical genomics laboratory to test TOPOLOGIX on a retrospective cohort of patient-derived resistance mutations where the outcome is known. The evaluation protocol will be pre-registered before the data is accessed. The deliverable is a publishable benchmark, a documented prototype, and an honest assessment of where the system fails. Why this project fits Transfyr: the fellowship's stated themes include scientific execution ontology, real-world evaluations, and biosecurity and bio-risk observability. Drug resistance is a bio-risk observability problem. A system that can predict resistance from sequence alone, without requiring a structure, is a monitoring tool for emerging threats. The scientific execution ontology theme maps directly to my data pipeline: every mutation in my benchmark will carry provenance metadata describing how it was measured, in what assay, under what conditions. The fellowship's emphasis on tacit expertise and physical-world AI maps to my clinical background. I have watched resistance emerge in patients. I know what the data does not capture. The risk I am most honest about: sequence-only prediction may hit a ceiling. Some resistance mechanisms are purely structural and may be invisible to sequence embeddings. My pre-registered evaluation plan includes a stopping rule. If TOPOLOGIX does not beat the structure-based baseline on the expanded benchmark at the end of phase two, I will report that result and pivot to a hybrid sequence-structure architecture. The fellowship funds research taste, and research taste includes knowing when a hypothesis is wrong. ESSAY: RESEARCH TASTE AND RIGOR The Transfyr fellowship lists exceptional research taste as a selection criterion. I interpret research taste as the ability to choose problems worth solving and to report results honestly, especially when the results are negative. My record demonstrates both. In 2025, I conducted a pre-registered, powered replication of a claim in the topological data analysis literature: that bipartite persistent homology of protein-ligand interfaces predicts hERG cardiotoxicity. The published literature suggested this was true. My replication found it was not. The topological features achieved an AUROC of 0.8426 against a plain descriptor baseline of 0.8782. The topological approach lost. I published the result. In the same period, I applied the same topological constructs to drug-resistance prediction. The result was worse: AUROC 0.425 and 0.485 on the Platinum benchmark. Interface geometry, measured this way, carries almost no signal for resistance. This negative result is what motivated TOPOLOGIX. I did not try to make the topology work harder. I changed the representation entirely, moving from structure to sequence. The ergofluids project is the clearest example of my reporting standards. I extended Koopman-operator methods with a Mori-Zwanzig memory kernel to model drug transport through tumor tissue. The synthetic-data validation gates passed. The first real-data gate, tested against digitized published figures, failed its primary pre-registered criterion. I reported the failure directly, in the project documentation, without reframing the success criterion. The project is methods-validation research, not a venture. There is no IP to protect and no product claim to defend. The Conjunctive Consolidation Threshold model, my main neuroscience line, shows the same discipline on the positive side. A 14-parameter Bayesian MCMC calibration with literature-elicited priors from a 1,847-record screen, all five pre-registered hypotheses confirmed, posterior super-additivity of 13 to 22 percentage points across model versions. Three sole-authored preprints, each in review at a peer-reviewed journal. Research taste is not about being right. It is about being honest about what the data says, choosing problems where the answer matters, and building systems that can be evaluated by others. That is what I do. ESSAY: MULTI-DOMAIN EXPERTISE AND AMBIGUITY The fellowship asks for comfort with ambiguity and curiosity about how science actually happens in physical environments. My background is unusual: I am a licensed pharmacist, a computational modeler, and a software engineer. I have worked in a clinical pharmacy dispensing medications, in a genomics surveillance pipeline tracking antimicrobial resistance, and in a computational lab building Bayesian models of brain circuits. This combination matters for the fellowship's mission. Scientific execution data, the actual records of how experiments are run, is messy. It comes from lab notebooks, instrument logs, protocol documents, and clinical records. It is incomplete, contradictory, and full of tacit knowledge that never gets written down. Most ML researchers never see this data. I have lived in it. As a clinical pharmacist at Ramset Pharmacy, I managed drug inventories and dispensed medications. I saw how resistance emerges in practice, not just in the literature. As a research assistant at GHRU-GSAR, I built surveillance pipelines for antimicrobial resistance genomics. I learned that the data pipeline is the science. If the metadata is wrong, the model is wrong, no matter how sophisticated the architecture. My multi-domain skills are concrete. I write production-grade Python with scipy, numpy, PyMC for Bayesian inference, and ODE solvers. I use Ripser and GUDHI for topological data analysis, NEURON and Brian2 for neural simulation, AlphaFold and RDKit for protein and drug representation, GROMACS and AutoDock for molecular dynamics. I deploy and operate my own infrastructure: Linux VPS, systemd, Caddy TLS, CI/CD, automated backups. I serve local LLMs with llama.cpp and swap models on demand. I have built four independent DuckDB-based ingest-to-analyze pipelines across life-sciences, tech/AI/security, and social-science domains. This breadth is not a distraction. It is the point. The fellowship wants researchers who can operate across the full arc of scientific execution, from the physical environment of the lab to the abstract space of the model. I have done both, and I know where the gaps are. ESSAY: COMMITMENT AND LOGISTICS The fellowship requires full-time commitment for 12 months and willingness to work in person in Boston or Cambridge. I confirm both. I am currently enrolled in the M.Sc. Digital Health program at Hasso Plattner Institute and University of Potsdam, starting Winter Semester 2026/27. I will self-certify leave for the fellowship year, as the fellowship's terms permit for current students. My employment as National Product Manager at Synthcare, which began in March 2026, will be paused for the fellowship period. I am a Nigerian citizen. I will require visa sponsorship to work in the United States. The fellowship indicates that visa support is available for international applicants. I have experience with international relocation, having moved from Nigeria to Germany for my graduate studies. My research infrastructure is portable. My code is on GitHub, my data pipelines are containerized, and my compute needs are modest. I can be productive from day one. CHECKLIST - [ ] Confirm eligibility for early-career researcher track with M.Sc. enrollment status - [ ] Verify visa sponsorship process for Nigerian citizens under fellowship terms - [ ] Confirm IP ownership clause details with fellowship administration before signing - [ ] Obtain self-certification of leave from Hasso Plattner Institute M.Sc. program - [ ] Confirm leave of absence from Synthcare National Product Manager position - [ ] Prepare TOPOLOGIX code repository for review, including pre-registration documents - [ ] Prepare Platinum and SKEMPI 2.0 benchmark evaluation scripts for reproducibility - [ ] Draft letter of recommendation request from Kent Berridge, University of Michigan - [ ] Draft letter of recommendation request from Samuel Gershman, Harvard University - [ ] Draft letter of recommendation request from Nathaniel Daw, Princeton University - [ ] Draft letter of recommendation request from Marcelo Mattar, NYU - [ ] Prepare publication list including three CCT preprints and Alcohol co-authored paper - [ ] Prepare documentation of negative results, hERG topology study and ergofluids gate failure - [ ] Submit application by August 15, 2026 deadline - [ ] Confirm in-person availability for Boston/Cambridge for 12 months EDITOR NOTES - Eligibility risk: the M.Sc. enrollment at HPI/Potsdam starts Winter Semester 2026/27, which overlaps with the fellowship year. The self-certification of leave must be confirmed in writing with both HPI and the fellowship administration before submission. This is the single biggest risk to the application. - Verification needed: the exact dates of the Synthcare employment and the leave-of-absence terms are not in the profile. The applicant must confirm that pausing employment for 12 months does not violate the employment contract. - Gap to fill: the profile does not specify which clinical genomics laboratory would partner for the phase three real-world evaluation. The applicant must either name a specific lab with a letter of intent or revise the research statement to describe the partnership as planned rather than confirmed. - IP ownership clause: the strategy notes flag the fellowship's IP ownership clause as a potential constraint. The applicant should review this clause carefully, especially given the independent research lines and the open-source nature of TOPOLOGIX. If the clause is unacceptable, this application should not proceed. - The chosen research line is TOPOLOGIX, consistent with the recommended framing angle. The negative topology results are presented as historical context that motivated TOPOLOGIX, not as current claims. The ergofluids project is described at its actual validation status, behind a failed real-data gate, and is not presented as validated IP or product.
Draft History
v2 — 2026-08-04 20:12 · 0 tokens · researcher
v1 — 2026-07-31 17:19 · 0 tokens · researcher