← Together AI Startup Accelerator MODERATE Startup
AI Draft — Together AI Startup Accelerator
For Eniola Olutogun, the strongest angle is the venture's core technology: predicting drug resistance mutations using ESM-2 protein language model. This directly aligns with Together AI's focus on AI compute and model training. Emphasize the need for GPU credits to fine-tune ESM-2 on SKEMPI 2.0 (3K mutations) to improve AUROC from 0.634 to ≥0.70, which is a concrete, measurable milestone. Highlight the venture's AI-first approach and the potential for pharma partnerships (e.g., Servier) as a path to real-world impact, making it an attractive candidate for compute support and VC network connections.
Full Research →
Model: deepseek/auto
Tokens: 0
Generated: 2026-08-04 20:54
Profile: startup
MOTIVATION LETTER The Together AI Startup Accelerator exists to fund teams where AI is the product, not a feature. My venture fits that definition exactly. I am Eniola Olutogun, a pharmacist and machine learning engineer, and I have built a model that predicts drug resistance mutations from protein sequence alone. No crystal structure required. The model uses ESM-2 protein language model delta-embeddings combined with ECFP4 drug fingerprints, classified by a Random Forest. On the Platinum benchmark, it achieves an AUROC of 0.804 plus or minus 0.025 under protein-grouped cross-validation across 553 mutations. That beats the published state of the art, mCSM-lig, which sits near 0.70. My model also covers 100 percent of mutations in the test set, where structure-limited tools cover roughly 18 percent. The compute credits and engineering support from Together AI would go directly to a specific, measurable milestone. My current model scores 0.634 AUROC on SKEMPI 2.0. The next step is fine-tuning ESM-2 on the SKEMPI 3K mutation set to push that score to 0.70 or higher. That fine-tuning run requires GPU capacity I do not currently have. This is a defined training job with a defined success threshold, not open-ended research. The accelerator's VC network matters for a different reason. I have an active pilot discussion with Servier in Suresnes, and existing relationships at Paris-Saclay, the Institut Pasteur, and Sanofi in Gentilly. A pharma partnership is the fastest path to revenue, but the compute support is what gets me to a pilot-ready model. Together AI's generalist portfolio is an advantage here, because the engineering team works with protein language models across domains, not just biotech. I am pre-seed, pre-incorporation, and the sole founder. The proof of concept is validated. The model works. What it needs now is scale in training data and a partner who can supply the infrastructure. That is the Build tier of this accelerator. I am applying for that tier. RESEARCH STATEMENT Drug resistance is the reason antibiotics fail, antivirals lose efficacy, and targeted cancer therapies stop working. The standard approach to predicting resistance mutations requires a crystal structure of the target protein. That structure does not exist for most clinically relevant proteins. My venture removes that dependency entirely. The method works as follows. I take a protein sequence and pass it through ESM-2, a protein language model trained on millions of sequences. I compute delta-embeddings, which capture the change in the model's internal representation when a mutation is introduced. I pair those embeddings with ECFP4 drug fingerprints, which encode the chemical structure of the drug in question. A Random Forest classifier then predicts whether that specific mutation confers resistance to that specific drug. The validation results are concrete. On the Platinum benchmark, which contains 553 mutations across 44 proteins, the model achieves an AUROC of 0.804 with a standard deviation of 0.025 under protein-grouped cross-validation. This grouping is important: it means the model is tested on proteins it has never seen during training, which is the realistic deployment scenario. The published state of the art, mCSM-lig, achieves roughly 0.70 on the same benchmark. My model also achieves 100 percent mutation coverage, because it only needs sequence data. Structure-based tools cover about 18 percent of mutations, because they require a resolved crystal structure. The gap is on SKEMPI 2.0, where my model scores 0.634 AUROC. That dataset contains binding affinity changes for 3,000 mutations, and it is the right training set to improve generalization. The roadmap is to fine-tune ESM-2 on those 3,000 mutations, which should raise SKEMPI performance to 0.70 or higher. That fine-tuning run is the immediate technical objective, and it requires GPU compute. The commercial path runs through pharma partnerships. I am in active discussion with Servier in Suresnes for a pilot. I have existing relationships at Paris-Saclay, the Institut Pasteur, and Sanofi in Gentilly. The product is a software tool that pharma companies use during preclinical development to identify which resistance mutations are likely to emerge for a candidate drug, before the drug enters the clinic. That information changes development decisions. The venture is pre-seed and pre-incorporation. The proof of concept is validated. The next milestone is the SKEMPI fine-tune, followed by the Servier pilot, followed by annual recurring revenue. Together AI's compute credits are the enabling resource for the first milestone. ESSAY: TECHNICAL FIT Together AI's selection criteria ask whether a startup will actively use the platform for training, fine-tuning, or inference. My venture will use it for all three. The core model is ESM-2, a 650-million-parameter protein language model. I need to fine-tune it on the SKEMPI 3K mutation set. That is a training job. The Random Forest classifier on top of the embeddings is a lightweight inference job. The full pipeline, from raw sequence to resistance prediction, will run repeatedly as new drug candidates are tested. The fine-tuning run is the gating item. My current SKEMPI 2.0 AUROC is 0.634. The target is 0.70. Fine-tuning ESM-2 on 3,000 mutations is a well-scoped job, but it requires sustained GPU access. I am a solo founder with no compute budget. The accelerator's credits close that gap directly. The engineering support is equally relevant. Together AI's team works with large language models daily. Protein language models are architecturally similar to text models, but they have domain-specific quirks in tokenization, sequence length, and embedding extraction. Having engineering guidance on efficient fine-tuning of ESM-2 would reduce the iteration time between training runs. The venture is AI-first by construction. The product is a model. The moat is the training methodology and the validation results. There is no wet lab component. This is a pure computational biology play, which means the accelerator's infrastructure is not a supplement to the product. It is the product. ESSAY: TEAM AND EXECUTION I am a pharmacist by training and a machine learning engineer by practice. That combination is the reason this venture exists. As a pharmacist, I saw drugs fail in the clinic because resistance emerged in ways that were not predicted during development. As an ML engineer, I saw that protein language models had reached the point where sequence-only prediction was viable. I built the model myself, as sole author, and validated it against public benchmarks. The execution record is specific. The model achieves 0.804 AUROC on Platinum, beating published state of the art. It covers 100 percent of mutations, where structure-based tools cover 18 percent. I have secured a meeting pipeline that includes SEMIA and Quest for Health, with WILCO One BioTech scheduled for October 2026. I have submitted applications to IncubAlliance and AI House. The EIC Accelerator and BPI i-Lab are future targets. The Servier pilot discussion is in progress. The accelerator's VC network is the specific value I need beyond compute. My partnerships are in France, at Servier, Sanofi, and the Institut Pasteur. The accelerator's network can introduce me to pharma and biotech investors who operate across Europe and the US. I am not looking for a generalist introduction. I am looking for investors who understand the difference between a model that predicts resistance from sequence and a model that requires a crystal structure. The risk in this venture is not technical feasibility. The proof of concept is done. The risk is execution speed: can I fine-tune the model, close the Servier pilot, and convert that pilot into recurring revenue before the compute runs out. The accelerator's credits and network directly address that risk. CHECKLIST - [ ] Confirm Together AI Startup Accelerator application portal access at the provided URL - [ ] Verify current Build tier eligibility, pre-seed stage, and no incorporation requirement - [ ] Prepare model validation metrics document with Platinum and SKEMPI 2.0 AUROC numbers - [ ] Draft one-page technical summary of ESM-2 delta-embedding methodology for engineering review - [ ] List specific GPU requirements for ESM-2 fine-tuning on SKEMPI 3K, including estimated hours - [ ] Prepare one-paragraph summary of Servier pilot discussion status, with named contact if permitted - [ ] Confirm SKEMPI 2.0 and SKEMPI 3K dataset access and licensing for commercial use - [ ] Prepare short founder bio emphasizing pharmacist-to-ML-engineer transition - [ ] Confirm WILCO One BioTech October 2026 timeline and any exclusivity clauses - [ ] Verify whether Together AI requires incorporation or bank account for credit disbursement EDITOR NOTES - Eligibility risk: The venture is pre-incorporation. Together AI's Build tier targets pre-seed startups, but the application may require a legal entity for the compute credit agreement. Verify before submitting. - Fact check needed: The SKEMPI 3K mutation count and the exact SKEMPI 2.0 AUROC of 0.634 should be re-verified against the latest model run before submission, as these are the two numbers the technical fit essay hinges on. - Gap to fill: The Servier pilot is described as "in progress" but no named contact or stage is provided. The applicant should insert the specific status, meeting date, or named individual if the application allows it. - The Platinum benchmark AUROC of 0.804 is the strongest number in the application and should be placed prominently in any verbal pitch or interview, as it directly beats the published state of the art. - The essay responses assume the accelerator asks for short answers on technical fit and team. If the actual application uses a single form with different questions, the content should be re-mapped to those exact prompts.
Draft History
v2 — 2026-08-04 20:22 · 0 tokens · startup
v1 — 2026-08-03 02:28 · 0 tokens · startup