MOTIVATION LETTER
The Platinum benchmark measures what matters in drug resistance prediction: can a model generalize to mutations it has never seen, across proteins it has never seen. My platform scores an AUROC of 0.804 plus or minus 0.025 on that benchmark using protein-grouped cross-validation, and it covers 100 percent of mutations in the test set. Structure-limited tools like mCSM-lig cover roughly 18 percent of those same mutations and score around 0.70 AUROC. That gap, 0.804 versus 0.70 with five times the coverage, is the commercial and scientific foundation of this venture.
Bpifrance's Start-up programme funds pre-seed deep-tech companies that can demonstrate a defensible technical advantage and a credible path to revenue in French strategic sectors. This venture sits at the intersection of two of those sectors, AI and health, and specifically addresses drug resistance, the single largest cause of treatment failure in oncology and infectious disease. The technology uses ESM-2 protein language model delta-embeddings combined with ECFP4 drug fingerprints, fed into a Random Forest classifier. No crystal structure is required. That design decision enables the 100 percent mutation coverage, because structure prediction fails precisely on the mutated, destabilized proteins where resistance emerges.
The roadmap is concrete. Fine-tune ESM-2 on the SKEMPI 2.0 mutation dataset, which contains over 3,000 experimentally measured binding affinity changes, to push AUROC to 0.70 or higher on that harder benchmark. Then run a pilot with Servier at their Suresnes site, using the platform to predict resistance mutations for a compound in their oncology pipeline. That pilot converts the platform from a benchmark result into a revenue-generating service. The target is annual recurring revenue from subscription access to the prediction API plus per-project consulting for preclinical resistance screening.
I am a pharmacist by training and a machine learning engineer by practice. I built this platform alone, from the ESM-2 embedding pipeline to the classifier evaluation harness. The proof of concept is complete and validated on public benchmarks. The next step is incorporation in Ile-de-France, which is already planned, followed by the Servier pilot and the first commercial contracts.
Bpifrance's selection criteria emphasize disruptive innovation, market potential, and team execution. The 0.804 AUROC with 100 percent coverage is the disruptive innovation. The global antimicrobial resistance market and the oncology resistance screening market are the commercial opportunity. The pharmacist-to-ML-engineer trajectory is the execution evidence. This application requests non-dilutive funding to incorporate, fine-tune the model on SKEMPI, and run the Servier pilot. The funds will be deployed within twelve months of disbursement.
RESEARCH STATEMENT
Drug resistance emerges through mutations that preserve or restore protein function while evading the drug's binding mechanism. Predicting which mutations will do this, before they appear in a patient, requires a model that can score any mutation in any protein against any drug. Structure-based tools cannot do this at scale because they require a crystal structure or a high-confidence homology model, and those structures are unavailable for most clinically relevant mutant proteins. My platform removes that dependency entirely.
The technical architecture is a three-part pipeline. First, ESM-2, a 650-million-parameter protein language model trained on 138 million sequences, generates embeddings for the wild-type and mutant protein sequences. Second, the delta between those embeddings, the change in the model's internal representation caused by the mutation, is computed. Third, that delta is concatenated with the ECFP4 fingerprint of the drug molecule and passed to a Random Forest classifier. The classifier outputs a probability that the mutation confers resistance to that drug.
Validation results are as follows. On the Platinum benchmark, 553 mutations with protein-grouped cross-validation, the platform achieves AUROC 0.804 with a standard deviation of 0.025. On SKEMPI 2.0, a harder benchmark measuring binding affinity changes across a wider range of protein-protein interactions, the platform achieves AUROC 0.634. The Platinum result exceeds the published state of the art for structure-based tools, mCSM-lig at approximately 0.70 AUROC, while covering 100 percent of mutations versus approximately 18 percent for those tools. The SKEMPI result is lower because that benchmark includes mutations far outside the training distribution of resistance-related changes, which is precisely the gap the fine-tuning roadmap addresses.
The next technical milestone is fine-tuning ESM-2 on the SKEMPI 2.0 dataset. SKEMPI 2.0 contains over 3,000 experimentally measured binding affinity changes from 319 protein-protein complexes. Fine-tuning the language model on this data will teach it the biophysical features that correlate with binding affinity changes, which are the same features that drive drug resistance. The target is AUROC 0.70 or higher on SKEMPI 2.0 after fine-tuning. This is an achievable target because the current 0.634 result was obtained with a frozen, off-the-shelf ESM-2; fine-tuning adds a training signal that the current pipeline lacks.
The commercial application is preclinical resistance screening. Pharmaceutical companies need to know, before clinical trials, which resistance mutations are likely to emerge for a candidate drug. This information guides combination therapy design, prodrug strategies, and patient stratification. The platform provides this as a service: a pharmaceutical partner submits a drug structure and a target protein sequence, and receives a ranked list of high-risk resistance mutations with confidence scores. The Servier pilot at Suresnes is the first deployment of this service.
The platform also addresses antimicrobial resistance surveillance. Clinical isolates can be sequenced and their resistance profiles predicted without waiting for phenotypic testing. This application is further down the roadmap, after the oncology and antiviral pilots, but it represents a larger addressable market.
The intellectual property position rests on the specific combination of delta-embeddings from ESM-2 with drug fingerprints for resistance prediction, the fine-tuned model weights, and the evaluation methodology. The model architecture itself is published, but the trained weights and the specific feature engineering are proprietary. A patent application covering the delta-embedding plus fingerprint method for resistance prediction is planned post-incorporation.
ESSAY: INNOVATION AND DEFENSIBILITY
The defensibility of this platform does not rest on a single algorithm. It rests on a data flywheel that competitors cannot replicate without the same combination of clinical partnerships and benchmark iteration. The core method, delta-embeddings from ESM-2 combined with ECFP4 fingerprints, is published and reproducible. What is not reproducible is the fine-tuned model that will emerge from training on SKEMPI 2.0 plus proprietary mutation-resistance data from the Servier pilot and subsequent pharmaceutical partners.
Each pilot generates a new dataset of experimentally confirmed resistance mutations. Each dataset fine-tunes the model further. Each fine-tuning round improves AUROC on the next benchmark. A competitor starting today would need to replicate the benchmark results, secure their own pharmaceutical partnerships, and iterate through the same fine-tuning cycles. That is an eighteen to twenty-four month head start, which in a pre-seed to seed stage company is the difference between market leadership and also-ran status.
The 100 percent mutation coverage is itself a defensible moat. Structure-based tools cannot cover mutations in proteins without solved structures, and most clinically relevant mutant proteins lack solved structures. My platform requires only a sequence, which is available for every protein in every clinical database. This means the platform can be deployed immediately on any new resistance problem without waiting for crystallography or cryo-EM. For a pharmaceutical partner, that speed is the value proposition.
The Random Forest classifier at the core is deliberately simple. This is a feature, not a limitation. It makes the model interpretable, which matters for regulatory acceptance and for pharmaceutical partners who need to explain predictions to their own internal review boards. The complexity lives in the embeddings and the fine-tuning, not in the final classifier. This design choice also makes the platform computationally lightweight, deployable on a single GPU server, which keeps the cost structure low for a pre-seed company.
ESSAY: MARKET AND COMMERCIAL STRATEGY
The immediate market is preclinical resistance screening for pharmaceutical companies developing oncology, antiviral, and antimicrobial drugs. The global antimicrobial resistance market was valued at approximately 8 billion USD in 2023 and is projected to grow as regulatory pressure increases for resistance data in drug approval dossiers. The oncology resistance screening segment is smaller but higher value per engagement, with a single pilot contract worth 50,000 to 150,000 euros.
The revenue model is a tiered subscription. Tier one is API access to the prediction platform, priced at 30,000 euros per year per partner. Tier two adds a dedicated analyst who runs custom resistance screens and delivers a written report, priced at 80,000 euros per year. Tier three is a joint research collaboration where the partner shares proprietary resistance data in exchange for co-developed model improvements, priced per project at 150,000 euros and up. The Servier pilot is structured as a tier three engagement.
The geographic strategy is France-first. Incorporation in Ile-de-France, proximity to Servier in Suresnes and Sanofi in Gentilly, and the existing relationship with Paris-Saclay's I2BC and the Institut Pasteur provide the scientific network needed for the first pilots. After the Servier pilot validates the commercial model, the next targets are mid-size European pharma companies and then the larger US and Japanese markets.
The cost structure is lean. The platform runs on a single GPU server, approximately 2,000 euros per month in cloud compute. The founder is the sole employee through the pilot phase. The Bpifrance Start-up grant covers incorporation costs, legal fees for the patent application, compute for the SKEMPI fine-tuning, and six months of runway to complete the Servier pilot. The total ask is within the 30,000 to 2.5 million euro range of the programme, with the specific amount to be determined in the application form based on the detailed budget.
CHECKLIST
- [ ] Confirm eligibility: verify Bpifrance Start-up programme requirements for non-French founders and planned incorporation timeline
- [ ] Prepare detailed budget breakdown: incorporation costs, legal fees, compute, runway, pilot expenses
- [ ] Draft one-page technical appendix with benchmark methodology and AUROC results
- [ ] Obtain letter of intent or expression of interest from Servier contact at Suresnes
- [ ] Obtain letter of support from Paris-Saclay I2BC or Institut Pasteur collaborator
- [ ] Prepare founder CV with pharmacy credentials and ML engineering portfolio
- [ ] Prepare pitch deck (10-15 slides) covering problem, technology, validation, market, roadmap
- [ ] Complete Bpifrance online application form at https://www.bpifrance.fr/start
- [ ] Verify incorporation timeline: confirm ability to incorporate in Ile-de-France before grant disbursement
- [ ] Prepare patent filing strategy document for the delta-embedding plus fingerprint method
- [ ] Confirm SKEMPI 2.0 dataset access and licensing terms for fine-tuning
- [ ] Prepare financial projections for 12, 24, and 36 months post-grant
EDITOR NOTES
- Eligibility risk: The applicant is not yet incorporated and is not French. Bpifrance Start-up typically requires a French-registered company. The application must explicitly state the commitment to incorporate in Ile-de-France and provide a timeline. Verify whether the programme allows grant disbursement to a company in formation or requires the incorporation to be completed first.
- Fact verification: The AUROC 0.804 on Platinum and 0.634 on SKEMPI 2.0 are from the applicant's own evaluation. The claim that mCSM-lig achieves approximately 0.70 AUROC needs a citation or verification against the published mCSM-lig paper. The Platinum benchmark's 553 mutations and protein-grouped CV protocol should be double-checked against the benchmark's official documentation.
- Servier pilot: The roadmap names Servier as a pilot partner, but no letter of intent or confirmed contact is listed. The application must include at least an expression of interest from Servier, or the roadmap should be revised to say "target pilot partner" rather than implying a confirmed engagement.
- Missing personal detail: The motivation letter and essays do not include the applicant's specific pharmacy background, years of experience, or ML engineering credentials beyond the one-line description. The applicant should insert a sentence in the motivation letter detailing their pharmacy practice context and their ML engineering training or portfolio, as this directly addresses Bpifrance's team execution criterion.
- SKEMPI 2.0 fine-tuning target: The roadmap claims AUROC 0.70 on SKEMPI 2.0 after fine-tuning is achievable. This is an assumption, not a validated result. The application should frame this as a target with a fallback plan if fine-tuning does not reach 0.70, such as ensemble methods or additional feature engineering, to avoid overpromising to the selection committee.