Programme Thesis
This NSF programme funds projects that unlock the value of existing datasets for AI-enabled scientific discovery, aiming to accelerate research by making data AI-ready, developing novel AI methods for data analysis, and fostering community standards. It exists to bridge the gap between the growing availability of large datasets and the AI tools needed to extract scientific insights from them.
Selection Criteria
- Intellectual merit: potential to advance knowledge in the field, novelty of the approach, and significance of the proposed contribution to AI-enabled discovery.
- Broader impacts: benefits to society, contributions to underrepresented groups, and advancement of scientific infrastructure (e.g., data sharing, community standards).
- Data value: clarity on how the project will unlock value from existing datasets, including data curation, annotation, or integration.
- AI methodology: soundness and innovation of the AI/ML methods proposed, including reproducibility and validation.
- Feasibility: qualifications of the PI, adequacy of resources, and realistic work plan.
- Alignment with NSF mission: fit with the programme's goals of enabling discovery through data and AI.
Past Winners / Cohort Profiles
Past winners are typically academic researchers (faculty, postdocs, or senior researchers) at US institutions, often with strong computational backgrounds and access to large datasets. They may be from fields like biology, chemistry, materials science, or social science, and their projects often involve creating shared datasets, benchmarks, or AI models that serve a broad community. Named examples are not available from the page content, but typical archetypes include data scientists, ML researchers, and domain scientists who propose to curate and analyze existing data to answer fundamental questions.
Ideal Candidate Fingerprint
The ideal applicant is a researcher with a proven track record in both domain science and AI/ML, who can articulate a clear vision for how a specific existing dataset can be transformed into a valuable resource for AI-driven discovery. They have strong technical skills in data engineering, ML, and reproducibility, and they propose a project with broad community impact, such as creating a benchmark or a shared model. They are likely affiliated with a US institution and have access to computational resources.
Recommended Framing
For Eniola, the strongest angle is to leverage the TOPOLOGIX project, which directly uses existing protein sequence and drug data (Platinum benchmark, SKEMPI 2.0) to build an AI model (ESM-2 + fingerprints + Random Forest) that predicts drug-resistance mutations. This aligns perfectly with the programme's goal of unlocking dataset value for AI-enabled discovery, as TOPOLOGIX demonstrates how to extract new predictive power from existing datasets, and its success (AUROC 0.804) shows a clear scientific contribution. Eniola should frame TOPOLOGIX as a case study for a broader methodology that can be applied to other datasets, emphasizing the potential for community-wide impact and the novel combination of protein language models with topological data analysis.
Watch Out
Eligibility: NSF grants typically require the PI to be affiliated with a US institution (university, non-profit, or for-profit). Eniola is an independent researcher and currently enrolled in a German university, which may make him ineligible as a PI. He would need a US-based collaborator or institution to submit the proposal. Additionally, the programme may have restrictions on funding for non-US citizens or residents, though NSF often allows foreign nationals if they are at US institutions. Competitive disadvantage: lack of a US institutional affiliation and limited prior NSF funding history may be a disadvantage. Also, the programme may prioritize projects with broader community engagement, which Eniola's independent work may not yet demonstrate.
Research History
2026-08-04 20:24 · medium confidence
2026-08-01 05:27 · medium confidence