An official website of the United States government

SPARK dbGaP — Task 1: Semantic Variable Normalization and Ontological Alignment
United States National Library of Medicine
15 Registrants
$125,000
Build AI that enables cross study biomedical discovery
View additional information at NIH Challenges.
 90 days left to register
Challenge Image

Overview

dbGaP contains vast genomic and phenotypic metadata written inconsistently across studies. This task asks solvers to build AI systems that map raw, unstructured variables to standardized ontology concepts, enabling precise, interoperable study discovery and cohort evaluation.

Timeline

  • Start: 9/15/2026 4:00 AM ET
    End: 1/2/2027 4:59 AM ET
  • 1/15/2027 5:00 AM ET
  • 2/1/2027 4:59 AM ET
  • 03/29/27
  • 03/30/27

Prize

Total cash prize: $125,000
Prize Description:

Prizes awarded under this Challenge will be paid by electronic funds transfer and may be subject to federal income taxes. HHS/NIH will comply with Internal Revenue Service withholding and reporting requirements, where applicable. Entities participating in this Challenge are encouraged, but not required, to obtain a free Unique Entity ID (UEI) via SAM.gov, as this will expedite prize payment. Additional information is available at sam.gov/content/entity-registration.

NIH/NLM reserves the right, in its sole discretion, to (a) cancel, suspend, or modify the Challenge, or any part of it, for any reason, and/or (b) not award any prizes if no submissions are deemed worthy.

NIH/NLM also reserves the right to validate submissions based on the Docker Images, repositories, and/or documentation provided by participants. Evidence of gaming the system will result in disqualification.

Use of Prize Funds
Participants are reminded that under this challenge announcement, NIH/NLM is awarding unrestricted cash prizes, not grants. The purpose of this Challenge is to reward innovation, not provide financial assistance, and NIH/NLM does not limit how winners may use cash prizes awarded to them.

For each individual task, a total pool of $125,000 will be split among the top two performing participants (whether an individual, team, or entity) using fixed percentage proportions:
Task 1
First Place (Winner): 80%--$100,000
Second Place: 20%----$25,000

Rules

Eligibility and Participation rules for challenges can be found here. For additional challenge-specific rules, download and view the following:

Terms & Conditions

Participants are solely responsible for ensuring their submissions and activities comply with all applicable laws, regulations, and policies. NIH cannot provide legal advice to third parties. Participants should consult their own legal counsel as necessary and appropriate.

Judging Criteria

Evaluation Metrics

Task 1: Semantic Variable Normalization and Ontological Alignment

An evaluation set comprising of 1,000 dbGaP variables will be provided to participants shortly before the end of the challenge period. Participants will apply their systems to the evaluation set and generate a submission file containing predicted mappings from those study variables to corresponding target source-vocabulary concepts nodes. Each submission will then be evaluated against a ground-truth set of mappings established by human subject-matter experts. Performance will be measured using a Mapping Alignment Accuracy (MAA) score, which is defined as the Percentage of dbGaP study variables for which the system identifies the expert-verified target concept(s).

MAA will serve as the primary metric for determining system ranking. Other evaluation criteria will also be considered:

•MAA score: Measures the Percentage of dbGaP study variables for which the system identifies the expert-verified target concept(s)

•Ease of Use: Assesses the effort required to deploy, configure, operate, and use the system in the standardized NIH cloud environment.

•Compute Time: Measures the computational time required to generate responses.

•Compute Effort: Measures the computational effort required to generate responses, including the resources and processing required by the system.

•Memory Footprint: Measures the memory and other computational resources required to operate the system and generate responses.

MAA will serve as the primary metric for determining final system rankings and identifying the overall winning systems. In the event of a tie in this metric, the additional evaluation criteria listed above may be used to differentiate systems and determine the final ranking.

Task 1 Evaluation Metrics - Algorithmic Evaluation
100 points

A Mapping Alignment Accuracy (MAA) score

Mapping Alignment Accuracy (MAA) score, which is defined as the Percentage of dbGaP study variables for which the system identifies the expert-verified target concept(s).

100 points
Total Score
100 points

Challenge Materials

Challenge Materials