An official website of the United States government

SPARK PubMed/PMC — Task 2: Open-Ended Exploratory Information Needs
United States National Library of Medicine
9 Registrants
$125,000
Build AI that transforms biomedical discovery
View additional information at NIH Challenges.
 90 days left to register
Challenge Image

Overview

Participants will build an end-to-end AI system that responds to complex biomedical questions by synthesizing evidence from PubMed® and PubMed Central® (PMC). Solutions must provide concise, objective, evidence-grounded responses with valid citations and relevant passages.

Timeline

  • Start: 9/15/2026 4:00 AM ET
    End: 1/2/2027 4:59 AM ET
  • 1/15/2027 5:00 AM ET
  • 2/1/2027 4:59 AM ET
  • 03/29/27
  • 03/30/27

Prize

Total cash prize: $125,000
Prize Description:

Prizes awarded under this Challenge will be paid by electronic funds transfer and may be subject to federal income taxes. HHS/NIH will comply with Internal Revenue Service withholding and reporting requirements, where applicable. Entities participating in this Challenge are encouraged, but not required, to obtain a free Unique Entity ID (UEI) via SAM.gov, as this will expedite prize payment. Additional information is available at sam.gov/content/entity-registration.

NIH/NLM reserves the right, in its sole discretion, to (a) cancel, suspend, or modify the Challenge, or any part of it, for any reason, and/or (b) not award any prizes if no submissions are deemed worthy.

NIH/NLM also reserves the right to validate submissions based on the Docker Images, repositories, and/or documentation provided by participants. Evidence of gaming the system will result in disqualification.

Use of Prize Funds
Participants are reminded that under this challenge announcement, NIH/NLM is awarding unrestricted cash prizes, not grants. The purpose of this Challenge is to reward innovation, not provide financial assistance, and NIH/NLM does not limit how winners may use cash prizes awarded to them.

For each individual task, a total pool of $125,000 will be split among the top two performing participants (whether an individual, team, or entity) using fixed percentage proportions:
Task 2
First Place (Winner): 80%--$100,000
Second Place: 20%----$25,000

Rules

Eligibility and Participation rules for challenges can be found here. For additional challenge-specific rules, download and view the following:

Terms & Conditions

Participants are solely responsible for ensuring their submissions and activities comply with all applicable laws, regulations, and policies. NIH cannot provide legal advice to third parties. Participants should consult their own legal counsel as necessary and appropriate.

Judging Criteria

Task 2: Open-Ended Biomedical Information Needs

Participants will build systems that answer broad, exploratory biomedical research questions by retrieving and synthesizing evidence from the challenge corpus of PubMed/PMC documents.

An evaluation set comprising approximately 1,000 natural-language questions will be provided to participants shortly before the end of the challenge period. Participants will apply their systems to the evaluation set and submit a response for each information need, including supporting citations and evidence passages.

Participant systems and responses will be evaluated against an expert-verified reference set of relevant documents and evidence.

This challenge task will use the following evaluation criteria:

•Grounded Nugget F1: Measures the precision and recall of correct, evidence-grounded nuggets.

•Conflict / Controversy Coverage: Measures whether systems appropriately identify and characterize conflicting findings, scientific uncertainty, competing hypotheses, or other areas of disagreement when present.

•Ease of Use: Assesses the effort required to deploy, configure, operate, and use the system in the standardized NIH cloud environment.

•Compute Time: Measures the computational time required to generate responses.

•Compute Effort: Measures the computational effort required to generate responses, including the resources and processing required by the system.

•Memory Footprint: Measures the memory and other computational resources required to operate the system and generate responses. T2QS will serve as the primary metric for determining final system rankings and identifying the overall winning systems. In the event of a tie, the additional evaluation criteria listed above may be used to differentiate systems and determine the final ranking.

Grounded Nugget F1 and Conflict / Controversy Coverage - Algorithmic Evaluation
100 points

Grounded Nugget F1: Measures the correctness and completeness of system responses with respect to the reference response and supporting evidence.

Conflict / Controversy Coverage: Measures whether systems appropriately identify and characterize conflicting findings, scientific uncertainty, competing hypotheses, or other areas of disagreement when present.

100 points
Total Score
100 points

Challenge Materials

Challenge Materials