A synthesis of current strategies in non-invasive prenatal testing, multi-cancer early detection and molecular residual disease.
Background: finding a small signal in a sea of similar DNA
Cell-free DNA (cfDNA) is the population of short, typically histone-bound DNA fragments circulating in plasma. It is now the substrate for prenatal screening, cancer detection and monitoring, transplant surveillance and infectious disease diagnosis, and in every one of those settings the analytical problem is the same: a diagnostically interesting minority population must be measured against an overwhelming and chemically identical background. The concentrations of cfDNA are low — 1 mL of plasma contains a median of 1,645 copies of fragmented genomic DNA (5.43 ng/mL) as determined by validated short-amplicon qPCR from 104 healthy donors [1].
Most cfDNA originates from dying cells. During apoptosis, chromatin is cut by nucleases at locations that are accessible between nucleosomes while the DNA wrapped around the histone core is shielded. As a result, what reaches the blood is a ladder of nucleosome-sized pieces. The dominant fragment length is about 167 base pairs, the length protected by a nucleosome (147 bp) plus the linker DNA associated with histone H1 (~20 bp). The fragments reflect the transcriptional state of the cell that they were released from and suggest tissue origin [2]. Sequencing across seven healthy donors shows that 67.5% to 80% of plasma cfDNA is mononucleosomal, that most of these fragments carry nicks, and that the cfDNA of a cancer patient is shifted toward shorter lengths (Figure 1) [3]. In healthy individuals, cfDNA originates from white blood cells (55%), erythrocyte progenitors (30%), vascular endothelial cells (10%) and hepatocytes (1%) [4]. Every clinical use of cfDNA therefore comes down to the same task: a minority population of DNA – placental in pregnancy or tumor-derived in cancer – must be measured against a chemically identical background that is overwhelmingly derived from blood cells.

The strategies now in clinical use differ mainly in which property of the minority population they read. Counting methods ask whether a chromosome contributes too many (or too few) fragments. Genotyping methods ask which alleles are present. Molecular counting methods ask how many molecules of each species were actually in the tube. Methylation methods ask which tissue the fragments came from and whether that tissue looks like cancer. The sections that follow take one example of each.
Non-invasive prenatal testing (NIPT): from counting to genotyping to counting molecules
Fetal DNA was demonstrated in maternal plasma in 1997 by PCR for a Y-chromosome sequence, found in 24 of 30 women carrying male fetuses and none of 13 carrying female fetuses [5]. A decade later two groups showed independently that shotgun sequencing of maternal plasma could detect fetal trisomy: reads are aligned, counted by chromosome, and a placental trisomy 21 appears as a small excess of chromosome 21 fragments over a reference distribution [6, 7]. A weakness in this approach is that fetal cfDNA comes from the placental tissue and not the fetus itself. If the mother is pregnant with twins, and one of those embryos was not viable due to aneuploidy or an embryonic lethal mutation, the twin will be resorbed but the placental tissue will still contain DNA from the ‘vanished twin’ for months, leading to false positives [8]. This counting approach, now run at shallow whole-genome depth, underlies many commercial NIPT products. Its strength is breadth, since it sees any chromosome. Its weakness is that counting precision is set by the number of fragments counted, and that it cannot tell a vanished twin or a triploid pregnancy from a fetal trisomy because it cannot determine the source of the DNA it is counting.
Natera’s Panorama test, published in 2012, asks a different question. Counting asks whether there is too much of chromosome 21 in the plasma; the SNP method asks which alleles the fetus inherited at each of thousands of polymorphic sites, and how many copies of each chromosome those alleles add up to. A single multiplex PCR amplifies about 11,000 single-nucleotide polymorphisms (SNPs) on chromosomes 13, 18, 21, X and Y, and the plasma is sequenced together with the mother’s own white-cell DNA. At a site where the mother is homozygous, any second allele in the plasma must be fetal and paternally inherited, and its fraction of the reads is half the fetal fraction; at a site where she is heterozygous, the fetal contribution tilts the 50:50 maternal ratio one way or the other. How far each ratio tilts, and in which direction, depends on whether the fetus carries one, two or three copies of that chromosome and on which parent the extra copy came from. The algorithm computes the likelihood of the observed allele ratios across all the SNPs on a chromosome under each of those hypotheses and reports the most likely one, together with a sample-specific confidence [9]. The answer is therefore a copy-number call for each of the five chromosomes and a fetal sex call, with an error estimate attached to each. In the original series of 166 samples, 145 passed a DNA quality filter and the algorithm made 725 of 725 correct chromosome calls among them, with an average calculated accuracy of 99.92%; the 21 samples that failed the filter received no call rather than a wrong one [9]. Because it genotypes rather than counts, the SNP method sees things counting cannot, such as a vanishing twin, an unrecognized twin or a diandric triploid [10].
The limit of allele-ratio genotyping appears when the question changes from aneuploidy to single-gene disease. Where the mother carries a recessive variant, the plasma variant allele fraction stays at 0.50 if the fetus is heterozygous like her, but rises only to 0.55 at 10% fetal fraction and 0.60 at 20% fetal fraction if the fetus is homozygous affected [11]. Distinguishing 0.50 from 0.55 requires knowing how many molecules of each allele were present, and read depth does not report that. PCR efficiency varies by locus, by sample and at random, so the relationship between molecules in and reads out is scrambled during library preparation; as the developers put it, because the major sources of noise are introduced during amplification and library preparation, increased sequencing depth does not necessarily result in improved accuracy [11].
BillionToOne’s answer, quantitative counting templates (QCTs), changes what amplicon sequencing can do: it turns an ordinary multiplex PCR run on an ordinary sequencer into an absolute molecular counter. A QCT is a synthetic double-stranded DNA molecule that shares the target’s primer-binding sites and amplicon length, so it co-amplifies at the same rate, but differs internally by a five-base identifier and a stretch of ten randomized bases, the ‘embedded molecular index’ or EMI. Roughly 100 to 1,000 QCT molecules are spiked into each PCR from a pool of up to 410 sequences, so two QCTs sharing an index is very unlikely. After sequencing, index sequences differing by two or fewer mismatches are clustered and the clusters counted, which recovers how many QCT molecules were actually pipetted into the tube without assuming a dilution. Two numbers now come out of the same sequencing run: the total number of reads from the QCT amplicon, and the number of QCT molecules that produced them, counted from their distinct EMI indices. Dividing the first by the second gives the number of reads that one input molecule generated in that reaction, the amplification factor. Because the QCT was built to amplify at the same rate as the target, that factor applies to the target as well, and dividing the target’s total reads by it gives the number of target molecules that were in the tube, expressed as haploid genome equivalents [11]. Chemists will recognize the similarity to an internal isotopic standard used in quantitative mass spectrometry: a chemically matched analyte, subject to the same losses, is spiked in at a known amount.
Natera and BillionToOne use the sequencing reads for different purposes. Both tests amplify their targets in one multiplex PCR reaction and sequence the products, and both avoid the adapter-ligation steps that lose most of a scarce sample and neither needs a paternal sample. Natera’s algorithm compares the two alleles at each SNP, which sit in the same amplicon and amplify at the same rate, so the ratio between them survives PCR bias even though the absolute read counts do not, and by combining thousands of such ratios along a chromosome it infers fetal copy number and, when more than one fetal genome is present, whose DNA is in the sample. The number of molecules behind those ratios is estimated statistically rather than measured. BillionToOne’s reads answer an absolute question. The QCT internal standard converts read counts into molecule counts, so the assay knows how many copies of each allele it actually sampled and can report a confidence for each call from counting statistics. In the validation work, the replicate-to-replicate variation in allele fraction was at the limit that the number of molecules allows [11]. This makes the small ratio shift of recessive single-gene testing measurable, where allele-ratio genotyping alone runs out of precision. Two further consequences follow from counting molecules. The assay reports its input in genome equivalents rather than nanograms, which matters because fragmented cfDNA yields far fewer amplifiable molecules per nanogram than its mass suggests, and the random index set in each reaction doubles as a sample fingerprint, so cross-contamination between samples is measured on every run rather than assumed from the indexing scheme [11].
The clinical consequence is a workflow that did not previously exist. Among 42,067 pregnant women screened for carrier status, 7,538 carriers (17.9%) proceeded to fetal testing without a partner sample for cystic fibrosis, spinal muscular atrophy, α-thalassemia and β-hemoglobinopathies. In 528 pregnancies with outcomes, including 25 affected, the fetal assay showed 96.0% sensitivity, 95.2% specificity and a negative predictive value of 99.8%, and every pregnancy assigned a fetal risk of 9 in 10 was affected [12].
Multi-cancer early detection: GRAIL’s bet on methylation
Multi-cancer early detection (MCED) asks whether any cancer signal exists in an asymptomatic person and, if so, what tissue it came from. GRAIL’s decision to read methylation rather than mutations was made empirically. In the first Circulating Cell-free Genome Atlas (CCGA) substudy, ten machine-learning classifiers were trained on the same plasma samples and validated independently, each built on a different cfDNA feature: whole-genome methylation, single-nucleotide variants with and without paired white-blood-cell background removal, copy-number changes, fragment length, fragment end positions, allelic imbalance, and a combination of all scores. With specificity fixed at 98%, the whole-genome methylation classifier was among the most sensitive. The mutation classifier matched it only when the patient’s white blood cell DNA had been sequenced and its variants removed. Methylation was also the best predictor of the tissue of origin [13]. Two conclusions follow. A methylation classifier matches a mutation classifier without sequencing the patient’s white cells, because clonal hematopoiesis is a mutation problem – blood-cell clones do not acquire a tumor’s methylation pattern [14]. The same study found that how well a classifier performed depended more on the fraction of cfDNA that came from the tumor (ctDNA) rather than on the cancer’s stage or type. Cancer stage matters only because later-stage tumors usually shed more DNA [13].
Methylation also suits the low-molecule regime. A tumor may lack a recurrent driver mutation, but it almost always carries thousands of altered CpG sites (cytosine-guanine dinucleotides, the positions at which mammalian DNA is methylated), and several CpGs are read on each fragment, so a single molecule can carry enough pattern to be assigned to a tissue. The Galleri assay uses bisulfite conversion, which deaminates unmethylated cytosine to uracil while leaving 5-methylcytosine intact, so that after amplification a methylated position reads as C and an unmethylated one as T. The converted DNA is sequenced across more than 100,000 informative methylation regions, and a classifier trained on cancer and non-cancer plasma returns a cancer-signal score and a predicted cancer signal origin [15]. In the second CCGA substudy of 6,689 participants, specificity was 99.3%; sensitivity in twelve pre-specified cancer types that account for about 63% of United States cancer deaths was 67.3% at stages I to III and rose from 39% at stage I to 92% at stage IV; and the tissue of origin was predicted for 96% of cancer-like signals and was correct in 93% of those [15].
What this approach changes is the scope of screening. Every established cancer screening test looks for one cancer in one organ, and recommended tests exist for only a handful of them: breast, cervix, colon and rectum, lung in heavy smokers, and prostate with reservations. Cancers of the pancreas, ovary, liver, esophagus, stomach and head and neck, and the blood cancers, have no screening test at all and are usually found once they cause symptoms, by which time most are advanced. A methylation classifier is not tied to any one organ. In the validation study it detected signals from more than 50 cancer types in people who had not yet been diagnosed, and when it found a signal it named the tissue of origin correctly in about nine of ten cases, which tells the physician where to look [16]. Sensitivity rises steeply with stage, from about one in six at stage I to nine in ten at stage IV, so the test is far better at finding a cancer that has begun to shed DNA than one that has not, but for the many cancers with no other screening route even that represents detection that would otherwise not happen [16].
The prospective studies bear this out. In PATHFINDER, the first deployment in adults over 50 with no signs or symptoms of cancer, three-quarters of the cancers the test found were of types for which no screening test is recommended, about half of the new diagnoses were stage I or II, and the predicted tissue of origin was correct in nearly all of them; the cost was that more people received a false signal than a true one, and false positives took a median of five months to resolve [17]. The randomized NHS-Galleri trial then asked, in 142,250 people over three annual rounds, whether adding the test to standard care shifts cancers to earlier stages at diagnosis. It did not meet its primary endpoint of fewer stage III and IV diagnoses across twelve pre-specified cancers, but stage IV diagnoses fell by 14% overall, and the effect grew with each round of screening, consistent with a test that misses a small cancer in one year and catches it in the next [18].
Molecular residual disease: how many loci, and whose tumor
Molecular residual disease (MRD) testing is a liquid biopsy – a blood test for tumor-derived DNA, applied after curative treatment. It asks whether tumor DNA persists, at a tumor fraction far lower than in advanced disease. The central design choice is whether to sequence the patient’s tumor first. Natera’s Signatera takes the tumor-informed route: the surgically removed tumor and matched normal are sequenced, a personalized multiplex PCR panel is built from clonal tumor mutations, and later plasma samples are interrogated only at those loci. In patients whose colorectal cancer had been removed with curative intent, a positive plasma result a month after surgery identified the patients who would go on to relapse, and a positive result during follow-up separated the two groups even more sharply; on average the blood test declared the recurrence about eight months before it could be seen on a scan [19]. In a cohort of more than a thousand patients, the same postoperative result did something more useful than predict relapse: it identified which patients with stage II or III disease went on to benefit from adjuvant chemotherapy, and by implication which ones were being treated for a cancer that was already gone [20]. That is the clinical meaning of residual-disease testing. It turns the decision to give or withhold chemotherapy after surgery from a judgment based on the tumor’s appearance under the microscope into one based on whether tumor DNA is still in the blood. The sensitivity that makes this possible comes from knowing which mutations to look for: the few thousand genome equivalents in a tube are the same whether one locus is examined or many, but each additional tracked mutation gives the assay another independent chance to catch a tumor molecule among them.
Guardant Health took the opposite route with their Guardant Reveal. Rather than building a personalized panel from each patient’s tumor, it reads a fixed panel of mutations and more than 20,000 methylation-informative regions directly from plasma, so the test can be run on any patient, from blood alone, within days of surgery, with no tissue sample and no waiting for a bespoke assay [21]. The cost of that convenience is sensitivity. Without knowing which mutations to look for, the assay must rely on methylation patterns shared across colorectal tumors, and although this approach caught about four in five recurrences of stage II or higher colon cancer over serial testing, with a lead over imaging of several months, it caught them unevenly: nearly every recurrence in the liver, but only about half of those in the lung and fewer than half in the peritoneum [21]. The unevenness is not a flaw in the chemistry. Tumors shed DNA into blood at different rates depending on where they grow, and a fixed panel with no patient-specific targets has less margin to detect a tumor that sheds little. A negative result from a tumor-naive test therefore means less for some recurrence patterns than for others, which is the trade a clinician accepts in exchange for a test that needs nothing but a tube of blood.
BillionToOne has carried its QCT counting chemistry into the same space. Its oncology assay, Northstar Response, is designed for patients who already have a diagnosed cancer and are on treatment. Its job is to answer, from one blood draw to the next, whether the amount of tumor DNA in the blood is going up or down. It applies QCTs to methylated tumor DNA so that changes in the number of methylated molecules, rather than variant allele fraction, are tracked during therapy; the reported coefficient of variation is under 10% at 1% tumor fraction, which is about half the variation of comparable tumor-naive panels reporting allele fraction. The assay reliably distinguished samples whose tumor fraction differed by 0.25 percentage points, across twelve solid tumor types [22]. The clinical cohort in that report is small and the assay is aimed at treatment-response monitoring rather than postoperative MRD, but it shows that absolute counting and methylation are complementary rather than competing choices.
Future directions of cell-free DNA diagnostics
The strategies that now define the field are answers to the two constraints set out at the beginning. Because molecules are scarce, the successful designs either count them directly (BillionToOne’s QCTs), pool the signal across many loci (Natera’s tumor-informed panels), or read a pattern rich enough that a few molecules suffice (GRAIL’s methylation classifier). Because the background is predominantly blood cell-derived DNA carrying its own mutations, successful designs either subtract that background by sequencing the patient’s white cells alongside the plasma or read a property the blood clones do not share. Methylation is one such property. Fragment length, illustrated in Figure 1, is another: a mutation changes a base call but does not move a fragment end, so fragment-based classifiers, though less sensitive than methylation alone, nicely complement it, and one company, Delfi Diagnostics, has developed a lung-cancer screening test on fragment length alone.
This article has focused on the best-known uses of cfDNA – NIPT, MCED, and MRD – but the strategies used can be applied wherever a minority population of DNA must be measured against the background of the body’s own genome. A transplanted organ is a genetically distinct tissue, and the same SNP genotyping that finds fetal cfDNA in maternal plasma can quantify a rising donor fraction when a transplant is rejected. Tumor genotyping using ctDNA can accurately capture a tumor’s genomic heterogeneity because ctDNA samples every shedding lesion at once and therefore may be used to find a druggable mutation shared by all of its clones. A pathogen carries a genome unlike any human sequence, so unbiased sequencing of plasma DNA can identify a bacterium, fungus, or virus without first guessing which one to look for. In every one of these settings the useful answer is the same – a count of molecules from one source against the background of all the others. The choices – what to sequence, how to count it accurately, and how much sample and depth the question truly demands – will decide what new applications reach the clinic.
References
- Meddeb R, Dache ZAA, Thezenas S, et al. Quantifying circulating cell-free DNA in humans. Sci Rep. 2019;9(1):5220. doi:10.1038/s41598-019-41593-4
- Snyder MW, Kircher M, Hill AJ, Daza RM, Shendure J. Cell-free DNA comprises an in vivo nucleosome footprint that informs its tissues-of-origin. Cell. 2016;164(1-2):57-68. doi:10.1016/j.cell.2015.11.050
- Sanchez C, Roch B, Mazard T, et al. Circulating nuclear DNA structural features, origins, and complete size profile revealed by fragmentomics. JCI Insight. 2021;6(7):e144561. doi:10.1172/jci.insight.144561
- Moss J, Magenheim J, Neiman D, et al. Comprehensive human cell-type methylation atlas reveals origins of circulating cell-free DNA in health and disease. Nat Commun. 2018;9(1):5068. doi:10.1038/s41467-018-07466-6
- Lo YMD, Corbetta N, Chamberlain PF, et al. Presence of fetal DNA in maternal plasma and serum. Lancet. 1997;350(9076):485-487. doi:10.1016/S0140-6736(97)02174-0
- Fan HC, Blumenfeld YJ, Chitkara U, Hudgins L, Quake SR. Noninvasive diagnosis of fetal aneuploidy by shotgun sequencing DNA from maternal blood. Proc Natl Acad Sci U S A. 2008;105(42):16266-16271. doi:10.1073/pnas.0808319105
- Chiu RWK, Chan KCA, Gao Y, et al. Noninvasive prenatal diagnosis of fetal chromosomal aneuploidy by massively parallel genomic sequencing of DNA in maternal plasma. Proc Natl Acad Sci U S A. 2008;105(51):20458-20463. doi:10.1073/pnas.0810641105
- Kleinfinger P, Luscan A, Descourvieres L, et al. Noninvasive prenatal screening for trisomy 21 in patients with a vanishing twin. Genes (Basel). 2022;13(11):2027. doi:10.3390/genes13112027
- Zimmermann B, Hill M, Gemelos G, et al. Noninvasive prenatal aneuploidy testing of chromosomes 13, 18, 21, X, and Y, using targeted sequencing of polymorphic loci. Prenat Diagn. 2012;32(13):1233-1241. doi:10.1002/pd.3993
- Curnow KJ, Wilkins-Haug L, Ryan A, et al. Detection of triploid, molar, and vanishing twin pregnancies by a single-nucleotide polymorphism-based noninvasive prenatal test. Am J Obstet Gynecol. 2015;212(1):79.e1-79.e9. doi:10.1016/j.ajog.2014.10.012
- Tsao DS, Silas S, Landry BP, et al. A novel high-throughput molecular counting method with single base-pair resolution enables accurate single-gene NIPT. Sci Rep. 2019;9(1):14382. doi:10.1038/s41598-019-50378-8
- Wynn J, Hoskovec J, Carter RD, Ross MJ, Perni SC. Performance of single-gene noninvasive prenatal testing for autosomal recessive conditions in a general population setting. Prenat Diagn. 2023;43(10):1344-1354. doi:10.1002/pd.6427
- Jamshidi A, Liu MC, Klein EA, et al. Evaluation of cell-free DNA approaches for multi-cancer early detection. Cancer Cell. 2022;40(12):1537-1549.e12. doi:10.1016/j.ccell.2022.10.022
- Razavi P, Li BT, Brown DN, et al. High-intensity sequencing reveals the sources of plasma circulating cell-free DNA variants. Nat Med. 2019;25(12):1928-1937. doi:10.1038/s41591-019-0652-7
- Liu MC, Oxnard GR, Klein EA, Swanton C, Seiden MV; CCGA Consortium. Sensitive and specific multi-cancer detection and localization using methylation signatures in cell-free DNA. Ann Oncol. 2020;31(6):745-759. doi:10.1016/j.annonc.2020.02.011
- Klein EA, Richards D, Cohn A, et al. Clinical validation of a targeted methylation-based multi-cancer early detection test using an independent validation set. Ann Oncol. 2021;32(9):1167-1177. doi:10.1016/j.annonc.2021.05.806
- Schrag D, Beer TM, McDonnell CH III, et al. Blood-based tests for multicancer early detection (PATHFINDER): a prospective cohort study. Lancet. 2023;402(10409):1251-1260. doi:10.1016/S0140-6736(23)01700-2
- GRAIL Inc. GRAIL reports full results from NHS-Galleri trial demonstrating substantial reduction in stage IV cancer diagnoses at 2026 ASCO Annual Meeting. Press release. June 2026. Available from GRAIL (sponsor release of a conference report; not peer reviewed at time of writing)
- Reinert T, Henriksen TV, Christensen E, et al. Analysis of plasma cell-free DNA by ultradeep sequencing in patients with stages I to III colorectal cancer. JAMA Oncol. 2019;5(8):1124-1131. doi:10.1001/jamaoncol.2019.0528
- Kotani D, Oki E, Nakamura Y, et al. Molecular residual disease and efficacy of adjuvant chemotherapy in patients with colorectal cancer. Nat Med. 2023;29(1):127-134. doi:10.1038/s41591-022-02115-4
- Nakamura Y, Tsukada Y, Matsuhashi N, et al. Colorectal cancer recurrence prediction using a tissue-free epigenomic minimal residual disease assay. Clin Cancer Res. 2024;30(19):4377-4387. doi:10.1158/1078-0432.CCR-24-1651
- Ye PP, Viens R, Shelburne KE, et al. Molecular counting enables accurate and precise quantification of methylated ctDNA for tumor-naive cancer therapy response monitoring. Sci Rep. 2025;15(1):5869. doi:10.1038/s41598-025-90013-3