8%
of the human genome — roughly 200 million base pairs — remained completely unsequenced from the "complete" 2003 Human Genome Project announcement until the T2T-CHM13 assembly finished the job in 2022.
— T2T Consortium, Science, 2022

When the Human Genome Project was announced complete in 2003, it was a genuinely historic achievement — and also, technically, an overstatement. The reference genome it produced only covered the euchromatic fraction of the genome — the "easier," less repetitive regions. Large stretches of heterochromatin, especially the dense repetitive sequences found at centromeres and near chromosome tips, were left as gaps. Not low-quality data — no data. The technology of the era simply couldn't assemble them.

Why these regions resisted sequencing for 20 years

The missing 8% wasn't randomly scattered — it was concentrated in the genome's most repetitive architecture: centromeric satellite DNA, recent segmental duplications, and the short arms of five acrocentric chromosomes (13, 14, 15, 21, and 22). These regions consist of long stretches of nearly identical repeating sequence, sometimes for millions of base pairs at a stretch. Short-read sequencing — reading the genome in 150-300 base pair fragments — simply cannot tell one copy of a repeat from another when reassembling the puzzle. It's the same fundamental limitation covered in our long-read vs. short-read piece, just at its most extreme.

What finally solved it

The Telomere-to-Telomere (T2T) Consortium combined ultra-long-read sequencing (PacBio HiFi reads capable of resolving highly repetitive stretches) with a clever biological shortcut: they sequenced a complete hydatidiform mole cell line, a rare cell type that's nearly completely homozygous — meaning both copies of each chromosome are virtually identical, removing the added complexity of assembling two different parental haplotypes at once.

The result, published in 2022, was T2T-CHM13: the first truly complete, gapless human genome assembly, comprising just over 3.05 billion base pairs with zero gaps in every chromosome except Y (which was added in a 2023 update).

What T2T added

Nearly 200 million base pairs of previously invisible sequence, including 1,956 new gene predictions — around 99-140 with real protein-coding potential.

What it corrected

Beyond filling gaps, T2T fixed structural errors present in the prior reference genome (GRCh38) that had persisted, uncorrected, for years.

Does this actually change anything for real diagnoses?

Yes — measurably. A 2025 study using T2T-CHM13 as the reference for real clinical samples found it improved alignment quality and variant calling confidence, including better detection of rare and deleterious variants specifically valuable for rare-disease diagnostics — the exact population most likely to have a disease-causing variant hiding in a previously unmappable region.

Why this matters for you, practically: Most clinical and consumer WGS pipelines still primarily use the older GRCh38 reference genome, not T2T-CHM13, for compatibility and validation reasons. That's not wrong — GRCh38 remains extremely well-validated for the vast majority of the genome. But as T2T-based pipelines become more standard, some patients whose disease-causing variant sits in a previously unreadable region may finally get answers that older pipelines structurally couldn't provide.

Sequencing technology keeps getting more complete

Your raw WGS data from Dante Labs can be reprocessed against newer reference genomes like T2T-CHM13 as tools mature — meaning a sequencing test you take today doesn't lock you into today's interpretation ceiling.

Get Your Whole Genome Sequenced → Use code GENOME for 10% off at Dante Labs

The Human Genome Project didn't lie in 2003 — the euchromatic genome really was complete. But "complete genome" and "complete human genome" turned out to be two very different claims, and it took nearly two more decades of technological progress to close that gap for good.