REVIEW 3 major objections 4 minor 4 references
GREGoR: Accelerating Genomics for Rare Diseases
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The GREGoR consortium has built an openly shared, multi-omics cohort of about 7,500 people from over 3,000 rare-disease families, most of them exome-negative, and argues this will accelerate discovery of missing genetic diagnoses.
desk verdict A well-organized consortium resource paper whose central data-availability claim overreaches; fix the abstract and it's a solid pointer for rare-disease genomics. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the GREGoR cohort and its data model: thousands of deeply phenotyped families, each classified with Human Phenotype Ontology terms, connected to pedigree structure, and measured across multiple assay layers, with every variant assigned a machine-readable identifier. The data model is modular, built on accepted ontologies and common standards, and is submitted through open workspaces that run automated quality control and then release data under controlled access to the broader community. The mechanism carrying the argument is ascertainment plus depth: families were chosen because standard testing failed, and the layered data—especially the integration of transcriptomic and epigenetic outliers with underlying genomic variants—provides the independent lines of evidence needed to nominate and validate hard-to-detect causes such as noncoding, structural, mosaic, and multi-locus variation.
What would settle it
An independent re-analysis of a random sample of the released families without consortium guidance, followed by a comparison of the resulting diagnoses against the consortium's reported labels and against public evidence standards, would test the resource's reliability; a study that finds many released families were actually diagnosable from their prior exomes alone would weaken the 'exome-negative, unsolved' premise.
Extended reading notes
Core claim
The paper's central claim is that GREGoR has created a shared resource for rare-disease genomics: an openly shared cohort of previously undiagnosed families—about 7,500 individuals from over 3,000 families in the described release—most of them exome-negative, with data spanning exome and short-read genome sequencing, long-read sequencing, transcriptomes, methylation, and, for a subset, a full matrix of chromatin-accessibility, proteomic, and metabolomic measurements. The paper asserts that this combination of scale, prior negativity, multi-omics depth, and rapid sharing lets the community extract diagnoses that standard single-assay pipelines miss. It reports that consortium-led or contributed studies have implicated 365 genes, more than a third as novel disease-gene discoveries or phenotypic expansions, and that over 400 families now have candidate diagnoses, with solved cases intended to serve as positive controls and unsolved cases as a discovery substrate. On the paper's own terms, the dataset and the data model are the achievement: they are designed so that external researchers can benchmark tools, reanalyze families, and develop standards for when emerging technologies such as long-read sequencing and multi-omics should be used.
Load-bearing premise
The resource's value depends on the assumptions that these families are genuinely unsolved—that prior clinical testing was thorough enough for 'exome-negative' to mean something—and that the released data and the consortium's self-reported diagnoses are accurate, consented, and usable by outsiders.
Editorial extensions
If this is right
- The released cohort gives researchers thousands of exome-negative families to mine for new disease genes and for variant classes that standard pipelines miss, including noncoding, tandem-repeat, structural, and mosaic variants.
- The pairing of solved-case labels with unsolved families lets the community benchmark new tools and measure diagnostic yield against a real, difficult cohort rather than synthetic or easy cases.
- The consortium's comparisons of exome versus short-read genome sequencing, and short-read versus long-read technologies, provide evidence for when each assay should be used, including the finding that short-read genomes add roughly eight percent diagnostic yield over exomes.
- The standardized data model and rapid release allow external researchers to combine GREGoR data with other population datasets, turning individual-level rare findings into a quorum of evidence for gene-disease relationships.
Reading between the lines
- If the released cohort is as clean and representative as claimed, it could become a community benchmark for variant interpretation, but that status depends on an independent audit of a sample of the consortium's solved-case labels against public curation standards.
- A testable extension of the sharing model is to measure the community solve rate: how many of the released 'unsolved' families receive diagnoses from external groups who never spoke to the consortium, using the public data alone.
- The paper's discovery count of 365 genes is self-reported across 83 publications; an independent replication study would need to re-adjudicate a random subset to separate novel gene discoveries from phenotypic expansions and to verify functional support.
- The multi-omics matrix on a subset of families invites a direct comparative test of which assay layer—methylome, transcriptome, chromatin accessibility, proteome, or metabolome—contributes the most diagnoses for a given phenotype class, an analysis the paper describes as ongoing but does not complete.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript describes the GREGoR Consortium, a US NHGRI-funded effort to study thousands of rare disease families who remain undiagnosed after standard clinical testing, with emphasis on exome-negative cases. The paper summarizes the consortium's data generation strategy (short-read and long-read genome sequencing, transcriptomics, methylation, and other -omics), its computational and functional validation approaches, its data-sharing infrastructure (AnVIL, dbGaP:phs003047, seqr, Matchmaker Exchange, a public variant browser), and its reported scientific output (83 papers, molecular diagnoses in 365 genes, >400 solved families). The central claim is that GREGoR has created a large, openly shared, multi-omics resource that will catalyze rare disease research and provide positive controls for benchmarking.
Significance. If the resource is as described, it would be a valuable community asset: a cohort of ~3,000 families enriched for exome-negative unsolved cases, with deep phenotyping and layered genomic data, plus structured sharing through AnVIL. The paper's concrete accession numbers, URLs, and descriptions of the data model and validation workflows are strengths that make the resource independently checkable. The reported functional validation of a large subset of discoveries and the emphasis on diverse ancestry are also positive features. However, the significance hinges on the accuracy of the data availability claims and the reliability of the self-reported discovery counts; both currently have internal inconsistencies that need correction before the resource claims can be fully accepted.
major comments (3)
- [Abstract vs. ACCELERATING DATA SHARING] The abstract states that 'all data generated...is rapidly made available to researchers worldwide via AnVIL', but the Data Sharing section reports that currently only DNA data on ~7,400 individuals and transcriptome data on >500 individuals are available, with long-read genomes, long-read RNA-seq, Fiber-seq, ATAC-seq, metabolomics, and proteomics 'planned releases' within the next year. This is a direct internal inconsistency: not all generated data are currently available, and the multi-omics component of the central 'foundational resource' claim is not yet realized. Please revise the abstract to accurately reflect current versus planned availability, or justify how 'planned' data count as 'rapidly made available'.
- [CONCLUSION] The conclusion states that the resource 'currently is supporting data for over 3500 families', while the abstract, the Data Sharing section, and Figure 2 report approximately 3,000 families (Figure 2 caption gives n=3,059, the Data Sharing section says 'over 3000 families'). This discrepancy in the primary cohort size is load-bearing for the central quantitative claim; the numbers must be reconciled and a single authoritative cohort count provided with the data release date.
- [Supplementary Table 1] The claim of '83 papers studying molecular diagnoses in 365 genes' relies on Supplementary Table 1, which lists many genes with the same PMID (34582790) that does not appear in the reference list, and the table contains more than 365 rows while several rows correspond to duplicate genes or to the same PMID repeated across many entries. The provenance and counting rule (papers vs. genes vs. diagnoses) is therefore unverifiable. Please clarify how the 83 papers and 365 genes were counted, correct the table's PMID/attribution errors, or state explicitly that these counts are self-reported and not independently audited.
minor comments (4)
- [ACCELERATING DATA SHARING] The phrase 'Data is shared prior to analysis' in Figure 2's caption is ambiguous: it could mean raw data are released before the consortium's own analysis, or that phenotype data are shared at submission; please clarify whether the 'solved' labels are added after the initial release and how users should interpret unsolved cases in the current version.
- [EVALUATING EMERGING METHODS FOR ASCERTAINING RARE DISEASE DIAGNOSES] The statement '83 papers studying molecular diagnoses in 365 genes with more than a third being novel disease gene discoveries' appears to double-count several genes (e.g., AHDC1, CDKL5, HECTD4 appear multiple times). The manuscript should define whether the 365 genes are unique genes or gene-paper pairs.
- [Box 1 and main text] Minor typographical issues include 'Dvision' in the author affiliation and 'till today' instead of 'to this day' in the 'Reference Genomes' section; these should be corrected in the final version.
- [ACCELERATING DATA SHARING] The text states 'nearly 200 participants with both exome and short-read genome data' but does not define the overlap with the transcriptome numbers; a simple Venn diagram or table of data type overlaps would improve usability for potential data users.
Circularity Check
No circular derivation; descriptive consortium overview with independently checkable data releases.
full rationale
This manuscript is a descriptive consortium overview rather than a derivation chain. It contains no equations, fitted parameters, or formal predictions whose conclusions are presupposed by their inputs. The central claim, creation of a large openly shared rare disease genomics resource, is independently checkable through the stated dbGaP accession (phs003047), AnVIL, and the public GREGoR variant browser. The catalogue of 83 publications and 365 gene diagnoses is presented with PMIDs and functional-work annotations in Supplementary Table 1, meaning the evidence consists of externally published, peer-reviewed outputs rather than parameters fitted inside this paper. The abstract's wording that 'all data generated' is 'rapidly made available' is in tension with the Data Sharing section's statement that long-read genomes, Fiber-seq, ATAC-seq, metabolomics, and proteomics are planned releases 'within the next year'; this is an internal consistency or accuracy concern, but it is not a circular reduction because no claim is defined in terms of another claim's truth. Self-citation is pervasive, with consortium members citing consortium-authored papers, but no load-bearing argument reduces to a self-citation chain: the resource claim is externally verifiable, the publication counts are externally published, and no uniqueness theorem or ansatz is imported from the authors' prior work to force the conclusion. Accordingly, there is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Rare diseases affect approximately one in twenty individuals worldwide.
- domain assumption More than half of individuals suspected to have a rare disease lack a genetic diagnosis.
- domain assumption The GREGoR cohort is composed of families that remained unsolved after prior clinical genetic testing, with most being exome-negative.
Cite this review
Pith. "Pith review of GREGoR: Accelerating Genomics for Rare Diseases." pith.science (2026). https://pith.science/paper/AKSWMTVN
@misc{pith2026241214338,
author = {Pith},
title = {Pith review of: GREGoR: Accelerating Genomics for Rare Diseases},
year = {2026},
howpublished = {\url{https://pith.science/paper/AKSWMTVN}},
note = {Machine review of arXiv:2412.14338}
}
read the original abstract
Rare diseases are collectively common, affecting approximately one in twenty individuals worldwide. In recent years, rapid progress has been made in rare disease diagnostics due to advances in DNA sequencing, development of new computational and experimental approaches to prioritize genes and genetic variants, and increased global exchange of clinical and genetic data. However, more than half of individuals suspected to have a rare disease lack a genetic diagnosis. The Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) Consortium was initiated to study thousands of challenging rare disease cases and families and apply, standardize, and evaluate emerging genomics technologies and analytics to accelerate their adoption in clinical practice. Further, all data generated, currently representing ~7500 individuals from ~3000 families, is rapidly made available to researchers worldwide via the Genomic Data Science Analysis, Visualization, and Informatics Lab-space (AnVIL) to catalyze global efforts to develop approaches for genetic diagnoses in rare diseases (https://gregorconsortium.org/data). The majority of these families have undergone prior clinical genetic testing but remained unsolved, with most being exome-negative. Here, we describe the collaborative research framework, datasets, and discoveries comprising GREGoR that will provide foundational resources and substrates for the future of rare disease genomics.
Reference graph
Works this paper leans on
-
[110]
Wenger, T. L. et al. SeqFirst: Building equity access to a precise genetic diagnosis in critically ill newborns. Preprint at https://doi.org/10.1101/2024.09.30.24314516 (2024). 111. Wojcik, M. H. et al. Genome Sequencing for Diagnosing Rare Diseases. N. Engl. J. Med. 390, 1985–1997 (2024). 112. Saad, A. K. et al. Biallelic in-frame deletion in TRAPPC4 in ...
-
[146]
Melnikov, A. et al. Systematic dissection and optimization of inducible enhancers in human cells using a massively parallel reporter assay. Nat. Biotechnol. 30, 271–277 (2012). 147. Scott, H. A. et al. A high throughput splicing assay to investigate the effect of variants of unknown significance on exon inclusion. Preprint at https://doi.org/10.1101/2022....
arXiv 2012
-
[183]
Gifford, C. A. et al. Oligogenic inheritance of a human heart disease involving a genetic modifier. Science 364, 865–870 (2019). 184. Bozkurt-Yozgatli, T. et al. Multilocus pathogenic variants contribute to intrafamilial clinical heterogeneity: a retrospective study of sibling pairs with neurodevelopmental disorders. BMC Med. Genomics 17, 85 (2024). 185. ...
work page 2019
-
[217]
Young, J. L. et al. Beyond race: Recruitment of diverse participants in clinical genomics research for rare disease. Front. Genet. 13, 949422 (2022). 218. Wojcik, M. H. et al. Rare diseases, common barriers: disparities in pediatric clinical genetics outcomes. Pediatr. Res. 93, 110–117 (2023). 219. Serrano, J. G. et al. Advancing Understanding of Inequiti...
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.