{"id":"f3de34ec-7265-4a68-bc46-4c85bf3c252d","arxiv_id":"2412.14338","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The GREGoR Consortium describes its publicly shared dataset of roughly 7,500 individuals from over 3,000 unsolved rare disease families, along with the multi-omics and analytical approaches it is applying to increase diagnostic yield.","lead":"GREGoR is a US consortium that has assembled and shared genomic and multi-omics data from about 7,500 individuals in over 3,000 families with rare diseases that remained undiagnosed after standard testing. This paper describes the consortium's data, tools, and discoveries, and argues the resource will speed up rare disease gene and variant discovery.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 'all data...rapidly made available' conflicts with Data Sharing section's planned releases of long-read, ATAC-seq, metabolomics, and proteomics, overstating the current multi-omics resource.","rationale":"The reader's CONDITIONAL verdict is well supported: the paper is a status report whose central claim is the existence and availability of a foundational multi-omics dataset. My stress-test isolates the availability claim, which is the most concrete and checkable component of that central claim. The internal inconsistency between the abstract's 'all data generated ... is rapidly made available' and the Data Sharing section's 'planned releases' of several multi-omics data types is a genuine overstatement, not a matter of external consensus. The reader flagged this same discrepancy in their rationale, but their stated weakest assumption emphasized consent, quality, and exome-negative status rather than the availability contradiction; hence partial agreement. The proposed test—inspecting the actual dbGaP/AnVIL contents—would decisively confirm or refute the overstatement. If the data are present, the concern is resolved; if absent, the abstract and the 'multi-omics foundational resource' language need revision, consistent with the reader's CONDITIONAL recommendation. I do not see a more load-bearing objection: the discovery claims are self-reported but not central to the resource's value, and the cohort description is plausible and auditable through the same accession.","tokens_in":29176,"tokens_out":5727,"duration_ms":47257,"concrete_test":"Enumerate all currently released data types and file counts under dbGaP accession phs003047 and the GREGoR AnVIL workspaces as of the manuscript submission date, and compare with the abstract's 'all data generated ... is rapidly made available' and the Data Sharing section's current vs planned releases. If the deposited filesets contain only exome/short-read genome DNA (plus a small transcriptome subset) and lack long-read genomes, long-read RNA-seq, Fiber-seq, ATAC-seq, metabolomics, and proteomics, then the abstract overstates current availability and the central 'multi-omics resource' claim should be downgraded to a planned capability. If all listed data types are already in AnVIL, the inconsistency is resolved and the abstract is accurate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the dataset being a large, openly shared, multi-omics resource. The abstract states 'all data generated, currently representing ~7500 individuals from ~3000 families, is rapidly made available to researchers worldwide via AnVIL.' The Data Sharing section ('ACCELERATING DATA SHARING') states 'Currently, DNA data on approximately 7400 individuals from over 3000 families is available with transcriptome data available for over 500 individuals' and 'Within the next year, planned releases include additional short and long read genomes, short and long read RNA-seq, Fiber-seq, ATAC-seq, metabolomics and proteomics.' These statements are in direct tension: if long-read, ATAC-seq, Fiber-seq, metabolomics, and proteomics are only planned, then not 'all data generated' are currently available, and the resource as presently released is mostly genomic DNA plus a modest transcriptome set. The multi-omics component, which is part of the foundational-resource claim, is thus not yet realized. This is an internal inconsistency, not a disagreement with field consensus, and it is the most load-bearing because a reader's decision to use the resource as a multi-omics benchmark depends on actual data availability. The self-reported discovery counts are secondary; the availability claim is directly about the central resource.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript describes the GREGoR Consortium, a US NHGRI-funded effort to study thousands of rare disease families who remain undiagnosed after standard clinical testing, with emphasis on exome-negative cases. The paper summarizes the consortium's data generation strategy (short-read and long-read genome sequencing, transcriptomics, methylation, and other -omics), its computational and functional validation approaches, its data-sharing infrastructure (AnVIL, dbGaP:phs003047, seqr, Matchmaker Exchange, a public variant browser), and its reported scientific output (83 papers, molecular diagnoses in 365 genes, >400 solved families). The central claim is that GREGoR has created a large, openly shared, multi-omics resource that will catalyze rare disease research and provide positive controls for benchmarking.","tokens_in":29366,"tokens_out":4497,"duration_ms":42049,"significance":"If the resource is as described, it would be a valuable community asset: a cohort of ~3,000 families enriched for exome-negative unsolved cases, with deep phenotyping and layered genomic data, plus structured sharing through AnVIL. The paper's concrete accession numbers, URLs, and descriptions of the data model and validation workflows are strengths that make the resource independently checkable. The reported functional validation of a large subset of discoveries and the emphasis on diverse ancestry are also positive features. However, the significance hinges on the accuracy of the data availability claims and the reliability of the self-reported discovery counts; both currently have internal inconsistencies that need correction before the resource claims can be fully accepted.","major_comments":[{"comment":"The abstract states that 'all data generated...is rapidly made available to researchers worldwide via AnVIL', but the Data Sharing section reports that currently only DNA data on ~7,400 individuals and transcriptome data on >500 individuals are available, with long-read genomes, long-read RNA-seq, Fiber-seq, ATAC-seq, metabolomics, and proteomics 'planned releases' within the next year. This is a direct internal inconsistency: not all generated data are currently available, and the multi-omics component of the central 'foundational resource' claim is not yet realized. Please revise the abstract to accurately reflect current versus planned availability, or justify how 'planned' data count as 'rapidly made available'.","section":"Abstract vs. ACCELERATING DATA SHARING"},{"comment":"The conclusion states that the resource 'currently is supporting data for over 3500 families', while the abstract, the Data Sharing section, and Figure 2 report approximately 3,000 families (Figure 2 caption gives n=3,059, the Data Sharing section says 'over 3000 families'). This discrepancy in the primary cohort size is load-bearing for the central quantitative claim; the numbers must be reconciled and a single authoritative cohort count provided with the data release date.","section":"CONCLUSION"},{"comment":"The claim of '83 papers studying molecular diagnoses in 365 genes' relies on Supplementary Table 1, which lists many genes with the same PMID (34582790) that does not appear in the reference list, and the table contains more than 365 rows while several rows correspond to duplicate genes or to the same PMID repeated across many entries. The provenance and counting rule (papers vs. genes vs. diagnoses) is therefore unverifiable. Please clarify how the 83 papers and 365 genes were counted, correct the table's PMID/attribution errors, or state explicitly that these counts are self-reported and not independently audited.","section":"Supplementary Table 1"}],"minor_comments":[{"comment":"The phrase 'Data is shared prior to analysis' in Figure 2's caption is ambiguous: it could mean raw data are released before the consortium's own analysis, or that phenotype data are shared at submission; please clarify whether the 'solved' labels are added after the initial release and how users should interpret unsolved cases in the current version.","section":"ACCELERATING DATA SHARING"},{"comment":"The statement '83 papers studying molecular diagnoses in 365 genes with more than a third being novel disease gene discoveries' appears to double-count several genes (e.g., AHDC1, CDKL5, HECTD4 appear multiple times). The manuscript should define whether the 365 genes are unique genes or gene-paper pairs.","section":"EVALUATING EMERGING METHODS FOR ASCERTAINING RARE DISEASE DIAGNOSES"},{"comment":"Minor typographical issues include 'Dvision' in the author affiliation and 'till today' instead of 'to this day' in the 'Reference Genomes' section; these should be corrected in the final version.","section":"Box 1 and main text"},{"comment":"The text states 'nearly 200 participants with both exome and short-read genome data' but does not define the overlap with the transcriptome numbers; a simple Venn diagram or table of data type overlaps would improve usability for potential data users.","section":"ACCELERATING DATA SHARING"}],"recommendation":"major_revision","confidential_remarks":"The paper is a descriptive resource report rather than a methods or hypothesis-testing paper, which is appropriate for its target journal. The main concern is not the absence of derivations but that the central data-availability claim is internally inconsistent and the discovery statistics are not readily auditable. Both are fixable with careful revision. I would not recommend rejection, as the consortium and resource are likely to be of genuine value, but the revised version must either correct the abstract and numbers or explicitly qualify the current status of each data type."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a consortium status report and data release, not a science paper. If you read it expecting a new method or a novel biological finding, you'll be disappointed. The value is in the resource: ~7,400 individuals from >3,000 families, mostly exome-negative, with deep phenotyping, transcriptomes for >500, and a public variant browser with >95M variants. The data model with VRS IDs and ClinGen allele IDs, plus the AnVIL/dbGaP accession (phs003047) and the explicit commitment to pre-analysis release, are genuinely useful infrastructure markers. I'd point a trainee to this as a model for how to write an honest consortium overview, except for one stumble.\n\nThe stumble is real. The abstract says 'all data generated... is rapidly made available.' The Data Sharing section says current DNA data on ~7,400 individuals and transcriptome data for >500, with long-read genomes, Fiber-seq, ATAC-seq, metabolomics and proteomics as 'planned releases within the next year.' Those two statements don't reconcile. The multi-omics resource is a stated pillar of the consortium's contribution, but as of the release described, it's mostly genomic DNA plus a modest transcriptome set. This overstates the current state and should be corrected, either by softening the abstract or by listing what's unavailable yet. The reader's stress-test note got this right.\n\nWhat's otherwise good: the problem framing (exome-negative, unsolved families) is sound and well-referenced. The 83 papers and 365 genes are self-reported but traceable to PMIDs in the supplement, and 44 of the 83 have functional work, so the 'discoveries' claim is not empty, it's just not audited in this paper. There's no math to check and no statistical claim that I can falsify. The citation pattern is heavy on consortium output, which is expected for a consortium review; the external citations to gnomAD, All of Us, UK Biobank, etc. are there.\n\nWho is this for: anyone deciding whether to use the GREGoR data as a benchmark or discovery cohort. That audience gets a clear route to the data. The paper deserves a serious referee, but mainly to enforce accuracy of the availability claims and make the timeline of releases explicit. I'd accept it for review with a request to fix the abstract and to distinguish 'currently available' from 'generated to date.'","headline":"A well-organized consortium resource paper whose central data-availability claim overreaches; fix the abstract and it's a solid pointer for rare-disease genomics.","tokens_in":30107,"tokens_out":2024,"would_cite":true,"duration_ms":18674,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The GREGoR consortium has built an openly shared, multi-omics cohort of about 7,500 people from over 3,000 rare-disease families, most of them exome-negative, and argues this will accelerate discovery of missing genetic diagnoses.","keywords":["rare disease genomics","GREGoR consortium","unsolved families","exome-negative","multi-omics","data sharing","long-read sequencing","diagnostic yield"],"falsifier":"An independent re-analysis of a random sample of the released families without consortium guidance, followed by a comparison of the resulting diagnoses against the consortium's reported labels and against public evidence standards, would test the resource's reliability; a study that finds many released families were actually diagnosable from their prior exomes alone would weaken the 'exome-negative, unsolved' premise.","tokens_in":28958,"feed_emoji":"🧬","tokens_out":9666,"duration_ms":79529,"temperature":0.7,"pith_summary":"This paper reports that the GREGoR consortium has assembled and rapidly released a deeply phenotyped, multi-omics dataset of roughly 7,500 people from more than 3,000 families with rare diseases, the large majority of whom had prior clinical genetic testing—usually exome sequencing—that found no cause. The central claim is that this cohort, precisely because it consists of 'unsolved' and mostly exome-negative families, plus layered genome, transcriptome, methylation, and other molecular data, will act as a substrate for finding new disease genes, new variant types, and better diagnostic methods. The paper argues that the consortium's standardized, machine-readable data model and open controlled-access release make the resource usable by researchers worldwide, not just by the centers that generated it. A sympathetic reader would care because more than half of suspected rare-disease patients still lack a molecular diagnosis, and a shared cohort of unsolved families with rich data is a direct way to attack that gap. The paper backs the resource's value with 83 consortium-linked publications implicating 365 genes and with candidate diagnoses for more than 400 families.","feed_headline":"7,500 unsolved rare-disease cases open to global researchers","feed_subtitle":"Consortium shares exome-negative families with genome, RNA, and epigenetic data so the world can hunt missing diagnoses.","key_machinery":"The load-bearing object is the GREGoR cohort and its data model: thousands of deeply phenotyped families, each classified with Human Phenotype Ontology terms, connected to pedigree structure, and measured across multiple assay layers, with every variant assigned a machine-readable identifier. The data model is modular, built on accepted ontologies and common standards, and is submitted through open workspaces that run automated quality control and then release data under controlled access to the broader community. The mechanism carrying the argument is ascertainment plus depth: families were chosen because standard testing failed, and the layered data—especially the integration of transcriptomic and epigenetic outliers with underlying genomic variants—provides the independent lines of evidence needed to nominate and validate hard-to-detect causes such as noncoding, structural, mosaic, and multi-locus variation.","core_discovery":"The paper's central claim is that GREGoR has created a shared resource for rare-disease genomics: an openly shared cohort of previously undiagnosed families—about 7,500 individuals from over 3,000 families in the described release—most of them exome-negative, with data spanning exome and short-read genome sequencing, long-read sequencing, transcriptomes, methylation, and, for a subset, a full matrix of chromatin-accessibility, proteomic, and metabolomic measurements. The paper asserts that this combination of scale, prior negativity, multi-omics depth, and rapid sharing lets the community extract diagnoses that standard single-assay pipelines miss. It reports that consortium-led or contributed studies have implicated 365 genes, more than a third as novel disease-gene discoveries or phenotypic expansions, and that over 400 families now have candidate diagnoses, with solved cases intended to serve as positive controls and unsolved cases as a discovery substrate. On the paper's own terms, the dataset and the data model are the achievement: they are designed so that external researchers can benchmark tools, reanalyze families, and develop standards for when emerging technologies such as long-read sequencing and multi-omics should be used.","pith_inferences":["If the released cohort is as clean and representative as claimed, it could become a community benchmark for variant interpretation, but that status depends on an independent audit of a sample of the consortium's solved-case labels against public curation standards.","A testable extension of the sharing model is to measure the community solve rate: how many of the released 'unsolved' families receive diagnoses from external groups who never spoke to the consortium, using the public data alone.","The paper's discovery count of 365 genes is self-reported across 83 publications; an independent replication study would need to re-adjudicate a random subset to separate novel gene discoveries from phenotypic expansions and to verify functional support.","The multi-omics matrix on a subset of families invites a direct comparative test of which assay layer—methylome, transcriptome, chromatin accessibility, proteome, or metabolome—contributes the most diagnoses for a given phenotype class, an analysis the paper describes as ongoing but does not complete."],"forward_implications":["The released cohort gives researchers thousands of exome-negative families to mine for new disease genes and for variant classes that standard pipelines miss, including noncoding, tandem-repeat, structural, and mosaic variants.","The pairing of solved-case labels with unsolved families lets the community benchmark new tools and measure diagnostic yield against a real, difficult cohort rather than synthetic or easy cases.","The consortium's comparisons of exome versus short-read genome sequencing, and short-read versus long-read technologies, provide evidence for when each assay should be used, including the finding that short-read genomes add roughly eight percent diagnostic yield over exomes.","The standardized data model and rapid release allow external researchers to combine GREGoR data with other population datasets, turning individual-level rare findings into a quorum of evidence for gene-disease relationships."],"supporting_citations":[{"why":"Documents that most patients referred for clinical whole-exome sequencing remain undiagnosed, establishing the diagnostic gap the consortium targets.","marker":"[6]"},{"why":"Shows that periodic reanalysis of clinical exome data yields additional diagnoses, supporting the consortium's reanalysis strategy for previously negative exomes.","marker":"[17]"},{"why":"Provides the framework for choosing technologies after panel or exome testing is inconclusive, guiding GREGoR's move to genomes and multi-omics.","marker":"[107]"},{"why":"Reports that rapid first-line genome sequencing diagnosed 50 percent of intervention infants versus 10 percent in conventional care, supporting first-tier genome sequencing claims.","marker":"[110]"},{"why":"Large study of 822 families quantifying that 72 percent of genome-derived diagnoses were exome-detectable and that genomes added over eight percent yield, the main quantitative evidence for short-read genome utility.","marker":"[111]"},{"why":"Describes RNU4-2, a noncoding RNA gene whose frequent neurodevelopmental syndrome was discovered with cases from multiple GREGoR sites, showing the cohort's power for noncoding discoveries.","marker":"[75]"},{"why":"Describes the web-based analysis tool used to query GREGoR data and link multi-omics outliers to genomic variants, supporting the data model's usability.","marker":"[137]"},{"why":"Describes the matchmaking exchange through which GREGoR candidate genes are shared with the global community, supporting the claim that data sharing drives discovery.","marker":"[227]"}],"fun_headline_variants":["Rare-disease consortium shares 7,500 exome-negative cases","Open rare-disease data yields 400 candidate diagnoses","3,000 exome-negative families now open to researchers","Genomics consortium opens exome-negative cases to world","Exome-negative families get open multi-omics resource"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The resource's value depends on the assumptions that these families are genuinely unsolved—that prior clinical testing was thorough enough for 'exome-negative' to mean something—and that the released data and the consortium's self-reported diagnoses are accurate, consented, and usable by outsiders.","fun_headline_variants_meta":{"raw":{"variants":["Rare-disease consortium shares 7,500 exome-negative cases","Open rare-disease data yields 400 candidate diagnoses","3,000 exome-negative families now open to researchers","Genomics consortium opens exome-negative cases to world","Exome-negative families get open multi-omics resource"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001553,"raw_usage":{"total_tokens":6235,"prompt_tokens":1003,"completion_tokens":5232,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":5151}},"tokens_in":619,"tokens_out":5232,"duration_ms":32407,"temperature":1.0,"reasoning_tokens":5151,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:18:51.008564+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent re-analysis of a random sample of the released families without consortium guidance, followed by a comparison of the resulting diagnoses against the consortium's reported labels and against public evidence standards, would test the resource's reliability; a study that finds many released families were actually diagnosable from their prior exomes alone would weaken the 'exome-negative, unsolved' premise.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports that rapid first-line genome sequencing diagnosed 50 percent of intervention infants versus 10 percent in conventional care, supporting first-tier genome sequencing claims."}],"review_version":1}