{"id":"169e95a6-313c-453b-a3ec-08ff25b41b6f","arxiv_id":"2412.06666","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Diff5T provides the first public 5.0T diffusion MRI dataset with raw k-space data, multi-shell diffusion images, and structural scans from 50 subjects.","lead":"This paper introduces Diff5T, a new open dataset of 5.0 Tesla human brain diffusion MRI scans from 50 healthy volunteers, including raw k-space data and reconstructed images. It is meant to support the development of image reconstruction, preprocessing, and diffusion modelling methods at an intermediate field strength.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Full Diff5T dataset is not yet accessible: Usage Notes state data 'will be gradually uploaded,' and the GitHub holds only example data, so the central open-benchmark claim is currently unverifiable.","rationale":"I read the strongest claim as asserting the existence and public availability of a first-of-kind 5T dMRI dataset with raw k-space. The single most load-bearing condition for that claim to hold is that the full 50-subject dataset is genuinely accessible. The paper's Usage Notes directly state that data will be uploaded only gradually and that only example data are currently on GitHub; the Data Records section offers no counts or manifest to substantiate the 50-subject claim. This makes the central claim currently unverifiable, matching the reader's weakest assumption. I considered whether an internal technical inconsistency might be more important, but none of the protocol numbers are contradictory: 90+90+90+21=291 volumes, and the 21 b0 images split as 15 PA plus 6 AP, consistent with the three-series description. The absence of quantitative reconstruction-quality metrics is a real secondary weakness, but it would not invalidate the dataset's existence if the data were accessible; therefore it is not the most load-bearing concern. Because the reader already made the verdict CONDITIONAL on data release, my read does not change that verdict; I agree with the reader's assessment and recommend keeping the conditional status until the full dataset is made available and its contents verified.","tokens_in":15416,"tokens_out":6342,"duration_ms":62600,"concrete_test":"At the time of review, visit the GitHub repository (https://github.com/ShoujunYu/Diff5T) and any linked data archive; attempt to download the complete Diff5T dataset. Verify whether all 50 subjects are present, whether each subject includes the 291 diffusion volumes (90 each at b=1000/2000/3000 plus 21 b0) and raw/reconstructed k-space files in .mat format, and whether the file sizes are consistent with a 176x176 matrix, 114 slices, and 48 channels. If only example data are available or the repository lacks a full inventory, the open-access claim is not met.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that Diff5T is an open 5.0T human brain diffusion MRI dataset containing raw k-space and reconstructed images for 50 subjects. For a dataset descriptor, the load-bearing requirement is that the data are actually available to the community. The manuscript's own Usage Notes undermine this: 'The dataset is publicly accessible, and the data along with the processing code will be gradually uploaded and updated in the coming period. The example data and processing code can be accessed at the GitHub repository.' Thus at publication time the full 50-subject dataset is not available; only example data are online. The Data Records section describes four data types but provides no file inventory, no per-subject manifest, no data sizes, and no repository DOI, so a reader cannot verify from the paper that 50 subjects' k-space and image data exist in a public archive. Without the full data, the first-of-its-kind open benchmark claim cannot be independently checked. A secondary but related weakness is that Technical Validation is purely qualitative (visual inspection of Fig. 2 and one RMS motion plot in Fig. 3b), with no quantitative agreement metrics between the offline reconstruction and vendor reconstruction; this weakens the 'benchmarking' utility claim but is less fundamental than the absence of the data itself.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript describes Diff5T, a 5.0 Tesla human brain diffusion MRI dataset collected from 50 healthy subjects on a United Imaging uMR Jupiter scanner. It includes raw k-space data, reconstructed k-space data, reconstructed images, and DICOM-to-NIfTI images for multi-shell, multi-direction dMRI (b = 1000, 2000, 3000 s/mm2, 90 directions per shell), together with 0.5 mm isotropic T1w and T2w structural images. The paper details the acquisition protocol, an offline reconstruction pipeline, a preprocessing pipeline, and example downstream analyses using DTI, NODDI, MSMT-CSD, and tractography. The stated goal is to provide the first open 5.0T brain dMRI dataset with raw k-space and to serve as a benchmarking resource for reconstruction and diffusion modeling.","tokens_in":15651,"tokens_out":6514,"duration_ms":69060,"significance":"If the full 50-subject dataset is actually released as described, Diff5T would fill a genuine gap: existing open dMRI datasets are predominantly 3T or 7T and typically do not include raw k-space. The 5T field strength, the multi-shell 90-direction acquisitions, and the paired k-space/images would be valuable for AI-based reconstruction, q-space processing, and artifact-correction research. The manuscript's strengths include detailed acquisition parameters, a step-by-step reconstruction and preprocessing pipeline, the use of publicly available software tools, and example outputs from standard diffusion models. However, the central open-access promise is currently not fulfilled: the Usage Notes say the data will be uploaded only gradually and that the GitHub repository contains only example data. In addition, the technical validation is qualitative, with visual inspection and a single motion plot rather than quantitative reconstruction or model-quality metrics. Both issues need to be resolved before the dataset descriptor can be accepted.","major_comments":[{"comment":"The central claim of an open benchmark dataset is not currently verifiable. The Usage Notes state that 'The dataset is publicly accessible, and the data along with the processing code will be gradually uploaded and updated in the coming period' and that only 'The example data and processing code can be accessed at the GitHub repository.' This directly contradicts the Data Records statement that the four data types 'are publicly available,' and it means readers cannot currently access the full 50-subject raw k-space and image data. For a dataset descriptor, the archived data are the load-bearing deliverable; the manuscript should not be accepted until the complete dataset is deposited in an archival repository with a stable identifier (e.g., a DOI) and the repository link is verified to contain all subjects and all four data types.","section":"Usage Notes; Data Records"},{"comment":"The Data Records section provides no file inventory, per-subject manifest, directory structure, file sizes, or total storage requirements. A reader cannot determine what files exist for each of the 50 subjects, how the raw k-space, reconstructed k-space, reconstructed images, and DICOM-to-NIfTI images are organized, or how k-space files correspond to reconstructed images. Add a data record table listing file naming conventions, formats, sizes, and subject/session identifiers, so that the claimed contents of the dataset can be checked against the archive.","section":"Data Records"},{"comment":"The technical validation is qualitative throughout. Reconstruction quality is assessed by 'visual inspection' of Fig. 2, with no quantitative agreement metrics (e.g., PSNR, SSIM, NRMSE, or ROI-based SNR) between the offline reconstruction and the vendor DICOM-to-NIfTI images. The motion assessment in Fig. 3b is presented as a single RMS plot without stating whether it is a representative subject or an aggregate, and without summary statistics or error bars across the 50 subjects. The microstructural modeling results in Fig. 4 are judged as 'visually consistent' without quantitative comparison to literature values or fitting-quality metrics. Since the paper's purpose is benchmarking, at least a small set of quantitative reconstruction and model-quality metrics should be reported.","section":"Technical Validation"}],"minor_comments":[{"comment":"The sentence 'current open-access dMRI datasets were primarily acquired on 3.0T or 7.0T MRI systems, with their own.' is a fragment and should be completed or rewritten.","section":"Background & Summary"},{"comment":"The Methods section says k-space data are saved in binary .bin format after extraction, while the Data Records section says the k-space data are in 'binary MAT-format (.mat)'. Please clarify the actual file container and naming convention.","section":"Methods; Data Records"},{"comment":"The note in Table 1 uses 'echo planer imaging'; this should be 'echo planar imaging'.","section":"Table 1"},{"comment":"There is a typo, 'Muti-channel combination,' which should be 'Multi-channel combination.'","section":"Methods; Data reconstruction"},{"comment":"The heading 'Filed bias correction' should be 'Bias field correction.'","section":"Methods; Data preprocessing"},{"comment":"The NORDIC PCA denoising step is described without its parameter settings (e.g., patch size, temporal window, rank threshold). Since the paper claims the reconstruction pipeline is fully documented for reproducibility, either report these parameters or point to the exact code version that contains them.","section":"Methods; Data reconstruction"},{"comment":"Please clarify in the Data Records section that T1w and T2w data are image-only and do not include raw k-space; the abstract and Background & Summary could otherwise be read as claiming k-space for all modalities.","section":"Data Records"}],"recommendation":"major_revision","confidential_remarks":"The main editorial risk is timing: a dataset descriptor whose central promise is open access cannot be accepted while the full dataset is still promised for 'the coming period.' The revision should be conditioned on the complete archive being available with a persistent identifier and a verifiable file inventory. Given that one author is affiliated with the scanner vendor, the editors may also wish to verify that the competing interests statement adequately covers this relationship."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuine first — an open 5.0T brain diffusion MRI dataset with raw k-space from 50 subjects — and the acquisition/reconstruction documentation is detailed enough to be useful. The main problem is that the full dataset is not actually public yet; the Usage Notes say it will be 'gradually uploaded,' and the GitHub only has example data. That makes the central benchmark claim conditional.\n\nWhat's new: no prior open dataset provides raw k-space for brain dMRI at 5T. The 3T and 7T datasets they cite (HCP, UK Biobank, AHEAD, etc.) are image-only or don't include raw k-space. So the resource fills a real gap for people developing reconstruction and artifact-correction methods at intermediate field strength. The 291 diffusion volumes with three shells, 21 b0s, plus 0.5mm T1w/T2w, is a solid protocol. The reconstruction chain (pre-whitening, NGC/GESTE, slice-GRAPPA, in-plane GRAPPA, POCS, adaptive combine, NORDIC) is described step by step, and the code is promised on GitHub. They also ran standard DTI, NODDI, MSMT-CSD, and tractography as sanity checks.\n\nSoft spots, in order:\n1. Availability. The Usage Notes say the data 'will be gradually uploaded and updated in the coming period.' As of now, only example data are on GitHub. For a dataset descriptor, the data being downloadable is the load-bearing requirement. The paper should state a release date, a DOI, and a file inventory, or the claims should be scaled back to 'announcement of a forthcoming dataset.'\n2. Technical validation is qualitative. They visually compare offline vs vendor reconstruction and show one RMS motion plot, but there are no quantitative metrics (e.g., NRMSE, SNR, QA maps) across the 50 subjects. For a benchmarking resource, that's a real omission, though not a fatal one.\n3. Minor narration issues: a couple of truncated sentences in the Background (e.g., 'with their own') and some typos. Not substantive.\n\nThe reader's stress-test note is right on point; I don't think it overstates.\n\nWho this is for: MRI reconstruction and diffusion imaging researchers who need raw k-space data at 5T. If the full release lands, this becomes a frequently cited resource. As it stands, it's a well-documented preview. I'd send it to peer review — a serious referee can push for the data release and quantitative validation — but I wouldn't cite the dataset itself until it's actually in a public archive.\n\nRecommendation: engage, but with a request for the data availability to be made concrete before acceptance.","headline":"Useful dataset descriptor for a genuine 5T k-space resource, but the open-access claim is currently ahead of the actual data release.","tokens_in":16221,"tokens_out":2555,"would_cite":false,"duration_ms":24413,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The Diff5T paper claims the first open 5.0T human brain diffusion MRI dataset with raw k-space and reconstructed images from 50 subjects.","keywords":["diffusion MRI","5.0 Tesla","k-space data","brain imaging","benchmark dataset","image reconstruction","tractography","microstructure"],"falsifier":"Check the public repository and data-hosting link listed in the Usage Notes: if raw k-space and reconstructed images for all 50 subjects are not downloadable, or if the released files do not match the stated 291 diffusion volumes and 0.5-mm structural scans, the open-benchmark claim fails.","tokens_in":15239,"feed_emoji":"🧠","tokens_out":11131,"duration_ms":99848,"temperature":0.7,"pith_summary":"Diff5T is presented as the first open 5.0-tesla human brain diffusion MRI dataset that contains raw k-space data alongside reconstructed images. It covers 50 healthy volunteers with 1.2-mm isotropic diffusion images acquired at three b-values ($b=1000$, $2000$, and $3000$ s/mm², 90 directions each) plus 0.5-mm T1w and T2w structural scans. The authors supply a documented offline reconstruction pipeline that converts raw scanner data into images, a preprocessing pipeline, and example fits of DTI, NODDI, MSMT-CSD, and tractography. If the full dataset is released as promised, it would give the MRI community real 5T k-space measurements for developing and benchmarking reconstruction and artifact-correction methods that currently rely heavily on synthetic k-space.","feed_headline":"First 5-tesla brain diffusion MRI dataset offers raw k-space","feed_subtitle":"Fifty healthy adults, three diffusion shells, 0.5-mm anatomy, and raw scanner data for benchmarking reconstruction.","key_machinery":"The central object is the Diff5T dataset, anchored by raw k-space data — the frequency-domain measurements from which MR images are computed by Fourier transformation. The mechanism that makes these raw data usable is the documented offline reconstruction pipeline: pre-whitening, Nyquist ghost correction, slice-GRAPPA and in-plane GRAPPA parallel imaging, POCS partial-Fourier recovery, adaptive coil combination, and NORDIC PCA denoising. The same pipeline also produces a reproducible baseline, since the authors describe their reconstruction as non-optimal and explicitly invite other researchers to refine it.","core_discovery":"Diff5T is claimed to be the first 5.0T brain imaging dataset to provide both raw k-space and reconstructed image data for diffusion MRI. It comprises 50 healthy subjects aged 18–38 scanned on a 5.0T system with 120 mT/m gradients and a 48-channel head coil: 1.2-mm isotropic diffusion volumes with 90 directions per shell at $b=1000$, $2000$, and $3000$ s/mm², 21 $b=0$ volumes, and 0.5-mm isotropic T1w and T2w structural images. The diffusion data are released as raw k-space, reconstructed k-space, reconstructed images, and converted scanner images, together with reconstruction and preprocessing scripts. To demonstrate usefulness, the authors fit DTI, NODDI, and multi-shell multi-tissue constrained spherical deconvolution and ran whole-brain tractography, reporting results visually consistent with human brain anatomy.","pith_inferences":["The authors stop at showing their pipeline works; combining Diff5T with 3T and 7T diffusion data from other studies would let the field separate field-strength effects from protocol effects — a comparison the paper itself does not run.","Because the raw k-space is empirical rather than synthesized from magnitude images, these data can serve as a testbed for AI reconstruction models that currently train on simulated k-space; the released pipeline gives such models a fixed, non-optimal baseline to beat.","The narrow inclusion criteria (healthy adults 18–38, one scanner) make Diff5T a clean method-development benchmark, but the same homogeneity limits any inference about aging, disease, or scanner variability."],"forward_implications":["Researchers can use the raw and reconstructed k-space to benchmark parallel-imaging and partial-Fourier reconstruction methods against a documented, non-optimal 5T pipeline.","The 90-direction, three-shell sampling supports DTI, NODDI, and MSMT-CSD in the same subjects, so microstructural metrics can be compared across models without scanner differences.","The 0.5-mm T1w and T2w volumes give registration targets and structural context for diffusion tractography.","The released processing scripts make the whole k-space-to-tractography chain reproducible and modifiable.","If the data behave as described, 5T diffusion MRI becomes a testable middle ground between 3T and 7T in resolution, SNR, distortion, and scan time."],"supporting_citations":[{"why":"Establishes the prior high-resolution whole-brain diffusion imaging that Diff5T extends to an intermediate 5.0T field with raw k-space.","marker":"[3]"},{"why":"Demonstrates 5.0T RF hardware and initial brain imaging, providing the scanner foundation for the dataset.","marker":"[27]"},{"why":"Precedent for an open raw-data MRI benchmark; motivates the gap that Diff5T fills for diffusion k-space.","marker":"[41]"},{"why":"Supplies the spherical-code method used to design the multi-shell, multi-direction diffusion protocol.","marker":"[50]"},{"why":"Supplies slice-GRAPPA, used to unalias the multi-band slices in reconstruction.","marker":"[58]"},{"why":"Supplies GRAPPA, used for in-plane parallel-imaging reconstruction.","marker":"[59]"},{"why":"Supplies POCS partial-Fourier reconstruction, used to recover missing k-space lines.","marker":"[60]"},{"why":"Supplies NORDIC PCA denoising, applied to the complex dMRI reconstruction.","marker":"[66]"},{"why":"Supplies topup, used to correct susceptibility-induced distortion from reversed phase-encoding b=0 pairs.","marker":"[69]"},{"why":"Supplies eddy, used for eddy-current and motion correction and for quantifying head motion.","marker":"[71]"}],"fun_headline_variants":["5T brain diffusion dataset adds raw k-space to public domain","First 5T dMRI dataset with raw k-space for open benchmarking","Diff5T: 50 brains, 3 shells, raw k-space at 5 tesla","Open 5T diffusion MRI dataset includes raw k-space and images"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the full 50-subject dataset will be made publicly available, but the Usage Notes only promise gradual uploads and currently offer example data.","fun_headline_variants_meta":{"raw":{"variants":["5T brain diffusion dataset adds raw k-space to public domain","First 5T dMRI dataset with raw k-space for open benchmarking","Diff5T: 50 brains, 3 shells, raw k-space at 5 tesla","Open 5T diffusion MRI dataset includes raw k-space and images"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000161,"raw_usage":{"total_tokens":1223,"prompt_tokens":920,"completion_tokens":303,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":221}},"tokens_in":536,"tokens_out":303,"duration_ms":3409,"temperature":1.0,"reasoning_tokens":221,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:24:30.511429+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the public repository and data-hosting link listed in the Usage Notes: if raw k-space and reconstructed images for all 50 subjects are not downloadable, or if the released files do not match the stated 291 diffusion volumes and 0.5-mm structural scans, the open-benchmark claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates 5.0T RF hardware and initial brain imaging, providing the scanner foundation for the dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the spherical-code method used to design the multi-shell, multi-direction diffusion protocol."},{"cited_title":"F., Polimeni J","cited_arxiv_id":null,"evidence_quote":"Supplies slice-GRAPPA, used to unalias the multi-band slices in reconstruction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies GRAPPA, used for in-plane parallel-imaging reconstruction."},{"cited_title":"M., Lindskogj E","cited_arxiv_id":null,"evidence_quote":"Supplies POCS partial-Fourier reconstruction, used to recover missing k-space lines."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies NORDIC PCA denoising, applied to the complex dMRI reconstruction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies topup, used to correct susceptibility-induced distortion from reversed phase-encoding b=0 pairs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies eddy, used for eddy-current and motion correction and for quantifying head motion."}],"review_version":1}