{"id":"fb87742c-e054-4a36-a513-eb4bc5e83fcc","arxiv_id":"2411.12091","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A standardized, open dataset of C. elegans calcium-imaging traces from 12 studies (about 900 worms, 250 labeled neurons) plus consensus connectome graphs from 4 publications.","lead":"This paper packages calcium-imaging recordings from 12 C. elegans studies (about 900 worms) and connectome graphs from 4 publications into uniformly formatted, openly available datasets. The value is that modelers can now build structure-function models of a full nervous system without re-processing each lab's proprietary format.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-lab comparability rests on an unvalidated label-identity assumption: Section 2.2.3(f) conflates unlabeled with unmeasured neurons, and no check shows that the same canonical name refers to the same cell across the 12 source studies.","rationale":"The reader and I converge on label identity as the weakest point. I do not think this warrants REJECT: if the audit passes, the resource is useful, and the paper is candid about many limitations. The missing URLs or commit hashes and the Algorithm 1 versus prose order mismatch are additional reproducibility defects, but they are secondary to label identity. A dataset paper's core promise is that others can safely compare across labs and animals; that promise is exactly what Section 2.2.3(f) and the absence of cross-lab validation leave unsecured. The verdict therefore stays CONDITIONAL: the central claim is plausible but must be backed by a concrete identity check before the released matrices should be used for the advertised cross-lab analyses.","tokens_in":14096,"tokens_out":6611,"duration_ms":74175,"concrete_test":"Run an independent label-identity audit: download the raw files for at least two overlapping datasets (e.g., Kato2015 and Yemini2021), re-extract each neuron trace using the original authors' published labels, and compare trace-by-trace to the matching column of the released preprocessed matrix. The audit passes only if every same-named column contains the same underlying raw trace after resampling and normalization, and no released column is a zero-filled or swapped placeholder. In the same pass, read the GitHub code for NeuralBasePreprocessor to determine whether normalization precedes resampling, then rerun the pipeline both ways and quantify how many traces change; if the implemented order differs from Section 2.2.3, the paper must state which order produced the released files.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the released artifacts are homogenized so that cross-lab and cross-animal comparisons are valid. That requires every column label (e.g., AVAL) to denote the same canonical neuron in every source dataset. The paper never validates this. Section 2.2.3(f) explicitly says unlabeled neurons and unmeasured neurons are used interchangeably, so a mask zero can mean either 'no trace' or 'trace without a name'; this already compromises coverage statistics and any pooled analysis. More importantly, no test is described showing that AVAL from Kato2015 is the same cell as AVAL from Yemini2021; the sources use different imaging, registration, and identification protocols, and a label shift or naming collision would silently corrupt the D-column matrices and the consensus connectome. A second inconsistency weakens the 'standardized' claim: Section 2.2.3(b/e) states resampling, then smoothing, then normalization as the final step, while Algorithm 1 normalizes before smoothing and resampling. If the released code follows Algorithm 1, the output traces differ from the published prose; if it follows the prose, the algorithm listing is wrong. Either way, unambiguous reproduction of the homogenized data is not currently possible from the paper alone. These issues do not disprove that the repository exists, but they do put the central usability claim at risk.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a preprocessing pipeline and two released datasets that aggregate calcium imaging neural activity from 12 C. elegans studies and connectome data from multiple electron microscopy and functional studies. The neural pipeline standardizes raw traces by organizing columns alphabetically by canonical neuron name, resampling to a common time step, optionally smoothing, z-scoring, and generating a binary mask of labeled neurons; the connectome pipeline builds graph tensors and a consensus connectome. The claimed contribution is a public, homogenized resource that supports cross-lab and cross-animal comparisons and downstream structure-function modeling.","tokens_in":14301,"tokens_out":6645,"duration_ms":69521,"significance":"If the identified issues are resolved, this would be a useful community resource: it consolidates 919 worms across 12 studies, provides open code and HuggingFace-hosted data, and standardizes connectomes into GNN-ready tensors. The paper is candid about several limitations, including calcium imaging's low-pass character, the mismatch between activity and connectome animals, and the subjective choices of delta-t and lambda. The main scientific risk is that the central usability claim depends on cross-lab label equivalence and on exact pipeline ordering, neither of which is currently established from the manuscript alone.","major_comments":[{"comment":"The cross-lab comparability claim is load-bearing, but the paper never validates the assumption that a canonical neuron name such as AVAL refers to the same cell across the 12 source studies. Section 2.2.3(f) explicitly states that unlabeled and unmeasured neurons are used interchangeably, so a mask entry of zero can mean either 'no trace' or 'trace without a name'. More importantly, no test is described showing that labels from different laboratories correspond to the same anatomical cells; different registration, imaging, and identification protocols could produce label shifts or naming collisions that would silently corrupt the D-column matrices and any pooled analysis. Please add a section reporting per-source label vocabularies and label-overlap statistics, and describe an independent validation (for example, registration to a common atlas or NeuroPAL-based identification) for at least a subset of neurons.","section":"Section 2.2.3(f) and Table 1"},{"comment":"The preprocessing order is described inconsistently. The prose in Section 2.2.3(b)-(e) states the main steps as resampling, optional smoothing, and then normalization, with normalization described as the final step in (e). Algorithm 1, however, normalizes the raw traces before smoothing and resampling. If the released code follows Algorithm 1, the output traces differ from the published prose; if it follows the prose, the algorithm listing is wrong. Either way, unambiguous reproduction of the homogenized data from the paper alone is currently impossible. Please correct the algorithm or the text and explicitly state which order was used to generate the released artifacts.","section":"Section 2.2.3(b)-(e) and Algorithm 1"}],"minor_comments":[{"comment":"The abstract says the connectivity dataset is compiled from 9 connectome annotations, while Section 2.1 says there are 10 distinct connectome source files and Appendix Table 5 lists 10 graph tensor files; please reconcile these counts.","section":"Abstract vs. Section 2.1 and Appendix Table 5"},{"comment":"Row 9 labels the source as Leifer2023 but the cited reference [15] is Randi et al.; please use a consistent author-year label that matches the bibliography.","section":"Table 1"},{"comment":"Equation (1) defines z-scoring with mu and sigma in R^D; please clarify that the subtraction and division are elementwise over the temporal dimension and specify whether sigma is the sample or population standard deviation.","section":"Section 2.2.3(e)"},{"comment":"The mask is defined as boolean(X_j != empty); since Section 2.2.3(f) conflates unlabeled and unmeasured neurons, please define exactly what non-empty means for each source file format.","section":"Algorithm 1"},{"comment":"The abstract says approximately 900 worms, but Table 1 sums to 919; please report the exact count where possible.","section":"Table 1 and Abstract"},{"comment":"Reference [35] is formatted as 'L. et al. Tian'; the standard form is 'Tian et al.'","section":"References"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers what it promises: a standardized release of C. elegans calcium imaging and connectome data, with common neuron ordering, resampling, masks, and a consensus connectome. That is a new artifact, not just a rehash of existing datasets, and it should save a lot of preprocessing pain. The metadata tables are careful, the limitation discussion is candid, and the code is structured so new datasets can be added. Credit where it is due: this is a solid, useful resource for anyone doing structure-function modeling in C. elegans.\n\nThe soft spots are real but not fatal. First, the label-identity assumption. The whole claim that columns named AVAL are comparable across labs rests on the same name denoting the same canonical cell in all 12 source studies. That is never validated. Section 2.2.3(f) explicitly conflates unlabeled with unmeasured neurons, which muddies the masks and any coverage statistics built on them. If labeling conventions differ across labs, pooled matrices and cross-lab comparisons fail even if the code runs. This is the load-bearing risk and it needs an explicit check, or at least an acknowledgment that the resource inherits whatever label heterogeneity exists in the sources.\n\nSecond, the pipeline order is inconsistent. The prose in 2.2.3 says resampling, then optional smoothing, then normalization as the final step. Algorithm 1 normalizes before smoothing and resampling. That is not a trivial typo; it changes the output traces and makes reproduction from the paper alone impossible. Either the code follows one path and the text is wrong, or vice versa. Also, the hosted artifacts are not pinned with URLs or commit hashes in the text, which further weakens reproducibility.\n\nThese are fixable. The resource itself is valuable; the paper is honest about limitations like calcium as an indirect measure, the connectome/activity animal mismatch, and the subjective lambda choice. There is no circularity burden because nothing is fitted to a target. The engineering is routine, but that is fine for a dataset paper.\n\nWho is this for? Computational modelers and C. elegans neuroscientists who want a ready-to-use data substrate. It deserves a serious referee, not a desk reject. My recommendation: send it to peer review, but require the authors to resolve the normalization/resampling order discrepancy, pin the repository versions, and either validate the label-identity assumption across source studies or explicitly scope the claim to avoid overstating cross-lab comparability.","headline":"A genuinely useful homogenized C. elegans data release, but the label-identity assumption is unvalidated and the pipeline description is internally inconsistent; both need fixing before the cross-lab comparability claim is safe.","tokens_in":14860,"tokens_out":914,"would_cite":true,"duration_ms":12634,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper releases a homogenized dataset of C. elegans neural activity and connectivity, combining 12 calcium-imaging studies and connectome annotations from 4 primary sources.","keywords":["C. elegans","calcium imaging","connectome","neural activity","data homogenization","whole-brain imaging","consensus connectome","structure-function modeling"],"falsifier":"Compare the same named neuron (say AWCR) in two laboratories' recordings under the same stimulus after running the pipeline: if its trace is systematically swapped with its contralateral partner or inconsistent across labs, the canonical-label assumption fails and pooled columns are not comparable; likewise, inspecting the source files for traces that were recorded but left unnamed would directly test the stated conflation of unlabeled with unmeasured neurons.","tokens_in":13849,"feed_emoji":"🧠","tokens_out":8170,"duration_ms":81095,"temperature":0.7,"pith_summary":"The paper's aim is to make C. elegans neuroscience data comparable across labs by releasing a homogenized collection of neural activity and connectivity measurements. It assembles calcium-imaging traces from 12 published studies, covering roughly 900 worms and about 250 uniquely labeled neurons, and standardizes them by resampling to a common time step, z-scoring each neuron trace, ordering neuron columns alphabetically, and attaching a binary mask that says which neurons were labeled. It also compiles connectome annotations from 4 primary electron-microscopy and signal-propagation studies into graph-structured files, plus a consensus connectome built by averaging connection weights across datasets. If the resource is sound, a researcher can pool recordings collected under different protocols and directly ask how synaptic wiring relates to neural dynamics in a small, fully mapped nervous system.","feed_headline":"900 worms of C. elegans brain recordings, standardized in one dataset","feed_subtitle":"Resampled calcium traces and consensus connectomes let labs compare activity across animals and studies.","key_machinery":"The load-bearing machinery is the canonical 300-neuron coordinate system. Every worm's fluorescence traces are placed into a $T_k \\times 300$ matrix with columns sorted alphabetically by standard C. elegans neuron names, and a companion binary mask $M^{(k)}$ records which neurons were labeled in that worm; the mask is kept separate from the matrix instead of zeroing unlabeled entries. On the connectivity side, the same coordinate system is used to build directed graph objects through dataset-specific preprocessor classes, with per-edge attributes for chemical synapses, gap junctions, and optional functional connectivity, and a consensus connectome is produced by averaging nonzero chemical and gap-junction weights across sources plus a weighted standard deviation.","core_discovery":"The central claim is that a heterogeneous collection of C. elegans recordings and wiring diagrams can be homogenized without destroying their scientific meaning. The released activity dataset represents each worm as a time-series matrix $X^{(k)} \\in \\mathbb{R}^{T_k \\times 300}$ with columns ordered alphabetically by canonical neuron name, resampled to $\\Delta t = 0.333$ s, z-scored across time per neuron, and paired with a binary mask $M^{(k)}$ of labeled neurons; the connectome half stores each source as a directed graph with chemical, gap-junction, and optional functional weights and also provides a consensus connectome formed by averaging nonzero weights across datasets. The paper's contribution is the resource and the pipeline, not a new biological mechanism: if the standardization holds, any downstream model can load these files and treat structure and function in a common neuron coordinate system.","pith_inferences":["The entire comparison rests on labels being canonical; a direct check would be whether the same named neuron responds consistently across two labs under matched stimuli, and the paper does not report such a validation.","Resampling to $\\Delta t=0.333$ s and z-scoring with full-trace statistics trade away fast dynamics and causal normalization, so users building real-time or spike-resolved models will need to redo preprocessing.","Because the paper explicitly treats unlabeled and unmeasured neurons as the same, a missing entry in the mask should not be interpreted as evidence that a neuron was silent.","A strong test of the resource would be to train a structure-function predictor on a subset of worms and test transfer to worms from a different laboratory dataset."],"forward_implications":["Any of the 12 activity datasets can be loaded in the same matrix format, so pooling worms across labs and protocols becomes a concatenation step rather than a format-reconciliation project.","Connectome files are already structured as graph tensors, so models built on graph neural networks can consume connectivity and calcium traces in a common interface.","The consensus connectome provides one wiring diagram with per-edge means and standard deviations, making cross-study variability in synapse counts visible.","The binary masks let users restrict analyses to neurons that were actually labeled, which matters for training and evaluating structure-function models.","A new imaging or connectome source can be added by writing one dataset-specific preprocessor subclass, so the collection can grow without restructuring the pipeline."],"supporting_citations":[{"why":"Supplies the original electron-microscopy wiring diagram that anchors the connectome graph and neuron list.","marker":"[37]"},{"why":"Provides the whole-animal connectome of both sexes used as a primary connectome source and for consensus averaging.","marker":"[4]"},{"why":"Adds developmental-stage connectomes as additional electron-microscopy sources for the consensus connectome.","marker":"[38]"},{"why":"Contributes the functional signal-propagation data, providing the optional functional edge weights and one activity dataset.","marker":"[15]"},{"why":"Contributes whole-brain calcium-imaging traces with neuron identification and one of the 12 activity datasets.","marker":"[39]"},{"why":"One of the 12 calcium-imaging studies whose traces are resampled and normalized into the activity matrices.","marker":"[10]"},{"why":"Another source calcium-imaging study supplying raw fluorescence traces for the homogenized activity dataset.","marker":"[11]"},{"why":"Sleep-state calcium-imaging study whose traces feed into the pooled activity matrices.","marker":"[13]"},{"why":"Largest single neural-activity contributor (577 worms) and the source for the amphid chemosensory consensus visualization.","marker":"[12]"}],"fun_headline_variants":["C. elegans neural data: 900 worms, standardized activity and wiring","One dataset, 900 worms: homogenized C. elegans brain recordings","Standardized C. elegans connectome and activity across labs","900 worms, 250 neurons: unified C. elegans neural dataset"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that neuron names used by different laboratories refer to the same canonical cells, so alphabetically ordered columns and masks make recordings comparable across studies; the paper also assumes unlabeled and unmeasured neurons can be treated interchangeably.","fun_headline_variants_meta":{"raw":{"variants":["C. elegans neural data: 900 worms, standardized activity and wiring","One dataset, 900 worms: homogenized C. elegans brain recordings","Standardized C. elegans connectome and activity across labs","900 worms, 250 neurons: unified C. elegans neural dataset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000349,"raw_usage":{"total_tokens":1910,"prompt_tokens":954,"completion_tokens":956,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":881}},"tokens_in":570,"tokens_out":956,"duration_ms":8542,"temperature":1.0,"reasoning_tokens":881,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:54:48.805184+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the same named neuron (say AWCR) in two laboratories' recordings under the same stimulus after running the pipeline: if its trace is systematically swapped with its contralateral partner or inconsistent across labs, the canonical-label assumption fails and pooled columns are not comparable; likewise, inspecting the source files for traces that were recorded but left unnamed would directly test the stated conflation of unlabeled with unmeasured neurons.","supporting_citations":[{"cited_title":"The structure of the nervous system of the nematode c","cited_arxiv_id":null,"evidence_quote":"Supplies the original electron-microscopy wiring diagram that anchors the connectome graph and neuron list."},{"cited_title":"Con- nectomes across development reveal principles of brain maturation in c","cited_arxiv_id":null,"evidence_quote":"Adds developmental-stage connectomes as additional electron-microscopy sources for the consensus connectome."},{"cited_title":"Neural signal propagation atlas of caenorhabditis elegans","cited_arxiv_id":null,"evidence_quote":"Contributes the functional signal-propagation data, providing the optional functional edge weights and one activity dataset."},{"cited_title":"Nested neuronal dynamics orchestrate a behavioral hierarchy across timescales","cited_arxiv_id":null,"evidence_quote":"One of the 12 calcium-imaging studies whose traces are resampled and normalized into the activity matrices."},{"cited_title":"A global brain state underlies c","cited_arxiv_id":null,"evidence_quote":"Sleep-state calcium-imaging study whose traces feed into the pooled activity matrices."},{"cited_title":"Functional imaging and quantification of multineuronal olfactory responses in c","cited_arxiv_id":null,"evidence_quote":"Largest single neural-activity contributor (577 worms) and the source for the amphid chemosensory consensus visualization."}],"review_version":1}