{"id":"2b213f00-b603-4c5f-ba2c-74d0885401d2","arxiv_id":"2608.07632","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A 116 GB compressed subset of JUMP supports reproducible benchmarking, with moderate JPEG XL compression preserving biological signal and stable model rankings.","lead":"The authors built JUMP-lite, a 116 GB subset of the 115 TB JUMP Cell Painting dataset, and show that JPEG XL compression preserves most biological signal for benchmarking cell-image representations. A smart generalist might read this because it offers an accessible, standardized way to compare classical and deep-learning profiling methods without needing petabytes of storage.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Head-to-head model ranking (Fig. 5) is confounded: CellProfiler features come from 6-9 sites per well while all deep-learning embeddings use 4 sites per well, so the top-tier claim is not yet supported.","rationale":"The reader's weakest assumption is exactly the one I would rank first. The paper's compression-fidelity result—HQ/MQ preserve PA/PC and recall within a few percent—is supported by mean-across-configuration deltas in Table 1 and Table S5 and does not depend on the CellProfiler comparison, so I do not see grounds to reject that result. The comparative benchmark, however, is a core contribution: the title promises benchmarking of cell representations, and Section 3.5's three-tier ranking is the paper's headline for that promise. Section S1.2.7 concedes that CellProfiler features were precomputed from 6-9 sites per well while all other methods use 4 sites per well. Since well-level aggregation over more sites reduces noise, any retrieval advantage for CellProfiler could be an artifact of input sampling rather than representation strength. The paper does not provide a matched-site comparison, so the ranking in Figure 5 and the claim that CellProfiler features and MorphEM lead are conditional on a confound the authors themselves identify. I also considered the test-metric sweep as a candidate concern, but Figure 5 averages over normalization configurations rather than reporting only the best selected configuration, which weakens that worry. The concrete test is a matched-site recomputation; if it confirms CellProfiler's rank, the benchmark stands, and if not, the model-comparison conclusions need revision. The reader's CONDITIONAL verdict is therefore appropriate and should be retained.","tokens_in":22406,"tokens_out":8250,"duration_ms":79938,"concrete_test":"Recompute the CellProfiler row of Figure 5 using cp measure on exactly the same four sites per well that define JUMP-lite, for a representative subset (e.g., the Target-2 pilot plates or a few hundred perturbations spanning CRISPR/ORF/Diverse/Bioactive), with the normalization sweep used for the deep-learning models. If CellProfiler's mean normalized score drops below MorphEM or into the mid-tier, the site-count confound drives the headline ranking; if it remains within the current top tier, the comparison is robust. A cheaper alternative: if per-site precomputed features are available in the Cell Painting Gallery, subsample to four sites per well and rerun the same aggregation and evaluation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central compression claim is not what breaks; Table 1 and Figure 4 support the HQ/MQ preservation result using mean-across-configuration deltas from deep-learning models. The load-bearing problem is the comparative benchmark, which is the paper's headline contribution. Section 3.5 and Figure 5 rank CellProfiler and MorphEM as the top tier, but Section S1.2.7 states that CellProfiler features were precomputed from the Cell Painting Gallery using six to nine imaging sites per well, while all deep-learning embeddings in JUMP-lite use exactly four sites per well. Aggregating more sites per well reduces noise in the well-level profile and can only help retrieval scores, so CellProfiler's leading NAP values (e.g., 0.815 vs. 0.777 on CRISPR PA) may reflect data quantity rather than representation quality. The paper acknowledges this in S1.2.7 but still presents the unadjusted comparison as the benchmark's main result, and the released JUMP-lite images alone cannot reproduce the CellProfiler row because those features were not derived from the JUMP-lite images. If the site-count advantage is substantial, the three-tier ranking—especially CellProfiler tied with MorphEM—is an artifact of the input sampling protocol, not a property of the representations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces JUMP-lite, a 116 GB curated and JPEG XL-compressed subset of the 115 TB JUMP Cell Painting dataset that retains roughly 24,401 perturbations, and Nahual, a Nix-based framework for deploying models with incompatible dependencies. The authors characterize the effects of lossy compression on segmentation, feature correlation, and downstream retrieval tasks (phenotypic activity, phenotypic consistency, and MOTIVE cross-modality recall), reporting that HQ compression yields a 1.2% mean performance change and MQ a 6.5% change while reducing storage by 97–98.8%. They benchmark five representation methods (CellProfiler, MorphEM, OpenPhenom, SubCell, DINOv2) plus simple baselines across eleven tasks and find that CellProfiler and MorphEM form a top tier, DINOv2/SubCell/OpenPhenom a middle tier, and that model ranking is stable across HQ/MQ compression.","tokens_in":22657,"tokens_out":9794,"duration_ms":86645,"significance":"If the central claims hold, JUMP-lite would be a valuable community resource: a roughly 1000-fold storage reduction with a multi-lab, multi-task evaluation suite, combined with Nahual's reproducibility infrastructure, could make systematic benchmarking of cell representations accessible to groups without petabyte-scale storage. The compression-fidelity analysis fills a real gap in the Cell Painting literature, where previous compact benchmarks such as RxRx3-core did not assess the impact of compression on downstream signal. The finding that engineered CellProfiler features remain competitive with modern self-supervised models is an important and falsifiable result, and the authors should be credited for their public code release, detailed curation documentation, and transparent reporting of the normalization sweep. However, the head-to-head model ranking is currently confounded by a site-count mismatch, and the selection of post-processing configurations on the evaluation metric leaves the absolute performance claims open to optimistic bias; these issues must be resolved before the benchmark's headline conclusions are reliable.","major_comments":[{"comment":"The headline ranking of representation methods is confounded by the input sampling protocol. Section S1.2.7 states that the precomputed CellProfiler features were derived from six to nine imaging sites per well, whereas all deep-learning embeddings in JUMP-lite are computed from exactly four sites per well. Because well-level aggregation over more imaging sites reduces noise, CellProfiler's top-tier NAP values (e.g., 0.815 vs. 0.777 on CRISPR PA in Figure 5) may reflect data quantity rather than representation quality. Moreover, the released JUMP-lite images contain four sites per well, so a user cannot reproduce the CellProfiler row from the released data. The acknowledgment in S1.2.7 does not justify the unadjusted comparison in Figure 5; the authors should either recompute CellProfiler features on the same four sites per well (or a matched subsample) or explicitly restrict the ranking claim to deep-learning models.","section":"3.5, Figure 5, S1.2.7"},{"comment":"The post-processing protocol is internally inconsistent and, under one reading, selects configurations on the evaluation metric. Section 3.1 states 'We report the configuration with the highest balanced PA/PC (rescaled variant)' and S1.2.6 says 'The post-processed profiles from the best-performing configuration were then used for downstream analyses,' but the Figure 5 caption reports 'the average across normalization configurations.' These are different quantities: the former is a maximum over 280–420 configurations, which optimistically biases all reported scores; the latter is a mean and gives some protection against overfitting. The authors must specify which quantity is used for each reported result, and for the best-configuration variant they should provide unbiased estimates (e.g., evaluation on a held-out split or at a fixed configuration). This is load-bearing for the quantitative claims in Table 1 and the tier assignment in Figure 5.","section":"3.1, S1.2.6, Figure 5 caption"},{"comment":"The three-tier ranking is presented without uncertainty quantification. Figure 5 shows only the mean normalized score per model, and Table S5 reports standard deviations across normalization configurations in percentage-change units; for example, the middle-tier scores (DINOv2 0.721, SubCell 0.715, OpenPhenom 0.688) are close, and the per-configuration spread in Table S5 is comparable to these differences. Without error bars or paired statistical tests on the normalized scores, the claim that these three models form a distinct middle tier is not supported. The authors should report the distribution across configurations for the normalized scores in Figure 5, or provide a statistical test of the ranking.","section":"Figure 5, Table S5"}],"minor_comments":[{"comment":"The caption states that the pooled distribution includes '48 normalization configurations each,' but S1.2.6 describes 280 configurations for engineered features and up to 420 for deep-learning embeddings; please reconcile these numbers.","section":"Figure 3c caption"},{"comment":"The main text says 'We evaluated four compression levels' (raw, HQ, MQ, d20), but Table S2 and Figure 2a present additional levels (zstd, jxl-effort-3, jxl-d2-e8, jxl-lq, jxl-d10, jxl-d15, jxl-d25); please clarify which levels are used in which analyses and why the additional levels are omitted from the main narrative.","section":"3.1, Table S2, Figure 2a"},{"comment":"S1.4.1 states that links to software, code, and repositories are unavailable due to anonymization, while the Data availability section provides GitHub and Zenodo URLs; in the final version these statements should be reconciled.","section":"S1.4.1, Data availability"},{"comment":"The 'Cell Count' baseline used in Figure 5 is not defined anywhere in the Methods or Supplementary Information; please specify how cell count is computed, aggregated to the well level, and normalized before being used as a profile.","section":"Figure 5, Methods"},{"comment":"The main text does not state whether the reported MOTIVE results use the 'full' or 'strict' annotation set described in S1.3.2; because the composition of the edge sets differ substantially (e.g., the full CC graph includes chemical-similarity RESEMBLES edges), this choice should be stated explicitly wherever MOTIVE recall values are reported.","section":"3.4, 3.5, S1.3.2"},{"comment":"The sentence 'appreciable degradation appearing only upon aggressive compression (d20)' is inconsistent with the MQ results in Table 1, where some tasks lose up to 15.4% (PC on Diverse); please qualify the claim to reflect the task-dependent MQ degradation.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The site-count confound is the decisive issue. If CellProfiler features are recomputed on the same four sites per well and the ranking persists, the paper would be a strong contribution; if the ranking shifts, the main comparative claim collapses. The normalization-selection inconsistency also needs to be resolved, and the authors should provide uncertainty estimates for the tier assignments. I recommend major revision rather than rejection because the compression-preservation core appears well-supported and the released resource is valuable to the community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The compression claim is the real contribution here, and it holds up. JUMP-lite is a genuinely useful resource — 116 GB, 24,401 perturbations, a thousandfold reduction from JUMP — and the JPEG XL fidelity assessment is the first systematic one for Cell Painting. Table 1 and Figure 4 support the headline numbers: at HQ the mean per-task change from raw is -1.2%, at MQ -6.5%, and d20 collapses signal as expected. The authors also check segmentation AP, feature correlations, and SSIM, which is the right set of probes. The rank-stability analysis (Figure S7) is the right way to ask whether compression changes conclusions, and the answer — rankings hold at HQ and MQ — is credible. Nahual looks like practical infrastructure for a real dependency-hell problem.\n\nThe soft spot is the model ranking, and the stress-test note is right about it. Figure 5 puts CellProfiler and MorphEM in the top tier, but CellProfiler features come from 6-9 imaging sites per well (precomputed from the Cell Painting Gallery) while every deep model sees exactly 4 sites. More sites per well means less noise in the well-level profile, so CellProfiler's lead could be data quantity rather than representation quality. The paper flags this in S1.2.7 and then presents the unadjusted ranking anyway. That is a load-bearing problem for the comparison claim, though not for the compression claim, which is built on the deep models alone. A matched-site comparison — even a subsampled sensitivity check — would settle it.\n\nTwo smaller things. First, the normalization configuration is selected on the evaluation metric itself: each model-compression pair gets its best balanced PA/PC from a 280-420 configuration sweep (S1.2.6). This inflates absolute scores. It probably does not change the ranking much since every method gets the same treatment, but it is an optimistic bias and the paper should be upfront that these are swept-best numbers, not out-of-sample ones. Second, the code availability statements contradict each other: S1.4.1 says repositories are unavailable due to anonymization, S1.2.5 says Nahual will be released after publication, and the Data Availability section gives GitHub and Zenodo links. Needs cleanup. Also worth noting: curation and evaluation draw on the same annotation sources (RefChem, MOTIVE). I do not read that as circular — the retrieval tasks are still nontrivial — but it does mean the benchmark covers well-annotated compounds by construction.\n\nWho this is for: anyone in image-based profiling who wants a tractable JUMP subset and a defensible answer on whether lossy JPEG XL is safe. The compression result deserves to be in the literature on its own. The ranking needs the matched-site fix before it can be taken at face value. This deserves a serious referee; the fixes are revision-scale, not rejection-scale.","headline":"The compression-fidelity result is solid and the dataset is a real contribution, but the headline model ranking is confounded by a site-count mismatch between CellProfiler and every deep model.","tokens_in":23218,"tokens_out":6212,"would_cite":true,"duration_ms":48012,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"JPEG XL compression at high and medium quality preserves biological signal in Cell Painting images, so a 116 GB subset of the 115 TB JUMP dataset can stand in for the full resource when benchmarking cell representations.","keywords":["image-based profiling","Cell Painting","JUMP dataset","lossy compression","JPEG XL","representation learning","benchmarking","phenotypic activity"],"falsifier":"Recompute the classical CellProfiler profiles from the same four imaging sites per well used for deep-learning embeddings and rerun the phenotypic activity and consistency tasks; if CellProfiler no longer ranks with MorphEM at the top, the three-tier ordering is an artifact of site count rather than of representation quality.","tokens_in":22209,"feed_emoji":"🔬","tokens_out":9312,"duration_ms":75003,"temperature":0.7,"pith_summary":"This paper attempts to make systematic benchmarking of cell-image representations feasible on ordinary hardware. It curates JUMP-lite, a 116 GB subset of the 115 TB JUMP Cell Painting dataset, and shows that lossy JPEG XL compression at high quality leaves mean downstream performance essentially unchanged (-1.2% relative to raw) while cutting storage by roughly 97%. At medium quality the dataset shrinks by ~98.8% at a -6.5% average cost, and model rankings remain stable across these compression levels. The authors also release an evaluation harness that deploys multiple deep-learning models with incompatible dependencies in the same reproducible pipeline. If the preservation claims hold, large-scale Cell Painting benchmarking stops being a privilege of groups that can store 115 TB.","feed_headline":"A 116 GB cell-image benchmark stands in for 115 TB of JUMP data","feed_subtitle":"Lossy JPEG XL keeps rankings stable in Cell Painting benchmarks, so labs with modest storage can compare models fairly.","key_machinery":"The machinery is a compression-fidelity ladder built on JPEG XL codec settings (near-lossless HQ, medium MQ, and aggressive d20), combined with a standardized evaluation harness. Ground truth comes from the intersection of RefChem records and MOTIVE drug-target graphs, and every representation is scored on phenotypic activity, phenotypic consistency, and cross-modality recall, normalized to NAP. The harness uses an orchestration pipeline for segmentation and featurization, and an inter-process bridge that runs each deep-learning model in an isolated, reproducible environment so all methods are processed identically.","core_discovery":"On its own terms, the paper's central discovery is that moderate JPEG XL lossy compression preserves enough biological signal in Cell Painting images to support fair benchmarking of representation methods. Across the full JUMP-lite dataset, high-quality compression changes mean per-task performance by -1.2% and stays within ±6% on every task, medium-quality compression costs -6.5% on average, and the relative ordering of models is preserved at both settings. Only the aggressive d20 setting destroys signal, with phenotypic activity dropping 20-34% and phenotypic consistency up to 31%. Using this evaluation, the paper reports three performance tiers among representations: CellProfiler features and MorphEM lead, DINOv2, SubCell, and OpenPhenom sit in the middle, and cell-count and random-weight ViT baselines trail.","pith_inferences":["If the compression-fidelity trade-off transfers to other fluorescence and bright-field assays, the same HQ/MQ recipe could shrink other large microscopy repositories, making them benchmarkable on commodity hardware.","The observation that some models score higher at intermediate compression than at raw hints that lossy compression may act as a mild denoiser, suppressing inter-site resolution noise; testing this directly would clarify when compression improves rather than degrades profiles.","Releasing JUMP-lite at four compression levels lets future work study how compression interacts with specific cell types, assays, and annotation schemes, and could guide compression choices for specialized benchmarks.","The same harness should make it straightforward to add new representation methods and re-run the entire benchmark without re-tuning the evaluation, which would accelerate comparisons as new foundation models appear."],"forward_implications":["A researcher with a modest server can run the full JUMP-lite benchmark at MQ compression and expect model rankings to match those from the uncompressed 115 TB dataset.","The ~97–99% storage reduction at HQ and MQ comes with average per-task performance changes of -1.2% and -6.5%, so conclusions about which representations are strong or weak do not hinge on using raw images.","Because aggressive d20 compression degrades PA by 20–34% and PC by up to 31%, the paper establishes a usable compression range and identifies settings that should be avoided.","CellProfiler features and MorphEM lead the benchmark, meaning classical engineered features remain a competitive baseline for image-based profiling, despite being roughly 200x slower than deep-learning pipelines."],"supporting_citations":[{"why":"Supplies the 115 TB JUMP Cell Painting dataset and its perturbation collection, the resource JUMP-lite samples from.","marker":"[12]"},{"why":"Provides the RefChem compound–target records used to select high-confidence compounds and to define phenotypic consistency ground truth.","marker":"[26]"},{"why":"Supplies MOTIVE's drug–target relationship graphs used for compound selection and the cross-modality retrieval evaluation tasks.","marker":"[2]"},{"why":"Defines the copairs retrieval metrics (PA/PC and NAP) that score every representation in the benchmark.","marker":"[27]"},{"why":"Establishes the prior compact benchmark, RxRx3-core, that JUMP-lite extends to Cell Painting at the scale of modern high-content screens.","marker":"[30]"},{"why":"Implements the cp measure feature extractor that produces the classical CellProfiler profiles benchmarked in the paper.","marker":"[33]"},{"why":"Provides the DINOv2 vision transformer used as one of the benchmarked deep-learning representations.","marker":"[34]"},{"why":"Orchestrates the ALIBY featurization pipeline that all models plug into for standardized image processing and profile generation.","marker":"[32]"},{"why":"Prior work on lossy image compression for microscopy that motivates the choice of JPEG XL and the compression-level sweep.","marker":"[52]"}],"fun_headline_variants":["JPEG XL shrinks 115 TB to 116 GB without shifting rankings","Moderate lossy compression preserves Cell Painting model rankings","1000x smaller cell-image benchmark keeps model rankings","Lossy compression keeps cell-image benchmarking fair at 116 GB"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's comparison of representation methods assumes that the classical feature pipeline, which uses 6-9 imaging sites per well, and the deep learning models, which use 4 sites per well, are directly comparable.","fun_headline_variants_meta":{"raw":{"variants":["JPEG XL shrinks 115 TB to 116 GB without shifting rankings","Moderate lossy compression preserves Cell Painting model rankings","1000x smaller cell-image benchmark keeps model rankings","Lossy compression keeps cell-image benchmarking fair at 116 GB"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001018,"raw_usage":{"total_tokens":4270,"prompt_tokens":891,"completion_tokens":3379,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":507,"completion_tokens_details":{"reasoning_tokens":3309}},"tokens_in":507,"tokens_out":3379,"duration_ms":25610,"temperature":1.0,"reasoning_tokens":3309,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:27:57.096883+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the classical CellProfiler profiles from the same four imaging sites per well used for deep-learning embeddings and rerun the phenotypic activity and consistency tasks; if CellProfiler no longer ranks with MorphEM at the top, the three-tier ordering is an artifact of site count rather than of representation quality.","supporting_citations":[{"cited_title":"Michael Ando, John Arevalo, Melissa Bennion, Nicolas Boisseau, Adriana Borowa, Justin D","cited_arxiv_id":null,"evidence_quote":"Supplies the 115 TB JUMP Cell Painting dataset and its perturbation collection, the resource JUMP-lite samples from."},{"cited_title":"Workflow for defining reference chemicals for assessing performance of in vitro assays.Altex, 36(2):261, 2018","cited_arxiv_id":null,"evidence_quote":"Provides the RefChem compound–target records used to select high-confidence compounds and to define phenotypic consistency ground truth."},{"cited_title":"Carpenter, and Shantanu Singh","cited_arxiv_id":null,"evidence_quote":"Supplies MOTIVE's drug–target relationship graphs used for compound selection and the cross-modality retrieval evaluation tasks."},{"cited_title":"Kalinin, John Arevalo, Erik Serrano, Loan Vul- liard, Hillary Tsang, Michael Bornholdt, Al ´an F","cited_arxiv_id":null,"evidence_quote":"Defines the copairs retrieval metrics (PA/PC and NAP) that score every representation in the benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the prior compact benchmark, RxRx3-core, that JUMP-lite extends to Cell Painting at the scale of modern high-content screens."},{"cited_title":"Mu ˜noz, Tim Treis, Alexandr A","cited_arxiv_id":null,"evidence_quote":"Implements the cp measure feature extractor that produces the classical CellProfiler profiles benchmarked in the paper."},{"cited_title":"DI- NOv2: Learning Robust Visual Features without Supervision, 2024","cited_arxiv_id":null,"evidence_quote":"Provides the DINOv2 vision transformer used as one of the benchmarked deep-learning representations."},{"cited_title":"Mu ˜noz.Phenotyping Single Cells of Saccharomyces Cerevisiae Using an End-to-End Analysis of High-Content Time-Lapse Microscopy","cited_arxiv_id":null,"evidence_quote":"Orchestrates the ALIBY featurization pipeline that all models plug into for standardized image processing and profile generation."},{"cited_title":"Deep-Learning- Based Image Compression for Microscopy Images: An Empir- ical Study.Biological Imaging, 4:e16, 2024","cited_arxiv_id":null,"evidence_quote":"Prior work on lossy image compression for microscopy that motivates the choice of JPEG XL and the compression-level sweep."}],"review_version":1}