{"id":"986604ad-8999-4cc2-93f5-1f59bc82301e","arxiv_id":"2505.10993","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review that taxonomizes over 150 generative-model papers in computational pathology into four generation-task domains and discusses datasets, evaluation, and clinical barriers.","lead":"This preprint surveys content generation models in computational pathology, covering more than 150 studies organized into image, text, molecular, and other generation tasks. It maps a fast-growing field and lays out open challenges, making it a possible entry point for those tracking AI in medicine.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survey's own scope definition is violated: Table V classifies encoder-only models (PLIP, Prov-GigaPath, GPFM) as generative, undermining the taxonomy's internal coherence.","rationale":"The reader's weakest assumption concerned the completeness of the literature search, pointing to the missing Supplementary Materials Section II. That is a legitimate concern but cannot be resolved from the preprint text. My review identified a different, more immediately checkable load-bearing issue: the survey's taxonomy is internally inconsistent with its own stated scope. The paper repeatedly emphasizes that only models with a decoder that synthesizes new images or text are in scope, yet Table V places encoder-only and contrastive models such as PLIP, Prov-GigaPath, and GPFM under 'Latent Representation Generation.' These models do not create new content in the sense the paper defines; they output feature vectors or embeddings. A reader relying on the taxonomy to understand the field of content generation would be misled about what counts as generative. This is a concrete, text-locatable inconsistency rather than an unverifiable external gap. It weakens the paper's central organizational contribution, so the conditional verdict is appropriate. However, the concern does not invalidate the survey's overall utility; it suggests the need for revision or clarification of the inclusion criteria and table entries rather than rejection. Thus I recommend keeping the verdict unchanged, but note that the basis for the conditional status should include this internal inconsistency. I disagree with the reader's choice of weakest assumption because the scope inconsistency is more specific and testable from the provided manuscript, whereas the literature-search completeness issue is deferred to missing supplementary material and cannot be evaluated here.","tokens_in":32818,"tokens_out":4560,"duration_ms":46026,"concrete_test":"Check every entry in Tables I–V against the paper's own definition of generative capacity. For each cited method, inspect the original publication's architecture to determine whether it contains a decoder that synthesizes new images or text. Pay particular attention to PLIP [177], Prov-GigaPath [181], and GPFM [182] in Table V: confirm whether they have any generative decoder or only encoders/contrastive heads. Also verify whether Redekop et al. [83] (an image generation method) belongs in the text generation table (Table III). If these models lack generative decoders, or if the Redekop entry is confirmed to be misclassified, the survey's taxonomy fails its own scope definition, and the paper should explicitly revise either the inclusion criteria or the table entries.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of the survey is to provide a comprehensive, organized review of content generation models, with a taxonomy as a core contribution. The paper explicitly defines its scope in Section I: 'The scope of this review is limited to models with generative capacity, defined by the presence of a decoder that can synthesize new images or text. In contrast, encoder-only or contrastive foundation models, such as CLIP, ALIGN, or DINOv2, are considered non-generative.' Yet Section III-D.3 and Table V classify several models that lack any generative decoder as 'Latent Representation Generation.' For example, PLIP [177] is a contrastive image-text model that learns joint embeddings for retrieval and classification; Prov-GigaPath [181] is a ViT encoder producing slide-level features; and GPFM [182] distills a universal feature backbone. None of these has a decoder that synthesizes new images or text. They generate feature representations, not content under the paper's own definition. This is not a borderline case; it directly contradicts the stated inclusion criterion. Additionally, Table III, devoted to text generation, lists Redekop et al. [83], a prototype-guided diffusion model for image synthesis, under Report Generation. That is a categorical misassignment. These internal inconsistencies call into question the rigor of the taxonomy, which is one of the survey's principal claimed contributions. If the boundary between generative and non-generative is applied inconsistently, then the survey is not strictly a survey of content generation models, and the comprehensiveness claim becomes difficult to interpret or verify. This concern is load-bearing because the survey's value rests on its organizational framework; if the framework's categories are not applied consistently, the reference utility for readers is weakened.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a survey of content generation models in computational pathology, claiming to be the first comprehensive review of the field. It organizes more than 150 publications into a taxonomy with four major domains: image generation, text generation, molecular profile–morphology generation, and other specialized generation tasks. The paper reviews underlying model families (VAEs, GANs, diffusion models, generative vision-language models), compiles commonly used datasets, and discusses capabilities, challenges, and future directions such as generative foundation models, agent-driven pathology, and the AI virtual cell. The central contribution is the proposed taxonomy and the synthetic organization of a rapidly growing literature.","tokens_in":33096,"tokens_out":6237,"duration_ms":60136,"significance":"If the taxonomy is internally consistent and the coverage is genuinely comprehensive, this survey would be a valuable reference for researchers entering the field. The paper is timely, covers a broad and rapidly expanding literature, and provides a balanced discussion of technical limitations, evaluation gaps, and clinical translation barriers. It is transparent about the absence of standardized benchmarks, which is an important and honest observation. The most valuable strengths are the breadth of the compiled literature, the structured tables organizing methods by task and architecture, and the practical compilation of datasets. However, the value of the survey hinges on the reliability of its taxonomy, and the internal inconsistencies detailed below currently weaken that contribution.","major_comments":[{"comment":"The paper explicitly defines its scope in Section I to include only models with generative capacity, i.e., those with a decoder that can synthesize new images or text, and it explicitly excludes encoder-only or contrastive foundation models such as CLIP, ALIGN, and DINOv2. Yet Section III-D.3 and Table V classify PLIP [177], Prov-GigaPath [181], and GPFM [182] under 'Latent Representation Generation' as if they were generative models. These are respectively a contrastive image-text model, a ViT encoder for slide-level features, and a knowledge-distilled feature backbone; none of them contains a decoder that synthesizes new images or text. The text in Section III-D.3 even presents these as examples of 'latent representation generation,' conflating feature extraction with content generation. This directly contradicts the paper's own stated inclusion criterion and undermines the internal coherence of the taxonomy, which is one of the principal claimed contributions. Please either remove these and similar encoder-only entries (e.g., Hu et al. [44], PRDL [178]) from the generative taxonomy, or substantially broaden the scope definition to explicitly cover representation generation, and then revise the Introduction and the category description to match.","section":"Section I vs. Section III-D.3 / Table V"},{"comment":"The entry 'Redekop et al. [83] Prototype Diffusion Synthesizing images guided by unsupervised prototypes for data-efficient SSL' is listed under Report Generation, but it is a conditional diffusion model for image synthesis, not a report generation method. This method is not discussed in the text of Section III-B.3, and its placement contradicts the task definition. This categorical misassignment indicates that the table curation needs a systematic pass against the textual descriptions and task definitions, since a similar error could affect other entries.","section":"Table III, Report Generation"}],"minor_comments":[{"comment":"The abstract describes 'over 150 representative studies' while the contribution list claims 'covering more than 150 publications up to July 2025,' and the Introduction calls the review 'comprehensive.' Please clarify whether the selection is exhaustive or representative, and define the inclusion criteria in the main text rather than only in supplementary material.","section":"Abstract / Section I (Contributions)"},{"comment":"The search strategy and the evaluation metrics are referenced as 'Supplementary Materials Section II' and 'Supplementary Materials Table I,' but the preprint as provided does not include the supplementary file. To support the comprehensiveness claim, please make the search strategy and the metrics table available, or summarize them in the main text.","section":"Section I / Section IV"},{"comment":"Several references lack complete bibliographic information; for example, [48] ('Diffusion-based generation of histopathological whole slide images at a gigapixel scale') and [78] ('Insmix') have no venue or arXiv identifier, and [115] ('Pathldm') has no venue. Please complete these entries for reproducibility.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The main blocker is the internal inconsistency between the scope definition and the inclusion of encoder-only models such as PLIP, Prov-GigaPath, and GPFM in the taxonomy. This is a fixable issue but one that requires a systematic revision of the taxonomy and the tables, not just a local edit. The survey is otherwise a useful reference, and I would look favorably on a revised version that resolves these inconsistencies and provides the missing supplementary search details."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good broad survey; the taxonomy has a real internal inconsistency that needs fixing before the 'comprehensive' claim is believable.\n\nThe paper does what a good survey should: it collects over 150 papers on generative modeling in computational pathology, arranges them by generation target (image, text, molecular-morphology, other), gives a decent historical arc from GAN/VAE to diffusion to VLMs, and includes a practical dataset table. I also appreciate the extended discussion of evaluation gaps and clinical deployment barriers; that section is honest and useful. For a newcomer, this is a reasonable map of the field.\n\nThe soft spot is the taxonomy. The authors explicitly define generative capacity as the presence of a decoder that synthesizes new images or text, and explicitly exclude contrastive models like CLIP and DINOv2. Then in Table V, under 'Latent Representation Generation,' they list PLIP (a contrastive image-text model), Prov-GigaPath (a ViT encoder), and GPFM (a distilled feature backbone). None of these has a decoder that synthesizes content. That is not a minor edge case; it is a direct violation of the stated inclusion criterion. Table III also has an obvious misassignment: Redekop et al., a prototype-guided diffusion model for image synthesis, appears under Report Generation in the text generation table. These errors undermine the taxonomy, which is one of the paper's main selling points.\n\nAlso, the 'comprehensive' claim depends on a search strategy that lives in missing supplementary material, so I cannot verify coverage. Several reference entries have incomplete author or venue data, which matters for a reference work. The discussion of 'no standardized benchmarks' is accurate and well-stated.\n\nOverall, the paper is useful and the core content is solid. But the scope definition and the tables disagree in places, and the authors need to either fix the tables or soften the scope claim. I would send it to peer review with a request for a careful pass over the taxonomy and reference list.","headline":"Useful broad survey with a genuine taxonomy flaw: encoder-only models are listed as generative, contradicting the paper's own scope.","tokens_in":33570,"tokens_out":2802,"would_cite":true,"duration_ms":27499,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims to be the first comprehensive map of content generation models in computational pathology, organizing more than 150 studies into four task families.","keywords":["computational pathology","generative models","diffusion models","generative adversarial networks","vision-language models","whole-slide images","synthetic data","literature survey"],"falsifier":"A reader could assemble an independent bibliography of content-generation papers in computational pathology published before July 2025 and check whether a substantial number of peer-reviewed works fall outside the four taxonomy categories or are absent from the survey's list of over 150 studies; finding such a body would refute the comprehensiveness claim.","tokens_in":32570,"feed_emoji":"🧫","tokens_out":5901,"duration_ms":57785,"temperature":0.7,"pith_summary":"This paper sets out to establish that content generation models have become a distinct, fast-growing methodological paradigm in computational pathology, and that the field is now mature enough to be surveyed under one taxonomy. It claims to be the first comprehensive survey of this area, covering more than 150 publications up to July 2025 and organizing them into four task families: image generation, text generation, molecular profile–morphology generation, and other specialized generation tasks. The review matters because pathology images are gigapixel, annotation-heavy, and heterogeneous, so the ability to synthesize realistic tissue, virtual stains, diagnostic reports, and molecular readouts could lower the cost of data, improve model robustness, and open archival slides to molecular analysis. The paper's contribution is the structure itself: a taxonomy, a dataset inventory, and a statement of the barriers that stand between current models and clinical deployment.","feed_headline":"Survey maps 150+ generative-model studies in computational pathology","feed_subtitle":"Image, text, and molecular-profile generation across 150+ papers, plus the barriers that block clinical use.","key_machinery":"The load-bearing device is the proposed taxonomy together with the definition of generative capacity. A model counts as a content generation model only if it has a decoder that can produce new images or text; encoder-only contrastive models are explicitly excluded. The taxonomy then sorts the field into four generation targets — image, text, molecular profile–morphology, and other — with sub-tasks under each, and the survey uses this grid to compare architectures (GAN/VAE, diffusion, LLM/VLM) across applications, datasets, and evaluation protocols. The taxonomy is what lets the review claim to be a reference framework rather than a list of papers.","core_discovery":"The central claim is that content generation in computational pathology is best understood as a family of tasks unified by a single criterion: the presence of a decoder that synthesizes new images or text, which separates generative models from encoder-only or contrastive foundation models. On that basis the survey proposes a four-part taxonomy — image synthesis (augmentation, mask-guided generation, artifact restoration, resolution scaling, text-to-image, stain synthesis), text generation (captioning, visual question answering, report generation, abstraction), molecular profile–morphology generation (virtual molecular profiling and reverse morphology synthesis), and other tasks (spatial layout, semantic outputs, latent representations, cell simulation). It traces an architectural arc from variational autoencoders and GANs through diffusion models to generative vision-language models, and argues that since roughly 2024 the field has shifted from proof-of-concept synthesis toward task-oriented and clinically motivated generation. The paper also claims that the main obstacles are now the fidelity of whole-slide image synthesis, the absence of pathology-specific evaluation metrics, computational cost, and unresolved ethical, legal, and regulatory questions.","pith_inferences":["If the taxonomy's categories are the right ones, testable predictions follow: methods developed for one sub-task in a family, such as cycle-consistency in stain transfer, should transfer to other sub-tasks in the same family, such as artifact restoration, with minimal modification.","The paper's own cost figures imply an equity consequence it states but does not develop: as generative foundation models grow, only well-resourced groups will train them, so democratization depends on open pretrained weights and efficiency research.","The emphasis on an 'AI virtual cell' suggests a concrete experiment the field could run: use a generative model to simulate the morphological effect of a genetic perturbation, then check the prediction against real spatial transcriptomics data; the paper stops at recommending the direction."],"forward_implications":["A newcomer can use the taxonomy to locate any pathology generation method and its nearest alternatives, which shortens the path from problem statement to candidate architecture.","The survey's division of labor — GANs for speed, diffusion for fidelity, VLMs for language — gives practitioners a first-pass selection rule when choosing a generation model for a task.","The identified gap in pathology-specific evaluation metrics implies that reported FID/SSIM improvements do not yet translate into clinical claims; standardized benchmarks are a prerequisite for comparing methods.","The proposed future of generative foundation models would unify image, text, and molecular generation in one system, making cross-modal tasks such as gene-to-image synthesis a single-model capability.","The catalogued datasets, including synthetic benchmarks like SNOW and PathGen-1.6M, provide the raw material for reproducing or extending most surveyed methods."],"supporting_citations":[{"why":"Establishes that existing surveys treat generation as one application among many, defining the gap this survey fills.","marker":"[23]"},{"why":"Shows that prior work focused on evaluation of foundation models rather than pathology-specific generation, reinforcing the need for a dedicated survey.","marker":"[24]"},{"why":"Supplies the variational autoencoder formalism that anchors the earliest model family in the timeline.","marker":"[28]"},{"why":"Introduces the adversarial training setup underlying early pathology synthesis methods.","marker":"[31]"},{"why":"Defines the denoising diffusion process that the survey identifies as the current fidelity leader.","marker":"[45]"},{"why":"Provides the latent diffusion and conditioning machinery used by many pathology generation pipelines.","marker":"[56]"},{"why":"Establishes visual instruction tuning, the basis of the generative vision-language lineage the survey traces.","marker":"[63]"},{"why":"Exemplifies the multimodal pathology assistant used to support text-generation applications and clinical copilot claims.","marker":"[18]"},{"why":"Supports the survey's claims about scalable mask-to-image synthesis via parallel patch diffusion.","marker":"[15]"},{"why":"Shows gigapixel whole-slide generation, the capability the survey flags as a key open challenge.","marker":"[48]"}],"fun_headline_variants":["Survey maps 150+ generative pathology studies into four tasks","Pathology generation: image, text, molecules in 150+ papers","From GANs to diffusion: taxonomy of pathology content generation","150+ studies on generative pathology models and their clinical gaps","Pathology content generation: a four-part framework and its limits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comprehensiveness claim rests on the literature search described in Supplementary Materials Section II, which is not included in this preprint; if that search missed a substantial body of relevant work, the survey's value as a complete reference would be weakened.","fun_headline_variants_meta":{"raw":{"variants":["Survey maps 150+ generative pathology studies into four tasks","Pathology generation: image, text, molecules in 150+ papers","From GANs to diffusion: taxonomy of pathology content generation","150+ studies on generative pathology models and their clinical gaps","Pathology content generation: a four-part framework and its limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000358,"raw_usage":{"total_tokens":1933,"prompt_tokens":933,"completion_tokens":1000,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":914}},"tokens_in":549,"tokens_out":1000,"duration_ms":9887,"temperature":1.0,"reasoning_tokens":914,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:59:00.729506+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could assemble an independent bibliography of content-generation papers in computational pathology published before July 2025 and check whether a substantial number of peer-reviewed works fall outside the four taxonomy categories or are absent from the survey's list of over 150 studies; finding such a body would refute the comprehensiveness claim.","supporting_citations":[],"review_version":1}