{"id":"1b695cd5-3274-4bed-a843-ff63ee6bc932","arxiv_id":"2412.15876","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"This viewpoint proposes an 'AI-in-the-loop' paradigm in which AI helps build and run biomedical visual analytics tools while human experts keep responsibility and control.","lead":"This paper lays out a vision for putting AI inside the loop of biomedical data-analysis tools while keeping human experts in charge. It describes how future AI assistants could help build, customize, and operate interactive data visualizations for medical and biological research.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central roadmap depends on generative AI mastering biomedical VA code and interaction, yet the paper itself identifies the scarcity of the very training data this requires; without a data strategy or benchmark evidence, Data2Code and related tasks remain unsupported conjectures.","rationale":"The reader's verdict is CONDITIONAL, which is the right level for a Visualization Viewpoint: the conceptual framework is coherent and the paper is unusually honest about its open challenges, including hallucination, legal constraints, and the scarcity of biomedical VA training data. No formal verification or benchmark can be demanded of a position paper. However, the roadmap's forward-looking thrust is load-bearing. The paper does not merely suggest using AI for assistance; it predicts that generative models will create the visualization pipeline itself. That prediction requires the very training resources the paper admits are missing. A single concrete benchmark of current LLM capability on biomedical VA code generation would not fully settle a future projection, but it would test whether the claimed trajectory has any existing empirical foothold. The comparative adoption-gap claim in the Conclusion is also unsupported: it cites a market-penetration statistic and asserts superiority of AI-in-the-loop over human-in-the-loop without comparative evidence. Both concerns point to the same underlying condition: the central vision's viability depends on future AI capability that is plausible but unestablished. Since the paper already frames itself as a viewpoint and the reader's CONDITIONAL verdict reflects this uncertainty, my analysis does not change the verdict; it sharpens the reason for the condition. If a future revision added a small feasibility study or explicitly reframed the roadmap as a research agenda rather than an expectation, the condition could be relaxed; if the benchmark failed badly, the roadmap would need substantial revision.","tokens_in":10841,"tokens_out":6322,"duration_ms":60172,"concrete_test":"Construct a held-out benchmark of biomedical VA tasks drawn from published open-source tools and papers (e.g., volume rendering of a specific modality, a multi-modal correlation view, a cohort-comparison dashboard). Prompt several current frontier LLMs to generate the corresponding visualization code or configuration, and have two expert visualization researchers independently score outputs for executability, correctness, and faithfulness to the data/task. If success rates are low even on familiar patterns and near zero on novel data types, the 'will be feasible' claim lacks current empirical support and the roadmap should be framed as explicitly speculative; if success is high, it would materially strengthen the paper's central vision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim in the paper is that a human-centered VA workflow with AI-in-the-loop will remain central for biomedical decision-making, and that the proposed workflow taxonomy (Data2Code, Data2Image, VisComponent2Tool, Task2Query, Task2Interact) will be a useful design framework. For this to hold, generative AI must become capable of producing correct, task-appropriate visualization components and interactions for complex biomedical data. The paper's own 'Challenges in Biomedical Data Analytics' section, items (5)-(7), identifies that biomedical training data are scarce, that most VA code is proprietary and therefore unavailable for training, and that expert analytical interactions are highly variable and difficult to validate. These are exactly the resources needed to train the proposed generative capabilities. The text asserts that Data2Code 'will be feasible' and that the pipeline 'will be part of our work environments in a few years' without offering a mechanism or evidence that this data/code/interaction bottleneck can be overcome. If it cannot, the engine of the Figure 2 roadmap stalls, even though the normative human-agency framing can survive. The conclusion's further claim that AI-in-the-loop 'will more effectively close' the radiology AI adoption gap than human-in-the-loop is an unsupported comparative prediction: reference [11] reports current low market penetration, not the relative effectiveness of either paradigm.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This Visualization Viewpoint argues that in biomedical data analytics, the increasing power of generative AI and LLMs will transform visual analytics workflows, but that agency and responsibility must remain with human experts. The authors propose an 'AI-in-the-loop' paradigm in which visual analytics serves as an interface integrating AI into human-centered workflows, in contrast to the traditional 'human-in-the-loop' framing. They present a workflow taxonomy (Data2Data, Data2Code, Data2Image, Data2VisComponent, Task2VisConfig, VisComponent2Tool, Task2Query, Task2Interact) and a spectrum from fully human-generated to end-to-end AI-generated VA tools, and discuss ethical/legal considerations and a set of research opportunities.","tokens_in":11071,"tokens_out":4167,"duration_ms":32979,"significance":"The central normative claim that human-centered VA should remain central in critical biomedical decision-making is coherent and consistent with current regulatory and ethical frameworks. The paper's explicit list of challenges (scarce biomedical training data, proprietary VA code, hard-to-validate expert interactions) is a useful contribution, and the proposed taxonomy may serve as a design framework. If the feasibility assumptions are met, the viewpoint provides a clear agenda for the visualization community. However, the paper does not provide evidence or a mechanism for the key capability assumptions underlying the roadmap, and one comparative prediction in the conclusion is unsupported. The contribution is a viewpoint worth publishing after revision, not a result with validated evidence.","major_comments":[{"comment":"The roadmap in Figure 2 rests on the assumption that generative AI will become able to produce correct, task-appropriate visualization components and interactions for complex biomedical data. The paper asserts this repeatedly, e.g., 'As LLMs evolve and specialize, direct generation of custom visualization code adapted to the current dataset will be feasible' and 'the process of creating a tool by mixing and linking different visualization components ... will be achievable in the near future', but the same paper's Section 'Challenges in Biomedical Data Analytics', items (5)-(7), identifies that biomedical training data are scarce, that most VA code is proprietary and unavailable for training, and that expert interactions are highly variable and difficult to validate. These are precisely the resources needed to train the proposed generative capabilities. The text does not offer a mechanism, a data strategy, or benchmark evidence that this bottleneck can be overcome. As written, this is a load-bearing unsupported conjecture rather than an argued roadmap. Please either provide evidence or a concrete strategy for overcoming the data/code/interaction bottleneck, or reframe these statements as open research questions and explicitly discuss the dependence of Figure 2 on this capability assumption.","section":"The future of VA workflows / Challenges in Biomedical Data Analytics"},{"comment":"The concluding claim that 'our proposed AI-in-the-loop approach ... will more effectively close this gap than human-in-the-loop approaches' is an unsupported comparative prediction. Reference [11] reports current market penetration of radiology AI products (~2%) and discusses barriers to adoption; it does not provide evidence comparing the effectiveness of AI-in-the-loop versus human-in-the-loop paradigms. The manuscript offers no data, case study, or analytical argument that would support this comparative claim. Please either remove the comparison or reframe it explicitly as an untested hypothesis, since as stated it exceeds what the cited material and the paper's own argument can support.","section":"Conclusion"}],"minor_comments":[{"comment":"In the 'Current VA status' subsection, 'thehuman-in-the-loop' should read 'the human-in-the-loop'.","section":"Visual Analytics Workflows: present and future"},{"comment":"In item (5), 'the the development' contains a duplicated 'the'; it should read 'the development'.","section":"Challenges in Biomedical Data Analytics"},{"comment":"The naming of the task is inconsistent: the list uses 'Data2Viscomponent' while Figure 2 and later text use 'Data2VisComp'; please standardize.","section":"The future of VA workflows"},{"comment":"The phrase 'Metas' llama3' should be written as 'Meta's Llama 3' for readability and correctness.","section":"Dangers, Ethical, and Legal Considerations"},{"comment":"The caption uses 'Task2Config' and 'Component2Tool' while the body uses 'Task2VisConfig' and 'VisComponent2Tool'; please align the shorthand with the taxonomy terminology.","section":"Figure 3 caption"}],"recommendation":"major_revision","confidential_remarks":"Given the venue (Visualization Viewpoint), the lack of empirical validation is acceptable for a position paper. However, the two load-bearing issues—the feasibility of the generative roadmap given the paper's own stated data bottlenecks, and the unsupported comparative prediction in the conclusion—should be resolved before publication. The manuscript is otherwise a clearly written and well-structured contribution to the ongoing discussion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a Viewpoint, not a research contribution, and it's a decent one. The \"AI-in-the-loop\" inversion of \"human-in-the-loop\" is a genuinely useful reframing—it keeps agency and responsibility with the human expert while treating AI as the integrated component inside the analytics loop. The workflow taxonomy (Data2Data, Data2Code, Data2Image, Data2VisComponent, Task2VisConfig, VisComponent2Tool, Task2Query, Task2Interact) is a reasonable organizing scheme, and Figure 3's spectrum from fully human-generated to end-to-end AI generation is helpful. The paper is also honest about many obstacles: scarce biomedical training data, proprietary VA code, variable expert interactions, hallucination, regulatory constraints. For a roadmap, that honesty is a real strength.\n\nWhere it wobbles: the central vision depends on generative AI becoming able to produce correct, task-appropriate visualization code and interactions for complex biomedical data. The paper asserts Data2Code \"will be feasible\" and that the pipeline \"will be part of our work environments in a few years,\" but it doesn't offer a mechanism or evidence for overcoming the data and code scarcity it itself lists as challenges (5)–(7). That's a load-bearing assumption, and the stress-test note is right: without a data strategy or benchmark evidence, these tasks remain conjectures. The framing about human agency can survive that failure, but the roadmap's engine stalls.\n\nAlso, the conclusion's claim that AI-in-the-loop \"will more effectively close\" the radiology AI adoption gap than human-in-the-loop is unsupported. Reference [11] gives current market penetration (~2%); it says nothing about the relative effectiveness of either paradigm. That sentence should be softened or backed with evidence.\n\nThese are soft spots proportionate to the genre. For a Viewpoint, the paper is well within bounds; it's clearly written, internally consistent, and grounded in the right literature. The citations are appropriate and not self-promotional.\n\nWho is this for? Visualization researchers, especially in biomedical domains, and anyone thinking about human-AI teaming in data analytics. It could organize a research agenda and is a reasonable discussion piece.\n\nRecommendation: send it to peer review as a Viewpoint. It deserves referee time; expect some revision on the feasibility claims and the comparative prediction.","headline":"A coherent, useful roadmap for biomedical VA with AI-in-the-loop, but its engine is a bet on generative AI capabilities that the paper itself admits lack training data; the comparative adoption claim is unsupported.","tokens_in":11610,"tokens_out":1884,"would_cite":true,"duration_ms":16919,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that biomedical visual analytics should put AI in the loop while keeping human experts in charge of agency and responsibility.","keywords":["AI-in-the-loop","visual analytics","biomedical visualization","large language models","generative AI","human-centered workflows","workflow taxonomy","explainability"],"falsifier":"A prospective comparison study would settle the claim: if an end-to-end AI-generated visualization tool with no human audit matches or exceeds the reliability of an AI-in-the-loop tool with a human expert in real biomedical decision tasks, for example in tumor-board treatment planning, then the paper's insistence on human agency and visual analytics as the integrating interface would lose its empirical basis.","tokens_in":10622,"feed_emoji":"🧠","tokens_out":6441,"duration_ms":51227,"temperature":0.7,"pith_summary":"This paper argues that the future of biomedical visual analytics should be organized around 'AI-in-the-loop' rather than 'human-in-the-loop': instead of treating humans as a feedback module inside an AI system, AI should be inserted as an assistive component inside human-centered visual analytics workflows. The authors propose a taxonomy of eight AI-supported steps, covering data processing, code generation, image creation, component configuration, dashboard assembly, natural-language queries, and intent-driven interaction, along with a spectrum ranging from fully human-generated tools to end-to-end AI generation. They contend that because biomedical decisions are high-stakes and the data are complex, volatile, and poorly standardized, human experts must keep agency and auditability, and visual analytics is the natural medium for that oversight. If the argument is right, future tools will be designed as modular, inspectable human-AI collaborations, and visualization researchers will shift from implementing routine code to designing new AI-assisted workflows and validation methods.","feed_headline":"AI-in-the-loop, not human-in-the-loop, for biomedical visual tools","feed_subtitle":"A roadmap for keeping expert agency while generative AI builds and powers biomedical visualizations.","key_machinery":"The paper's central organizing device is the 'AI-in-the-loop' workflow model, together with an eight-task taxonomy of AI-supported pipeline steps and an automation spectrum from fully human-generated to fully AI-generated tools. The taxonomy (Data2Data, Data2Code, Data2Image, Data2VisComponent, Task2VisConfig, VisComponent2Tool, Task2Query, Task2Interact) defines where AI enters the classic visual analytics pipeline, while the spectrum shows how these steps can combine into systems of varying automation. This machinery carries the argument by giving researchers a concrete language for describing and building future AI-assisted biomedical VA tools, and by making explicit that a fully end-to-end AI tool would require regeneration rather than reconfiguration when intent changes.","core_discovery":"The central claim is that agency and responsibility must remain with human experts in biomedical decision-making, and that visual analytics therefore should be used as a tool for integrating AI into human-centered workflows, which is the reverse of the common 'human-in-the-loop' framing that inserts humans into AI systems. On this view, AI becomes an integral part of the visual analytics loop: it assists with data processing (Data2Data), generates visualization code and components (Data2Code, Data2VisComponent), creates images directly from data (Data2Image), configures and assembles tools (Task2VisConfig, VisComponent2Tool), and mediates natural-language queries and interactions (Task2Query, Task2Interact). The paper argues that such AI-supported workflows can be mixed at various automation levels, and that even near-future generative systems, despite hallucinations and reliability limits, will make low-code biomedical VA tool development feasible while human experts audit each result. The claimed consequence is that biomedical visual analytics will remain central, and that the approach could help close the gap between approved AI radiology products and their actual 2% market penetration.","pith_inferences":["The taxonomy could be turned into a benchmark: each of the eight arrows names a capability that can be evaluated separately, giving the community measurable milestones for how close generative AI is to realizing the roadmap.","The 'regeneration problem' the paper identifies, where intent-driven changes to an end-to-end tool require full regeneration, suggests a design principle that component-based, configurable AI tools will likely dominate in practice over monolithic end-to-end generation.","The AI-in-the-loop framing may generalize beyond biomedicine to other regulated, high-stakes domains such as financial auditing or public-safety analytics, where explainability and human sovereignty are similarly mandated.","A direct test of the roadmap would be to implement a small biomedical VA scenario, for example a multi-modal tumor-board view, and measure whether current LLMs can produce the Data2Code and Task2Query stages with acceptable reliability; the paper does not perform this test."],"forward_implications":["Tool builders will be able to mix and match AI-supported stages, for instance a human-written component with an AI-generated configuration, or an AI-generated view with a human-crafted dashboard.","For high-stakes uses, every AI-generated component must remain inspectable and editable because the human expert is ultimately responsible, and closed end-to-end generation is the least desirable end of the spectrum.","Task2Query and Task2Interact will let users steer complex biomedical data with natural language and predicted intentions, but only when the underlying AI can reliably interpret high-level structures such as anatomy.","VA researchers will spend less time on routine code generation and more on new algorithms, interface designs, and validation tools, since standard components can be produced by assistive AI.","The biomedical field will adopt AI-assisted VA more slowly than recreational AI applications because training data and code are scarce and mostly proprietary."],"supporting_citations":[{"why":"Documents that seamless integration of deep learning into VA tools has a long path ahead, motivating the need for a new paradigm.","marker":"[1]"},{"why":"Defines visual analytics and the human-in-the-loop principle that the paper reframes into AI-in-the-loop.","marker":"[5]"},{"why":"Shows a current LLM-based tool that can generate only simple charts from data, establishing the gap the roadmap aims to close.","marker":"[6]"},{"why":"Supplies evidence that LLMs still hallucinate and give unreliable explanations, motivating the human-agency and auditability requirement.","marker":"[7]"},{"why":"Presents the 'generalist medical AI' vision that the paper argues against full automation and for human oversight.","marker":"[8]"},{"why":"Reports the roughly 2% market penetration of approved radiology AI products, motivating the claim that human-centered integration will close the adoption gap.","marker":"[11]"}],"fun_headline_variants":["AI-in-the-loop: humans stay in charge of biomed visuals","Reverse the loop: AI serves, humans command biomed visuals","Human agency remains: AI joins the biomedical visual loop","AI in the loop, humans in charge for biomedical visual tools"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The roadmap assumes that generative AI and large language models will become capable of correctly and reliably producing custom visualization code, direct data-to-image transformations, natural-language queries, and intent-driven interactions for complex biomedical data, despite the paper's own list of obstacles (scarce training data, proprietary VA code, hallucination, and hard-to-validate expert interactions).","fun_headline_variants_meta":{"raw":{"variants":["AI-in-the-loop: humans stay in charge of biomed visuals","Reverse the loop: AI serves, humans command biomed visuals","Human agency remains: AI joins the biomedical visual loop","AI in the loop, humans in charge for biomedical visual tools"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000863,"raw_usage":{"total_tokens":3766,"prompt_tokens":990,"completion_tokens":2776,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":2706}},"tokens_in":606,"tokens_out":2776,"duration_ms":16222,"temperature":1.0,"reasoning_tokens":2706,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:59:56.187148+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A prospective comparison study would settle the claim: if an end-to-end AI-generated visualization tool with no human audit matches or exceeds the reliability of an AI-in-the-loop tool with a human expert in real biomedical decision tasks, for example in tumor-board treatment planning, then the paper's insistence on human agency and visual analytics as the integrating interface would lose its empirical basis.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents that seamless integration of deep learning into VA tools has a long path ahead, motivating the need for a new paradigm."},{"cited_title":"Abramson, J","cited_arxiv_id":null,"evidence_quote":"Defines visual analytics and the human-in-the-loop principle that the paper reframes into AI-in-the-loop."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows a current LLM-based tool that can generate only simple charts from data, establishing the gap the roadmap aims to close."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies evidence that LLMs still hallucinate and give unreliable explanations, motivating the human-agency and auditability requirement."},{"cited_title":"Cui and T","cited_arxiv_id":null,"evidence_quote":"Reports the roughly 2% market penetration of approved radiology AI products, motivating the claim that human-centered integration will close the adoption gap."}],"review_version":1}