Pith. sign in

REVIEW 4 major objections 5 minor 87 references

Medillustrator: Improving Retrospective Learning in Physicians' Continuous Medical Education via Multimodal Diagnostic Data Alignment and Representation

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Medillustrator is a visual analytics system that aligns MRI images, diagnostic text, and lab indicators so novice physicians can screen, analyze, and revisit high-value teaching cases more effectively than with their standard hospital…

desk verdict Solid system paper whose main claim about learning is untested; worth reviewing, but the evaluation needs a reframing or a retention measure. read the letter →

arxiv 2411.15593 v1 pith:WKLBWQMR submitted 2024-11-23 cs.HC

classification cs.HC
keywords RetrospectiveLearningofPhysiciansVisualAnalyticsMultimodalDataAlignmentContinuousMedicalEducationImagingCase-BasedNoviceUserStudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that retrospective learning—physicians reviewing past patient cases to improve their skills—can be made substantially easier for novices by presenting multimodal diagnostic data as one aligned, navigable whole. Medillustrator combines an overview of patient mentions and embedding projections, a detail view where diagnostic text is overlaid on the corresponding areas of MRI images, reference ranges for lab indicators, and a record view for saving insights. In a controlled study with 13 novice physicians, the group using Medillustrator selected more of the ten benchmark cases a senior panel had judged valuable and wrote more accurate rationales than the group using the hospital information system, and they also rated the system higher on usability. If the result holds, hospitals and continuing medical education programs could move from senior-physician intuition toward a structured, tool-assisted workflow for turning routine patient records into learning material.

What carries the argument

The carrying object is the Medillustrator system itself, specifically its multimodal alignment and representation pipeline. The pipeline has four load-bearing parts: a text-to-image grounding step that connects diagnostic phrases to bounding boxes and then to pixel-level segmentation in MRI images; an embedding step that projects image, text, and indicator features into a shared low-dimensional space where each patient is drawn as a glyph whose three segments encode the overlap similarity of nearest-neighbor sets across modality pairs; a detail view with a practice phase (raw image, user annotations, imaging indicators on parallel axes) and a learning phase (aligned diagnostic text layered onto the image by a force-directed layout); and a record view that saves analyzed cases as cards for later review.

What would settle it

Check the rating scales with a different panel: have an independent set of senior physicians label the same 50-case set, measure inter-rater reliability, and rerun the task; if the Medillustrator advantage disappears against a differently labeled gold standard, the result is an artifact of the chosen ten. Alternatively, give both groups a delayed retention test or follow-up supervised case discussion weeks later; if the Medillustrator group shows no better recall or diagnostic performance, the "learning" improvement is task-specific, not retrospective learning.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that a visual analytics approach built around semantic alignment of multimodal data supports novice physicians' retrospective learning better than the conventional record browser. The system's design follows a discovery-learning structure: start with an overview to find high-value cases, analyze a case in detail with practice-then-learning phases, then record and revisit conclusions. Its evaluation shows the Medillustrator group outperforming the baseline on both measured report dimensions—completeness (median 5.71 vs 5.57 on a 10-point scale) and accuracy (4.16 vs 3.5)—as well as on perceived ease of use, screening help, and confidence, with significance on a non-parametric rank test.

Load-bearing premise

The effectiveness claim rests on treating the ten cases chosen by two senior physicians as the ground truth for what is worth learning, and treating how many of those cases a novice selects (plus the precision of the written rationale) as the measure of learning, without measuring the reliability of that ground truth or showing that the task predicts real clinical improvement.

Editorial extensions

If this is right

  • If the measured advantage is real, novice physicians can use the embedding glyphs to detect cases where image and indicator signals disagree with the text diagnosis—precisely the cases most likely to be misdiagnosed and most valuable to study.
  • The practice-then-learning split lets trainees commit to their own reading of an image before the system reveals the physician-aligned diagnosis, turning passive review into an active retrieval exercise.
  • Because the alignment pipeline runs on physician annotations, the same workflow could be rebuilt for other image-heavy specialties by repeating the annotation and fine-tuning process.
  • Recorded case cards give trainees a persistent artifact for retrospection, which the formative interviews identified as a bottleneck in current continuing medical education practice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The study's completeness scores are close (medians 5.71 vs 5.57), so the larger, more robust benefit may be screening efficiency and user confidence rather than raw identification; a longer-term test with retention measures would clarify whether the learning gain is durable.
  • A natural testable extension is to replace the senior-physician gold standard with outcome-based labels—for example, cases that later involved diagnostic revision or adverse events—and see whether the system helps novices discover those cases too.
  • The text-to-image grounding component could be reused as a clinical explanation aid beyond training, for example to highlight what a diagnostic sentence refers to in a scan during case conferences or decision support.
  • Tracking physicians' record-view entries over many sessions would turn the tool into a portfolio that measures learning curves, something a single one-hour experiment cannot show.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents Medillustrator, a visual analytics system intended to support novice physicians' retrospective learning from multimodal diagnostic data (MRI images, diagnostic text, and laboratory indicators). The system includes a data processing pipeline (image annotation, Grounding DINO fine-tuning, SAM segmentation), a modeling engine (ResNet-50 and BERT embeddings fused and projected with UMAP, k-NN neighborhood and Jaccard-based glyphs), and a five-view interface organized around discovery-learning levels (overview, detail, retrospection). The authors report a formative study with six physicians, a case study with one intern, and a controlled in-lab between-subject study with 13 participants comparing Medillustrator to a hospital HIS baseline. The user study reports significantly higher scores for the Medillustrator group on report completeness (U=3, p<0.01) and accuracy (U=3.5, p<0.05), as well as on several usability and effectiveness Likert items.

Significance. If the central claim were fully established, the system would address a real and underserved need in continuous medical education: helping novices identify high-value cases and interpret multimodal data in a unified view. The paper's strengths include a transparent description of the modeling pipeline, a realistic baseline comparison (the hospital HIS), a formative study that grounds design requirements, and a case study that illustrates the intended workflow in detail. The statistical reporting is transparent in the sense that U and p values are given. However, the evaluation as designed measures immediate screening and analysis performance, not learning outcomes. There is no retention test and no transfer task, and the ground-truth benchmark selection is not validated with inter-rater reliability. These issues are load-bearing because the abstract and conclusion frame the result as demonstrating enhanced retrospective learning. The system itself is plausible and the qualitative insights are useful, but the evidence does not yet support the broadest claims.

major comments (4)
  1. [§5.2 (Procedure) and §7 (Conclusion)] The study's dependent variables are the number of benchmark cases selected during a one-hour session and the precision of the written rationales. These measure immediate screening and analysis performance, not the acquisition, retention, or transfer of knowledge. The Abstract's claim that Medillustrator 'enhance[s] physicians' retrospective learning processes' and the Conclusion's statement that 'user evaluations demonstrate Medillustrator's effectiveness in aiding novice physicians to efficiently identify and analyze research cases' are not supported by the same evidence: the former is a learning claim, the latter is a performance claim. A delayed retention test or a transfer task on unseen cases would be required to support the learning claim; alternatively, the claims should be restricted to immediate performance.
  2. [§5.2 (Procedure)] The 10 benchmark cases are described as 'curated by two senior physicians' with no operational definition of what makes a case valuable and no inter-rater reliability reported for this selection. Because the completeness score is defined as the number of selected cases matching this ground truth, the validity of the completeness measure depends entirely on the reliability and validity of that judgment. The ICC value reported (above 0.75) is for the two senior raters who scored the reports, not for the benchmark selection itself. The paper should report the explicit selection criteria (e.g., diagnostic discrepancy, atypical presentation, educational value) and an inter-rater statistic such as Cohen's kappa for the benchmark selection. If the same two senior physicians both selected the benchmarks and scored the reports, the report scoring is also at risk of criterion contamination.
  3. [§5.2 (Procedure)] Participants in the Medillustrator condition received a 5-minute system tutorial followed by 10 minutes of free exploration before the one-hour task, while the baseline condition did not receive an equivalent orientation period. The measured effect may therefore be attributable to additional time on task or differential attention rather than to the system's design. The authors should either give the baseline group an equivalent familiarization activity (e.g., a structured orientation to the HIS features relevant to the task) or otherwise control for time and attention across conditions, and should discuss this confound in Section 6, which currently does not acknowledge it.
  4. [Table 2 and Appendix B] The median completeness scores are 5.71 (Medillustrator) versus 5.57 (baseline) on a 10-point scale. Under the scoring rubric in Appendix B, both values fall in the same band (5-6: 'Identifies 70-84% of cases'), corresponding to the same number of benchmark cases (7 of 10). The reported p<0.01 with U=3 (n1=7, n2=6) indicates a significant rank difference, but the absolute difference of 0.14 points is small and its practical significance is unclear. The paper should report effect sizes (e.g., rank-biserial correlation) and discuss whether the observed difference is educationally meaningful.
minor comments (5)
  1. [§2.2] The citation 'Duanm et al.' should read 'Duanmu et al.' (reference [19]).
  2. [§4.2.2] The model names 'resnet50' and 'bert-base-uncased' should be typeset consistently as 'ResNet-50' and 'BERT-base-uncased'.
  3. [§4.3.2] The three-set similarity formula JA,B,C = |A∩B∩C| / |A∪B∪C| is not the standard Jaccard index for three sets; either define the intended normalized intersection measure explicitly or use the standard formulation.
  4. [§5.2 (Participants)] The compensation amount '20 compensation' lacks a currency unit; please specify the currency and amount.
  5. [§5.1 Case Study Part II] In the description of step 13, 'c2 also be examined as abnormal' should read 'p2 was also examined as abnormal'.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: system outputs are computed independently of the outcome measure, and the evaluation uses an external baseline and externally curated benchmark ground truth.

full rationale

The paper contains no derivation chain that reduces to its own inputs. Medillustrator's embeddings, modality alignments, glyphs, and detail-view annotations are computed from the curated hospital dataset through pretrained and fine-tuned models (ResNet-50, BERT, UMAP, Grounding DINO, SAM), and none of these components is fitted to the study's outcome measure. The user study compares 13 participants against an external baseline (the HIS system) and against a ground truth of 10 senior-curated benchmark cases that was constructed independently of the system's outputs, so the completeness and accuracy scores are not forced by construction. The same-group self-citation (e.g., [68]) appears only as related-work support and is not load-bearing for the central claim; there is no invoked uniqueness theorem and no ansatz smuggled in by citation. The strongest legitimate concerns are construct-validity and experimental-confound issues (measuring one-session screening and rationale quality rather than a retention or transfer test, extra tutorial time in the Medillustrator arm, and the benchmark's unspecified labeling criteria), and these belong to correctness risk rather than circularity. Score 2 reflects the presence of minor non-load-bearing self-citations while the central empirical claim remains externally grounded.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new entities, particles, forces, or conserved quantities. Its free parameters are the standard fitting choices of a machine-learning and visualization pipeline: the fine-tuned alignment model, UMAP parameters, the k in kNN, and fusion weighting. The axioms are the evaluation assumptions: ground-truth case selection, baseline fairness, report score as a learning proxy, transferability of pretrained models, and annotation quality. The most consequential unexamined assumptions are the ground-truth definition and the lack of a learning measure beyond case-selection performance.

free parameters (4)
  • Grounding DINO fine-tuned alignment model = mAP 0.53 on 150 test pairs
    The model is trained on roughly 600 annotated image-text pairs and fine-tuned to perform the alignment that anchors text to image regions. The 50 epochs and the resulting mAP are reported, but no error analysis or downstream impact is given. This is a fitted component that the Detail View's central functionality depends on.
  • UMAP embedding parameters = Not reported
    UMAP is applied to image, text, indicator, and concatenated embeddings; the metric, n_neighbors, min_dist, and random seed are not reported. These choices affect the Embedding View layout and all kNN computations.
  • k in k-nearest neighbors = k = 5 shown in the case study
    The kNN sets and Jaccard similarities in the glyph design depend on the choice of k. The paper states k = 5 in the case study but does not specify the k used in the system or in the user study, or justify the choice.
  • Number of UMAP dimensions and fusion weights = Not reported
    The image, text, and indicator embeddings are concatenated to form the fused representation; no weighting or normalization scheme is reported. The relative scale of 2048-dimensional image features, 768-dimensional text features, and low-dimensional indicator features determines the fusion result.
assumptions (5)
  • domain assumption The hospital's HIS system is an appropriate baseline for a retrospective-learning tool.
    The user study compares Medillustrator against HIS because it is the tool physicians currently use for data browsing. If HIS is not a fair baseline (e.g., it was not designed for learning tasks), the comparison measures task-fit rather than educational effectiveness. This is stated in Section 5.2, Baseline System.
  • domain assumption The 10 'textbook-deviation' cases curated by two senior physicians are the correct ground truth for valuable learning cases.
    The evaluation assumes that cases deviating from typical symptom-description patterns are the valuable ones. No operational definition, selection protocol, or inter-rater reliability is given for this curation. This is the central measurement assumption of the user study (Section 5.2, Procedure).
  • domain assumption Report completeness and accuracy scores measure retrospective learning quality.
    The paper equates selecting the right cases and writing precise rationales with learning. No test of knowledge retention, transfer, or diagnostic skill is administered. The report scores are a proxy for learning, not a direct measure (Section 5.2, Measurement).
  • domain assumption Pretrained models (Grounding DINO, SAM, ResNet50, BERT-base-uncased) transfer to cervical spine MRI with only fine-tuning on ~600 pairs.
    The alignment pipeline assumes that a general open-set detector and a general vision encoder work on medical images with limited fine-tuning. The reported mAP of 0.53 indicates substantial residual error, and the downstream effect of this error on the learning outcome is not assessed.
  • domain assumption The annotation of MRI regions by senior physicians is accurate and consistent.
    Image-text alignment training depends on manual bounding-box annotations of cervical spine regions. No annotation consistency statistics are reported, and the limitation section notes that manual annotation may limit data diversity and introduce bias.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Medillustrator: Improving Retrospective Learning in Physicians' Continuous Medical Education via Multimodal Diagnostic Data Alignment and Representation." pith.science (2026). https://pith.science/paper/WKLBWQMR

@misc{pith2026241115593,
  author       = {Pith},
  title        = {Pith review of: Medillustrator: Improving Retrospective Learning in Physicians' Continuous Medical Education via Multimodal Diagnostic Data Alignment and Representation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WKLBWQMR}},
  note         = {Machine review of arXiv:2411.15593}
}
read the original abstract

Continuous Medical Education (CME) plays a vital role in physicians' ongoing professional development. Beyond immediate diagnoses, physicians utilize multimodal diagnostic data for retrospective learning, engaging in self-directed analysis and collaborative discussions with peers. However, learning from such data effectively poses challenges for novice physicians, including screening and identifying valuable research cases, achieving fine-grained alignment and representation of multimodal data at the semantic level, and conducting comprehensive contextual analysis aided by reference data. To tackle these challenges, we introduce Medillustrator, a visual analytics system crafted to facilitate novice physicians' retrospective learning. Our structured approach enables novice physicians to explore and review research cases at an overview level and analyze specific cases with consistent alignment of multimodal and reference data. Furthermore, physicians can record and review analyzed results to facilitate further retrospection. The efficacy of Medillustrator in enhancing physicians' retrospective learning processes is demonstrated through a comprehensive case study and a controlled in-lab between-subject user study.

Figures

Figures reproduced from arXiv: 2411.15593 by the authors.

Figure 1
Figure 1. Medillustrator is composed of a data process module, a modeling engine, and a visualization interface designed to align with the discovery-driven learning workflow. After gathering discussions and design requirements from the formative study, we developed Medillustrator, an interactive visu￾alization system tailored to assist physicians in efficient and effec￾tive learning and retrospection using diagnostic data. Th… view at source ↗
Figure 2
Figure 2. A Glyph design in fusion modal and connections across fusion modal and unimodal. A1 Current glyph design. A2 - A3 Alternatives based on the rose chart and box plot. comprehension. For the second alternative ( [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Case Study Part I: Exploring high-value patient cases [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: In Case Study Part II, the physicians’ interaction workflow comprises: [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Participants were randomly assigned to two groups: one using [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: In-task survey results from usability and effectiveness aspects in [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

87 extracted references · 44 canonical work pages

  1. [1]

    Active Learning Strategies to Improve Progression from Knowledge to Action

    2020. Active Learning Strategies to Improve Progression from Knowledge to Action. 46, 1 (Feb 2020), 1–19. https://doi.org/10.1016/j.rdc.2019.09.001

  2. [2]

    Hermann Ebbinghaus (1885). 2013. Memory: A Contribution to Experimental Psychology. Annals of Neurosciences 20 (2013), 155 – 156. https://doi.org/10. 5214%2Fans.0972.7531.200408

  3. [3]

    Yura Ahn, Gil-Sun Hong, Kye Jin Park, Choong Wook Lee, Ju Hee Lee, and Seon-Ok Kim. 2021. Impact of diagnostic errors on adverse outcomes: learning from emergency department revisits with repeat CT or MRI. Insights into Imaging 12, 1 (2021), 160. https://doi.org/10.1186/s13244-021-01108-0

  4. [4]

    Rao, and David C

    Kofi-Buaku Atsina, Laurence Parker, Vijay M. Rao, and David C. Levin

  5. [5]

    Bernard, Florian Jung, Jörn Kohlhammer, Thorsten May, Kathrin Scheckenbach, and Stefan Wesarg

    Andreas Bannach, J. Bernard, Florian Jung, Jörn Kohlhammer, Thorsten May, Kathrin Scheckenbach, and Stefan Wesarg. 2017. Visual analytics for radiomics: Combining medical imaging with patient data for clinical research. 2017 IEEE Workshop on Visual Analytics in Healthcare (VAHC) (2017), 84–91. https: //doi.org/10.1109/V AHC.2017.8387545

  6. [6]

    Mazmanian

    Georges Bordage, Brian Carlin, and Paul E. Mazmanian. 2009. Continuing medical education effect on physician knowledge: effectiveness of continuing medical education: American College of Chest Physicians Evidence-Based Educational Guidelines. Chest 135 3 Suppl (2009), 29S–36S. https://doi.org/10.1378/chest.08- 2515

  7. [7]

    Richard E Boyatzis. 1998. Transforming qualitative information: Thematic analy- sis and code development. sage

  8. [8]

    Boyatzis

    Richard E. Boyatzis. 1998. Transforming Qualitative Information: Thematic Analysis and Code Development. https://api.semanticscholar.org/CorpusID: 60526017

Show all 87 references
  1. [9]

    J. B. Brooke. 1996. SUS: A ’Quick and Dirty’ Usability Scale. https://api. semanticscholar.org/CorpusID:107686571

  2. [10]

    Jérôme Seymour Bruner. 1960. The Process of Education. https://api. semanticscholar.org/CorpusID:177285798

  3. [11]

    Caban and David Gotz

    Jesus J. Caban and David Gotz. 2015. Visual analytics in healthcare - opportunities and research challenges. Journal of the American Medical Informatics Association : JAMIA 22 2 (2015), 260–2. https://doi.org/10.1093/jamia/ocv006

  4. [12]

    Zhihao Chen, Yangqiaoyu Zhou, Anh Tu Tran, Junting Zhao, Liang Wan, Gideon Su Kai Ooi, Lionel T. E. Cheng, Choon Hua Thng, Xinxing Xu, Yong Liu, and H. Fu. 2023. Medical Phrase Grounding with Region-Phrase Context Contrastive Alignment. ArXiv abs/2303.07618 (2023). https://doi...

  5. [14]

    Victoria Clarke and Virginia Braun. 2017. Thematic analysis. The journal of positive psychology 12, 3 (2017), 297–298. https://doi.org/doi/10.1037/13620-004

  6. [15]

    Coburn, Keith T

    Can Cui, Haichun Yang, Yaohong Wang, Shilin Zhao, Zuhayr Asad, Lori A. Coburn, Keith T. Wilson, Bennett A. Landman, and Yuankai Huo. 2022. Deep multimodal fusion of image and non-image data in disease diagnosis and prognosis: a review. Progress in biomedical engineering (Brist...

  7. [16]

    Zhou, and Huamin Qu

    Weiwei Cui, Yingcai Wu, Shixia Liu, Furu Wei, Michelle X. Zhou, and Huamin Qu. 2010. Context preserving dynamic word cloud visualization. 2010 IEEE Pacific Visualization Symposium (PacificVis)(2010), 121–128. https://doi.org/10. 1109/PACIFICVIS.2010.5429600

  8. [17]

    Rebecca Donkin, Heather Yule, and Trina Fyfe. 2023. Online case-based learning in medical education: a scoping review. BMC Medical Education 23, 1 (2023),

  9. [18]

    Norbert Donner-Banzhoff. 2018. Solving the Diagnostic Chal- lenge: A Patient-Centered Approach. The Annals of Family Medicine 16, 4 (2018), 353–358. https://doi.org/10.1370/afm.2264 arXiv:https://www.annfammed.org/content/16/4/353.full.pdf

  10. [19]

    Hongyi Duanmu, Pauline Boning Huang, Srinidhi Brahmavar, Stephanie Lin, Thomas Ren, Jun Kong, Fusheng Wang, and Tim Q. Duong. 2020. Prediction of Pathological Complete Response to Neoadjuvant Chemotherapy in Breast Cancer Using Deep Learning with Integrative Imaging, Molecular...

  11. [20]

    DynaMed. 2024. DynaMed. https://www.dynamed.com/

  12. [21]

    Sandra G Hart and Lowell E Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. In Advances in psychology. V ol. 52. Elsevier, 139–183. https://doi.org/10.1016/S0166-4115(08) 62386-9

  13. [22]

    Partridge, Habib Rahbar, Debosmita Biswas, Christoph I

    Greg Holste, Savannah C. Partridge, Habib Rahbar, Debosmita Biswas, Christoph I. Lee, and Adam M. Alessio. 2021. End-to-End Learning of Fused Image and Non- Image Features for Improved Breast Cancer Classification from MRI. 2021 IEEE/CVF International Conference on Computer Vi...

  14. [23]

    Shih-Cheng Huang, Anuj Pareek, Saeed Seyyedi, Imon Banerjee, and Matthew P. Lungren. 2020. Fusion of medical imaging and electronic health records using deep learning: a systematic review and implementation guidelines. NPJ Digital Medicine 3 (2020). https://doi.org/10.1038/s41...

  15. [24]

    P. Jaccard. 1912. THE DISTRIBUTION OF THE FLORA IN THE ALPINE ZONE.1. New Phytologist 11 (1912), 37–50. https://doi.org/10.1111/j.1469- 8137.1912.tb05611.x

  16. [25]

    Berg, Wan-Yen Lo, Piotr Dollár, and Ross B

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross B. Girshick. 2023. Segment Anything. 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (...

  17. [26]

    Malcolm S Knowles. 1975. Self-directed learning: A guide for learners and teachers. (1975)

  18. [27]

    Yingshu Li, Yunyi Liu, Zhanyu Wang, Xinyu Liang, Lingqiao Liu, Lei Wang, Leyang Cui, Zhaopeng Tu, Longyue Wang, and Luping Zhou. 2023. A Compre- hensive Study of GPT-4V’s Multimodal Capabilities in Medical Imaging. ArXiv abs/2310.20381 (2023). https://doi.org/10.1101/2023.11.0...

  19. [28]

    Paul Pu Liang, Yiwei Lyu, Gunjan Chhablani, Nihal Jain, Zihao Deng, Xingbo Wang, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2022. MultiViz: Towards Visualizing and Understanding Multimodal Models. In International Conference on Learning Representations. https://doi.org/...

  20. [29]

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chun yue Li, Jianwei Yang, Hang Su, Jun-Juan Zhu, and Lei Zhang. 2023. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. ArXiv abs/2303.05499 (2023). https://doi.org/10....

  21. [30]

    Tzu-Hung Liu and Amy M Sullivan. 2021. A story half told: a qualitative study of medical students’ self-directed learning in the clinical setting. BMC Medical Education 21, 1 (2021), 1–11. https://doi.org/10.1186/s12909-021-02913-3

  22. [31]

    Lu, Tiffany Y

    Ming Y . Lu, Tiffany Y . Chen, Drew F. K. Williamson, Melissa Zhao, Maha Shady, Jana Lipková, and Faisal Mahmood. 2020. AI-based pathology predicts origins for cancers of unknown primary. Nature 594 (2020), 106 – 110. https: //doi.org/10.1038/s41586-021-03512-4

  23. [32]

    Henry B Mann and Donald R Whitney. 1947. On a test of whether one of two random variables is stochastically larger than the other.The annals of mathematical statistics (1947), 50–60. https://doi.org/stable/2236101

  24. [33]

    Wilson, Bimal H

    Spyridon Marinopoulos, Todd Dorman, Neda Ratanawongsa, Lisa M. Wilson, Bimal H. Ashar, Jeffrey L. Magaziner, Redonda G. Miller, Patricia A. Thomas, Gregory Prokopowicz, Rehan Qayyum, and Eric B. Bass. 2007. Effectiveness of continuing medical education. Evidence report/technol...

  25. [34]

    Leland McInnes, John Healy, and James Melville. 2018. Umap: Uniform man- ifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426 (2018)

  26. [35]

    Susan F McLean. 2016. Case-based learning and its application in medical and health-care fields: a review of worldwide literature. Journal of medical education and curricular development 3 (2016), JMECD–S20377. https://doi.org/10.4137/ JMECD.S20377

  27. [36]

    Medscape. 2024. Medscape. https://www.medscape.org/

  28. [37]

    Wagner-Larsen, Erlend Hodneland, Camilla Krakstad, In- gfrid S

    Eric Mörth, Kari S. Wagner-Larsen, Erlend Hodneland, Camilla Krakstad, In- gfrid S. Haldorsen, Stefan Bruckner, and Noeska N. Smit. 2020. RadEx: Integrated Visual Exploration of Multiparametric Studies for Radiomic Tumor Profiling. Computer Graphics Forum 39 (2020). https://do...

  29. [38]

    Nguyen, Maura Borrego, Cynthia J

    Kevin A. Nguyen, Maura Borrego, Cynthia J. Finelli, Matt DeMonbrun, Caroline Crockett, Sneha Tharayil, Prateek Shekhar, Cynthia Waters, and Robyn Rosenberg

  30. [39]

    Fred Paas, Alexander Renkl, and John Sweller. 2003. Cognitive Load Theory and Instructional Design: Recent Developments. Educational Psychologist 38 (2003), 1 – 4. https://doi.org/10.1207/S15326985EP3801_1

  31. [40]

    Walsh, and Samir C

    Chandni Pattni, Michael Scaffidi, Juana Li, Shai Genis, Nikko Gimpaya, Rishad Khan, Rishi Bansal, Nazi Torabi, Catharine M. Walsh, and Samir C. Grover

  32. [41]

    Yu Qi Qiao, Jun Shen, Xiao Liang, Song Ding, Fang Yuan Chen, Li Shao, Qing Zheng, and Zhi Hua Ran. 2014. Using cognitive theory to facilitate medical education. BMC Medical Education 14, 1 (2014), 79. https://doi.org/10.1186/ 1472-6920-14-79

  33. [42]

    Ziyuan Qin, Huahui Yi, Qicheng Lao, and Kang Li. 2022. Medical Image Un- derstanding with Pretrained Vision Language Models: A Comprehensive Study. ArXiv abs/2209.15517 (2022). https://doi.org/10.48550/arXiv.2209.15517

  34. [43]

    Einck, Anna Vilanova, and Eduard Gröller

    Renata Georgia Raidou, Oscar Casares-Magaz, Artem Amirkhanov, Vitali Moi- seenko, Ludvig Paul Muren, John P. Einck, Anna Vilanova, and Eduard Gröller

  35. [44]

    Lior Rokach. 2010. Ensemble-based classifiers. Artificial Intelligence Review 33 (2010), 1–39. https://doi.org/10.1007/s10462-009-9124-7

  36. [45]

    Marianne Rosendal, Dorte Ejg Jarbøl, Anette Fischer Pedersen, and Rikke Sand Andersen. 2013. Multiple perspectives on symptom interpretation in primary care research. BMC Family Practice 14, 1 (2013), 167. https://doi.org/10.1186/1471- 2296-14-167

  37. [46]

    Bryan C Russell, Antonio Torralba, Kevin P Murphy, and William T Freeman

  38. [47]

    Yuval Shahar, Dina Goren-Bar, David Boaz, and Gil Tahan. 2006. Distributed, intelligent, interactive visualization and exploration of time-oriented clinical data and their abstractions. Artificial intelligence in medicine 38 2 (2006), 115–35. https://doi.org/10.1016/j.artmed.2...

  39. [48]

    Shaik Shehanaz, Ebenezer Daniel, Sitaramanjaneya Reddy Guntur, and Sivaji Satrasupalli. 2021. Optimum weighted multimodal medical image fusion using particle swarm optimization. Optik 231 (2021), 166413. https://doi.org/10.1016/j. ijleo.2021.166413

  40. [49]

    T. Shimizu. 2023. Twelve tips for physicians’ mastering expertise in diagnostic excellence. MedEdPublish 13 (2023), 21. https://doi.org/10.12688/mep.19618.1 Version 1; Peer review: 2 approved with reservations, 1 not approved

  41. [50]

    Ben Shneiderman. 1996. The eyes have it: a task by data type taxonomy for infor- mation visualizations. Proceedings 1996 IEEE Symposium on Visual Languages (1996), 336–343. https://doi.org/10.1016/B978-155860915-0/50046-9

  42. [51]

    Shrout and Joseph L

    Patrick E. Shrout and Joseph L. Fleiss. 1979. Intraclass correlations: uses in assessing rater reliability. Psychological bulletin 86 2 (1979), 420–8. https: //doi.org/doi/10.1037/0033-2909.86.2.420

  43. [52]

    Muhlbaier, Elaine Hart-Brothers, Glenda M

    Mina Silberberg, Lawrence H. Muhlbaier, Elaine Hart-Brothers, Glenda M. Small, Arwen E. Bunce, Rupal Patel, Seronda Robinson, and Sherman A. James. 2021. Melding Multiple Sources of Knowledge: Using Theory and Experiential Knowl- edge to Design a Community Health Intervention ...

  44. [53]

    Stephenson, Sara L

    Christopher R. Stephenson, Sara L. Bonnes, Adam P. Sawatsky, Lukas W. Richards, Cathy D. Schleck, Jayawant N. Mandrekar, Thomas J. Beckman, and Christopher M. Wittich. 2020. The relationship between learner engagement and teaching effectiveness: a novel assessment of student e...

  45. [54]

    Nicole Sultanum, Farooq Naeem, Michael Brudno, and Fanny Chevalier. 2022. ChartWalk: Navigating large collections of text notes in electronic health records Medillustrator: Improving Retrospective Learning in Physicians’ Continuous Medical Education via Multimodal Diagnostic D...

  46. [55]

    John Sweller. 1988. Cognitive Load During Problem Solving: Effects on Learning. Cogn. Sci. 12 (1988), 257–285. https://doi.org/10.1016/0364-0213(88)90023-7

  47. [56]

    John Sweller, Jeroen J. G. van Merrienboer, and Fred Paas. 1998. Cognitive Architecture and Instructional Design. Educational Psychology Review 10 (1998), 251–296. https://doi.org/10.1023/A:1022193728205

  48. [57]

    Masami Tagawa. 2008. Physician self-directed learning and education. The Kaohsiung Journal of Medical Sciences 24, 7 (2008), 380–385. https://doi.org/10. 1016/S1607-551X(08)70136-0

  49. [59]

    VisualDx. 2024. VisualDx. https://www.visualdx.com/

  50. [60]

    Guotai Wang, Xiangde Luo, Ran Gu, Shuojue Yang, Yijie Qu, Shuwei Zhai, Qianfei Zhao, Kang Li, and Shaoting Zhang. 2022. PyMIC: A deep learning toolkit for annotation-efficient medical image segmentation. Computer methods and programs in biomedicine 231 (2022), 107398. https://...

  51. [61]

    Syeda-Mahmood

    Hongzhi Wang, Vaishnavi Subramanian, and Tanveer F. Syeda-Mahmood. 2021. Modeling Uncertainty in Multi-Modal Fusion for Lung Cancer Survival Analysis. 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI) (2021), 1169–1172. https://doi.org/10.1109/ISBI48211.2021.9433823

  52. [63]

    Shanshan Wang, Cheng Li, Rongpin Wang, Zaiyi Liu, Meiyun Wang, Hongna Tan, Yaping Wu, Xinfeng Liu, Hui Sun, Rui Yang, Xin Liu, Jie Chen, Hui-Chong Zhou, Ismail Ben Ayed, and Hairong Zheng. 2020. Annotation-efficient deep learning for automatic medical image segmentation. Natur...

  53. [64]

    Xingbo Wang, Jianben He, Zhihua Jin, Muqiao Yang, and Huamin Qu. 2021. M2Lens: Visualizing and Explaining Multimodal Models for Sentiment Analysis. IEEE Transactions on Visualization and Computer Graphics PP (2021), 1–1. https://doi.org/10.1109/TVCG.2021.3114794

  54. [65]

    Yunhai Wang, Xiaowei Chu, Chen Bao, Lifeng Zhu, Oliver Deussen, Baoquan Chen, and Michael Sedlmair. 2018. EdWordle: Consistency-Preserving Word Cloud Editing. IEEE Transactions on Visualization and Computer Graphics24 (2018), 647–656. https://doi.org/10.1109/TVCG.2017.2745859

  55. [66]

    Hilde Worum, Daniela Lillekroken, Birgitte Ahlsen, Kirsti Skavberg Roaldsen, and Astrid Bergland. 2019. Bridging the gap between research-based knowledge and clinical practice: a qualitative examination of patients and physiotherapists’ views on the Otago exercise Programme. B...

  56. [67]

    Fan, and Qu Li

    Yuchen Wu, Yuansong Xu, Shenghan Gao, Xingbo Wang, Wen gang Song, Zhi- heng Nie, X. Fan, and Qu Li. 2023. LiveRetro: Visual Analytics for Strategic Retrospect in Livestream E-Commerce. IEEE Transactions on Visualization and Computer Graphics 30 (2023), 1117–1127. https://doi.o...

  57. [68]

    Ouyang Yang, Yuchen Wu, Hengbao Wang, Chenyang Zhang, Furui Cheng, Chang Jiang, Lixia Jin, Yuanwu Cao, and Qu Li. 2023. Leveraging Historical Medical Records as a Proxy via Multimodal Modeling and Visualization to Enrich Medical Diagnostic Learning. IEEE transactions on visual...

  58. [69]

    Lu Ying, Tan Tang, Yuzhe Luo, Lvkeshen Shen, Xiao Xie, Lingyun Yu, and Yingcai Wu. 2021. GlyphCreator: Towards Example-based Automatic Generation of Circular Glyphs. IEEE Transactions on Visualization and Computer Graphics PP (2021), 1–1. https://doi.org/10.1109/TVCG.2021.3114877

  59. [70]

    Haipeng Zeng, Xingbo Wang, Yong Wang, Aoyu Wu, Ting-Chuen Pong, and Huamin Qu. 2022. GestureLens: Visual Analysis of Gestures in Presentation Videos. IEEE Transactions on Visualization and Computer GraphicsPP (2022), 1–1. https://doi.org/10.1109/TVCG.2022.3169175

  60. [71]

    Xiner Zhu, Yichao Wu, Haoji Hu, Xianwei Zhuang, Jincao Yao, Di Ou, Wei Li, Mei Song, Na Feng, and Dong-Guo Xu. 2022. Medical lesion segmentation by combining multi-modal images with modality weighted UNet. Medical physics (2022). https://doi.org/10.1002/mp.15610 Chinese CHI 20...

  61. [78]

    The system is easy to use

  62. [79]

    I would like to recommend this system to others and use it in the future

  63. [80]

    The system is helpful for screening valuable patient cases for learning

  64. [81]

    The system is helpful for performing patient case analysis

  65. [82]

    I am satisfied when using the system

  66. [83]

    Effectiveness

    I feel distracted when using the system. Effectiveness

  67. [84]

    I can search and filter my concerned patient cases conveniently

  68. [85]

    I can compare and identify consistencies in cases across modalities efficiently

  69. [86]

    I can view multimodal diagnostic data and contextual information of case accessibly

  70. [87]

    I can record my analysis insights for further reviews and revisions conveniently

  71. [88]

    I am satisfied with the quality of the selected cases and related analyses in my report

  72. [89]

    I am confident with my analysis

  73. [90]

    Table 3: In-task survey for participants in 7-point Likert scale(1: Strongly Disagree, 7: Strongly Agree)

    My entire analysis process is efficient. Table 3: In-task survey for participants in 7-point Likert scale(1: Strongly Disagree, 7: Strongly Agree). Medillustrator: Improving Retrospective Learning in Physicians’ Continuous Medical Education via Multimodal Diagnostic Data Align...

  74. [564]

    https://doi.org/10.1186/s12909-023-04520-w

  75. [2008]

    https://doi.org/10.1007/s11263- 007-0090-8

    LabelMe: a database and web-based tool for image annotation.International journal of computer vision 77 (2008), 157–173. https://doi.org/10.1007/s11263- 007-0090-8

  76. [2018]

    Computer Graphics Forum 37 (2018)

    Bladder Runner: Visual Analytics for the Exploration of RT-Induced Bladder Toxicity in a Cohort Study. Computer Graphics Forum 37 (2018). https://doi.org/10.1111/cgf.13413

  77. [2020]

    American Journal of Roentgenol- ogy 214, 1 (2020), W55–W61

    Advanced Imaging Interpretation by Radiologists and Nonradiol- ogist Physicians: A Training Issue. American Journal of Roentgenol- ogy 214, 1 (2020), W55–W61. https://doi.org/10.2214/AJR.19.21802 arXiv:https://doi.org/10.2214/AJR.19.21802 PMID: 31691611

  78. [2021]

    International Journal of STEM Education 8, 1 (2021), 9

    Instructor strategies to aid implementation of active learning: a systematic literature review. International Journal of STEM Education 8, 1 (2021), 9. https: //doi.org/10.1186/s40594-021-00270-7

  79. [2023]

    PLOS ONE 18, 7 (07 2023), 1–15

    Video-based interventions to improve self-assessment accuracy among physicians: A systematic review. PLOS ONE 18, 7 (07 2023), 1–15. https: //doi.org/10.1371/journal.pone.0288474

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.