Pith. sign in

Paper Citation Record · LEDGER

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination

As of 10 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2506.01902.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01902 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:36:12.724334Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 633b41cb-0744-44e5-8de2-3f7b3138447b · outbound

This paper cites Deep learning in medical image analysis: A third eye for doctors,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Deep learning in medical image analysis: A third eye for doctors,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:15.334429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:10.829997Z digest=sha256:5d97688dedeed26592df5a0a17faebf2147afce8e7fd5504471ec229007c3b77

Observation 9d55cf2f-cb9d-4278-9994-242b2ea03d3a · outbound

This paper cites Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.926927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.926927Z digest=sha256:e52d9156ff28a1437ed3bec9dae549502353e2140651423cfb825807b00ea3b0

Observation 23eae31c-6844-4a6e-b694-2791d5a5904e · outbound

This paper cites Learning transferable visual models from natural language super- vision,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Learning transferable visual models from natural language super- vision,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:15.150660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:11.037765Z digest=sha256:c16aac35566a745d931fc48ca7536d2bc65f99fb3ea1aa21fdb3b76671cc2921

Observation 3f1463d5-58d9-4fe9-843b-83ceea490472 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:14.990715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:11.147365Z digest=sha256:ab8743dafff8fc49cba7068e2860a35ea0385e72691f11e58bf86b37602b6d52

Observation 4adabc46-6f02-414f-81ed-c3081dcb8229 · outbound

This paper cites Making the Most of Text Semantics to Improve Biomedical Vision–Language Processing,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Making the Most of Text Semantics to Improve Biomedical Vision–Language Processing,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:11.257865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:11.257865Z digest=sha256:2821edc97c271ee0d94c4c11715fd926a33c5f37ddc9eb780b03927d1c9ee685

Observation ebf3608c-470c-4af6-b008-8ffeaf72f98e · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:14.805910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:11.369209Z digest=sha256:56e1ec315c8a2364b37b1b3aa90cb4df2fa3c69b96f0f243fb8e91e0d4c75949

Observation 3a2fe4b2-438d-4525-b328-e9e470a3d137 · outbound

This paper cites Lessons from Natural Language Inference in the Clinical Domain.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Lessons from Natural Language Inference in the Clinical Domain

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:11.452383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:11.452383Z digest=sha256:d54932bf2c1aad044f2609e1b264ae8f2dd50684c1e816be072ef015679e5915

Observation 3228846b-5eb7-4284-a517-7ce2fe87e845 · outbound

This paper cites Improving factual completeness and con- sistency of image-to-text radiology report generation,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Improving factual completeness and con- sistency of image-to-text radiology report generation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:14.663254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:11.552685Z digest=sha256:1ef312d94f1ea925c931fb252f6a8fd286cccc95e29a8d0dd2602a412bfbabec

Observation b90f7903-7ad9-4b22-a44b-9bc28b5b2175 · outbound

This paper cites Chexpert: A large chest radiograph dataset with uncertainty labels and ex- pert comparison,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Chexpert: A large chest radiograph dataset with uncertainty labels and ex- pert comparison,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:14.481434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:11.786585Z digest=sha256:1bae5cb4153f21b587892853cd496bdcc02ee77eb7a170717524b5fb55d7cf50

Observation e64449a7-797f-4298-991b-da2c3db4d5a4 · outbound

This paper cites Contrastive learning of medical visual representations from paired images and text,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Contrastive learning of medical visual representations from paired images and text,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:14.317898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:11.885502Z digest=sha256:e6e51cd3c163ba8f947bb34ecfe234ae129e6bd87e865e0615c0e79c6f2e5d51

Observation 4285469d-accf-4316-a9f5-f740ad6f4c00 · outbound

This paper cites Joint learning of localized representations from medical im- ages and reports,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Joint learning of localized representations from medical im- ages and reports,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:14.141935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:11.963517Z digest=sha256:a34756510a753d4abee1e6fddb511508e9989c4ed8fc8695afa04c173432c430

Observation bb87e176-0ef2-41f6-97f3-7a66dfbe47fc · outbound

This paper cites MedCLIP: Contrastive Learning from Unpaired Medical Images and Text.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination MedCLIP: Contrastive Learning from Unpaired Medical Images and Text

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:12.053292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:12.053292Z digest=sha256:4850f64d18159ed649c41a5727ffb625cc704945708401520fb090b8759f2f19

Observation e5159f0c-b5fa-4e4a-877c-c9917cdf326b · outbound

This paper cites Expert-level detection of patholo- gies from unannotated chest x-ray images via self- supervised learning,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Expert-level detection of patholo- gies from unannotated chest x-ray images via self- supervised learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:13.956412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:12.126466Z digest=sha256:6514eaed6ebfd6a178e20a181723c021783516126aa04076ef0b55fbfc0136f3

Observation 90938677-ade9-4073-b59f-d009f93e7bc3 · outbound

This paper cites Pubmed data download,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Pubmed data download,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:13.777040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:12.171617Z digest=sha256:8f51083f3c400dadbed9dd98f6023c3575c99791567e3ddfa17c5c9783e87f17

Observation 78c2f522-1df3-440c-b82e-9f582de5a44e · outbound

This paper cites MIMIC-III, a freely accessible critical care database,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination MIMIC-III, a freely accessible critical care database,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:12.263639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:12.263639Z digest=sha256:c2768b6ea9aae47c5aad81fdbc4105eedbfd65117d9255e6ea75c2645529548c

Observation 5dfa1b2d-c567-4320-a70f-0c99040516cd · outbound

This paper cites Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:13.624497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:12.369487Z digest=sha256:4e0595193d884d57562cf668969448f137b7dab7dccf72318e0ec0d161da8579

Observation 5a6286b7-0585-4073-b95b-a4973f05ef7c · outbound

This paper cites Spacy 2: Natural lan- guage understanding with bloom embeddings, convolu- tional neural networks and incremental parsing. neural machine translation,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Spacy 2: Natural lan- guage understanding with bloom embeddings, convolu- tional neural networks and incremental parsing. neural machine translation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:13.436265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:12.437052Z digest=sha256:8a35fb3fe19544882bc2ad8dc0c588847da79556dfcd9da9a9b34c093502b05c

Observation a2c5e2d7-7b63-4fdb-8904-ec6b09a0ca2f · outbound

This paper cites What context features can transformer language models use?.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination What context features can transformer language models use?

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:13.180575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:12.561961Z digest=sha256:96b212c8142a104f065342e6828c40509650e82f3d5f64c87d0e1be7ddf1b5b8

Observation 8ed3cef4-4c0f-4adc-81be-85dacf0c6b0f · outbound

This paper cites Design and development of a multimodal biomedical information retrieval system,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Design and development of a multimodal biomedical information retrieval system,

Reference 19

Resolution
verified exact
doi, observed 2026-08-07T11:36:12.952389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:36:12.656437Z digest=sha256:11ed8fb5b168f8fd87a5c3f9467c26489dbe35d556264bfaf8d690d6032ed805

Observation c26caedf-6e7e-453b-829f-8bb8686e4f5e · outbound

This paper cites An overview of gradient descent optimization algorithms.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination An overview of gradient descent optimization algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:12.724334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:12.724334Z digest=sha256:5847fb308990fcd1ab3c057b59da69a67b404dbec710d0d3aa3528e5c80f9a34

Observation fec5e85a-9246-4461-aedc-c8233b3bac18 · outbound

This paper cites Improving Factual Completeness and Consistency of Image-to-Text Radiology Report Generation.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Improving Factual Completeness and Consistency of Image-to-Text Radiology Report Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:11.685118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:11.685118Z digest=sha256:43277b4be6997e6827f30b7664f93ff3d93927a97e66205a44a778071ef28166

Pith citing papers

No inbound Pith citation observations are available.