Pith. sign in

Paper Citation Record · LEDGER

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination

As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2506.01902.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01902 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:36:12.724334Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 633b41cb-0744-44e5-8de2-3f7b3138447b · outbound

This paper cites Deep learning in medical image analysis: A third eye for doctors,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Deep learning in medical image analysis: A third eye for doctors,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:15.334429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:10.829997Z digest=sha256:d98c1ff841c78fd6eb4cbeaa8b0474a8d33fbf0bbca338b19c58c8aa20e7cfae

Observation 9d55cf2f-cb9d-4278-9994-242b2ea03d3a · outbound

This paper cites Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.926927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.926927Z digest=sha256:5305802a2e752319d90a01436f667a24c824a068822a37e82d8ccff8e897f9eb

Observation 23eae31c-6844-4a6e-b694-2791d5a5904e · outbound

This paper cites Learning transferable visual models from natural language super- vision,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Learning transferable visual models from natural language super- vision,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:15.150660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:11.037765Z digest=sha256:46ae68fb364a85f754f755962769b13ec2e5678bfbc4195f946e5b38148dcc3e

Observation 3f1463d5-58d9-4fe9-843b-83ceea490472 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:14.990715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:11.147365Z digest=sha256:26f33063e4e9f0f2434d8a6633865aa3fe1c072af8ecc922abf8a728c3086e3f

Observation 4adabc46-6f02-414f-81ed-c3081dcb8229 · outbound

This paper cites Making the Most of Text Semantics to Improve Biomedical Vision–Language Processing,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Making the Most of Text Semantics to Improve Biomedical Vision–Language Processing,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:11.257865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:11.257865Z digest=sha256:990123955b25a22b2d2ad1750f27042f3e362be97aa0472d1d625fcc2662c4d7

Observation ebf3608c-470c-4af6-b008-8ffeaf72f98e · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:14.805910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:11.369209Z digest=sha256:399cfe2e03f24e0a59a4828c1d0ed0615997f8ebf56ccab442797e44070cef9b

Observation 3a2fe4b2-438d-4525-b328-e9e470a3d137 · outbound

This paper cites Lessons from Natural Language Inference in the Clinical Domain.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Lessons from Natural Language Inference in the Clinical Domain

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:11.452383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:11.452383Z digest=sha256:832ee1c8595a2d4fa6356d248058367c05b93962d95bc72d3fe6068a0493987f

Observation 3228846b-5eb7-4284-a517-7ce2fe87e845 · outbound

This paper cites Improving factual completeness and con- sistency of image-to-text radiology report generation,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Improving factual completeness and con- sistency of image-to-text radiology report generation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:14.663254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:11.552685Z digest=sha256:b19f31e7c57ce6e201e5717961d68105e40aa205bb6d527e63bad09bb891f9f3

Observation b90f7903-7ad9-4b22-a44b-9bc28b5b2175 · outbound

This paper cites Chexpert: A large chest radiograph dataset with uncertainty labels and ex- pert comparison,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Chexpert: A large chest radiograph dataset with uncertainty labels and ex- pert comparison,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:14.481434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:11.786585Z digest=sha256:d69b03a180a2ae032bd888399464157bab2372a58553439e443d255f494147da

Observation e64449a7-797f-4298-991b-da2c3db4d5a4 · outbound

This paper cites Contrastive learning of medical visual representations from paired images and text,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Contrastive learning of medical visual representations from paired images and text,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:14.317898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:11.885502Z digest=sha256:9b4b5117d9582cf3f13f23c6f93b3c8fdb4ca6e3e667eb9e26ce46fa8ebde47d

Observation 4285469d-accf-4316-a9f5-f740ad6f4c00 · outbound

This paper cites Joint learning of localized representations from medical im- ages and reports,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Joint learning of localized representations from medical im- ages and reports,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:14.141935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:11.963517Z digest=sha256:07a96e8eba211f571c720aff4ff1fa972be330d2c94e96a3294b48200d62f22e

Observation bb87e176-0ef2-41f6-97f3-7a66dfbe47fc · outbound

This paper cites MedCLIP: Contrastive Learning from Unpaired Medical Images and Text.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination MedCLIP: Contrastive Learning from Unpaired Medical Images and Text

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:12.053292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:12.053292Z digest=sha256:7e258e716e30480b140b034d5b8d8814d73ec40f065a7d41093d2965d10a086c

Observation e5159f0c-b5fa-4e4a-877c-c9917cdf326b · outbound

This paper cites Expert-level detection of patholo- gies from unannotated chest x-ray images via self- supervised learning,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Expert-level detection of patholo- gies from unannotated chest x-ray images via self- supervised learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:13.956412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:12.126466Z digest=sha256:72806e842dece5f888417274efe8570b10278c60e14ee3c1c55cc870c36e7e88

Observation 90938677-ade9-4073-b59f-d009f93e7bc3 · outbound

This paper cites Pubmed data download,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Pubmed data download,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:13.777040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:12.171617Z digest=sha256:ad80d884d34e0f41c53371bffd2a227f8592b82fb26b9ee29002fefe24abbaa0

Observation 78c2f522-1df3-440c-b82e-9f582de5a44e · outbound

This paper cites MIMIC-III, a freely accessible critical care database,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination MIMIC-III, a freely accessible critical care database,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:12.263639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:12.263639Z digest=sha256:0a71e2c12d97efc595fc99bef0e228f09ac9914e455392607b0a74b00f6ef88d

Observation 5dfa1b2d-c567-4320-a70f-0c99040516cd · outbound

This paper cites Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:13.624497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:12.369487Z digest=sha256:d94235a84c979c8d82de9a19a69b4763c231e23c351f9f62e815fd53681166a2

Observation 5a6286b7-0585-4073-b95b-a4973f05ef7c · outbound

This paper cites Spacy 2: Natural lan- guage understanding with bloom embeddings, convolu- tional neural networks and incremental parsing. neural machine translation,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Spacy 2: Natural lan- guage understanding with bloom embeddings, convolu- tional neural networks and incremental parsing. neural machine translation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:13.436265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:12.437052Z digest=sha256:2cff6f02714da15f758d29e45799714a18985f7caa077a7eaa79a8609dca29ad

Observation a2c5e2d7-7b63-4fdb-8904-ec6b09a0ca2f · outbound

This paper cites What context features can transformer language models use?.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination What context features can transformer language models use?

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:13.180575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:12.561961Z digest=sha256:cdc1f82ddb34e98943391c68937b177a94d743316ca27dcdb9bee3c76b0931fb

Observation 8ed3cef4-4c0f-4adc-81be-85dacf0c6b0f · outbound

This paper cites Design and development of a multimodal biomedical information retrieval system,.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Design and development of a multimodal biomedical information retrieval system,

Reference 19

Resolution
verified exact
doi, observed 2026-08-07T11:36:12.952389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:36:12.656437Z digest=sha256:2b694e8c8cd7fccc981de48d66cf2347853db75b54982047a98737b010ed263a

Observation c26caedf-6e7e-453b-829f-8bb8686e4f5e · outbound

This paper cites An overview of gradient descent optimization algorithms.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination An overview of gradient descent optimization algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:12.724334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:12.724334Z digest=sha256:72940931eaab315ab9f046ceebaa4b504d7314bc72c6ee40f7892634fe149237

Observation fec5e85a-9246-4461-aedc-c8233b3bac18 · outbound

This paper cites Improving Factual Completeness and Consistency of Image-to-Text Radiology Report Generation.

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination Improving Factual Completeness and Consistency of Image-to-Text Radiology Report Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:11.685118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:11.685118Z digest=sha256:e1c36fc69c4405982a5f22a5fc7d71cc77a196d7d30bedcd0cbdeedadbe0f169

Pith citing papers

No inbound Pith citation observations are available.