Pith. sign in

Paper Citation Record · LEDGER

Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2112.13906.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.13906 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:59:24.095085Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:44.358634Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9a3440f3-bbb9-4565-b977-3fd864f3f479 · inbound

BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs cites this paper.

BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:42:22.455703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T10:42:22.378367Z digest=sha256:7425a097c2b7d36c99ce8e3041cd07352e10775a38a86885a3d7b5c623daf7aa

Observation 1253e8ae-703d-481f-a8ee-ad8737c6c709 · inbound

Data-Centric Foundation Models in Computational Healthcare: A Survey cites this paper.

Data-Centric Foundation Models in Computational Healthcare: A Survey Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:13:53.045700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-24T04:13:05.328492Z digest=sha256:814173ea8ac9eaab2dfe1e215dadf270bb1917a3f84bb25dd2456ed59916ba43

Observation ea0e080f-3aff-4fc5-8ea6-e45229723100 · inbound

MCP-MedSAM: A Powerful Lightweight Medical Segment Anything Model Trained with a Single GPU in Just One Day cites this paper.

MCP-MedSAM: A Powerful Lightweight Medical Segment Anything Model Trained with a Single GPU in Just One Day Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:16:48.565097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:16:48.565097Z digest=sha256:af7cede1e0d107b72e8f636bfabf09389174874b90c2ca842c69473e528ace37

Observation 04bb37c1-f7d1-4c44-9418-1504f70e9d46 · inbound

DiffCLIP: Few-shot Language-driven Multimodal Classifier cites this paper.

DiffCLIP: Few-shot Language-driven Multimodal Classifier Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:11:02.330619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:11:02.330619Z digest=sha256:09843427b7c0d162a35b7a9eca03f18b0781a8684572cfdf035a73cda1dbc35b

Observation 033cae36-d1eb-47eb-9047-9c3962c0daac · inbound

MedCoT: Medical Chain of Thought via Hierarchical Expert cites this paper.

MedCoT: Medical Chain of Thought via Hierarchical Expert Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:34.845840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:53:34.845840Z digest=sha256:4f29ad19537a95a8e783d2a1117c7305d9ec8ecc9d456dbe1fe94932157023b6

Observation aa81f00a-7300-4779-83d9-6321e4a77720 · inbound

KPL: Training-Free Medical Knowledge Mining of Vision-Language Models cites this paper.

KPL: Training-Free Medical Knowledge Mining of Vision-Language Models Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:35:58.075715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:35:58.075715Z digest=sha256:3138fffef42b761fa0be24750ea591a2161ec194c6db76cc6552364fb0712fda

Observation 969c01af-5edb-4cb7-b6e5-d359eb2c90ad · inbound

Chest X-ray Foundation Model with Global and Local Representations Integration cites this paper.

Chest X-ray Foundation Model with Global and Local Representations Integration Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T20:10:52.333997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:10:52.333997Z digest=sha256:1310e38d8a7454cc91f85271f9318327973b5fbb1ac2d8e8942b84152b7ef9e2

Observation 92c40f9c-aa93-431e-8779-1596923ebaad · inbound

Exploring Language Patterns of Prompts in Text-to-Image Generation and Their Impact on Visual Diversity cites this paper.

Exploring Language Patterns of Prompts in Text-to-Image Generation and Their Impact on Visual Diversity Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:59:24.095085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:59:24.095085Z digest=sha256:365da9e8ee2592cb2213ed82c5bb69d2763a4443c755fcd4e15778e4a46267be

Observation a18ae110-5f86-45ee-b14d-4347e8e05511 · inbound

Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis cites this paper.

Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:37:15.143976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T10:34:28.875334Z digest=sha256:48ce3b95d92c15e3a1af1cc154a18a0b7936dd2060e95880c791b72cceda8c81

Observation da3006bc-2a3e-4827-a30f-e1cb6c25853c · inbound

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning cites this paper.

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:58:06.852295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:58:06.852295Z digest=sha256:fc442961ad56498191b74f3b20068cf7058c368d1a9da8f5cc7e944de8f3eb13

Observation ccb3c126-1df2-4150-a4e6-468ddb3e62d0 · inbound

VSF-Med:A Vulnerability Scoring Framework for Medical Vision-Language Models cites this paper.

VSF-Med:A Vulnerability Scoring Framework for Medical Vision-Language Models Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.632061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.632061Z digest=sha256:8ff1a3d65b556427132dbde13f9171c17535303cec5508ea6fb4bdd6906ef3b3

Observation fca317d1-0493-4faa-b198-6f911d0aaf6e · inbound

Prompt Mechanisms in Medical Imaging: A Comprehensive Survey cites this paper.

Prompt Mechanisms in Medical Imaging: A Comprehensive Survey Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:25.038638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:02:25.038638Z digest=sha256:e82a6dac9793670187809ed63c07501322374befc844d822dd974dc783443229

Observation df982fbe-0fad-4113-b25f-48f58704bd8e · inbound

LangMamba: A Language-driven Mamba Framework for Low-dose CT Denoising with Vision-language Models cites this paper.

LangMamba: A Language-driven Mamba Framework for Low-dose CT Denoising with Vision-language Models Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:13.302139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:13.302139Z digest=sha256:0c73af84a8b8efc45102b3fc459317f1eaf02d0ec757bb77ecd0ac571639d8df

Observation b4f7a10a-0182-41a5-a26b-c32e6ebd09de · inbound

BiomedAP: A Vision-Informed Dual-Anchor Framework with Gated Cross-Modal Fusion for Robust Medical Vision-Language Adaptation cites this paper.

BiomedAP: A Vision-Informed Dual-Anchor Framework with Gated Cross-Modal Fusion for Robust Medical Vision-Language Adaptation Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:23:41.047579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T19:18:54.747828Z digest=sha256:6fbfcd0b0c8d44ffbf264cf54b8ecd4044eceb3c234b9c7a33b306e62f81ab25

Observation d48379c2-451b-4338-902e-0921ee07d167 · inbound

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain cites this paper.

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:46:46.326764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T06:40:00.643198Z digest=sha256:ee9fa7fb6dcbe893c4493a8d5a5657230f2efffe53d6251916ae574171cc7fde

Observation 8654be57-58df-4174-af77-925585a79198 · inbound

Geometry-Aware Distillation for Prompt Tuning Biomedical Vision-Language Models cites this paper.

Geometry-Aware Distillation for Prompt Tuning Biomedical Vision-Language Models Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.146118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T06:34:42.631503Z digest=sha256:977189206d3bb2163d818b6490fa6de28a91cc5613afad551e49c28902ae42e9

Observation f311d815-f819-45e7-a1cc-a7113c79f018 · inbound

Brain-Adapter: A Dual-Stream Vision-Language MIL Framework for Comprehensive 3D CT Diagnosis of Acute Intracranial Pathologies cites this paper.

Brain-Adapter: A Dual-Stream Vision-Language MIL Framework for Comprehensive 3D CT Diagnosis of Acute Intracranial Pathologies Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.360421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T08:53:14.124335Z digest=sha256:3f01d3f07327a8f4e993c6e3980599d4dbf3d20313b119b2e341bb3cd1cd32a4