Pith. sign in

Paper Citation Record · LEDGER

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine

As of 16 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2508.12108.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12108 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:30:01.202763Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved36
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f4146d61-2732-4516-b624-c2b7aeb8fb25 · outbound

This paper cites Gloria: A multimodal global- local representation learning framework for label-efficient medical image recognition.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Gloria: A multimodal global- local representation learning framework for label-efficient medical image recognition

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:02.073438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:00.875731Z digest=sha256:b68ffd0b6a546563f6be549b680729f8b17b5ca6ab52897298eee03a4a8db460

Observation 8e0f9474-ea4f-4127-9d48-13630f15a833 · outbound

This paper cites Medclip: Contrastive learning from unpaired medical images and text.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Medclip: Contrastive learning from unpaired medical images and text

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:02.059231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:00.881084Z digest=sha256:c13925a3328258c6f5e502285ec35d6d89a664968eaa307e36ee7df6369040eb

Observation 1ca8d35d-1b76-42a3-a122-00b868588b1d · outbound

This paper cites MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.885386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.885386Z digest=sha256:ca5f9348a62cc6d32c938ab9dbb5df02ff8d3e346d127929b414c6557a712f40

Observation aeb99c0f-ea84-4b85-8a41-045a9e128482 · outbound

This paper cites Towards unifying medical vision-and-language pre-training via soft prompts.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Towards unifying medical vision-and-language pre-training via soft prompts

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:02.044294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:00.890714Z digest=sha256:fc0d5bacb35679b7812e5ff03cad083f7c0aa890eed4b6f54ca64fd0710d389e

Observation 0da28f04-e7b2-43ac-8fe4-e3b7a886af18 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Learning transferable visual models from natural language supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.896331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.896331Z digest=sha256:e63e1fd5338075f83850c5c9e297153389f93d129db5b7b9d6214e530657755b

Observation 31fd3ddf-0019-47de-87cc-ccc666725f89 · outbound

This paper cites Conditional prompt learning for vision- language models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Conditional prompt learning for vision- language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.901789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.901789Z digest=sha256:9747d2a65238c369dd390e0dcabb148e4a753237f3c8c28c4691f14f14431de7

Observation a82cf4d4-0d5e-44a8-a9cc-09e4c9d47e98 · outbound

This paper cites Learning to prompt for vision-language models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Learning to prompt for vision-language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.907408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.907408Z digest=sha256:79f26b0198a16b695e330ad15c7d7bc702d81b0c4c3e95518ad0518ac8b7be5c

Observation 3d24152b-dd7f-4b0f-b4a2-dc688ba3e347 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Groupvit: Semantic segmentation emerges from text supervision

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:02.002510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:00.911719Z digest=sha256:78530eb034cac71cf5f15065c78ba456f508c245fd719d875843cfcad4b78399

Observation 31fd943f-85c2-41fd-80ff-adf890337742 · outbound

This paper cites Learning to exploit temporal structure for biomedical vision-language processing.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Learning to exploit temporal structure for biomedical vision-language processing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.986018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:00.916634Z digest=sha256:9979ddfb00c7467bd7c52596efcb4c4e3b797b70d8667f8b6a395d23f8a8780c

Observation 1b345788-b913-4dd1-9761-227c17a82168 · outbound

This paper cites Cplip: zero-shot learning for histopathology with comprehensive vision-language alignment.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Cplip: zero-shot learning for histopathology with comprehensive vision-language alignment

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.969720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:00.921164Z digest=sha256:5dd5cc2a47aaff5b9be123183fdd3d533a8ecfac761c00f37c92e74204316bfe

Observation 86beeb90-f59c-497a-bb3a-c9b04b201fe4 · outbound

This paper cites Lu, Bowen Chen, Andrew Zhang, Drew F.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Lu, Bowen Chen, Andrew Zhang, Drew F

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.954878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:00.926099Z digest=sha256:a1d62976c207e436e1ab85f6f87ada3d4330239a07de88d10a0da0fe037b7b8a

Observation 68253beb-c63f-4592-be58-8e38d3c53ec3 · outbound

This paper cites Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.940037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:00.930789Z digest=sha256:9e35443ec72d55bb17b148fa188a37854cb7f2c1c21eb03ec680020cc0e13437

Observation 75d4139c-ac4d-494b-8a81-927e794d3f02 · outbound

This paper cites Quilt-1M: One Million Image-Text Pairs for Histopathology.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Quilt-1M: One Million Image-Text Pairs for Histopathology

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.935549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.935549Z digest=sha256:f27f95c17d8cddfd3fb0642b120f19a2b886b162fb472b5a40e6cc54650f35b9

Observation 0e1c04ed-6500-4676-b46e-c47e48b92be3 · outbound

This paper cites Towards generalist foundation model for radiology.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Towards generalist foundation model for radiology

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.925831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:00.942063Z digest=sha256:3482ff95aa53288174893cc2a3f574dc5edeaaba0c20600c585b293139eed155

Observation 1c8df07b-4282-4af2-aa22-24f49dff84f6 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.948324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.948324Z digest=sha256:37341cdacf6eddb64e8ea8f6d2b1a2a379242188ad913701a9303a39272ecafe

Observation 49a9153b-bea6-4ccc-a6a9-ce080367e416 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Align before fuse: Vision and language representation learning with momentum distillation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.954157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.954157Z digest=sha256:69d4933f314ee26b0b603366dc0e6bada59634f57fabaceff5f95de01139f94d

Observation 005f1bc3-deb3-41eb-92ac-a0cec3d855d0 · outbound

This paper cites Bert: Pre-training of deep bidi- rectional transformers for language understanding.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Bert: Pre-training of deep bidi- rectional transformers for language understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.959794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.959794Z digest=sha256:5ce7276f71df758b49ba1cbf89439d032f7a9fec0066a7060e21b18292ce8558

Observation 496cac27-372c-429e-8797-9669ab826acc · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Momentum contrast for unsupervised visual representation learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.964572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.964572Z digest=sha256:e36923ab6e331472e04f36645a55ae274cf1ad4b634c1092b38c8fc4f3cc4c67

Observation 68afe0c2-a45f-4924-881b-173de1dd5779 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine A simple framework for contrastive learning of visual representations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.969520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.969520Z digest=sha256:a6bba19c6a352a6eecd79dee0fa64543f8b1ce50d422db8254ac22ee6d54a83e

Observation 37c0d002-2bd2-46df-85b9-868fba55a5fc · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Bootstrap your own latent-a new approach to self-supervised learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.974254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.974254Z digest=sha256:76d1b879aa530afaec0d369a38a8737759b3ae1fc7cd08bc15cb7600614174a3

Observation 5fca479a-bce1-4cd0-9d94-f14a8b21dab4 · outbound

This paper cites Masked autoencoders are scalable vision learners.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Masked autoencoders are scalable vision learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.979116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.979116Z digest=sha256:28d568ea11b0f01b3dd29e655317e2bb3fa7d03ce226442f77c258e5dbaef25e

Observation 4a61a804-ebb8-4aeb-ab57-d0578605cec3 · outbound

This paper cites Generative pretraining from pixels.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Generative pretraining from pixels

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.984629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.984629Z digest=sha256:0b05f770f59153ac1fd848ddbd5f005767f808f8cf55dfb86ef785d88766cd82

Observation 06722dc4-a9ea-472f-beba-ce93f12056e7 · outbound

This paper cites Unsupervised learning of visual representations by solving jigsaw puzzles.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Unsupervised learning of visual representations by solving jigsaw puzzles

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.989479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.989479Z digest=sha256:e7130c96340853570ec3607c50fd161af44a471a2a2be21afefdf8feb5b7959e

Observation 63e30dd3-80c0-4541-b335-0e6e2a759c11 · outbound

This paper cites Colorization as a proxy task for visual understanding.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Colorization as a proxy task for visual understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.842145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:00.994362Z digest=sha256:b4d0bebf20eaceac6e9f233ec9018b5bf9fb8dad9dfceed15d9239dc16cee678

Observation 1b4a9d23-d3ff-4d1a-ab2a-89b4ac9b4460 · outbound

This paper cites Self-supervised representation learning by rotation feature decoupling.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Self-supervised representation learning by rotation feature decoupling

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.826670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:00.998873Z digest=sha256:529a3bf679c255b069bd2d31304cb7bd19efa91fae110908030ab488f470ce21

Observation 0090d0f4-6a32-498e-bf27-7f1752ca3a0b · outbound

This paper cites What makes instance discrimination good for transfer learning?.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine What makes instance discrimination good for transfer learning?

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:30:01.367879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.003745Z digest=sha256:7b1659065e2baa216cc919a71bc825ab55426b62f28efb69652a50e2e0eca3b9

Observation d4d30880-b7a7-407a-a716-b7c9dd9545f4 · outbound

This paper cites Scale-mae: A scale-aware masked autoencoder for multiscale geospatial representation learning.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Scale-mae: A scale-aware masked autoencoder for multiscale geospatial representation learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.812120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.009754Z digest=sha256:5caa55a0707e41838fd3b0e120e42239f3fbd2a69890bb4b6566deef743c3c11

Observation d8a30164-549d-4257-b6ae-c5d0719a8c23 · outbound

This paper cites Attention is all you need.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Attention is all you need

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.014686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.014686Z digest=sha256:9a3b2fe811a192a67be4648cae8775fa48e74467d58d09dc974c806a1a7fb787

Observation bff68324-f635-4ec9-9917-50600d642e81 · outbound

This paper cites Language models are few-shot learners.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Language models are few-shot learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.019806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.019806Z digest=sha256:698356ea4826561185f0112169069954c3878bd6ba7acec54ef1397a69e3c29d

Observation 742dac01-52cc-4341-9119-f9f86ba32dc5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine LLaMA: Open and Efficient Foundation Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.024294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.024294Z digest=sha256:9da016b1426a6d6b96c4ee0ae5e8354d5a1a6101cf84d90c1915238fc8b0f8f1

Observation 5967b54d-c824-42c1-b5bc-2245925714b9 · outbound

This paper cites Med3D: Transfer Learning for 3D Medical Image Analysis.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Med3D: Transfer Learning for 3D Medical Image Analysis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.029512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.029512Z digest=sha256:b2753d411d4c2c56278ea24e5f2e33bcf920b6bf801830a1687078fe80603f1b

Observation 91e4ed54-2f39-4b4a-9c5e-f58f477b6416 · outbound

This paper cites Models genesis.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Models genesis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.780656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.034025Z digest=sha256:afd1de9b11bfe5f557f01624d2599b46497a6c884bf2bd2f5ff14266c2531846

Observation 724bb59d-0b2c-45c6-99ab-a656e72b9bd5 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.038158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.038158Z digest=sha256:20e3ce0d913206d41da1c9aaf0d2d949880bd22d83b00e50583c16ee12bdeb3d

Observation bac93bcb-f821-48f2-8ff5-f2c549c8a74f · outbound

This paper cites LXMERT: Learning Cross-Modality Encoder Representations from Transformers.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.043454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.043454Z digest=sha256:645d791b84cff86d79b0a176ac00c16e7c35f3fac80438a522c2465417f0e480

Observation d0d91a95-a7b6-46a2-ac37-2c3d84f12e06 · outbound

This paper cites Uniter: Universal image-text representation learning.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Uniter: Universal image-text representation learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.757157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.048116Z digest=sha256:341b879603446889de973adc68bfc60b4dd16a58f7e4f07063a416bb30414986

Observation f6b30fb9-6e03-4c3a-add7-a7fcd66b4145 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.743642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.053213Z digest=sha256:3435cb4e5eae0f17d3f8aa631c42eb5a69da75808805131359bd6c2fca0ecba5

Observation a5ce9f45-656d-4090-a193-445a6a8a5b1b · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Vinvl: Revisiting visual representations in vision-language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.057338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.057338Z digest=sha256:36379062b3adbb703226dc113e3727b45e9ef58aed243b093fefbf0cb69c32a9

Observation edff5fdb-e67d-4bf3-a270-1ef58d6df557 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models, 2022.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Coca: Contrastive captioners are image-text foundation models, 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.062249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.062249Z digest=sha256:7cf51f8a2aa1ccd6abf300506cfdf2e377eb9b61463304bda3dd429d5b9ebb75

Observation 55eeb42c-144f-4743-adf4-d5521a0efe40 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.066457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.066457Z digest=sha256:1af139cad516405bef6cc2ce60f719f3c0e2b8703a09e619a51931adb479d429

Observation 65a2707a-9740-4ef6-b6b8-a553c277d1ab · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.071454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.071454Z digest=sha256:beb8fc886795ce682fe33e98b7a42591330e2f053a926a3c57f3a95fc3f37f82

Observation 6df0aa4f-a65e-41d4-a663-6f98aff101bf · outbound

This paper cites Multi-modal understanding and generation for medical images and text via vision-language pre-training.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Multi-modal understanding and generation for medical images and text via vision-language pre-training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.694139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.075870Z digest=sha256:eb5b6572e6027cfb6e15cabb18e7c60e54ccd4afb9103807d3150c44165409fa

Observation ff93f537-8c21-40d3-a176-520255c67df5 · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.080927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.080927Z digest=sha256:160a0c865efd5dba1fb68872accd37e19361adb956ae022e68462dc10bfad778

Observation fa7cdee5-57b6-4bea-8382-678291a191c2 · outbound

This paper cites Slip: Self-supervision meets language- image pre-training.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Slip: Self-supervision meets language- image pre-training

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.678960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.085437Z digest=sha256:e915e1d2388e57d81a9e11688bcfd75b769f0c13346f2c320789d5d8016666f8

Observation 30c66838-3b1e-4226-9be9-6e0896c0c7c8 · outbound

This paper cites Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm, 2022.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm, 2022

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.664933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.089740Z digest=sha256:d883fdf311bc8a6c806cb6b4f3d087cb9b1f5c5f45f35bb046f22582581a4f4a

Observation e0504519-2856-4003-a9bc-3a893792e187 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Swin transformer: Hierarchical vision transformer using shifted windows

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.094663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.094663Z digest=sha256:e270d42714f83ecd0c73691e90aa51f3380a21ce234cc680e6dc45230963fcf9

Observation 35a0962c-c2a3-47dc-a28a-8495e21b17fd · outbound

This paper cites Self-supervised pre-training of swin transformers for 3d medical image analysis.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Self-supervised pre-training of swin transformers for 3d medical image analysis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.099089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.099089Z digest=sha256:39675b26c67e167751afb84459a1edfee5e88b4d6d349330653490b36e27f3fa

Observation d36c25de-c9bd-462e-9e61-ddc1f30d4e7f · outbound

This paper cites Masked image modeling advances 3d medical image analysis, 2022.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Masked image modeling advances 3d medical image analysis, 2022

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.633279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.103864Z digest=sha256:be47dbef513485cce88ddab4baf809d29097f62a479272f1d4de4f0ca377fc02

Observation a1337afe-146a-4800-bfa3-7a63abe36566 · outbound

This paper cites V oco: A simple-yet-effective volume contrastive learning framework for 3d medical image analysis.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine V oco: A simple-yet-effective volume contrastive learning framework for 3d medical image analysis

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.619401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.108158Z digest=sha256:ceaf48c14506bfa552a71537185c62537a164a529a557d352e9b3c05ecbd3d8b

Observation 39c3030c-6260-4ec2-bd81-ceaee3ac0f57 · outbound

This paper cites Context Encoders: Feature Learning by Inpainting.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Context Encoders: Feature Learning by Inpainting

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.113224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.113224Z digest=sha256:7a2201c1efb319cdd72a2899b029c0d48f95b525414caa41dacee3ab42b8f9a8

Observation d126cba3-332c-4507-9895-8e13a2929756 · outbound

This paper cites Unsupervised Representation Learning by Predicting Image Rotations.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Unsupervised Representation Learning by Predicting Image Rotations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.118205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.118205Z digest=sha256:1ea7fffbfdfbd0f49ab905a89edfbed79cdff2091bd7ab21bd085be5cc870680

Observation 18bc6f70-c1eb-4cf3-8c96-c2732c5712ac · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Representation Learning with Contrastive Predictive Coding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.122950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.122950Z digest=sha256:a7774594641c5f45f14af063decac7881cce6917a751283c20b8a10076df9b7d

Observation 0f456fc9-a844-4b98-8b47-c3c06dd978f3 · outbound

This paper cites Fast WordPiece Tokenization.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Fast WordPiece Tokenization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.128489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.128489Z digest=sha256:1d2be0c68064f20a37eefc1cd4bc4504eeeb6572dce7640a47d3812182d12f81

Observation e9a74340-58c9-4891-88f7-ce9118f5469c · outbound

This paper cites Zero-shot text-to-image generation.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Zero-shot text-to-image generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.604840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.133289Z digest=sha256:2be98867bde58dac89fb6690ddce2f9e64977144da4198644cfa0f1919fd1649

Observation 35db7e1a-a920-464b-accd-d190d1e28d70 · outbound

This paper cites Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.138329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.138329Z digest=sha256:f4064646f65552b6075fae8f311ff15e68e52c3776fc3159fc45a90c56593c23

Observation 5231ee45-4b53-4ccf-bf1c-777711b32a05 · outbound

This paper cites Abdomenct-1k: Is abdominal organ segmentation a solved problem? IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):6695–6714, 2022.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Abdomenct-1k: Is abdominal organ segmentation a solved problem? IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):6695–6714, 2022

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.581111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.142837Z digest=sha256:24d25347302f65a029124396dd29765da04da3174f9c7a00c0bd85e084a7a028

Observation ee2307a5-8a02-497e-bdf9-dbeedff77eff · outbound

This paper cites Ct-org, a new dataset for multiple organ segmentation in computed tomography.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Ct-org, a new dataset for multiple organ segmentation in computed tomography

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.566777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.148002Z digest=sha256:d0ce32891b80c7f3a3bc55f043ec67d29bf39ec45d081727cfba39fc835fcc54

Observation 517d1b0f-fc32-4797-9814-9daa32543141 · outbound

This paper cites Ledsam, and Olaf Ronneberger.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Ledsam, and Olaf Ronneberger

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.553169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.152348Z digest=sha256:9e1c28c0ad3e195767af8c4fde703da0766df31963cc01f4b28285377aae7af6

Observation 0419b32a-6f0d-4f36-92ed-505400bf3365 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Bleu: a method for automatic evaluation of machine translation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.539265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.157227Z digest=sha256:bf8072267893620d66b9f1f2e5417bbe4b1b019a594c0bd8e7ea1b1fd771155a

Observation 6787a5d8-402d-4bd3-92a5-6ae847698a9b · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Rouge: A package for automatic evaluation of summaries

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.524066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.161539Z digest=sha256:69a4762777a56694e618bcb458a9ae57c4a6f29e89e0155a34098afc9ef8b412

Observation 759c3292-5139-4c5c-b49f-4d11bdce6f35 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine BERTScore: Evaluating Text Generation with BERT

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.166429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.166429Z digest=sha256:d9df9b0ba7af6ad7f916405c476d24288294748e59846a6912e42c060652a8ed

Observation b728e4d2-5e7e-44e3-bf75-bf7ca624c657 · outbound

This paper cites Decoupled weight decay regularization, 2019.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Decoupled weight decay regularization, 2019

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.170711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.170711Z digest=sha256:76592ed54668ecd359a9642e348e64df231425f2821af0d878de2f7783ccb0ef

Observation 77ac4f29-0498-4704-82b8-6e3cb1cf4238 · outbound

This paper cites https://huggingface.co/ContactDoctor/Bio- Medical-Llama-3-8B, 2024.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine https://huggingface.co/ContactDoctor/Bio- Medical-Llama-3-8B, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.500789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.175442Z digest=sha256:772241031a20e6d24705776bda9e0cc9221d0626f6fb9ab361313fbe67dd987e

Observation 0e4abb36-f73f-4b20-98c6-4184023dd95c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Lora: Low-rank adaptation of large language models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.179662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.179662Z digest=sha256:f688521ea8d7d0a9f0261f2135c655bfdafafa4ce8712339ef3ea67655c67bf8

Observation ff94ef9a-2544-4cf9-8091-f84bc85c13bd · outbound

This paper cites Segment anything.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Segment anything

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.184855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.184855Z digest=sha256:9ec56195f3602e6fe4d5af2de8a85db5033a11184591da54064fe365c873f4dc

Observation 807c9353-43ba-4bfa-8ccb-a1f472aea0fc · outbound

This paper cites Segvol: Universal and interactive volumetric medical image segmentation.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Segvol: Universal and interactive volumetric medical image segmentation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.467821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.189214Z digest=sha256:bf4894d8261442991a726ae70cbefcf16bcf47b8f80961ec66b3cc47844532f6

Observation 70085536-0c4e-4680-b283-50c93e0ab6bd · outbound

This paper cites Segment anything in medical images.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Segment anything in medical images

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.453474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.193525Z digest=sha256:cedfeb155cabd45642f0c80169bfcd1d775218549e6c5c9f3c0439adb7c06f10

Observation d7497ca5-54bc-45f8-9024-6fce54bf1fea · outbound

This paper cites Pmc- clip: Contrastive language-image pre-training using biomedical documents.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Pmc- clip: Contrastive language-image pre-training using biomedical documents

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.438351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.198485Z digest=sha256:b78c3bd33575f92da3c1c6c4c3c5c7171c14156679451fb2e4248f9828e2649a

Observation f449df5a-836c-4ed5-8d9d-055db0c5d757 · outbound

This paper cites Accelerate: Training and inference at scale made simple, efficient and adaptable.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Accelerate: Training and inference at scale made simple, efficient and adaptable

Reference 68

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T17:30:01.424454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:30:01.202763Z digest=sha256:c8c8b7f8c16221169872cfbd2e5c477c92cc22173d588789aa5f47f074fb3285

Pith citing papers

No inbound Pith citation observations are available.