Pith. sign in

Paper Citation Record · LEDGER

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine

As of 21 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2508.12108.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12108 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:30:01.202763Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved36
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f4146d61-2732-4516-b624-c2b7aeb8fb25 · outbound

This paper cites Gloria: A multimodal global- local representation learning framework for label-efficient medical image recognition.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Gloria: A multimodal global- local representation learning framework for label-efficient medical image recognition

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:02.073438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:00.875731Z digest=sha256:6d845b9018ab7dc7d314eb20ff0a86041760ef46306945db735b44fa83f6bc9d

Observation 8e0f9474-ea4f-4127-9d48-13630f15a833 · outbound

This paper cites Medclip: Contrastive learning from unpaired medical images and text.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Medclip: Contrastive learning from unpaired medical images and text

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:02.059231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:00.881084Z digest=sha256:1790fbad6eb6b009f113435cfe9a9e8e96c54af4deb4692c7d51128093fb9d66

Observation 1ca8d35d-1b76-42a3-a122-00b868588b1d · outbound

This paper cites MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.885386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.885386Z digest=sha256:ac15ded93e5d9075ef4837ca1a452b35c3cce51e6d61d979ace9bf7b51db12a2

Observation aeb99c0f-ea84-4b85-8a41-045a9e128482 · outbound

This paper cites Towards unifying medical vision-and-language pre-training via soft prompts.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Towards unifying medical vision-and-language pre-training via soft prompts

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:02.044294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:00.890714Z digest=sha256:3ed92ba6745091068a1cba275e355d7470c26056b6763918f7327a19da07bfdc

Observation 0da28f04-e7b2-43ac-8fe4-e3b7a886af18 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Learning transferable visual models from natural language supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.896331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.896331Z digest=sha256:8b0009d2aeecb56b1eb49595d3431ab7c2ef0c2fa3d6599336dba19c4c090d12

Observation 31fd3ddf-0019-47de-87cc-ccc666725f89 · outbound

This paper cites Conditional prompt learning for vision- language models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Conditional prompt learning for vision- language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.901789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.901789Z digest=sha256:faeafd09cf4d45177ba89b1b7220f5612d28605763b77cd67119d2e8fab17d90

Observation a82cf4d4-0d5e-44a8-a9cc-09e4c9d47e98 · outbound

This paper cites Learning to prompt for vision-language models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Learning to prompt for vision-language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.907408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.907408Z digest=sha256:7324ad04b9bb5b15e9455a3704169a36080069ec9fbb66a4ed07e4d8d742e185

Observation 3d24152b-dd7f-4b0f-b4a2-dc688ba3e347 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Groupvit: Semantic segmentation emerges from text supervision

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:02.002510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:00.911719Z digest=sha256:a7d13cc309c46602464f52768735747bbd435d42a2c3a75044e98a8c110be9e8

Observation 31fd943f-85c2-41fd-80ff-adf890337742 · outbound

This paper cites Learning to exploit temporal structure for biomedical vision-language processing.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Learning to exploit temporal structure for biomedical vision-language processing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.986018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:00.916634Z digest=sha256:58bb435d50b58ad071d8eb6c2caec976cbb42a2145a1d70519238523e885b184

Observation 1b345788-b913-4dd1-9761-227c17a82168 · outbound

This paper cites Cplip: zero-shot learning for histopathology with comprehensive vision-language alignment.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Cplip: zero-shot learning for histopathology with comprehensive vision-language alignment

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.969720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:00.921164Z digest=sha256:34cffdb604dda764ec3fa527f126f7c465724f89610407132f811bf1a70c6754

Observation 86beeb90-f59c-497a-bb3a-c9b04b201fe4 · outbound

This paper cites Lu, Bowen Chen, Andrew Zhang, Drew F.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Lu, Bowen Chen, Andrew Zhang, Drew F

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.954878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:00.926099Z digest=sha256:0dd7987830381b63b80d9ad68323da63169c9168923dd971eb30f1ac4d38da3a

Observation 68253beb-c63f-4592-be58-8e38d3c53ec3 · outbound

This paper cites Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.940037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:00.930789Z digest=sha256:b900431a3ed6733e59ae2ad354ef18335882b83a67681c23226a5c6ebe6d3e9f

Observation 75d4139c-ac4d-494b-8a81-927e794d3f02 · outbound

This paper cites Quilt-1M: One Million Image-Text Pairs for Histopathology.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Quilt-1M: One Million Image-Text Pairs for Histopathology

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.935549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.935549Z digest=sha256:f2138bf372a44fd5aeaefb42fe707da4134249aaf7d77cb4f4fe033bfba082e4

Observation 0e1c04ed-6500-4676-b46e-c47e48b92be3 · outbound

This paper cites Towards generalist foundation model for radiology.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Towards generalist foundation model for radiology

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.925831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:00.942063Z digest=sha256:54f32ba8d390291f5a37c0cbd2de1823d895758e4f7ecb4acc5de1be7b3f597b

Observation 1c8df07b-4282-4af2-aa22-24f49dff84f6 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.948324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.948324Z digest=sha256:4a7d64a8f23ef221a2fe3ef95a41ec1145decc66b9e381a8882a96d252bc6f9f

Observation 49a9153b-bea6-4ccc-a6a9-ce080367e416 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Align before fuse: Vision and language representation learning with momentum distillation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.954157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.954157Z digest=sha256:53ecb4997d8021f0f6befe727aa9f64a57c5f783971bc08be67d7093ffdd0aad

Observation 005f1bc3-deb3-41eb-92ac-a0cec3d855d0 · outbound

This paper cites Bert: Pre-training of deep bidi- rectional transformers for language understanding.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Bert: Pre-training of deep bidi- rectional transformers for language understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.959794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.959794Z digest=sha256:7ad68a401d98651139b32b2b27d54ce36f86ae16bbb55aee72735ba267c6f740

Observation 496cac27-372c-429e-8797-9669ab826acc · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Momentum contrast for unsupervised visual representation learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.964572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.964572Z digest=sha256:a363f4b3e247e8fcd8484ce03afb525e2e942d230ffb8c8385f90fc15d69d03c

Observation 68afe0c2-a45f-4924-881b-173de1dd5779 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine A simple framework for contrastive learning of visual representations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.969520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.969520Z digest=sha256:5163de9565d05384f9d135079a90af21ba02fe25313a4ba51ef83f79f629a8c4

Observation 37c0d002-2bd2-46df-85b9-868fba55a5fc · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Bootstrap your own latent-a new approach to self-supervised learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.974254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.974254Z digest=sha256:67f9d4ca23c8c6d345c900da898ca7344d98c11cd59504531716af58b451e9d3

Observation 5fca479a-bce1-4cd0-9d94-f14a8b21dab4 · outbound

This paper cites Masked autoencoders are scalable vision learners.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Masked autoencoders are scalable vision learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.979116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.979116Z digest=sha256:84e54d1ae61da3dda6003436be3bff69f2022b36ac513fb3985992d20dbbff2d

Observation 4a61a804-ebb8-4aeb-ab57-d0578605cec3 · outbound

This paper cites Generative pretraining from pixels.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Generative pretraining from pixels

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.984629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.984629Z digest=sha256:83f155b4e331cc92356dd43cae57585d7a51263a84b699a2c2855de7a5d54cd0

Observation 06722dc4-a9ea-472f-beba-ce93f12056e7 · outbound

This paper cites Unsupervised learning of visual representations by solving jigsaw puzzles.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Unsupervised learning of visual representations by solving jigsaw puzzles

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:00.989479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:00.989479Z digest=sha256:86c4cf618df542638290021152e7dc49b3c88af2a9a10d1d5cd46dca10ff9b85

Observation 63e30dd3-80c0-4541-b335-0e6e2a759c11 · outbound

This paper cites Colorization as a proxy task for visual understanding.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Colorization as a proxy task for visual understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.842145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:00.994362Z digest=sha256:c1df2ba48a4557139d38fb422293c3e74dbd37dd49ebe6020f7cb494fc072672

Observation 1b4a9d23-d3ff-4d1a-ab2a-89b4ac9b4460 · outbound

This paper cites Self-supervised representation learning by rotation feature decoupling.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Self-supervised representation learning by rotation feature decoupling

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.826670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:00.998873Z digest=sha256:8a2b2ccbc4bf1382dea1666a9e99cb1c6b2889cc00e0d8b348d0134c8da9446d

Observation 0090d0f4-6a32-498e-bf27-7f1752ca3a0b · outbound

This paper cites What makes instance discrimination good for transfer learning?.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine What makes instance discrimination good for transfer learning?

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:30:01.367879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.003745Z digest=sha256:9d527f9c535cd022a366a35c59bc203f5a65bd1c28f2650a5e37617cc235525e

Observation d4d30880-b7a7-407a-a716-b7c9dd9545f4 · outbound

This paper cites Scale-mae: A scale-aware masked autoencoder for multiscale geospatial representation learning.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Scale-mae: A scale-aware masked autoencoder for multiscale geospatial representation learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.812120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.009754Z digest=sha256:c1b74ffa256f7ff5ae7be4a042bb45a47c8d8a12f5dcd28daa4e5a94544ec607

Observation d8a30164-549d-4257-b6ae-c5d0719a8c23 · outbound

This paper cites Attention is all you need.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Attention is all you need

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.014686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.014686Z digest=sha256:78de82a2d70b22911264534ce8cf0908a3093d6babb90eed3e3519bab30de9a2

Observation bff68324-f635-4ec9-9917-50600d642e81 · outbound

This paper cites Language models are few-shot learners.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Language models are few-shot learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.019806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.019806Z digest=sha256:cd004b5083fea7500b89cef94a1e3ed00dafe597549b65f83fa12374963db307

Observation 742dac01-52cc-4341-9119-f9f86ba32dc5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine LLaMA: Open and Efficient Foundation Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.024294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.024294Z digest=sha256:318168f02fdfe7f485bc809b8250731162842937a82b609b1f40c5e25ccbf73b

Observation 5967b54d-c824-42c1-b5bc-2245925714b9 · outbound

This paper cites Med3D: Transfer Learning for 3D Medical Image Analysis.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Med3D: Transfer Learning for 3D Medical Image Analysis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.029512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.029512Z digest=sha256:e8cfbd260e73955004d406ff4b4cbafe20071413721f9f02f7f7a6a65a7e05e7

Observation 91e4ed54-2f39-4b4a-9c5e-f58f477b6416 · outbound

This paper cites Models genesis.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Models genesis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.780656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.034025Z digest=sha256:a5c1c345d5062461ac3a5e14ae14b9962c4ec57419f86742e0759a69216b9b9b

Observation 724bb59d-0b2c-45c6-99ab-a656e72b9bd5 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.038158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.038158Z digest=sha256:bd087a6314f627e3a131636835c8e96cd205d4d58c4576a0507a38076e795285

Observation bac93bcb-f821-48f2-8ff5-f2c549c8a74f · outbound

This paper cites LXMERT: Learning Cross-Modality Encoder Representations from Transformers.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.043454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.043454Z digest=sha256:00499b6c52aa66250b2ed118ee3e6767cf478cf9c2302e90b75cd14cebb31029

Observation d0d91a95-a7b6-46a2-ac37-2c3d84f12e06 · outbound

This paper cites Uniter: Universal image-text representation learning.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Uniter: Universal image-text representation learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.757157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.048116Z digest=sha256:92903e236d7ab73a8c3a0889b895a312d3da39679699c84301bce6be1e43ce00

Observation f6b30fb9-6e03-4c3a-add7-a7fcd66b4145 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.743642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.053213Z digest=sha256:14871725660250d1c7abe003a4d15a1becd999790a43db5e03e31faceb0bba93

Observation a5ce9f45-656d-4090-a193-445a6a8a5b1b · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Vinvl: Revisiting visual representations in vision-language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.057338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.057338Z digest=sha256:d0ed4907f7347ffcc0b83e0b1d9d0dbdaf8030d02f3daf7eda787438ffab529e

Observation edff5fdb-e67d-4bf3-a270-1ef58d6df557 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models, 2022.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Coca: Contrastive captioners are image-text foundation models, 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.062249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.062249Z digest=sha256:2c9beece27285a10ae0ff102aea60ce19af32afbae66ac15669374958c848f0b

Observation 55eeb42c-144f-4743-adf4-d5521a0efe40 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.066457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.066457Z digest=sha256:4b3bf408c6b877114172ef1e25e7815b7fdc4d249a5911007fd2cbe93483e827

Observation 65a2707a-9740-4ef6-b6b8-a553c277d1ab · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.071454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.071454Z digest=sha256:62ffc1d5981e194f6793ca74e49fb9b8f4e0a077d112b323683c9b13278bec2f

Observation 6df0aa4f-a65e-41d4-a663-6f98aff101bf · outbound

This paper cites Multi-modal understanding and generation for medical images and text via vision-language pre-training.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Multi-modal understanding and generation for medical images and text via vision-language pre-training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.694139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.075870Z digest=sha256:e7035f8a3fcebdff420174fe4166948b701e661a2c1575703939f3bead59a902

Observation ff93f537-8c21-40d3-a176-520255c67df5 · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.080927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.080927Z digest=sha256:638c3ecaaaecbd7a4a88bf281787c1f5a81ccce12b28d6980e4a2a8370b3ae47

Observation fa7cdee5-57b6-4bea-8382-678291a191c2 · outbound

This paper cites Slip: Self-supervision meets language- image pre-training.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Slip: Self-supervision meets language- image pre-training

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.678960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.085437Z digest=sha256:c93820b86726697e9bc749e00df58e4934860201c801c2d188cc0fc20192d89b

Observation 30c66838-3b1e-4226-9be9-6e0896c0c7c8 · outbound

This paper cites Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm, 2022.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm, 2022

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.664933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.089740Z digest=sha256:98781f7156079e1991a967b657698989ee6e46c8edfe664c942e3cddd2728ea9

Observation e0504519-2856-4003-a9bc-3a893792e187 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Swin transformer: Hierarchical vision transformer using shifted windows

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.094663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.094663Z digest=sha256:dd6da4584d1062a5a0d87d2a2f81300d964fc00e724571088165ea107b8e0375

Observation 35a0962c-c2a3-47dc-a28a-8495e21b17fd · outbound

This paper cites Self-supervised pre-training of swin transformers for 3d medical image analysis.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Self-supervised pre-training of swin transformers for 3d medical image analysis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.099089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.099089Z digest=sha256:132790a3067bee0f0915f58076dc71ad5af9a43a4a31c458d102fb70d8d4a91a

Observation d36c25de-c9bd-462e-9e61-ddc1f30d4e7f · outbound

This paper cites Masked image modeling advances 3d medical image analysis, 2022.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Masked image modeling advances 3d medical image analysis, 2022

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.633279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.103864Z digest=sha256:99bc82b9ba23ee76a44a72dcf9777ff6276433e5c851e5b682d3f7e79cba66f4

Observation a1337afe-146a-4800-bfa3-7a63abe36566 · outbound

This paper cites V oco: A simple-yet-effective volume contrastive learning framework for 3d medical image analysis.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine V oco: A simple-yet-effective volume contrastive learning framework for 3d medical image analysis

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.619401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.108158Z digest=sha256:90415ab35370c3091c278c8497e5baba9fb1c1f43991095bd73099d0c36f2ea8

Observation 39c3030c-6260-4ec2-bd81-ceaee3ac0f57 · outbound

This paper cites Context Encoders: Feature Learning by Inpainting.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Context Encoders: Feature Learning by Inpainting

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.113224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.113224Z digest=sha256:66d734d7c1115ea306358ea746e62aca7166955366e97eee49dbc2dc2fc513b7

Observation d126cba3-332c-4507-9895-8e13a2929756 · outbound

This paper cites Unsupervised Representation Learning by Predicting Image Rotations.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Unsupervised Representation Learning by Predicting Image Rotations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.118205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.118205Z digest=sha256:310fdb19d2c1459de54ce039d25eff9c0792c69c5cd00a524ab7fb95429c07c4

Observation 18bc6f70-c1eb-4cf3-8c96-c2732c5712ac · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Representation Learning with Contrastive Predictive Coding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.122950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.122950Z digest=sha256:6620151e29e7d0297bcdbcbcc58825c378e67a3e1338f2f9e60ec9fcfbcc2060

Observation 0f456fc9-a844-4b98-8b47-c3c06dd978f3 · outbound

This paper cites Fast WordPiece Tokenization.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Fast WordPiece Tokenization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.128489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.128489Z digest=sha256:4352786bb3e1511c83723f2a3af110784e8bd591cd70bc71da321fe762e63fca

Observation e9a74340-58c9-4891-88f7-ce9118f5469c · outbound

This paper cites Zero-shot text-to-image generation.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Zero-shot text-to-image generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.604840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.133289Z digest=sha256:053a0bba47f7e483f7424acc25f03f03421d24130b104dd90442b9db757de20e

Observation 35db7e1a-a920-464b-accd-d190d1e28d70 · outbound

This paper cites Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.138329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.138329Z digest=sha256:06c137f28dda76a3818362c486f7c1a0ec503a0d0dbdcde19823dec1ecc9dc08

Observation 5231ee45-4b53-4ccf-bf1c-777711b32a05 · outbound

This paper cites Abdomenct-1k: Is abdominal organ segmentation a solved problem? IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):6695–6714, 2022.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Abdomenct-1k: Is abdominal organ segmentation a solved problem? IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):6695–6714, 2022

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.581111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.142837Z digest=sha256:6d2be1a211d0d28fc6c8ed3ac9aa1d3577e897e92e416440a2369e007921034d

Observation ee2307a5-8a02-497e-bdf9-dbeedff77eff · outbound

This paper cites Ct-org, a new dataset for multiple organ segmentation in computed tomography.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Ct-org, a new dataset for multiple organ segmentation in computed tomography

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.566777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.148002Z digest=sha256:85a7c04e2cdece06edabc151d32cf95618cc217a50bb1e9a1305369dcb58fc63

Observation 517d1b0f-fc32-4797-9814-9daa32543141 · outbound

This paper cites Ledsam, and Olaf Ronneberger.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Ledsam, and Olaf Ronneberger

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.553169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.152348Z digest=sha256:06b5f56677fb65dfed610e42724e83e5b73c62c4a0c7409a1c3f6a2c02c4d260

Observation 0419b32a-6f0d-4f36-92ed-505400bf3365 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Bleu: a method for automatic evaluation of machine translation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.539265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.157227Z digest=sha256:62b63ebf945efaf136f4ba31b5f921fe68bab6101bc87c3abd3f3829ab522a8a

Observation 6787a5d8-402d-4bd3-92a5-6ae847698a9b · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Rouge: A package for automatic evaluation of summaries

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.524066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.161539Z digest=sha256:b1c652b6d8bd25b00acf4ef714243299b930b83a6c407e8467a66696889d68d1

Observation 759c3292-5139-4c5c-b49f-4d11bdce6f35 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine BERTScore: Evaluating Text Generation with BERT

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.166429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.166429Z digest=sha256:687e269d93f97b51cf7a1862acebae285768a32d8ba9a0a99f0739d0f8d9801a

Observation b728e4d2-5e7e-44e3-bf75-bf7ca624c657 · outbound

This paper cites Decoupled weight decay regularization, 2019.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Decoupled weight decay regularization, 2019

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.170711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.170711Z digest=sha256:78539b563a7f7b08fd7312d6f2c50f2c49873a2b9b164eb92a4ca8fd3ff825c5

Observation 77ac4f29-0498-4704-82b8-6e3cb1cf4238 · outbound

This paper cites https://huggingface.co/ContactDoctor/Bio- Medical-Llama-3-8B, 2024.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine https://huggingface.co/ContactDoctor/Bio- Medical-Llama-3-8B, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.500789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.175442Z digest=sha256:34b29541dcffe8396d54fcd982e23ca38b093eb92e94a9519187be93b3442bfd

Observation 0e4abb36-f73f-4b20-98c6-4184023dd95c · outbound

This paper cites Lora: Low-rank adaptation of large language models.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Lora: Low-rank adaptation of large language models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.179662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.179662Z digest=sha256:6c5951ac509aed68f224b3a5d0f1b7c9a98e42ea80dcfb59a725e0329e0180cd

Observation ff94ef9a-2544-4cf9-8091-f84bc85c13bd · outbound

This paper cites Segment anything.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Segment anything

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:01.184855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:30:01.184855Z digest=sha256:b2a4d866c0f9d06bacc393c604caf0e7cfa576a4b937f0b4a9e7a250dfb16866

Observation 807c9353-43ba-4bfa-8ccb-a1f472aea0fc · outbound

This paper cites Segvol: Universal and interactive volumetric medical image segmentation.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Segvol: Universal and interactive volumetric medical image segmentation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.467821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.189214Z digest=sha256:b25a47cc941456821f037dde3f35ffceb10a75e07464d25326ac20f5a1c99a73

Observation 70085536-0c4e-4680-b283-50c93e0ab6bd · outbound

This paper cites Segment anything in medical images.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Segment anything in medical images

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.453474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.193525Z digest=sha256:9fb1a05685054f94a6412f023160f9317dc782b30be73047e983e3c6eb38f5f6

Observation d7497ca5-54bc-45f8-9024-6fce54bf1fea · outbound

This paper cites Pmc- clip: Contrastive language-image pre-training using biomedical documents.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Pmc- clip: Contrastive language-image pre-training using biomedical documents

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:30:01.438351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.198485Z digest=sha256:6de1b513438f4e05c9d79ad3925e415e7ba0c949e812634e94344cf26eb2c97f

Observation f449df5a-836c-4ed5-8d9d-055db0c5d757 · outbound

This paper cites Accelerate: Training and inference at scale made simple, efficient and adaptable.

VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine Accelerate: Training and inference at scale made simple, efficient and adaptable

Reference 68

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T17:30:01.424454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:30:01.202763Z digest=sha256:aa991649faa2959d7bde33aab1fff61442875e5eab9c9317f29a86d7d780d60b

Pith citing papers

No inbound Pith citation observations are available.