Pith. sign in

Paper Citation Record · LEDGER

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment

As of 21 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2608.04472.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04472 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:45:56.691342Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c7b839a-5f6a-46f7-926e-88ccdc5516a4 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment Representation Learning with Contrastive Predictive Coding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:45:56.664842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:45:56.664842Z digest=sha256:d87b707a281261e21fa1e3555f7241c6c5c74071674d5fa31c7263af5326c5e0

Observation ea2445f4-c9e6-4f0c-956c-8c79e69e8dd4 · outbound

This paper cites Qwen3 Technical Report.

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment Qwen3 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:45:56.687064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:45:56.687064Z digest=sha256:dbd1b2dd8120b8b49eced2a340321f31ae2ab2b7f234c76797aa4066d9f1b794

Observation 9cb512ef-835b-42e4-8c5d-1012110b70cf · outbound

This paper cites BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:45:56.691342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:45:56.691342Z digest=sha256:7372d9e3c769238a33b693688ea1b050fb8e03f2beb0e7636de1001c4f365a73

Observation 64c8acf4-c212-4730-aaea-96e47a917db7 · outbound

This paper cites DINOv3.

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment DINOv3

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-15T14:45:56.678372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:45:56.678372Z digest=sha256:6a2ae1b1e8489383cafc82d4dc7adf3a88a97992ff377c185b998f03ec35011c

Observation 9ab9a584-4293-4151-9189-77e8358e8c17 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment On the Opportunities and Risks of Foundation Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-15T14:45:56.639316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:45:56.639316Z digest=sha256:2d515c0a542fe8de0d05fefe179a7b7129d1c45c1bd4204dfb8943262dfb759a

Observation 9a763d7e-a112-4b96-b613-4a89530f22cc · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment DINOv2: Learning Robust Visual Features without Supervision

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T14:45:56.669203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:45:56.669203Z digest=sha256:08b5de432c10724c42552892ac66c814fbdc0fe92e1f86585c66b98100019ea8

Observation 03ac2f40-6232-4ac3-9d10-5397d805875b · outbound

This paper cites EndoDINO: A Foundation Model for GI Endoscopy.

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment EndoDINO: A Foundation Model for GI Endoscopy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T14:45:56.646310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:45:56.646310Z digest=sha256:f5e2107f4155a33c2a81224d9b59fa19cb470a02e59002da9c648b75e5ab6bae

Observation c900c5eb-72cf-41cc-8856-d5c940aeb14f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T14:45:56.656052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:45:56.656052Z digest=sha256:81274dfd0d8bbe16466801356981ea0c783c9b4a7a6854694d14adc1100937cd

Observation f2959058-8677-4505-ba65-375a5f7d8584 · outbound

This paper cites A benchmark for endoluminal scene segmentation of colonoscopy images.

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment A benchmark for endoluminal scene segmentation of colonoscopy images

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:45:56.819851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:45:56.682540Z digest=sha256:140b65a0ac56a3b52e1ac59800807dfc98de534b13220e1180569e290a1f97ca

Observation 0763bb18-70fb-4868-b129-5e583f9242b4 · outbound

This paper cites Tumorchain: Interleaved multimodal chain-of-thought reasoning for traceable clinical tumor analysis.

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment Tumorchain: Interleaved multimodal chain-of-thought reasoning for traceable clinical tumor analysis

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:45:56.831698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T14:45:56.660693Z digest=sha256:ea8bfe1e1e837136738187d533a500c2c69ba57869a3f3d105e9071fd455182a

Observation 7c653307-518c-40f0-9a8f-0fe6ee6cbba9 · outbound

This paper cites Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding.

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T14:45:56.673709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:45:56.673709Z digest=sha256:0ad79ffb063027c13c3876c1f10cb045b49ba0749bcdbb51933a2b88658b5280

Observation ca0a022a-e05b-4742-8c6c-d43052d58f91 · outbound

This paper cites Polyp-PVT: Polyp Segmentation with Pyramid Vision Transformers.

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment Polyp-PVT: Polyp Segmentation with Pyramid Vision Transformers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T14:45:56.651117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:45:56.651117Z digest=sha256:ba299828e2b2d45270724215a2d7cbd694025a3b35fa370d832794877a698a91

Pith citing papers

No inbound Pith citation observations are available.