Pith. sign in

Paper Citation Record · LEDGER

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2507.07006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07006 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:55:42.909470Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy12
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9932535-b6c4-47bc-a3f4-c420b54bb7be · outbound

This paper cites Inference of captions from histopatho- logical patches,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Inference of captions from histopatho- logical patches,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.712752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.712752Z digest=sha256:16184e65f6d73913f2530b256bb0a4e69efce19ca3267ddb72bb5eb3de9fbbea

Observation 739cf3c5-fac6-455f-a33e-610924a0e4a3 · outbound

This paper cites Artificial intelligence for digital and compu- tational pathology,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Artificial intelligence for digital and compu- tational pathology,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.438872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.719966Z digest=sha256:e2cd95ea7e76029e793b44d576f2f93e043753977928e825d79b493f8fac760b

Observation 2139d102-de3e-4115-acf9-04a56fdd76db · outbound

This paper cites PathAlign: A vision-language model for whole slide images in histopathology.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning PathAlign: A vision-language model for whole slide images in histopathology

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.731503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.731503Z digest=sha256:188161ec7c6b6b1616d387efbb54ebaa1761d2793b2553c043c69c46a29f7c98

Observation 7befc8a2-bffe-4772-b5eb-76a5584d584c · outbound

This paper cites Attention-based deep multiple instance learning,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Attention-based deep multiple instance learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.422584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.741430Z digest=sha256:4415abc5ec9c89aa67e1e7f537a5672052a5f9e0f246c7c18941f464a9c6e67a

Observation 728e977a-1ac4-40dd-ad78-842a2b4fe611 · outbound

This paper cites Transmil: Transformer based correlated multiple instance learning for whole slide image classification,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Transmil: Transformer based correlated multiple instance learning for whole slide image classification,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.403249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.749496Z digest=sha256:8dc4e91cb5b21b1fb1ee23fba50d1ee26ee52182c8713ce2b06c6aac6726cb56

Observation 6f2a7822-8f9b-4d26-8dac-0a14827ebf07 · outbound

This paper cites A survey on graph-based deep learning for computational histopathology,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning A survey on graph-based deep learning for computational histopathology,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.384796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.757346Z digest=sha256:adda721dea67f7b1005be2cf4f8a47c25edd9adb533f8d1cab0d070f2d7d01a4

Observation cf583242-523d-4b31-a16a-a31f1ad19426 · outbound

This paper cites Enhanced descriptive captioning model for histopathological patches,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Enhanced descriptive captioning model for histopathological patches,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.368148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.765499Z digest=sha256:74f96321e531f988dad9ce3314e89a50e315e463ff5911cb16e2cffb8dd42663

Observation 7b837ee8-7723-49bf-8f44-306d52212c12 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning LLaMA: Open and Efficient Foundation Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.773499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.773499Z digest=sha256:2b75c8b69814d02920b7562861afbb1a1a7537ab95ad4190adffe396c8815df7

Observation 02665d5e-0043-4810-b3c6-686424d71821 · outbound

This paper cites Clinicalt5: A generative language model for clinical text,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Clinicalt5: A generative language model for clinical text,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.780443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.780443Z digest=sha256:689f0e33159e52727ef8a5eac6037edc5fcf582fbb02101ce33e8c4c02b0a970

Observation b8c755d5-f623-452f-9812-c7db0bba8579 · outbound

This paper cites Biogpt: generative pre-trained transformer for biomedical text genera- tion and mining,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Biogpt: generative pre-trained transformer for biomedical text genera- tion and mining,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.787842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.787842Z digest=sha256:aeef8bf85f5bbbfe53e4031a86ec0ed8d50c7bdba1245c3cc04608da03241d44

Observation 6efab8ef-5069-41e8-84b1-25c65161e0c8 · outbound

This paper cites A generalist vision–language foundation model for diverse biomedical tasks,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning A generalist vision–language foundation model for diverse biomedical tasks,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.326042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.793771Z digest=sha256:6dfe2ac38ff2a5e889dccbc30d5d4caee7dfba66b485de3562d103de2c361d29

Observation 65261c9b-15de-480b-a638-794e88d896c9 · outbound

This paper cites Imagenet large scale visual recognition challenge,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Imagenet large scale visual recognition challenge,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.799479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.799479Z digest=sha256:55196bfda18c48e8e63b2dc6f986523599f893825738299a533754bfe6f704a9

Observation e266c0f8-8819-4563-a81e-561eb8c19ea5 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Learning transferable visual models from natural language supervision,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.807114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.807114Z digest=sha256:84decf21fa171dbd8e96c6a99a138889eaf56d08a5673097802c09fe4ca5928d

Observation 48256ba8-a2cf-45c3-8012-278cf66882fb · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.812086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.812086Z digest=sha256:e07f7f07efa0e69923cce7a81da9f9031031c52c2b4f5956b568a991928fd781

Observation 495defc2-7de4-4268-b6b0-fb972fd55273 · outbound

This paper cites PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:55:43.039262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.817005Z digest=sha256:6cdaa994988729c275a7e76c475d96da536e741326628430e6d59bfbe89288d2

Observation f7dc63f0-fc8b-47cf-8f76-817d84553ec2 · outbound

This paper cites Dual-stream multiple instance learn- ing network for whole slide image classification with self-supervised contrastive learning,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Dual-stream multiple instance learn- ing network for whole slide image classification with self-supervised contrastive learning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.827274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.827274Z digest=sha256:8bf0c5782312eab71d442f79d20807e9658c1a03c84e7dac0161ece532adc53c

Observation b4a3ca2f-297d-4129-bd08-f09cd1b4717c · outbound

This paper cites Dtfd-mil: Double-tier feature distillation multiple instance learning for histopathology whole slide image classification,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Dtfd-mil: Double-tier feature distillation multiple instance learning for histopathology whole slide image classification,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.833688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.833688Z digest=sha256:711d43c2414272c044c3342d2cea8748344ff1347155ba52902b34ed6773aa75

Observation 97216124-38e4-4580-af9d-2cfcef796500 · outbound

This paper cites Visual language pretrained multiple instance zero-shot transfer for histopathology images,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Visual language pretrained multiple instance zero-shot transfer for histopathology images,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.244937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.839603Z digest=sha256:f2770661362925e256fec2c91964fa16df2a1f12903bd7b8d5011b829f932443

Observation 5109a3e4-46d6-4c05-8712-f734fc9d1172 · outbound

This paper cites CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.848632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.848632Z digest=sha256:d23f477df700dc9be7b0e9f7100669b988142ab26b23473bd442455bc7844c16

Observation 2658958b-a776-418b-bdd7-870f4b9e42d3 · outbound

This paper cites Large language models in healthcare and medical domain: A review,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Large language models in healthcare and medical domain: A review,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.221309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.856878Z digest=sha256:ecaea1fc203521b1dfc6de8d6ea4618598ff9958bda8fc2f16fdbf8bbb763e44

Observation 02e84621-ff0b-464c-b8bd-93d432a6204c · outbound

This paper cites What a Whole Slide Image Can Tell? Subtype-guided Masked Transformer for Pathological Image Captioning.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning What a Whole Slide Image Can Tell? Subtype-guided Masked Transformer for Pathological Image Captioning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:55:42.992013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.863470Z digest=sha256:ca2a0b912e7047fd67293909b8c1bdce87a2797059795ad2381aea8be200bfa3

Observation 1de277b8-78c8-49df-8c52-4116cf10e648 · outbound

This paper cites Patch-based convolutional neural network for whole slide tissue image classification,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Patch-based convolutional neural network for whole slide tissue image classification,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.197120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.869493Z digest=sha256:88c29a74bc6971a1b7af65798cf4bc8faefda2ffe0b41f8ff64acf33de25c1f0

Observation 68dadbd5-0921-47af-b719-846e4ebc45d4 · outbound

This paper cites Unsupervised deep embedding for clustering analysis,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Unsupervised deep embedding for clustering analysis,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.173062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.874812Z digest=sha256:f7340309435703b015bcdd5d13c9255778aba15a2dce3a828a4279e71873939d

Observation 424ea3b9-5158-4c0b-b301-59e52794ab7d · outbound

This paper cites Attention is all you need,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Attention is all you need,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.879763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.879763Z digest=sha256:a90810d3d3b3c13986da8a514654758742dbfcec199969a01f79b1bcf0527d94

Observation 3b64879b-1ced-4796-877f-520141799fea · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Categorical Reparameterization with Gumbel-Softmax

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.885997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.885997Z digest=sha256:6858619f17caef5bb78b28d17176db661e283874ff78428acc43377caeddd4ed

Observation 7e0bb6d2-6446-4a3e-9a87-5c3cbc466655 · outbound

This paper cites Graph attention networks,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Graph attention networks,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.893311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.893311Z digest=sha256:6195d631efed9db3c578b41be88c4df1d997313bfd1a20b7055fe0717891c591

Observation 18ce9874-2ea8-4468-a38a-e9a3918433be · outbound

This paper cites Rectifier nonlinearities improve neural network acoustic models,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Rectifier nonlinearities improve neural network acoustic models,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.898607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.898607Z digest=sha256:3cae9ea1e3bfbe17d0157445ec8ee4de3fec20216150bc2058a74f3095cf5a75

Observation 54c78102-d8bc-4517-8bc0-9e5faaa68d60 · outbound

This paper cites A dataset for breast cancer histopathological image classification,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning A dataset for breast cancer histopathological image classification,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.122386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.904383Z digest=sha256:1c217521b55b7ab30c35117eefcc010062fa51faa00082028c0daf655393d203

Observation be848d36-dd34-4587-b09c-cf81813c94c5 · outbound

This paper cites A survey of evaluation metrics used for nlg systems,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning A survey of evaluation metrics used for nlg systems,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.099400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:55:42.909470Z digest=sha256:0e9a68f0c384fd775b347898fd27b97d372202a9f56ac0a82e57da706b68db73

Pith citing papers

No inbound Pith citation observations are available.