Pith. sign in

Paper Citation Record · LEDGER

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation

As of 15 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2506.19665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19665 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:32:12.368913Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:54:41.933749Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T14:04:44.782937Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 367f08d5-a61b-4022-9354-8d86cc2fbe2c · outbound

This paper cites Generating Radiology Reports via Memory-driven Transformer,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Generating Radiology Reports via Memory-driven Transformer,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:13.168414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.134889Z digest=sha256:fd8ed8c3e13db03ce01a031d12b1c267f21e4a059148c1049f23ca50170015dd

Observation 28fb95c4-5a4d-4f79-9133-897db0be870e · outbound

This paper cites Ratchet: Med- ical transformer for chest x-ray diagnosis and reporting,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Ratchet: Med- ical transformer for chest x-ray diagnosis and reporting,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:13.153635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.140671Z digest=sha256:1b05190d8dc9d62d27bcd72a8e8db8a6be972b078072f39386e1d0522839f65f

Observation 1969a769-dc09-4cff-9912-09564feac391 · outbound

This paper cites XrayGPT: Chest radiographs summarization using large medical vision-language models,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation XrayGPT: Chest radiographs summarization using large medical vision-language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:13.136583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.145862Z digest=sha256:b5dc572c53659ebb696320b8234352b905e470845f25f688c39352b0e54b0f27

Observation 180e9ed2-6054-4456-8ba0-2e7398734ee2 · outbound

This paper cites Automated pul- monary nodule detection in ct images using deep convolutional neural networks,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Automated pul- monary nodule detection in ct images using deep convolutional neural networks,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:13.119633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.150610Z digest=sha256:3aa9710d547cc5365fc545f26e9483dfafaefca859ff18fed600c6c84a22c13f

Observation da48fa2a-df88-4086-bfed-01a3911aadd3 · outbound

This paper cites Unet++: Redesigning skip connections to exploit multiscale features in image segmentation,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Unet++: Redesigning skip connections to exploit multiscale features in image segmentation,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.155593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.155593Z digest=sha256:e95c0ffd9a571d8cbaeba2d442f11730dd281e3bd4238db7a320bc6f5d345978

Observation 125d1a6d-9d89-4ebb-9403-9ec3eb97f8c5 · outbound

This paper cites Unet 3+: A full-scale connected unet for medical image segmentation,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Unet 3+: A full-scale connected unet for medical image segmentation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:13.095917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.160776Z digest=sha256:b689b40475bd7eacddbdb5a23ce29f85d0cfd09ca84f023fd910b398eea92ac4

Observation 9cd81f06-d292-4024-bde1-30e12691d85f · outbound

This paper cites Sanet: A slice-aware network for pulmonary nodule detection,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Sanet: A slice-aware network for pulmonary nodule detection,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:13.078691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.165822Z digest=sha256:7cca06e3e559ecfd99214c7ac2800ba53ec8fc065577c7b3fb7139a1f46b30fa

Observation cd436671-ec74-430e-98a9-970cdf21850d · outbound

This paper cites Lung nodule classi- fication of ct images based on the deep learning algorithms,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Lung nodule classi- fication of ct images based on the deep learning algorithms,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:13.061052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.170558Z digest=sha256:204abd0da8df8a60c8f0978a4656ecaf3e9203a3085900b5088649e88a684ee9

Observation d58d259d-d401-4a1b-a1e5-f3e757615fc7 · outbound

This paper cites Semi-supervised deep transfer learning for benign-malignant diagnosis of pulmonary nodules in chest ct images,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Semi-supervised deep transfer learning for benign-malignant diagnosis of pulmonary nodules in chest ct images,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:13.044603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.175080Z digest=sha256:758591db89cffde3a548e7d817f6920dd18f57089986d02f30bb37dd1302b14a

Observation a47d74a5-4f11-4fe8-9561-17c82b38dec0 · outbound

This paper cites Medical-vlbert: Medical visual language bert for covid-19 ct report generation with alternate learning,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Medical-vlbert: Medical visual language bert for covid-19 ct report generation with alternate learning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:13.027359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.179933Z digest=sha256:04d8c342fbbb4a0fffb45ae74a8673aa320aa9296c95e2ecb1cb35da09eba803

Observation cbba84c2-3204-4ade-821f-ea9ee0d6bdbd · outbound

This paper cites Swin-unet: Unet-like pure transformer for medical image segmentation,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Swin-unet: Unet-like pure transformer for medical image segmentation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:13.011806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.184755Z digest=sha256:ca3663cbb3c3c18d452a110181ac3b7a48e1a73fa12b0cce4b5483a5f63ced6a

Observation 96de92c1-6407-4ab0-a594-3305520295d2 · outbound

This paper cites A multilevel transfer learning tech- nique and lstm framework for generating medical captions for limited ct and dbt images,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation A multilevel transfer learning tech- nique and lstm framework for generating medical captions for limited ct and dbt images,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.994063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.189513Z digest=sha256:797667198762ec8db75f3b0bfef9d8e4b503de93f140add642ed047620d5f265

Observation 64adb230-7e95-43b0-9dea-13c51be53730 · outbound

This paper cites 3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation 3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.193990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.193990Z digest=sha256:a638babc7eb3cbe41012e92f657e2793e0c82db9e732259f163b4a4723a0920d

Observation 56244bfb-ed67-45a2-8988-2f80e98ebebd · outbound

This paper cites Dia-LLaMA: Towards Large Language Model-driven CT Report Generation.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Dia-LLaMA: Towards Large Language Model-driven CT Report Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.198866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.198866Z digest=sha256:93dc0f9f8ee8a3dac71f333600c6af76cae23530c83af99d3f493a87e2bb8899

Observation 798aebe8-d62b-4fb3-a650-7e1e5543e530 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.203692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.203692Z digest=sha256:b2f3334945597cd2622f6bd07317eace988b361f8535df6d717b14d17024cf7c

Observation c59326e6-4a23-44a0-9c05-3ed2fb40b3f9 · outbound

This paper cites CT2Rep: Automated radi- ology report generation for 3d medical imaging,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation CT2Rep: Automated radi- ology report generation for 3d medical imaging,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.978635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.209920Z digest=sha256:9d4f3a599a1bb8738714c56e65e8b2044eb063831cc6affe641e0cf05a5d5c96

Observation af163ed8-d53e-4773-99a3-e5323732fd2d · outbound

This paper cites Generative adversarial network with robust discrimina- tor through multi-task learning for low-dose ct denoising,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Generative adversarial network with robust discrimina- tor through multi-task learning for low-dose ct denoising,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.964142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.214814Z digest=sha256:699669cfc3763fe9d22facbb37d8d36c73d832bb162e63eedd05924b5621d27a

Observation f627e1b8-aa3f-480c-a2f7-fb4346f16c3a · outbound

This paper cites Ct-agrg: Automated abnormality-guided report gen- eration from 3d chest ct volumes,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Ct-agrg: Automated abnormality-guided report gen- eration from 3d chest ct volumes,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.218634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.218634Z digest=sha256:93980f6ecce8e59eefdc9ad119a77462da36d7f1fa4a1882fc376e95b0ab1979

Observation fc50a145-4d10-4287-994e-6ca53b752b03 · outbound

This paper cites Explainable artificial intelligence (xai) in deep learning-based medical image analysis,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Explainable artificial intelligence (xai) in deep learning-based medical image analysis,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.949968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.223133Z digest=sha256:a78ce9a9780078ce2fee93f8ab6219a1900a95053d8358066f93d722bcaa36e3

Observation 76f4dd0a-0b75-410a-8c9e-d1100f39cb3e · outbound

This paper cites Granularity matters: pathological graph-driven cross-modal alignment for brain ct re- port generation,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Granularity matters: pathological graph-driven cross-modal alignment for brain ct re- port generation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.936427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.227035Z digest=sha256:80c0a3e0c6b6ca416cd45e825937675641d2ababbc7ee0cc564f9b9e3ae568b4

Observation 19bcddfe-177b-4ab6-8a62-b77db346b3d3 · outbound

This paper cites Co-occurrence rela- tionship driven hierarchical attention network for brain ct report generation,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Co-occurrence rela- tionship driven hierarchical attention network for brain ct report generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.923687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.231662Z digest=sha256:57747e6afe9babd35c870182e3ff8bfbb6b48b521a0821a1707947cf674cdc47

Observation da325d47-66f0-4861-b1e8-ed5de4f01339 · outbound

This paper cites MedXChat: A Unified Multimodal Large Language Model Framework towards CXRs Understanding and Generation.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation MedXChat: A Unified Multimodal Large Language Model Framework towards CXRs Understanding and Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.237088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.237088Z digest=sha256:1d9a79aa14c1de8d3f5a49665e7f8605e4adb0b21adebb3717de328b7e2cd15e

Observation 649d54e8-ed4f-48ad-8ca4-503837bd8ee2 · outbound

This paper cites M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.241470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.241470Z digest=sha256:90ddd71810d1a1753ca8d637393c2eb3b526a32166764fcc6838677ca1e14c16

Observation a2d7e8e9-b6a6-45b2-a811-2f070e7df6ef · outbound

This paper cites 3d u-net: learning dense volumetric segmentation from sparse annotation,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation 3d u-net: learning dense volumetric segmentation from sparse annotation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.245487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.245487Z digest=sha256:b0d592b929569e3d612e4ec40ee49930c75f199796e587173242667fe4c5cf6c

Observation 5757f897-6832-4329-b9cb-3e663e37f500 · outbound

This paper cites Unetr: Transformers for 3d medical image segmentation,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Unetr: Transformers for 3d medical image segmentation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.900222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.249438Z digest=sha256:c41fd72dcbf199a4d1e7ccf4b157e95088d192bdad61d516ca1b72b1283a3e79

Observation fc3708e2-e4ff-417b-a52c-63add6bed9fc · outbound

This paper cites Swin unet3d: a three-dimensional medical image segmentation network combining vision transformer and convolution,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Swin unet3d: a three-dimensional medical image segmentation network combining vision transformer and convolution,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.884895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.253850Z digest=sha256:917d292c63b5892b6aaccb881fc38a74fddafe082f0a06eccfb2e29295c82af1

Observation 81006820-01af-4204-b949-980076001d63 · outbound

This paper cites AutoRG-Brain: Grounded Report Generation for Brain MRI.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation AutoRG-Brain: Grounded Report Generation for Brain MRI

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.258658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.258658Z digest=sha256:b21408aa6024bf13f9eb38554bbced9d0d4e28ba8eff36639a4df8d46594d007

Observation a548735d-fc80-40d2-ab10-601b05493a4e · outbound

This paper cites Unetr++: delving into efficient and accurate 3d medical image segmentation,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Unetr++: delving into efficient and accurate 3d medical image segmentation,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.263691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.263691Z digest=sha256:6d5d32e5a059f1106b692b9371d7e6f72e1fb56e1ff020b522d9f28ee844a6db

Observation 41e7a6e5-aa89-40a0-a3c0-fc494ad37cf0 · outbound

This paper cites Generalized zero-shot chest x-ray diagnosis through trait-guided multi-view semantic embedding with self- training,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Generalized zero-shot chest x-ray diagnosis through trait-guided multi-view semantic embedding with self- training,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.859997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.269093Z digest=sha256:18fbf5befe7258eeab3cc02f998ec92ad3b1eafce241e806076378e3944e4f4f

Observation 071448da-0079-4beb-93d2-e8d96be952a9 · outbound

This paper cites A Systematic Review of Deep Learning-based Research on Radiology Report Generation.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation A Systematic Review of Deep Learning-based Research on Radiology Report Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.273289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.273289Z digest=sha256:331602dc939720c1f6d49f2a4b8d067abd1d7af65cc5d4a5c84ef72e641c8a24

Observation 3ba19b73-e693-430b-a337-0640381fc447 · outbound

This paper cites Diffusion networks with task-specific noise control for radiology report generation,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Diffusion networks with task-specific noise control for radiology report generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.842508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.277980Z digest=sha256:c340d58bbd453c0441558c888fee37d9cc0b2001b3779321c243ebd0e3cf0586

Observation b4cad25a-6fb7-4a63-81d4-39219785f6ea · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.284220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.284220Z digest=sha256:d8357796954975cbd5a32d53a5cf8f3e58bbfec73a41d659758e1df031fc0214

Observation 81ab1e01-997c-4b21-a498-da0787ed8ff0 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.289541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.289541Z digest=sha256:2b2df57ca8488d391b72e142aa2b1b3f41663b8c419feed25a19df5347e77a2d

Observation a0eda7da-c545-45be-8f70-d3240eb44e52 · outbound

This paper cites Attention is All You Need,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Attention is All You Need,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.825075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.294311Z digest=sha256:9505d0423d41d36da3068165f0fb2bf9a4389a9bb878a02964f53c84a57eed58

Observation 9a924e7b-e134-4523-a7a8-bdb8b3ce4d2c · outbound

This paper cites Attention on Atten- tion for Image Captioning,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Attention on Atten- tion for Image Captioning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.807991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.299524Z digest=sha256:f109cde723f11a504357d132aa65f261d644873e8d40b7d813b38a0a41b0da40

Observation 5aef8628-f471-4698-91cf-bd08b2c018c8 · outbound

This paper cites Supertagging Combinatory Categorial Grammar with Attentive Graph Convolutional Networks,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Supertagging Combinatory Categorial Grammar with Attentive Graph Convolutional Networks,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.305076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.305076Z digest=sha256:1a3989814450505c11d1fade5cab0c755ff53f98b2f1cd193c2850fed0c10dad

Observation 1cbfe6b6-ef2d-4a1b-aef6-0bdaf524b62d · outbound

This paper cites Learning multimodal contrast with cross-modal memory and reinforced contrast recognition,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Learning multimodal contrast with cross-modal memory and reinforced contrast recognition,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.782532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.310747Z digest=sha256:8997ae0f136bf7f3089bf9dfafb71e293d96ebe8cd15ff7de081a20a00800bf0

Observation 75b37336-6549-4e4c-9a10-6193af4b7727 · outbound

This paper cites Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.315843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.315843Z digest=sha256:6bf925fdd4466437925794d8e2f6ff8ce8d4df1da181cc1269266e07fe362db4

Observation 9340a571-be55-44bc-9d91-42260ff55e62 · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Efficient Estimation of Word Representations in Vector Space

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.320567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.320567Z digest=sha256:8a094f12f701d1e91eac25e4a38f3da400414f153213c1de0d1499c64f2810c6

Observation a7bf9c1b-a9d5-4814-b749-cb4ca7ca8aee · outbound

This paper cites Learning Word Representations with Regularization from Prior Knowledge,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Learning Word Representations with Regularization from Prior Knowledge,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.769189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.325404Z digest=sha256:60d6c87b8e16f415e0c6c1cf9a4a539e25c5d6daf127e62df0ca8455a3fc440f

Observation 967e2c36-9cfe-4a0c-b389-0cca2b88e994 · outbound

This paper cites Joint Learning Embeddings for Chinese Words and Their Components via Ladder Structured Networks,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Joint Learning Embeddings for Chinese Words and Their Components via Ladder Structured Networks,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.755551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.329299Z digest=sha256:3e44e67ff2dacc43d13f0c1e34d607850133dcaa3865b22ad2db4b5686cc16ba

Observation cd0c300e-3cb2-49dd-91fb-3527ceaa4a56 · outbound

This paper cites BERT: Pre- training of Deep Bidirectional Transformers for Language Un- derstanding,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation BERT: Pre- training of Deep Bidirectional Transformers for Language Un- derstanding,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.742569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.333480Z digest=sha256:f3686bee542ffdbf5642c265fe1267c7d919fe51778945462de34b92f97528a2

Observation 4748c049-45e7-4336-abd9-d105b77bcd31 · outbound

This paper cites BART: Denoising sequence-to-sequence pre-training for natural language genera- tion, translation, and comprehension,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation BART: Denoising sequence-to-sequence pre-training for natural language genera- tion, translation, and comprehension,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.728219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.337350Z digest=sha256:7bc991cf847b533557d63dfa3b99fb9c31adae7b6559ffd3faa0df0c548dcd7d

Observation d3d8c2a8-66b8-49c7-b3ec-31807c5dc49e · outbound

This paper cites Llava-Med: Training a large language-and-vision assistant for biomedicine in one day,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Llava-Med: Training a large language-and-vision assistant for biomedicine in one day,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.343975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.343975Z digest=sha256:ef5cc03bd14cf72096086f9e9d09dfebd07350b13b779371bda1bb985cfb2b5d

Observation b39ee149-08f4-471a-8ef0-e2dcbde7fd37 · outbound

This paper cites ChiMed-GPT: A Chinese medical large language model with full training regime and better alignment to human preferences,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation ChiMed-GPT: A Chinese medical large language model with full training regime and better alignment to human preferences,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:32:12.704547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T18:32:12.348069Z digest=sha256:122de5f60eb89f433ba233a5c1c0eed1f4f6bdd3a8976718ad6a3e2fcd21d807

Observation a71f6404-88c7-41b3-9eb4-6791e729f2a2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.352437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.352437Z digest=sha256:893dbacead410adc984caa6d39dff9b7c215b1b6fab1a4214a37b2fb7524e701

Observation d615b444-1c9a-43a4-9e3f-2fb4d4f395af · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.355978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.355978Z digest=sha256:29c7c7c012b6ca90115af81c9d03aa9ca617d8f34694f84955b99ff2e4fdfd77

Observation 96cdba50-ffd2-4e65-83aa-38ea572e8f9c · outbound

This paper cites Representation Decomposition for Learning Similarity and Contrastness Across Modalities for Affective Computing.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Representation Decomposition for Learning Similarity and Contrastness Across Modalities for Affective Computing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.359954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.359954Z digest=sha256:a4864bde7b6c834a640c6afd9432f0f56bb7377a663122b4fdab93b3296ba73a

Observation cdc9654a-2f32-4f42-985d-8243acb6a2c7 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.364655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.364655Z digest=sha256:34ce24ca15b76bb539a50a6dda7642539739aa05e7526d830bcf33e868c75905

Observation 00197f94-90f0-428f-891d-e79198f9ffc8 · outbound

This paper cites Vivit: A video vision transformer,.

Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation Vivit: A video vision transformer,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:12.368913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:12.368913Z digest=sha256:7c0b4e04d4db7b51ad8b0a4363fd655ef159adad713aa178885bee51d897fc65

Pith citing papers

Observation 867720fa-cb2a-42bc-b8c2-e00437be774d · inbound

Computed Tomography Visual Question Answering with Cross-modal Feature Graphing cites this paper.

Computed Tomography Visual Question Answering with Cross-modal Feature Graphing Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:41.933749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:41.933749Z digest=sha256:7738e302693205141ebe739b8221b50cc7e89698d6038db998e77d5100834458

Observation 74970003-7cf3-4931-a301-d5b8be4844fc · inbound

SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation cites this paper.

SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:04:44.784670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T13:57:09.053990Z digest=sha256:94b68984a382984ae47719a7645d34ca557b6a977ec12cc1f3b2169c8d931da9