Pith. sign in

Paper Citation Record · LEDGER

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding

As of 18 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2506.19288.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19288 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:59.323308Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:19:24.701282Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy30
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3528a83-a7ee-40c9-8223-310854cfe123 · outbound

This paper cites A data set for airborne maritime surveillance environments,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding A data set for airborne maritime surveillance environments,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:06.929626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:54.380544Z digest=sha256:40fdbc73008d94b80cfffc251cbf0b26d7c31ed42343eef0125e013ab3b56a58

Observation 27844cef-f1f5-49bf-89ae-bd96dc7abad6 · outbound

This paper cites Asy-vrnet: Waterway panoptic driving perception model based on asymmetric fair fusion of vision and 4d mmwave radar,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Asy-vrnet: Waterway panoptic driving perception model based on asymmetric fair fusion of vision and 4d mmwave radar,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:06.623178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:54.500267Z digest=sha256:197f441b29c54a97d971003b4ca97a06fa226b0e0dc296ce9bbd636dae23816e

Observation d187f314-9867-4cd8-92e4-20129d9884a0 · outbound

This paper cites Usvtrack: Usv-based 4d radar-camera tracking dataset for autonomous driving in inland waterways,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Usvtrack: Usv-based 4d radar-camera tracking dataset for autonomous driving in inland waterways,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:06.412323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:54.565772Z digest=sha256:052d9d7c691927145e1e97129f9f692094ba510ba0ebc47c1af46d1c98e833b6

Observation 58f25f48-1449-4749-baeb-add1f0d6227c · outbound

This paper cites Are we ready for unmanned surface vehicles in inland waterways? the usvinland multisensor dataset and benchmark,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Are we ready for unmanned surface vehicles in inland waterways? the usvinland multisensor dataset and benchmark,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:06.214649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:54.693238Z digest=sha256:85ba79f54af652f863a8b28a09eec933106c3c77f4729b4c52299a493bd068fb

Observation 2c3b71b9-e01b-4816-bcc3-40a397e49e88 · outbound

This paper cites Achelous: A fast unified water-surface panoptic perception framework based on fusion of monocular camera and 4d mmwave radar,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Achelous: A fast unified water-surface panoptic perception framework based on fusion of monocular camera and 4d mmwave radar,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:06.026304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:54.824633Z digest=sha256:124ae27fd4084c105db642560350e00a4000c898610e0bf85c78dd63fcec6697

Observation 79cf704c-20eb-4558-a91b-4e448572d1cc · outbound

This paper cites Achelous++: Power-Oriented Water-Surface Panoptic Perception Framework on Edge Devices based on Vision-Radar Fusion and Pruning of Heterogeneous Modalities.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Achelous++: Power-Oriented Water-Surface Panoptic Perception Framework on Edge Devices based on Vision-Radar Fusion and Pruning of Heterogeneous Modalities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:54.910034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:54.910034Z digest=sha256:15b6e95529a0648a3b1a2eba9da5535caddb9d2f07fbdc7f262337780e1363ab

Observation 0576f8b1-8e07-4bea-8e14-96544a8f7a1f · outbound

This paper cites Watervg: Waterway visual grounding based on text-guided vision and mmwave radar,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Watervg: Waterway visual grounding based on text-guided vision and mmwave radar,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:05.860607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:55.019750Z digest=sha256:ea15e3491553524d07b3c07208177b572e638f88f49413eca9d900e535528caa

Observation 7f149857-30d3-4942-a92c-43a369d6c990 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:55.097495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:55.097495Z digest=sha256:014980493b808dcd7236a050f8e82b2c3a79254213682da117c6ea6d29b9e8f4

Observation 431507a3-2ede-4617-b2b5-ea98e04d71f4 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:05.709641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:55.164644Z digest=sha256:a30ed4e07a69a719b8531f47cd355a6e18489a3ec9e32c1a4ce2b491cf522225

Observation 7334db13-52ec-418f-87ab-77411f5f0147 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:05.546474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:55.220892Z digest=sha256:97bf1deee9c360b58503481c0e2175a8ba23af3af270595a0c8f97c367c52372

Observation 601eddcf-5424-4c4d-a648-0289034f962a · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Im2text: Describing images using 1 million captioned photographs,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:05.365422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:55.270503Z digest=sha256:61af7407dffb41dffd9587578d7b7693c7665d996afcec476c8d9365c9f02a89

Observation da03011a-f677-40e5-b2df-e04dad77a848 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Vizwiz grand challenge: Answering visual questions from blind people,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:05.181603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:55.337293Z digest=sha256:d49a1c249b4d837df554124c578d4960865d23bc0adde1a32e9ae6bcdc455eee

Observation e1d16594-950c-45b2-9659-bccd1f4ee483 · outbound

This paper cites Learning deep represen- tations of fine-grained visual descriptions,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Learning deep represen- tations of fine-grained visual descriptions,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:04.957616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:55.390881Z digest=sha256:bdf3f941469b28de66725fa0b428989f4ab30ee9f8ab78d35783d43900031dc1

Observation 2cb29d9f-7359-49d0-8fde-146b32fce0fe · outbound

This paper cites Fashion captioning: Towards generating accurate descrip- tions with semantic rewards,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Fashion captioning: Towards generating accurate descrip- tions with semantic rewards,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:04.757576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:55.446462Z digest=sha256:3d11b661eacfcdab1063417d3ec1b12a7fb13c3f718d77d4371491c8957e8884

Observation 4022af13-e19d-43dc-918c-2f1726fbd04d · outbound

This paper cites Break- ingnews: Article annotation by image and text processing,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Break- ingnews: Article annotation by image and text processing,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:04.514738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:55.500107Z digest=sha256:69cb13c5410aa6768fc1f83d3dcba1d33866c19346db9963eda77d00bfb96c46

Observation 75037dd3-a52a-4951-b344-c6a3f57d8f03 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Textcaps: a dataset for image captioning with reading comprehension,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:04.264772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:55.578720Z digest=sha256:8c0445719a839aaf0f2956822ded2511970d33916257dab29a494d0d1b98a5b7

Observation 329a93c2-2979-4e66-b914-a3310f0f000e · outbound

This paper cites Deep learning approaches on image captioning: A review,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Deep learning approaches on image captioning: A review,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:04.048418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:55.637589Z digest=sha256:778ea87fe624137fc04464f652067ba3df2441718d1a08cbbbb45708aac9a515

Observation 5ebd6dec-ff85-48d5-ba1d-f98c01e19e20 · outbound

This paper cites Waterscenes: A multi-task 4d radar-camera fusion dataset and benchmarks for autonomous driving on water surfaces,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Waterscenes: A multi-task 4d radar-camera fusion dataset and benchmarks for autonomous driving on water surfaces,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:03.821704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:55.692726Z digest=sha256:e647f3124593007a51de347b4bc068e843369d33a8cbe1787e9c35071df32069

Observation 8a1acd54-b7bc-4ea2-8881-6048d843805c · outbound

This paper cites an unresolved cited work.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:13:03.672962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:55.735704Z digest=sha256:35b3fc0a16f9acd4b2fb4c1ef8679e37051dc5d58907bde45a8280e8f6df07ec

Observation c198008a-e337-447b-8511-0a07ceeeca1c · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:55.793428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:55.793428Z digest=sha256:1a3f63113e74f3bc5154034ec6658f549f1d4bd7f3b6de3f902e3196ad8cd713

Observation dfaff453-8f72-497a-862d-28849e747998 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Learning transferable visual models from natural language supervision,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:55.845170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:55.845170Z digest=sha256:df62ab633223b5e2cfd7ee6b233fcb8317e5ea4bdf9716568f4f9877aa7e24ee

Observation a0c4bf0c-1a5d-4af1-aaf4-e86de4796e50 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Align before fuse: Vision and language representation learning with momentum distillation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:55.902203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:55.902203Z digest=sha256:21c8c0aecfcaecf584ca0e929f77736c0c7fa76e42452bcee22999cf1eeb140f

Observation 615558af-e3db-48dd-94ec-a9b5986425d3 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:55.972188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:55.972188Z digest=sha256:5ef647c369f781f8166b9d92287e89547fe2560c1e4243738923877c4524ecd9

Observation 8d91ef8b-87e7-4f03-b628-b2e580385b65 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:56.022808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:56.022808Z digest=sha256:64d9387c265e12d3cbc929f554f2f1baaa37215a57d97d8ad5dd75d3f01bdeaf

Observation 811bea4f-c2ec-423b-9955-40aacfb6ac4f · outbound

This paper cites Visual instruction tuning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Visual instruction tuning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:56.086648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:56.086648Z digest=sha256:617005bde877a8f9e4516acacfc9ef20208faacfe3cc76367f5e1de2125b9dad

Observation 7e759b75-38a1-4828-a366-2cae7131c08d · outbound

This paper cites Qwen2. 5 technical report,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Qwen2. 5 technical report,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:03.389808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:56.155414Z digest=sha256:cbf73b513aebf989cf582991f92d52ee9cf1609924d6e9844cdffb57dc2d28f6

Observation 3174bdf5-d3e6-42cf-a494-b862fee32ba3 · outbound

This paper cites Task-adaptive attention for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Task-adaptive attention for image captioning,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:56.258705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:56.258705Z digest=sha256:f7524bead8a3b31967e2efb1511a1cd8a60d77aa466b88beed357d57665e1aca

Observation f5901940-4ad3-43da-97f9-3d76d4d4e981 · outbound

This paper cites Vision-enhanced and consensus- aware transformer for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Vision-enhanced and consensus- aware transformer for image captioning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:03.193848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:56.355939Z digest=sha256:78718c0e112d003003d7af7116892f3a032d6ad106ce7e32812edb224a90649f

Observation b1cb9332-8d86-436e-8410-4221571f69a7 · outbound

This paper cites Adaptive path selection for dynamic image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Adaptive path selection for dynamic image captioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:02.958702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:56.434598Z digest=sha256:3f86deeda502635cadec78dd9d833059b1d1f5706d153546823fd67fae906ece

Observation c81c2a21-60cd-4941-b0fa-f8d52be3c83e · outbound

This paper cites Double-stream position learning trans- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 14 former network for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Double-stream position learning trans- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 14 former network for image captioning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:02.744958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:56.534168Z digest=sha256:7d42bc808991b60ac743fc651258d02b7af1a49e4e8c60c2442f2383c4eb4050

Observation 4ebe8d1e-bf2b-449c-8e2e-150ce11c4437 · outbound

This paper cites A comprehen- sive survey of 3d dense captioning: Localizing and describing objects in 3d scenes,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding A comprehen- sive survey of 3d dense captioning: Localizing and describing objects in 3d scenes,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:02.556920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:56.603663Z digest=sha256:c14d43690d9fea044eee184ef1b311e6ece3531267958b5a511f4556be7264b2

Observation 31978f6c-b1d4-4a7c-a8c2-79ae2febe81b · outbound

This paper cites Spt: Spatial pyramid transformer for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Spt: Spatial pyramid transformer for image captioning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:02.375032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:56.743713Z digest=sha256:65ac51da2a7c16abfe4c84e311274fe558d848e7426e5b4007a349e9348f38d2

Observation c05fa789-e8b7-46ba-b3e6-d75d53bcf8d7 · outbound

This paper cites Multimodal transformer with multi- view visual representation for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Multimodal transformer with multi- view visual representation for image captioning,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:56.928074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:56.928074Z digest=sha256:9ec6a279afb75c741d63bdf141119e4530d784e7fe2167ee0d49d2b21c748862

Observation ccabe1f6-81d6-4c7e-9296-bd3d434b7c5b · outbound

This paper cites MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:56.998779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:56.998779Z digest=sha256:efe0825f330311d475c196bd2228a4acd72a2e135f00c5b8f5d709e7874db33c

Observation c5e489b2-72b8-42ec-b315-0eda9e0dcf3a · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:02.224750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:57.061781Z digest=sha256:f980211501a4535dbc2618280dac55c5115eb9fced8ed6072c80b7912bb766fa

Observation 500fdbab-b469-46d2-af42-cf3f814ed487 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:57.142952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:57.142952Z digest=sha256:4dddaba41816dbe1a9b52f4c11587866091163a923e7c40bed12a09b283fb3f1

Observation edbf4cf7-badf-40a5-b9ba-4172a7c575d7 · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:57.273153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:57.273153Z digest=sha256:ac50e577f05f6294328a0bc5fa60b918abbb9f88d434db052536dadc5bcaae34

Observation aa0ae569-7531-44af-8e29-9d3b379ef496 · outbound

This paper cites Mask-vrdet: A robust riverway panoptic perception model based on dual graph fusion of vision and 4d mmwave radar,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Mask-vrdet: A robust riverway panoptic perception model based on dual graph fusion of vision and 4d mmwave radar,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:01.997200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:57.364159Z digest=sha256:b6e883f222815d7c075e2389c87489ea9579fa343f1ac265006773c03fd497a2

Observation af2b816b-3b25-4b3c-b476-0dcaae6dfda0 · outbound

This paper cites NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave Radar.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave Radar

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:59.949657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:57.454821Z digest=sha256:27817868f41cd8f7b4d1b45de3317514a0ec8713e902a4ca92e9ef46c7535641

Observation 6cab96c3-08cf-4e27-9026-62cc3e78ca61 · outbound

This paper cites GPT-4 Technical Report.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding GPT-4 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:57.527835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:57.527835Z digest=sha256:fe4474a592c7394e2ba38fb78040a607968727d46652dba8b3f1dbfea0e787ca

Observation 081bb90d-1ec1-4d9d-bdec-b68e315506ea · outbound

This paper cites DeepSeek-V3 Technical Report.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding DeepSeek-V3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:57.627616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:57.627616Z digest=sha256:aa4d4bb079ec06f135814f89e5bea3c00ed325df06fd9e8b7ddc6cb68ce9a7f1

Observation 2111c1d1-a8db-441d-b219-6cc38c9c150c · outbound

This paper cites Mobileclip: Fast image-text models through multi-modal reinforced training,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Mobileclip: Fast image-text models through multi-modal reinforced training,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:01.851137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:57.757659Z digest=sha256:a2e23b55c4d4c571ccc2a7af649138db76c57cd31995241c327f3f538c2e249d

Observation 0c5fce76-134f-4e97-a4e0-9e70b75f8a98 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Coca: Contrastive captioners are image-text foundation models,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:57.830163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:57.830163Z digest=sha256:6e6aa92c8de9e2c4ba3c5ba6c870a17025e75b268ee58f165ad5e794e6784f9c

Observation d61b9442-439c-4671-96b3-98a58e1265df · outbound

This paper cites Sigmoid loss for language image pre-training,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Sigmoid loss for language image pre-training,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:57.922031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:57.922031Z digest=sha256:5bd6fea3edba45332fcc0e18e4fa8b39a6c34feee114eb2465292c8f49913a41

Observation 7bd0ff1f-2783-425c-8f03-d20c7b35d236 · outbound

This paper cites Agent attention: On the integration of softmax and linear attention,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Agent attention: On the integration of softmax and linear attention,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:01.649115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:58.027415Z digest=sha256:bc36e46130129a86a662e13b72a9e3589415545b5afec112b8a7518a43ab0d1a

Observation abb2ff5e-e04c-496d-9c7a-99f5e45d3fe4 · outbound

This paper cites Dual-level collaborative transformer for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Dual-level collaborative transformer for image captioning,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:01.360157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:58.117598Z digest=sha256:865df7e98cc484b11ab6ba5997c779679d831c2142f7ace2684e8105626f280c

Observation 810a7100-652f-4612-bdfb-ef681593103d · outbound

This paper cites Comprehending and ordering seman- tics for image captioning,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Comprehending and ordering seman- tics for image captioning,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:01.148678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:58.239788Z digest=sha256:cd4d22c0c3b7d736a2300a62182e722c59a949012eb2e43c9d37dc94a3aa485b

Observation c1a5c21e-a9df-45cd-b725-08b75466a7fa · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Oscar: Object-semantics aligned pre-training for vision-language tasks,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:58.335734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:58.335734Z digest=sha256:0e84f2db3f567bb81fcafff794e4f967ebf0d221594a71f8ffb4490dc0d2841f

Observation 145b0466-f1ea-4ed3-9f35-537488884268 · outbound

This paper cites an unresolved cited work.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:58.477687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:58.477687Z digest=sha256:43bb204b6b0e7f6f78ec71c6befb932ab0b32ecdcfd1cdd93ee188a3ec7e5a14

Observation 666e3eec-e2e9-46bd-8dd8-258b15ae58fe · outbound

This paper cites Qwen2.5-VL Technical Report.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Qwen2.5-VL Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:58.578001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:58.578001Z digest=sha256:5618cec34ff02f8001ef40021ea1fae540b4fba3ccd44108a53f44a14009423b

Observation 427794b7-6134-4495-99f0-437837990883 · outbound

This paper cites MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:59.596017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:58.675967Z digest=sha256:18789371dd8e10e9ecab546ef185ee22c91bf4db7287428909d5bada217cc873

Observation 44a722cf-0c9f-4a35-94af-64a5b07bfb53 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Rouge: A package for automatic evaluation of summaries,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:58.785881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:58.785881Z digest=sha256:9080b71e79de5a5cd48b32b9b0e90e9da7545e8d89e127f4a0c79cb273286abb

Observation 976e26ed-e61f-4926-a8a6-318f66543f8d · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Bleu: a method for automatic evaluation of machine translation,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:58.879043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:58.879043Z digest=sha256:766e6c985e8d1b38cbeee4cefc8a7e1e3d5b78c7b3195ea01fc75f9e2328ce78

Observation 410c64ab-1944-44a5-91c3-6d9a08ea2364 · outbound

This paper cites Lavie and A.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Lavie and A

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:00.851698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:59.001834Z digest=sha256:55824f74dbf8319cc0fd42c15c9bd23376fad0efe4e08f9a6fbde846bf6c5d23

Observation 7157aa19-ca8c-463a-8a93-0f58f8e2020a · outbound

This paper cites Cider: Consensus- based image description evaluation,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Cider: Consensus- based image description evaluation,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:59.106904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:59.106904Z digest=sha256:0234e7d2f911de837d7114656a17b933055f782d6270c4dd93c27d92469ea35f

Observation 1f34d844-ea05-4991-95c6-f0c3e2dfafa7 · outbound

This paper cites Drivelm: Driving with graph visual question answering,.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Drivelm: Driving with graph visual question answering,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:00.559148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:59.233791Z digest=sha256:744dabfa8cf90f1e68abc8adedd5c339214c0793a1d2cc1d91b05782efc384a6

Observation 53883036-0972-4f8f-85de-8dcd25d81229 · outbound

This paper cites an unresolved cited work.

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:13:00.327800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T23:12:59.323308Z digest=sha256:246be3c1aadaa7f9a2aea8ac8b5d669b7af2361d15a0412aaf461230a7a9bf2f

Pith citing papers

Observation 9fdddfd6-4ace-4f4d-9d43-1eb7af353c64 · inbound

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey cites this paper.

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T19:19:24.701282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:19:24.701282Z digest=sha256:af295de453807f370aca03091f7c682e1ce594934cbed0b082c6ad904e24bf72